Friday, November 30, 2012

Dashboards and Big Data: Before Fruit Ninja, Cybernetics


Will Wiles, The New York Times, November 29, 2012

THIS spring, Fraser Nelson, editor of The Spectator, chastised David Cameron, Britain’s prime minister, for spending too much time on his iPad. “One of his senior advisers says the P.M. spends ‘a crazy, scary amount of time playing Fruit Ninja,’ ” Mr. Nelson wrote in The Daily Telegraph. Such is Mr. Cameron’s devotion that he ordered the creation of an iPad app that would allow him to monitor the British economy.

According to the Web site of the Cabinet Office, the application is now complete. Called the No. 10 Dashboard, after the prime minister’s residence, it gives “the prime minister, other ministers, and senior Whitehall officials an at-a-glance overview of everything that’s happening in government and elsewhere” — stock prices, housing and jobs data, information on the performance of government departments, as well as the “political context”: polls, commentary and indications of the national mood from sources like Twitter. The prime minister liked having “quotable facts about what was going on,” said Alice Newton, one of the app’s developers.

The site included an example of the kind of data the app offers — a graph of Britain’s stuttering gross domestic product — something you’d expect Mr. Cameron to be able to see whenever he closes his eyes, let alone on his iPad.

The No. 10 Dashboard could be the White House Dashboard; Mr. Cameron plans to show the app to President Obama at the Group of 8 summit meeting.

The whole story is bathed in the white heat of 21st-century digital technology. “Ours is the first generation of digital natives coming through to work in government,” Ms. Newton said. But there’s nothing new about it. If Mr. Obama follows in Mr. Cameron’s footsteps, he should know that he’s also on the trail of another leader — the ill-fated socialist president of Chile, Salvador Allende.

What’s the connection between men so widely separated by ideology, geography and time? After assuming power in 1970, Allende’s administration was faced with economic paralysis, inequality and civil strife. The usual socialist prescription of nationalization of industry was applied, but Allende also looked to an unexpected quarter: the relatively new and niche science of cybernetics. Cybernetics is the study of control and communication in large, complex systems, be they organisms, machines or organizations. It spans management theory, information technology, psychology, biology and sociology. The Chilean government approached a British cybernetician named Stafford Beer and asked him to build a cybernetic hub for the management of the country’s economy — something that had never been attempted before, or since.

A recent history, “Cybernetic Revolutionaries,” by Eden Medina, gives the subsequent project, Cybersyn, the serious attention it deserves. The system, once up and running, would have channeled data from Chile’s nationalized industries into an operations room in Santiago, where Allende’s ministers would have made informed decisions in chairs with control panels built into the armrests. Photos of this room still retain a veneer of giddy futurity. White surfaces, pared-down interfaces, rounded corners — a dash of Apple in the mix.

Ms. Medina contends that rather than a tool of technocratic centralization, Cybersyn was democratic in design and intent. Workers would have access to the heart of government and could inform decision makers. The aim was transparency and economic and social homeostasis, a natural equilibrium. Mr. Beer called the system the “Liberty Machine.” But Allende’s regime, overthrown by a C.I.A.-backed military coup in 1973, didn’t last long enough to see Cybersyn operate.

The No. 10 Dashboard taps into the same desire to master available information — a desire that has only grown as the amount of information in circulation has increased. Where Cybersyn needed dedicated national infrastructure and rooms full of equipment, the app runs on a hand-held device. And yet, the dashboard is actually less sophisticated.

It is not truly cybernetic because it lacks a mechanism to translate all that data into action. It can display information; it cannot consult and control. Less a driver’s dashboard, it is more a window out of which a passenger can observe the national scenery speeding past. Government might be given the illusion of hand-held, one-stop manageability, but no actual managing is going on.

The app could thus be an apt metaphor for politicians reduced to spectators by the surges and shocks of the globalized world. Mr. Cameron should remember that there’s at least one other instance of government-by-app: the team that worked on restructuring Greece’s debt used iPads too, equipped with an app purpose-built for the job.

However that turns out, we can at least say this: in terms of distractions, these apps are marginally more useful than Fruit Ninja.

Will Wiles is the author of the novel “Care of Wooden Floors.”



Thursday, November 29, 2012

Nate Silver: In Silicon Valley, Technology Talent Gap Threatens G.O.P. Campaigns


Nate Silver, The New Yorks Times Five Thirty Eight Blog, November 28, 2012

SAN FRANCISCO – I live in Brooklyn, where President Obama won 81 percent of the vote this month. It’s hard to find anywhere in the country that is more Democratic-leaning.

But San Francisco qualifies. Here, Mr. Obama won 84 percent of the vote, while Mitt Romney took just 13 percent. Even John McCain, who won 14 percent of the vote four years ago, performed slightly better than Mr. Romney did.

And unlike the New York metropolitan area, where Long Island, the borough of Staten Island and many suburbs in New York and New Jersey remain competitive in presidential elections, it is hard to find any significant pockets of support for Republican candidates in the nine counties that make up the San Francisco Bay Area.

Instead, Mr. Obama won the nine counties of the Bay Area by margins ranging from 25 percentage points (in Napa County) to 71 percentage points (in the city and county of San Francisco). In Santa Clara County, home to much of the Silicon Valley, the margin was 42 percentage points.

Over all, Mr. Obama won the election by 49 percentage points in the Bay Area, more than double his 22-point margin throughout California.

Although San Francisco, Oakland and Berkeley have long been liberal havens, the rest of the region has not always been so. In 1980, Ronald Reagan won the Bay Area vote over all, along with seven of its nine counties. George H.W. Bush won Napa County in 1988.

Republicans have lost every county in the region by a double-digit margin since then. But Democratic margins have become more and more emphatic. Mr. Obama’s 49-point margin throughout the Bay Area this year was considerably larger than Al Gore’s 34-point win in 2000, for example, or Bill Clinton’s 31-point win in 1992.

Even without the Bay Area’s vote, Democrats would still be favored to win California by solid margins. So why does any of this matter?


The reason is that Democrats’ strength in the region is hard to separate out from the growth of its core industry — information technology – and the advantage that having access to the most talented individuals working in the field could provide to Democratic campaigns.

Companies like Google and Apple do not have their own precincts on Election Day. However, it is possible to make some inferences about just how overwhelmingly Democratic are the employees at these companies, based on fund-raising data. (The Federal Election Commission requires that donors to presidential campaigns disclose their employer when they make a campaign contribution.)

Among employees who work for Google, Mr. Obama received about $720,000 in itemized contributions this year, compared with only $25,000 for Mr. Romney. That means that Mr. Obama collected almost 97 percent of the money between the two major candidates.

Apple employees gave 91 percent of their dollars to Mr. Obama. At eBay, Mr. Obama received 89 percent of the money from employees.

Over all, among the 10 American-based information technology companies on Fortune’s list of “most admired companies,” Mr. Obama raised 83 percent of the funds between the two major party candidates.

Mr. Obama’s popularity among the staff at these companies holds even for those which are not headquartered in California. About 81 percent of contributions at Microsoft, which is headquartered in Redmond, Wash., went to Mr. Obama. So did 77 percent of those at I.B.M., which is based in Armonk, N.Y.

It does not require an algorithm to deduce that the sort of employees who may be willing to donate substantial money to a political campaign may also be those who would consider working for it.

Since Democrats had the support of 80 percent or 90 percent of the best and brightest minds in the information technology field, it shouldn’t be surprising that Mr. Obama’s information technology infrastructure was viewed as state-of-the-art exemplary, whereas everyone from Republican volunteers to Silicon Valley journalists have criticized Mr. Romney’s systems. Mr. Romney’s get-out-the-vote application, Project Orca, is widely viewed as having failed on Election Day, perhaps contributing to a disappointing Republican turnout.

This is not intended to absolve Mr. Romney and his campaign entirely. There were undoubtedly many bright and talented information technology professionals who worked for Mr. Romney, and who might have fielded a better product given better management.

Even if only 10 percent or 20 percent of elite information technology professionals would consider working for a Republican like Mr. Romney, this is still a reasonably large talent pool to draw from.

But Democrats are drawing from a much larger group of potential staff and volunteers in Silicon Valley.

Perhaps a different type of Republican candidate, one whose views on social policy were more in line with the tolerant and multicultural values of the Bay Area, and the youthful cultures of the leading companies here, could gather more support among information technology professionals.

Ron Paul, the libertarian-leaning Republican, raised about $42,000 from Google employees, considerably more than Mr. Romney did.




Wednesday, November 28, 2012

The Five Forces Shaping the 21st Century


Vivek Ranadivé, Forbes, November 27, 2012

The 21st century didn’t start in the year 2000. It started in 2010, the same way the 20th century began in 1908 with the advent of the automobile. It became the century of highways and freeways, the century of the auto—the American century. Similarly, if you look at what happened a couple of years ago, there were all kinds of crossover points that happened around the same time: more cell phones than landlines, more laptops than desktops, more debit cards than credit cards, more farmed fish than wild fish, more girls in college than boys.


I am dedicated to the belief that if you get the right information to the right place at the right time and in the right context, you can make the world a better place. This is something I call the two-second advantage. In order to achieve that, you need to understand five forces shaping the 21st century.

The first is the massive explosion of data. Look at all the data that was created from the beginning of mankind until a couple of years ago. Since then, ten times as much data has been created. Think about it: 10 times as much data just in the last two years then in all of history. The amount of video content that will go up on YouTube today will be far more than all that Hollywood has created since its inception.

The second force is the rise of mobility. It took 100 years for there to be a billion landlines, 10 years for there to be a billion cell phones and just one year for there to be a billion smart cell phones. We live in a time where everyone on the planet will have one of these smart cell phones.

The third is the emergence of platforms—social, cloud and so on. It used to be that if you wanted to reach an audience of millions or tens of millions, you had to be a large corporation. Today, platforms like YouTube, the iPhone app store and Facebook, allow individuals to reach massive global audiences. Just recently a Korean rapper put his song up on YouTube and it became a global phenomenon within days. That’s the third factor to consider… how do you leverage the platforms that are out there?

The fourth is the rise of Asia. A few hundred years ago, India and China were about two-thirds of the world’s economy. Most economists predict that at some point in the next century, we will revert to that same state. That is because anything that can be done in India and China will be done in India and China. Any 21st-century strategy must take that into account.

The fifth and final force is that Math is trumping Science. If the 20th century was the century of Science, I believe the 21st century will be the century of Math. When I say “Math trumping Science,” I mean that you no longer have to know the why of something, you have to know the what. You simply have to know that if A and B happen, then C will happen; you have to find the pattern. For years, AIDS researchers tried to find the secret of how the AIDS virus mutated and they were not able to. About a year ago, they converted it into a Math problem and put it into a game called Foldit. Within a week, gamers had found the answer—something scientists had not been able to find for years.

Anyone can build their business by understanding and harnessing these five forces and utilizing the two-second advantage. Think outside the box, be passionate, innovative and creative, and you, too, can make the world a better place.


Immelt: The Future of the Internet is Intelligent Machines


Jeff Immelt, GigaOm, November 28, 2012 

Thanks to the internet, almost anything consumers might want is just a click away. But for businesses, the gains have been much less dramatic. That is about to change, with the arrival of the Industrial Internet.

As we know it today, the internet has been largely about connecting people to information, people to people, and people to business. Monetization strategies range as widely as the options available, and for all the success, there are more failures. While many of the advancements have been extraordinary – even unthinkable a short time ago – too often we’re still left asking, “to what end?”

The internet can give consumers nearly anything with just a click, but global economies remain challenged.  The internet has become the biggest library in the world, but education is just now beginning to take advantage and change.  The internet can provide businesses with unprecedented data, but true insight remains contentious and change is slow.

The real opportunity for change is still ahead of us, surpassing the magnitude of the development and adoption of the consumer internet. It is what we call the “Industrial Internet,” an open, global network that connects people, data and machines. The Industrial Internet is aimed at advancing the critical industries that power, move and treat the world.

There are now many millions of machines across the world, ranging from simple electric motors to highly advanced MRI machines. There are tens of thousands of fleets of sophisticated machinery, ranging from power plants that produce electricity to aircraft that move people and cargo around the world. There are thousands of complex networks ranging from power grids to railroad systems, which tie machines and fleets together.

This vast physical world of machines, facilities, fleets and networks can more deeply merge with the connectivity, big data and analytics of the digital world. This is what the Industrial Internet Revolution is all about.

Productivity Revolution
The Industrial Internet leverages the power of the cloud to connect machines embedded with sensors and sophisticated software to other machines (and to us) so we can extract data, make sense of it and find meaning where it did not exist before. Machines – from jet engines to gas turbines to CT scanners – will have the analytical intelligence to self-diagnose and self-correct. They will be able to deliver the right information to the right people, all in real time. When machines can sense conditions and communicate, they become instruments of understanding. They create knowledge from which we can act quickly, saving money and producing better outcomes.

As an example, we have pushed the boundaries of physical and material sciences in our aircraft engines to the point where these engines are more powerful and efficient than ever. We will continue to improve them physically, but at the same time we can use software, monitoring and big data analytics to attack the $284 billion in annual waste in the airline industry that is caused by fuel inefficiency, unscheduled aircraft maintenance, and delayed flights.

Consider that just a one percent improvement in aircraft engine maintenance efficiency can reduce related costs by $250 million annually. A similar one percent fuel savings in power generation could add more than $4 billion annually to the global economy.

Whether in terms of operations, performance or maintenance excellence, all industries are looking for their next major productivity gains. Health care is burdened by a system where doctors and caregivers have to go searching for vital information; it is inefficient at best and life threatening at worst. We need to make the data more intelligent and integrated, more predictive and proactive, so information finds the doctor instead of the other way around.

Intelligent data flows speed up care delivery and can prevent chronic conditions by getting the treatment right the very first time. Similarly, in terms of health management costs, “intelligent” hospitals are deploying systems that behave like air traffic control for medical staff and devices, and provide a full detailed view of hospital resources. Better utilization cuts capital expenses. Better asset location leaves nurses more time to focus on patients. Better management improves patient flow, cuts operating costs, and saves hospitals millions.
There are similar scenarios in every other major industry, and the economic benefit can be huge.  Assuming growth similar to what prevailed during the internet boom, the Industrial Internet revolution will add about $15 trillion to global GDP by 2030. That’s the equivalent of adding another U.S. economy to the world.

The amazing aspect of this growth is that it stems from what appears to be minor productivity improvements. At GE, we have 5,000 software engineers and another 9,000 IT engineers. We’re focused on mining for just one percent gains in productivity. The potential is irresistible.

Roadmap to the Revolution
In the near future, I expect nothing short of an open, global fabric of highly intelligent machines that connect, communicate and cooperate with us. This Industrial Internet is not about a world run by robots, it is about combining the world’s best technologies to solve our biggest challenges. It’s about economically and environmentally sustainable energy, curing the incurable diseases, and preparing our infrastructure and cities for the next 100 years.

To do this industry and government need to work together on two critical areas: standardization and security.  We need to establish common standards so that innovative minds can develop the best solutions for the machines and systems that move our world. Just as the advancement of mobile devices and operating systems have brought forth a prosperous  “app” economy, a standard language for machines will unleash waves of innovation that will truly change how the world works. This is a critical step and needs government policies that favor advancement.

Attaining the vision set forth for the Industrial Internet will also require an effective internet security regime. Cyber security should be considered in terms of both network security (a defense strategy specific to the Cloud) and the security of devices that are connected to the network. We need industry to effectively secure facilities and networks and governments to enforce a regulatory regime that promotes innovative solutions and international standards.

The Industrial Internet era has already begun. And during a time when the global economy is recovering but remains volatile and where resources are constrained for people, governments, and companies, what we need most is to not lose sight of a real opportunity to create meaningful change around the world.
After all, this is what revolutions are all about.

Jeff Immelt is the chairman and CEO of GE. 

Tuesday, November 27, 2012

Big Data Won't Save Pharma, But Smart Data Might


Data analytics has the potential to do much more if applied across the pharmaceutical enterprise.

Guy Cavet, Genetic Engineering and Biotechnology News, November 2012

Intelligent use of large-scale data has become fundamental to other industries: finance, insurance—even sports. But despite its importance in areas of research, data analytics has the potential to do much more if applied across the pharmaceutical enterprise.

In the past 15 years, biology has been transformed by the availability of large-scale genetic and genomic data. The first ten years of work on the human genome yielded one draft genome. The last ten years have yielded over ten thousand. Advances in technology have enabled high-throughput gene expression profiling, cancer genome analysis, and other disciplines to change the way biology is studied. Cheminformatics allows companies like Numerate to screen millions of compounds for activity by purely computational prediction. However, there are much greater opportunities for data-driven transformation across the broader pharmaceutical enterprise.

These opportunities arise, in part, because of the broad trend toward data being tracked and recorded in new and far-reaching ways. Importantly, many of these are outside the pharma industry. Medical records are collected electronically on an unprecedented scale, driven in part by federal “meaningful use” programs. These records reveal how diseases manifest and how treatments are used in the real world. Social media also contains vast amounts of information on real patient experiences with both diseases and treatments. And in the sales and marketing of drugs, data on program effectiveness is collected in real-time by reps, and companies like Aktana are interpreting it to understand where physicians perceive value.

Getting data is only the first step. The true value arises from analytics that generate actionable insights. In many cases, this means predictive modeling: developing algorithms that reveal what drives an outcome of interest (such as response to therapy or drug choice) and allowing that outcome to be predicted in the future. The data scientists that can carry out this type of analysis are multidisciplinary experts with skills from statistics, computer science, biology, chemistry and other fields, and they are highly sought after.

The conventional ways to engage data analysts involve building internal teams of scientists or buying time from consultants. However, data analytics is also particularly well-suited to crowdsourcing, which opens up a problem for many people to address. It’s inevitable that most of the world’s experts in any domain are outside any single pharma company. Even with strong internal teams, as Bill Joy of Sun Microsystems insightfully noted, “Most of the smartest people work for someone else.” Crowdsourcing allows those people to be tapped in a highly flexible and cost-effective manner. A team of experts can coalesce around a problem, working on it only as long as necessary, and then move on.

In 2006, Netflix used crowdsourcing to improve their ability to suggest movies to their customers. Rather than just inviting people to work in isolation, they set up an online competition in which people submitted entries in real-time and vied to come up with the best solution. This is a particularly effective approach to predictive modeling analytics. Seeing their rivals above them on a leaderboard drives people to continuously generate better results. In the Netflix competition, the company’s internal method was surpassed within six days, and the eventual winner was more than 10% better.

The competition approach is equally applicable to pharmaceutical industry problems. For example, Boehringer Ingelheim sponsored a contest to develop methods to predict small molecule safety that resulted in a 25% improvement over an industry standard approach. In the Heritage Health Prize competition, methods are being developed to predict which patients will require hospitalization, and for how long, over the next twelve months. Other competitions have been used to predict patient outcomes, sales patterns, and clinical outcomes. In each case, the results were better than any methods that had previously existed.

The full potential of data analytics requires accessing and using data in creative ways. For example, after a drug launches, information about the drug is rapidly generated in the outside world through patient and physician experiences. This information is currently largely untapped. It is entered into electronic medical records, tweeted, posted on Facebook, and entered into community sites such as Patients Like Me. This data is often unstructured and very noisy, but companies such as Israeli startup Treato are beginning to systematically organize it. Despite the complexities of working with data like this, skilled data scientists can extract meaningful patterns about drug-drug interactions, what drives patients to start and stop medications, or which patients will not adhere to their prescriptions, to name a few.

Predictive models even have the potential to tackle some of the most critical decisions in drug development, such as whether a clinical trial will be successful or whether a licensing deal will eventually lead to a drug. Billions of dollars rest on these decisions, but it is rare that all available relevant data is systematically employed to predict the probability of success. Of course, no algorithm can make such predictions with perfect accuracy, and no computation can replace a clinical trial. However, for an organization deciding between multiple costly development programs, having any improvement in ability to predict results is immensely valuable.
Putting data beyond the company firewall for outside experts to use may not be a natural step for organizations that are accustomed to carefully protecting their sensitive information. However, with the appropriate steps, the confidentiality and privacy of pharmaceutical and medical data can be carefully preserved. For example, when Boehringer Ingelheim sponsored a competition to predict small molecule activity, neither the structures of the molecules nor the specifics of the activity were revealed. In a competition to identify patients with type 2 diabetes using electronic medical records, the data was carefully de-identified to meet HIPAA standards. Privacy and confidentiality concerns can also be addressed by restricting access to trained and trusted individuals.

With drug development costs rising and approvals declining, new approaches are sorely needed. It’s too simplistic to see “big data” as a knight in shining armor, but the intelligent use of rich data, regardless of size, has the potential to help dramatically with problems from basic research to commercial operations.

Wednesday, November 21, 2012

Campaigns’ Use of Supporters’ Data Worries Privacy Advocates


Craig Timberg, The Washington Post, November 20, 2012


Shortly before Election Day, a Stanford graduate student reported that the campaign Web sites of both President Obama and Republican Mitt Romney were “leaking” personal information about their supporters through careless data handling.

Had it been Facebook and Google, a federal investigation might have ensued, and the companies could have suffered significant public relations setbacks and perhaps fines. But the Federal Trade Commission, the government agency most focused on personal privacy, has no jurisdiction over campaigns or political groups.
That is a small example of what privacy advocates say is a big problem with efforts to protect personal information in the United States: The politicians are not guarding the chicken coop. They are the foxes.

Obama’s sophisticated use of Big Data gave him a crucial edge in what, based on popular support alone, should have been a close election. Republicans are desperate to catch up. But it’s not clear who is positioned to protect the rights of voters at a time when politicians from both parties increasingly build their campaigns on the insights that commercial data brokers provide.

Washington has a community of professional privacy advocates at places such as the ACLU, the Electronic Privacy Information Center and the Center for Digital Democracy. Jeff Chester, executive director of the Center for Digital Democracy, said he approached lawmakers from both parties to express his concerns long before the election. But he got nowhere.

“Maybe we’re digital Don Quixotes,” Chester said. “There was a lack of interest, not surprisingly.”

People routinely tell pollsters that they’re concerned about online privacy, and Chester and his colleagues in the field count some allies on Capitol Hill and in the White House. The FTC under Chairman Jon Leibowitz and David Vladeck, head of its Bureau of Consumer Protection, have made the agency far more aggressive on consumer privacy generally — even if political campaigns are beyond their reach.

Yet overall the laws in the United States are much less strict than in Europe, where there are tight limits on what personal information can be collected and how long it can be kept. Companies caught crossing the line can provoke furious backlashes among their users.

The American political landscape, by comparison, is amorphous when it comes to privacy. There are widespread concerns on both the right and left but no single, coherent constituency demanding greater protections.

For all the talk in recent years about online privacy, data-hungry Google remains the most popular search engine and data-hungry Facebook the most popular social media site. Both worked closely with the campaigns and also have growing lobbying operations in Washington. Google’s Executive Chairman Eric Schmidt was a regular visitor at Obama’s Chicago campaign headquarters, say those who worked there, offering advice to the campaign’s data-savvy technologists.

Privacy advocates say the tide will eventually turn, when Americans truly understand the extent to which their information is collected and traded. A recent poll by the University of Pennsylvania’s Annenberg School for Communications found that nearly two-thirds of people would be less likely to support a candidate who bought data about voters’ online activities and used it to tailor political ads.

“People still don’t quite understand this stuff,” said Joseph Turow, the lead researcher on the Annenberg poll. He said politicians are “hoping people will, quote-unquote, get used to it.”

Tuesday, November 20, 2012

Who Are the Doctors Most Trusted by Doctors? Big Data Can Tell You.


Ki Mae Heussner, GigaOm, November 16, 2012

ZocDocHealthgradesVitalsYelp and other sites can tell you what patients think of their doctors. But finding out in any aggregate way what doctors think of their peers has been much harder, if not near impossible, for patients — up until now.

By accessing information in government databases through FOIA (Freedom of Information Act) requests, healthcare innovators are now able to share connections between doctors that are based on millions of physician referrals — a valuable indicator of who doctors hold in esteem.

Last month, Fred Trotter, a self-identified “hacktivist,” revealed that he had obtained a dataset of Medicare physician referrals through a FOIA request and was making the initial data available to those who supported a Medstartr crowdfunding campaign meant to build out his “DocGraph” and make it freely available. This week, he announced that he not only blew past his $15,000 funding goal, but was launching a second campaign to integrate his current data with an additional dataset.

HealthTap, a Palo Alto-based startup that connects patients with an online network of 17,000 doctors, also this week launched a new feature based partly on Trotter’s data. Called “DOConnect,” it combines Trotter’s Medicare data with physician data from its own site and other sources to give patients a new window into their doctors’ networks.

“This isn’t just friendships and business connections. This is who doctors trust,” said HealthTap co-founder and CEO Ron Gutman. “If you could know who your doctor’s doctor is, if you knew who they would choose, this lets you see that for the first time.”

The new tool, which reflects 25 million doctor referral connections, enables patients to see how many doctors are linked to a particular doctor, as well as their locations. As patients search for new physicians and specialists, being able to see who their current doctors are linked with could help them decide who to visit.  
It also gives doctors an opportunity to build online networks that reflect their offline networks, Gutman said. In a post about his “DocGraph” project, Trotter said that his data wasn’t strictly a “referral” data set because, in some cases, doctors might be linked through a patient they both happened to see at the same time, not through an active referral. But Gutman emphasized that HealthTap’s DOConnect considered more than Medicare referrals in mapping connections between doctors.

In releasing the dataset, Trotter said his main goal was to create doctor-rating algorithms that “patients find useful and doctors find fair.” But he also hoped that academics, health policy wonks, entrepreneurs and others would use it to bring more transparency to health care overall.

Todd Park, the U.S. Chief Technology Officer, has frequently talked up the value of “setting data free” and has backed hackathons, “datapaloozas” and other open data initiatives to highlight the need for innovators to use government data for the public good — this is a great example of that vision and, hopefully, points to more similar projects in the future.

“Our goal is to empower the patient, make the system transparent and accountable, and release this d

Monday, November 19, 2012

Crovitz: Obama's 'Big Data' Victory


Marketing politicians is now like selling drinks. It involves filtering policies and voters through algorithms.

L. Gordon Crovitz, The Wall Street Journal, November 18, 2012

When the Obama campaign emailed supporters to join a $40,000-a-ticket dinner in June at the New York home of actress Sarah Jessica Parker, journalists at ProPublica noticed something odd. They uncovered seven versions of the email solicitation for the fundraiser, some mentioning a second fundraiser that night, a concert by Mariah Carey, others that Ms. Parker is a mother, and still others that Vogue editor Anna Wintour would be at the dinner.


Who got which email depended on "big data"—information about each fundraising prospect and how different people react to different messages. In this year's election, it looks as if the Obama team's use of such data was one of its biggest edges over the Romney effort.

Some uses of big data were known before the election—for instance, the Obama website was even more assiduous than online retailers like Best Buy about dropping "cookies" on users' computers to gather information about their online habits. Reporting since the election makes clear just how important the role of data was in deciding the election.

Campaign manager Jim Messina pledged to "measure every single thing in this campaign" and built an analytics department five times the size of the 2008 effort. A Time magazine reporter got access to the data scientists in the campaign's Chicago headquarters on the condition that the reporter would keep mum until after the election. "What they revealed as they pulled back the curtain," Time recently reported, "was a massive data effort that helped Obama raise $1 billion, remade the process of targeting TV ads and created detailed models of swing-state voters that could be used to increase the effectiveness of everything from phone calls and door knocks to direct mailings and social media."

According to the magazine, the campaign created a "single massive system that could merge the information collected from pollsters, fundraisers, field workers and consumer databases as well as social-media and mobile contacts with the main Democratic voter files."

The campaign's "chief scientist," Rayid Ghani, had been at Accenture, where he co-wrote an academic paper describing work helping companies that "analyze large amounts of transactional data but are unable to systematically 'understand' their products." For example, Mr. Ghani helped grocers figure out why people bought orange juice by reducing the product to attributes that could be analyzed by algorithms—"Brand: Tropicana, Pulp: low, Fortified with: Vitamin-D, Size: 1 liter, Bottle type: plastic."

Marketing politicians is now like selling drinks. It involves filtering polices and voters through algorithms.

The Obama campaign focused on data showing the "persuadability" of voters. Multivariate tests identified issues and positions that could move undecided voters, ProPublica said: "The persuasion scores allowed the campaign to focus its outreach efforts—and their volunteer calls—on voters who might actually change their minds as the result. It also guided them in what policy messages individual voters should hear."

Big data give incumbents a big advantage, which seems to have surprised the Romney team. The Obama campaign has used cookies to track its supporters online since the 2008 election. It spent the past 18 months creating a new, unified database, factoring in some 80 pieces of information about each person, from age, race and sex to voting history. (The campaign denied reports that it tracked visits to pornography sites in its outreach algorithms.) The Romney campaign says it tried to match the Obama campaign's collection and analysis of data but had to start from scratch and had just seven months after the primaries.

What does this mean for you? Voters need to develop buyer-beware habits. The era of politicianssaying the same thing to all voters is over. Campaigns aim to tell voters exactly what each wants to hear: data-driven pandering.

Another consequence is that efforts by the Federal Trade Commission and other agencies to regulate data mining in the name of privacy are destined to collapse. Last month, Sen. Jay Rockefeller (D., W.Va.) sent a letter to the top "information broker" companies, accusing them of being "elusive" about what data they collect. Companies such as Acxiom and Experian replied that much of their information comes from government databases. They should also point out that political campaigns are among the most sophisticated users of the consumer data they collect.

The Obama campaign deserves credit for its big win through the sophisticated use of big data. As for regulators, they should understand that the information genie will not go back into the bottle—whether consumer information is used to sell orange juice or politicians.

A version of this article appeared November 19, 2012, on page A17 in the U.S. edition of The Wall Street Journal, with the headline: Obama's 'Big Data' Victory.

WSJ - CEO Council on Big Data


Big Data: Opportunities and Risks Co-Chairs
Tim Armstrong, Chairman and CEO, AOL Inc.
Dominic Barton, Global Managing Director, McKinsey & Co.
R. Marcelo Claure, Chairman, President and CEO, Brightstar Corp.

Subject Expert
David J. Rothkopf, President and CEO, Garten Rothkopf

Big Data: The Top Four Recommendations
1. Big Data Is Opportunity
Industry already recognizes that the advent of Big Data is a potential engine of significant economic growth. Policy makers must address issues such as security and privacy but not restrict this opportunity. Consumers, government and companies all have a role to play in defining policy. Countries need to recognize that the treatment of data is ultimately an issue of national competitiveness.


2. Create Rules of the Road
The government should define which activities concerning data are legal or illegal. For activities that are legal, consumers should then define what is public or private in their own cases through opting in or out in terms of what they choose to reveal about themselves.


3. Enact Cybersecurity Standards
The private sector, government, the military and consumers should jointly develop detailed standards and incidence-reporting practices for cybersecurity, building on existing industry best practices.


4. Create a Global Data Treaty
The U.S. should play a leadership role in coordinating with international bodies to pass a treaty to establish clear, universal standards on data privacy and ownership.


It's a phenomenon commonly referred to as Big Data, and it has generated widespread debate over a host of issues. What should companies do with such data? How can it be used profitably? How should it be treated? Should there be laws governing its use? Internationally recognized safeguards for consumer privacy? And how do U.S. companies stay competitive as more countries learn how to process and interpret such data?

The Wall Street Journal's John Bussey moderated a task-force discussion about just such issues. Here are edited excerpts of their presentations to the CEO Council:

Opportunity First
JOHN BUSSEY: Our group ended up with essentially two principles and two action items. Dominic, if you could take our first, please?

DOMINIC BARTON: Our group felt very strongly that this is a huge opportunity, and we shouldn't focus on how to protect or regulate Big Data before we recognize how important it is. Companies that use Big Data effectively get about a 6% productivity improvement versus others.

We're at the early stages. It's in every sector. And we all felt it can be the next wave of productivity growth if we use this data effectively.

Depending on how you protect or use data, it can actually lead to a country's competitive advantage. We've got small city-states like Singapore and Abu Dhabi that are allowing, for example, consumer medical information to be publicly available to get innovation.
So let's not lose sight of how important it can be for productivity improvement as we think about the protection.

MR. BUSSEY: Marcelo, our second principle.
MARCELO CLAURE: First of all, we had a fascinating group. We had a tremendous amount of interaction. And we quickly realized that all of us live in a hyper-connected world. By that I mean we're passing a huge amount of data every day through a connected car, a connected cellphone, a connected tablet, a connected home.

Big Data: The Top Four Recommendations
1. Big Data Is Opportunity
Industry already recognizes that the advent of Big Data is a potential engine of significant economic growth. Policy makers must address issues such as security and privacy but not restrict this opportunity. Consumers, government and companies all have a role to play in defining policy. Countries need to recognize that the treatment of data is ultimately an issue of national competitiveness.


2. Create Rules of the Road
The government should define which activities concerning data are legal or illegal. For activities that are legal, consumers should then define what is public or private in their own cases through opting in or out in terms of what they choose to reveal about themselves.


3. Enact Cybersecurity Standards
The private sector, government, the military and consumers should jointly develop detailed standards and incidence-reporting practices for cybersecurity, building on existing industry best practices.


4. Create a Global Data Treaty
The U.S. should play a leadership role in coordinating with international bodies to pass a treaty to establish clear, universal standards on data privacy and ownership.


Companies around the world increasingly collect and process vast amounts of customer data, particularly data tracking consumer behavior on the Web.

 We see it as a tremendous opportunity, but what is the government role?

Some members of our group argued that government shouldn't be allowed to regulate or do anything related to Big Data. But the rest of the group agreed that we've got to define the basics of what the government is going to do.

And where we came out at the end was, government should have the role of defining what is legal and what is illegal in the use of Big Data. And that's where the government should stop. We think it should be left up to the consumer to define what he or she wants to share and doesn't want to share, what is public and what is private.

Limiting government control to what's legal and illegal will allow the consumer to make more choices. As a consumer, you should be able to determine how much of your Facebook profile you want to share, or whether you want to share information about where you shopped last.

Security and Treaty
MR. BUSSEY: And Tim has our two action items.

TIM ARMSTRONG: The first one is probably more important than it seems on the surface, especially for private companies. I think we started this as kind of a government conversation, but mostly it's private companies that actually own the infrastructure that makes the company or the country work.

So cybersecurity is something that really needs to be dealt with. And specifically, having standards around this is important in two different directions.

One is standards of what happens when something gets attacked. How do you report it, and what is the infrastructure to help companies with that?

The second piece which is important and which came up in our discussion is when your data lives outside the country, or when you're going to do things with data in other countries. There are groups like CFIUS [the U.S. government's Committee on Foreign Investment in the U.S., which reviews foreign investment deemed to have an effect on national security] which are inside the government, which, if you haven't dealt with them, you will, over these type of data assets. Cybersecurity is really important.

The last action item is coming up with a global data treaty. The U.S. is probably the largest economy involved in Big Data right now, and we think it's important for the U.S. to take a leading role in defining what some of the Big Data policies and standards should be in such a treaty. We would hope that we, the U.S., could define a basic treaty to start off, and then other countries could either adopt or augment the treaty over time.

Sunday, November 18, 2012

You Can't Say That on the Internet


Evgeny Morozov, The New York Times, November 16, 2012

A BASTION of openness and counterculture, Silicon Valley imagines itself as the un-Chick-fil-A. But its hyper-tolerant facade often masks deeply conservative, outdated norms that digital culture discreetly imposes on billions of technology users worldwide.

What is the vehicle for this new prudishness? Dour, one-dimensional algorithms, the mathematical constructs that automatically determine the limits of what is culturally acceptable.

Consider just a few recent kerfuffles. In early September, The New Yorker found its Facebook page blocked for violating the site’s nudity and sex standards. Its offense: a cartoon of Adam and Eve in the Garden of Eden. Eve’s bared nipples failed Facebook’s decency test.

That’s right — a venerable publication that still spells “re-elect” as “reëlect” is less puritan than a Californian start-up that wants to “make the world more open.”

And fighting obscenity can be good for business. Impermium, a Silicon Valley company that helps Web sites deal with unwanted reader comments, has begun marketing technology that identifies “all kinds of harmful content — such as violence, racism, flagrant profanity, and hate speech — and allows site owners to act on it in real-time, before it reaches readers.” Impermium will police the readers — but who will police Impermium?

Apple, too, has strayed from its iconoclastic roots. When Naomi Wolf’s latest book, “Vagina: A New Biography,” went on sale in its iBooks store, Apple turned “Vagina” into “V****a.” After numerous complaints, Apple restored the title, but who knows how many other books are still affected?

True, these books are still on sale. Unlike the good old United States Post Office, which once confiscated “Lady Chatterley’s Lover” and other books it deemed too lewd, Silicon Valley does not engage in direct censorship. What it does, though, is present ideas and terms that have gained public acceptance as something to be ashamed of. Silicon Valley doesn’t just reflect social norms — it actively shapes them in ways that are, for the most part, imperceptible.

The proliferation of the Autocomplete function on popular Web sites is a case in point. Nominally, all it does is complete your search query — on YouTube, on Google, on Amazon — before you’ve finished typing, using an algorithm to predict what you’re most likely typing. A nifty feature — but it, too, reinforces primness.

How so? Consider George Carlin’s classic comedy routine “Seven Words You Can Never Say on Television.” See how many of those words would autocomplete on your favorite Web site. In my case, YouTube would autocomplete none. Amazon almost none (it also hates “penis” and “vagina”). Of Carlin’s seven words, Google would autocomplete only “piss.”

Until recently, even the word “bisexual” wouldn’t autocomplete at Google; it’s only this past August that Google, after many complaints, began to autocomplete some, but not all, queries for that term. In 2010, the hacker magazine 2600 published a long blacklist of similar words. While I didn’t verify all 400 of them on Google, a few that I did try — like “swastika” and “Lolita” — failed to autocomplete. Is Nabokov not trending in Mountain View? Alas, these algorithms are not particularly bright: unable to distinguish between Nabokov’s novel and child pornography, they assume you want the latter.

Why won’t tech companies let us freely use terms that already enjoy wide circulation and legitimacy? Do they fashion themselves as our new guardians? Are they too greedy to correct their algorithms’ mistakes?
Thanks to Silicon Valley, our public life is undergoing a transformation. Accompanying this digital metamorphosis is the emergence of new, algorithmic gatekeepers, who, unlike the gatekeepers of the previous era — journalists, publishers, editors — don’t flaunt their cultural authority. They may even be unaware of it themselves, eager to deploy algorithms for fun and profit.

Many of these gatekeepers remain invisible — until something goes wrong. Thus, in early September, the online livestream from the Hugo Awards, the Oscars of the science fiction world, was interrupted with a cryptic copyright warning, right before the popular author Neil Gaiman was to deliver an acceptance speech.
Apparently, Ustream — the site streaming the ceremony — was using the services of another company to determine whether its streamed videos violated any copyrights. The partner company draws on a very large video archive to see, in real time, if what’s being streamed matches anything in its collection. Somehow, the celebratory video that preceded Mr. Gaiman’s speech tripped a copyright match, and the feed was cut off, even though the organizers had all the requisite permissions (and, under the doctrine of fair use, probably didn’t need them anyway).

The limitations of algorithmic gatekeeping are on full display here. How do you teach the idea of “fair use” to an algorithm? Context matters, and there’s no rule book here; that’s why we have courts. From the perspective of sticky, amorphous human culture, semi-automation — pairing up humans with algorithms — beats full automation. Sometimes, gaps are productive. But will profit-driven Silicon Valley ever acknowledge this insight?

Our reputations are increasingly at the mercy of algorithms, too. No one knows this better than Bettina Wulff, the former German first lady who has sued Google for autocompleting searches for her name with words like “escort” and “prostitute.” Ms. Wulff insists that Google’s algorithms spread false rumors about her; Google says that the suggested terms are just an “algorithmically generated result of objective factors, including the popularity of the entered search terms.”

Google’s defense would sound tenable if its own algorithms weren’t so easy to trick. In 2010, the marketing expert Brent Payne paid an army of assistants to search for “Brent Payne manipulated this.” Soon anyone typing “Brent P” into Google would see that phrase in their autocomplete suggestions. After Mr. Payne publicized his experiment, Google removed that particular suggestion, but how many similar cases have gone undetected? What is “objective” about such algorithmic “truths”?

Quaint prudishness, excessive enforcement of copyright, unneeded damage to our reputations: algorithmic gatekeeping is exacting a high toll on our public life. Instead of treating algorithms as a natural, objective reflection of reality, we must take them apart and closely examine each line of code.

Can we do it without hurting Silicon Valley’s business model? The world of finance, facing a similar problem, offers a clue. After several disasters caused by algorithmic trading earlier this year, authorities in Hong Kong and Australia drafted proposals to establish regular independent audits of the design, development and modifications of computer systems used in such trades. Why couldn’t auditors do the same to Google?
Silicon Valley wouldn’t have to disclose its proprietary algorithms, only share them with the auditors. A drastic measure? Perhaps. But it’s one that is proportional to the growing clout technology companies have in reshaping not only our economy but also our culture.

Obviously, Silicon Valley won’t develop or embrace similar norms overnight. However, instead of accepting this new reality as a fait accompli, we must ensure that, in pursuing greater profits, our new algorithmic gatekeepers are forced to accept the idea that their culture-defining function comes with great responsibility.
The author of the forthcoming book “To Save Everything, Click Here: The Folly of Technological Solutionism.”

Thaler: Applause for the Numbers Machine


Richard H. Thaler, The New York Times, November 18, 2012

THE biggest winners on Election Day weren’t politicians; they were numbers folks.

Computer scientists, behavioral scientists, statisticians and everyone who works with data should be proud. They told us who was going to win, but they also helped to make many of those victories happen.

Three groups of geeks deserve the love they rarely receive: people who run political polls, those who analyze the polls and those who figure out how to help campaigns connect with voters.

Many people doubted the accuracy of political polling this year. Part of the skepticism was based on the wide range of predictions, with some showing President Obama in the lead, and others Mitt Romney. But there were additional, structural reasons to worry whether pollsters would be able to find representative samples of voters.

One problem is that people are harder to reach on the telephone these days. About a third of voters no longer have a land line, and many of those who have them don’t pick up calls from strangers. So modern polling companies have to work harder to find voters willing to answer questions, then have to guess which of these respondents will actually show up and vote.

So it may come as a surprise that, collectively, polling companies did quite well during this election season. Although there was a small tendency for the pollsters to overestimate Mr. Romney’s share of the vote, a simple average of the polls in swing states produced a very accurate prediction of the Electoral College outcome. Notably, the most accurate polls tended to be done via the Internet, many by companies new to this field. That’s geek victory No. 1.

This relatively accurate polling data provided the raw material for the second group of election pioneers: poll analysts like Nate Silver, who writes the FiveThirtyEight blog for The New York Times, as well as Simon Jackman at Stanford, Sam Wang at Princeton and Drew Linzer at Emory University.

What do poll analysts do? They are like the meteorologists who forecast hurricanes. Data for meteorologists comes from satellites and other tracking stations; data for the poll analysts comes from polling companies. The analysts’ job is to take the often conflicting data from the polls and explain what it all means.

Worry about the reliability of the polling data led to widespread skepticism, or even outright hostility, toward poll analysts. The phrase “garbage in, garbage out” was one of the more polite criticisms bouncing around the Internet in the days before the election.

Because the polls were not, in fact, garbage, the first job of a poll analyst was quite easy: to average the results of the various polls, weighing more reliable and recent polls more heavily and correcting for known biases. (Some polls consistently project higher voter shares for one party or the other.)

A harder but more valuable task is to help readers translate the polling data into forecasts of the probability of victory. In Florida, where the final polls showed essentially a tie, according to Mr. Silver’s weighting method, it’s easy to see why he said the chance of either candidate winning the state was 50 percent. Ultimately, President Obama would very narrowly carry the state.

But what about North Carolina, where Mr. Silver projected that Mitt Romney would get 50.6 percent of the vote and President Obama, 48.9 percent? Looking at that very small difference, what probability would you have assigned to a Romney victory in that state?

Most people would guess something very close to 50-50. But not a good numbers guy. By looking back at previous elections with polling data this close, Mr. Silver estimated that Mr. Romney’s chances of winning North Carolina were 74 percent, a number that may seem surprisingly high. (Mr. Romney won the state.)

The slightly larger but still seemingly tiny lead that the president held in Ohio, another swing state, led poll analysts to predict that the chance of an Obama victory in Ohio was around 90 percent. And because Mr. Romney would have to win several such states with small Obama leads in order to prevail in the Electoral College, the analysts ended up with similarly high degrees of confidence in an overall Obama victory. They ended up predicting the Electoral College outcome almost exactly right, especially if you consider the final outcome in Florida to be a virtual tie, as they had projected.

Pundits making forecasts, some of whom had mocked the poll analysts, didn’t fare as well, and many failed miserably. George F. Will predicted that Mr. Romney would win 321 electoral votes, which turned out to be very close to President Obama’s actual total of 332. Jim Cramer from CNBC was nearly as wrong in the opposite direction, projecting that the president would win 440 electoral votes.

There is a lesson here. When it comes to assessing the chances of some complicated combination of events, gut feelings are pretty much useless. Pundits are no better at forecasting election outcomes than they would be at predicting the final path of a hurricane. Smart pundits should consider either abandoning this activity, or consulting with the geeks before rendering their guesses.

The third set of folks who deserve recognition in this election cycle were a group of young people working in a windowless room at Obama headquarters, affectionately known as the cave. They were part of the effort by the numbers-oriented campaign manager, Jim Messina, to maximize turnout.

THERE are two basic parts of an election campaign. The first comes under the category of messaging — deciding what a candidate should say and what ads to run. Most of the commentary we read about elections focuses on this component.

The second part is turnout, and in some ways is even more important. Here is a simple bit of math that you don’t have to be a geek to understand: It doesn’t matter which candidate a person prefers unless that person shows up and votes.

Pundits will debate for eternity which campaign did a better job of communicating its message, but there is no doubt which campaign won the turnout contest. Young, black and Hispanic voters all turned out in higher numbers than expected, and they often supported President Obama.

Much was made of the big Obama advantage in field offices in swing states. But those field offices would have been little good to the campaign without modern tools to find potential voters, have them register and encourage them to vote. In the weeks leading up to the election, the Obama canvassers had accurate lists of potential voters and field-tested scripts for their contacts with voters. This explains in part why Democrats were such heavy users of early voting.

By contrast, Project Orca, a get-out-the-vote computer program for the Romney campaign that wasn’t designed to be used until Election Day, reportedly had some bugs.

There should be something reassuring about this Obama campaign efficiency to all Americans, even those who supported Mr. Romney based on his success in business. When it came to the business of running a campaign, it was the former professor and community organizer who had the more technologically savvy organization and made more effective use of its resources, including geek power.

Richard H. Thaler is a professor of economics and behavioral science at the Booth School of Business at the University of Chicago. He was an informal adviser to the Obama campaign.