Wednesday, August 29, 2012

Facebook: The Real Presidential Swing State



David Talbot, Technology Review, July 20, 2012

The outcome of the 2012 campaign could have less to do with grand vision than with online data analytics and peer-to-peer voter targeting. 

Facebook and Internet campaign strategies grew up at the same time. In 2003 and early 2004, when Facebook was a new dorm-room plaything, Howard Dean's presidential campaign pioneered Internet fund-raising. By 2008, Facebook had crossed the 100-million-user mark and was coming to dominate online social networking; that year, Barack Obama's campaign wielded a custom social-networking site that helped win the White House (see "How Obama Really Did It"). A Facebook cofounder, Chris Hughes, helped build that site.

Now, in 2012, Facebook is central to the upcoming presidential election. Both Obama and his Republican opponent, Mitt Romney, are well aware that half or more of the electorate is on Facebook. Both campaigns' websites are entwined with Facebook pages; visitors are encouraged to log in with their Facebook accounts and then post messages supporting the candidates for their friends to see. What Facebook also gives the candidates is an arena for testing, analyzing, and distributing precisely targeted political advertising. Both campaigns can also use Facebook to urge their supporters to vote and, potentially, to lobby their undecided friends in swing states. That means this is where the 2012 election might be won or lost—even if far more money will be spent elsewhere, especially on TV ads.

Making use of social connections can lead to the ideal form of marketing: individual messages of persuasion delivered by trusted friends. You can see the president's campaign reaching for this goal with Obama 2012, an app that his supporters can use to integrate their Facebook accounts with the campaign's website. The app's avowed task is to give people a quick and easy way to access the volunteering and organizing functions that worked so well for Obama in 2008. But the permission screen that comes with the app makes clear that it has another purpose as well. When I installed the app, I noticed that it said it would grab information about my friends: their birthdates, locations, and "likes."

Facebook's policies require that such data be used in only the context of the app itself, but even so, the campaign should be able to create tools that prompt supporters to approach voting-age friends in swing states and craft personalized appeals based on what the campaign can infer about those friends' interests and views. Similar tools are coming from other quarters, too. In July NGP VAN, a company in Somerville, Massachusetts, that maintains a database on all registered U.S. voters and helps Democratic candidates access the data, released a Facebook app called Social Organizing. The app lets Democratic volunteers log in with Facebook and match their friends with voters in the database. Like the Obama app, NGP VAN's makes it possible for candidates to execute a peer-to-peer persuasion strategy using Facebook.

So don't be surprised—especially if you live in a state that is considered up for grabs, such as Ohio or Florida—if you hear from an old college friend with a political pitch based on what the campaign thinks is important to you, as suggested by your Facebook data. If you've "liked" a page blaming Obama for high gas prices, you might be reminded about his pro-drilling positions.

Don't be surprised if you hear from an old friend with a pitch based on what the campaign thinks is important to you. If you've "liked" a page blaming Obama for high gas prices, you might be reminded about his pro-drilling positions.

The Obama campaign didn't respond to requests for an interview about its plans, but Joe Trippi, the Democratic strategist who pioneered Internet fund-raising for Dean in 2003 and 2004, expects that the campaign will use sophisticated methods to determine how and when to encourage peer-to-peer appeals in the final weeks of the race. "What's most important in terms of being able to reach people is to know not only that the voter is undecided—and also what issues, what is holding them up from crossing the line—but who their friends are in the network that might be able to talk to them," he says. "And then get those friends the information that says, 'We need you to talk to your friend in Pennsylvania about these three issues that matter most to them.' This is a field organizer's dream." Certainly it is more than Trippi could have dreamed of as a $15-a-day campaign worker knocking on doors in Jones County, Iowa, for Senator Edward Kennedy in 1979, carrying shoeboxes of index cards indicating whether voters said they supported Kennedy for the next year's Democratic presidential nomination.

The Romney campaign's website also encourages supporters to log in using Facebook, but it requests permission only to view the individual user's information—not information about the user's friends "right now," says the Romney campaign's digital director, Zac Moffatt. The same is true for the Republican National Committee's Facebook app. This may change, though, because the Republicans share Trippi's view. "I think you will start to see, on our side, that app permissions will get changed," says a Republican digital strategist who spoke on condition of anonymity. "Republicans are working on apps that take advantage of all the things in the Facebook social graph."

How much information can the campaigns glean this way? Consider that the average friend count on Facebook is 190. As of early August, more than 150,000 people were using the Obama 2012 app. Multiply those numbers and you get more than 28 million people. Now, surely many friend lists overlap, and many of those people aren't even voters. And some users block the ability of apps like Obama's to gather information about them when their friends install the programs (a Consumer Reports study, however, found that only 37 percent of users touch app settings). But even if these factors make 90 percent of Obama supporters' friends useless to the campaign, the president's campaign app would still have intelligence on 2.8 million American voters who didn't necessarily take any explicit action to share it.

Persuading just a small percentage of those people could be crucial. In 2000, the contested election that put George W. Bush in office was determined largely by 537 votes in Florida, out of six million cast in that state. And in 2004 Bush beat John Kerry by fewer than 120,000 people out of 5.6 million who voted in Ohio. (Facebook is the virtual battleground within that battleground state. In 2012, just over five million account holders of voting age lived in Ohio—out of a total voting-age population of 8.8 million, according to Well & Lighthouse, a Democratic consulting firm.) Given math like that, the right peer-to-peer and message targeting strategies "could be the difference in swing states," Trippi says.

Fast and on target

In addition to any peer-to-peer strategies they might employ, the candidates are already waging online advertising campaigns that are more scientifically designed and demographically precise than the ones Obama and John McCain deployed in 2008. Political operatives can now rapidly test ad copy across multiple demographics, getting strategic insights within hours. They can even keep track of exactly which ads individual computer owners have clicked on.

These abilities were brought to bear in an ad campaign that rolled out in March of 2010, when President Obama signed the Patient Protection and Affordable Care Act—so-called Obamacare.

The midterm elections were just eight months away, and the president was concerned for a vulnerable ally, Harry Reid of Nevada, the Senate majority leader. On the health care issue alone, Reid's online strategist, Jon-David Schlough, developed 18 sets of targeted advertisements for people in different demographic groups. For example, the version geared to students pointed out that the legislation would let them keep their parents' insurance until age 26; the one for the elderly focused on what it would do to close a Medicaid benefit gap known as the doughnut hole.

Then, for each of the 18 campaigns, different versions were tested on Facebook. Schlough says the site gave him access to a wide range of demographic groups, made it possible to place small ads at low rates, and offered easy ways to experiment rapidly with different combinations of headline, image, and text. The versions that generated the most clicks would get wider distribution on multiple websites.

Eventually, the campaign could be sure that, say, an ad about being able to stay on parental insurance plans would be shown to a specific 25-year-old four times a day for two weeks. It's called nanotargeting, and "it's now a component of all campaigns," says Schlough, the founder of Well & Lighthouse. "Political types are used to large data-set analysis on things like polling data and turnout data. But the fact that so much more data is available, so much faster, is allowing us to innovate a lot quicker."

For Reid, such innovations might have been decisive. Consider that his opponent, Tea Party favorite Sharron Angle, spent about as much as Reid and was ahead in the polls in the weeks leading up to the election. In the end, Reid won by more than 5 percentage points.

David Talbot is Technology Review's chief correspondent.

This article was revised on August 15, 2012.

Tuesday, August 28, 2012

Gartner Says Big Data Makes Organizations Smarter, But Open Data Makes Them Richer

Open Data on the Agenda for Gartner Symposium/ITxpo, October 21-25, Orlando, Florida

STAMFORD, Conn., August 22, 2012—

Whereas "big data" will make organizations smarter, open data will be far more consequential for increasing revenue and business value in today's highly competitive environments, according to Gartner, Inc.

"Big data is a topic of growing interest for many business and IT leaders, and there is little doubt that it creates business value by enabling organizations to uncover previously unseen patterns and develop sharper insights about their businesses and environments," said David Newman, research vice president at Gartner. "However, for clients seeking competitive advantage through direct interactions with customers, partners and suppliers, open data is the solution. For example, more government agencies are now opening their data to the public Web to improve transparency, and more commercial organizations are using open data to get closer to customers, share costs with partners and generate revenue by monetizing information assets."

Gartner analysts believe an open data strategy should be a top priority for any organization that uses the Web as a channel for delivering goods and services. Open data strategies support outside-in business practices that generate growth and innovation. Enterprise architects help their organization connect independent open data projects by creating actionable deliverables and information-sharing practices that generate business-focused outcomes for achieving strategic customer growth and retention objectives.

Gartner analysts said that any business that has a data warehouse should consider how it can use data as a strategic asset and revenue generator. Maturing technologies for data quality and data anonymization can help mitigate regulatory restraints and risk factors. Open data APIs provide simple, Web-oriented means for data exchange, and linked data techniques are effective for generating big datasets. When considering the long-term benefits of an open data strategy, organizations should investigate the types of data exchange now emerging where information producers and consumers share data for profit.

Emerging data marketplaces are also places for organizations to open their data — potentially turning their "data into dollars." The challenge is to keep the barriers to entry low to enable participation by different types of business and streamlined processes for adding and vetting data sources. Monetizing data is a technological and operational challenge. If an enterprise's goal is to unlock its data's full revenue potential, it needs to be able to reach all possible data buyers efficiently.

"With tight budgets and continued economic uncertainty, organizations will need leaders who can craft breakthrough strategies that drive growth and innovation," said Mr. Newman. "As change agents, enterprise architects can help their organizations become richer through strategies such as open data."

Although openness is a pervasive and persistent issue in IT, there is very little agreement about exactly what "open" means. According to Gartner analysts, an informal definition of openness is a level playing field where everyone can play a game that can evolve. There is a positive relationship between the openness of information goods (for example, code, data, content and standards) and information services (for example, services that offer information goods, such as the Internet, Wikipedia, OpenStreetMap and GPS) and the size and diversity of the community sharing them. From the viewpoint of enterprise information architects, this is known as the information-sharing network effect: the business value of a data asset increases the more widely and easily it is shared.

Open data APIs are a lightweight approach to data exchange. Their use is now considered a best practice for opening data and functionality to developers and other businesses. Organizations use APIs to generate new sources of revenue, spur innovation, increase transparency and improve brand equity.

"The challenge for organizations is to determine how best to use APIs and how an open data strategy should align with business priorities," Mr. Newman said. "This is where enterprise architects can help. While some internal IT functions may be using APIs to fulfil local or specific application needs, the enterprise architecture process harvests and elevates good works as first-class strategic priorities that create business-focused outcomes. As a strategic enabler, APIs are a powerful means with which to build an ecosystem, and a first step toward monetizing data assets."

Additional information is available in the Gartner report "Open for Business: Learn to Profit by Open Data." The report is available on Gartner's website at http://www.gartner.com/resId=1947015.

About Gartner Symposium/ITxpo

Gartner Symposium/ITxpo is the world's most important gathering of CIOs and senior IT executives. This event delivers independent and objective content with the authority and weight of the world's leading IT research and advisory organization, and provides access to the latest solutions from key technology providers. Gartner's annual Symposium/ITxpo events are key components of attendees' annual planning efforts. IT executives rely on Gartner Symposium/ITxpo to gain insight into how their organizations can use IT to address business challenges and improve operational efficiency.

Additional information for Gartner Symposium/ITxpo in Orlando, Florida, October 21-25, is available at http://www.gartner.com/us/symposium. Members of the media can register for the event by contacting Janessa Rivera at janessa.rivera@gartner.com.

Additional information from the event will be shared on Twitter at http://twitter.com/Gartner_inc and using #GartnerSym.

Friday, August 24, 2012

Don't Build a Database of Ruin


Paul Ohm, Harvard Business Review, August 23, 2012

Many businesses today find themselves locked in an arms race with competitors to see who can convert customer secrets into the most pennies. To try to win, they are building perfect digital dossiers, to use a phrase coined by Daniel Solove, massive data stores containing hundreds, if not thousands or tens of thousands, of facts about every member of our society. 
In my work, I've argued that these databases will grow to connect every individual to at least one closely guarded secret. This might be a secret about a medical condition, family history, or personal preference. It is a secret that, if revealed, would cause more than embarrassment or shame; it would lead to serious, concrete, devastating harm. And these companies are combining their data stores, which will give rise to a single, massive database. I call this the Database of Ruin. Once we have created this database, it is unlikely we will ever be able to tear it apart.

I have become convinced that my earlier, bleak predictions about the Database of Ruin were in fact understated, arriving before it was clear how Big Data would accelerate the problem. Consider the most famous recent example of big data's utility in invading personal privacy: Target's analytics team can determine which shoppers are pregnant, and even predict their delivery dates, by detecting subtle shifts in purchasing habits. This is only one of countless similarly invasive Big Data efforts being pursued. In the absence of intervention, soon companies will know things about us that we do not even know about ourselves. This is the exciting possibility of Big Data, but for privacy, it is a recipe for disaster.

If we stick to our current path, the Database of Ruin will become an inevitable fixture of our future landscape, one that will be littered with lives ruined by the exploitation of data assembled for profit. But we can chart a different course, in various ways. I think our brightest engineers can develop innovative privacy-enhancing technologies which will enable new techniques for data analytics that minimize costs to privacy. I hope that public institutions and industry, through self-regulation, will devise ways to better balance the burdens on privacy and the benefits of Big Data. If nothing else, I anticipate that society will slowly develop new norms for engaging with the massive amount of information collected about us, creating informal rules governing when and how it is appropriate to release, collect, and use data, the way minors have learned to speak and listen carefully on social networks.

But every one of these correctives requires the same thing: time. We need to slow things down, to give our institutions, individuals, and processes the time they need to find new and better solutions. The only way we will buy this time is if companies learn to say, "no" to some of the privacy-invading innovations they're pursuing. Executives should require those who work for them to justify new invasions of privacy against a heavy burden, weighing them against not only the financial upside, but also against the potential costs to individuals, society, and the firm's reputation. Companies should do this not only as matter of good corporate social responsibility, but also because it will likely square with the government's recommendations for protecting privacy, which seem to advise caution and deliberation, under the banner of "context."

Earlier this year, Federal government officials released two privacy reports — the White House's White Paper and the FTC's Final Privacy Report — that together describe a national privacy policy for the foreseeable future. Although the two reports vary on some particulars, they both point to context as a central, important, and fundamental measuring stick we should use to assess decisions that bear on personal privacy.

The FTC report offers three broad recommendations: Privacy by Design, Simplified Choice for Businesses and Consumers, and Greater Transparency. In discussing the second recommendation — a call for simplified and more transparent choice — the FTC suggests a carve out. "Companies do not need to provide choice before collecting and using consumer data for practices that are consistent with the context of the transaction or the company's relationship with the consumer, or are required or specifically authorized by law." Under this standard, it might be "consistent with the context," for a company in a direct business relationship with a customer to use that customer's information to deliver ads for its other services, but it might be inconsistent with the context — thus requiring notice and choice — to sell that information to third-party advertisers, the FTC explains.

Similarly, the White House white paper defines a "Consumer Privacy Bill of Rights," which would protect, among other things, "Respect for Context." "Consumers have a right to expect that companies will collect, use, and disclose personal data in ways that are consistent with the context in which consumers provide the data," the paper explains.

These parallel pronouncements mean that companies that deal with personal information (meaning all companies, really) need to focus much more often than they have on the history of privacy practices in their industries. Although neither report defines in depth what it means by the word "context," to me the message seems to be: do not push the privacy envelope. Companies that use personal information in ways that go well beyond the practices of their competitors risk crossing the line from responsible steward to reckless abuser of consumer privacy.

The lesson is plain: compete vigorously and beat your competitors in every legitimate way, except when it comes to privacy invasion. Too many companies have learned this lesson the hard way, launching invasive new services that have triggered class action lawsuits, Congressional inquiries, and media firestorms. These companies knew that they were treading where others had feared to go. This may have felt like an exciting opportunity. It should have felt instead like perilous risk-taking, because it meant hurtling beyond the contextual borderlands defined by past practice.

At the Intersection of Big Data and Healthcare: What 7.2 Million Medical Records Can Tell Us

Kenneth Hines, Computing Computer Consortium Blog, August 23, 2012

We’ve featured lots of stories about Big Data over the last several months, but here’s a fascinating new one that illustrates the value of Big Data analytics in addressing important national priorities. Researchers at SENSEable City Lab – a new research initiative of the Massachusetts Institute of Technology — together with colleagues at GE Healthymagination have analyzed data from over 7 million electronic medical records, illustrating in a powerful visual the (sometimes surprising) relationships between medical conditions on the basis of the frequency of co-occurrences. They’re calling this extensive disease network the “Health  InfoScape.”

When you have heartburn, do you also feel nauseous? Or if you’re experiencing insomnia, do you tend to put on a few pounds, or more? By combing through 7.2 million of our electronic medical records, we have created a disease network to help illustrate relationships between various conditions and how common those connections are…

We often have a tendency to think of illness as an isolated event, but our first analysis details the numerous (sometimes unexpected) associations that exist around any given condition. This gives us new insight as to how closely connected some seemingly un-related health conditions might be. Such results force us to re-examine conventional categories of disease classification, as the boundaries between traditional disease categories are thoroughly blurred.

Our initial results are a mix of the expected and the unexpected — simultaneously challenging and reaffirming our preconceptions of health pattern, within individuals and across the U.S.
Click here to “take a look by condition or condition category and gender to uncover interesting associations.”

And there’s yet more opportunity for researchers here:

Now that we have a succinct picture of the human health network in the country, we will continue our investigation by delving deeper into how the environments around us factor into these results.

The Health InfoScape constitutes a perfect example of the role of Big Data science and engineering into the future.

Thursday, August 23, 2012

Computational social science: Making the links

From e-mails to social networks, the digital traces left by life in the modern world are transforming social science.
Jim Giles, Nature, August 22, 2012


Jon Kleinberg's early work was not for the mathematically faint of heart. His first publication1, in 1992, was a computer-science paper with contents as dense as its title: 'On dynamic Voronoi diagrams and the minimum Hausdorff distance for point sets under Euclidean motion in the plane'.

That was before the World-Wide Web exploded across the planet, driven by millions of individual users making independent decisions about who and what to link to. And it was before Kleinberg began to study the vast array of digital by-products generated by life in the modern world, from e-mails, mobile phone calls and credit-card purchases to Internet searches and social networks. Today, as a computer scientist at Cornell University in Ithaca, New York, Kleinberg uses these data to write papers such as 'How bad is forming your own opinion?'2 and 'You had me at hello: how phrasing affects memorability'3 — titles that would be at home in a social-science journal.


“I realized that computer science is not just about technology,” he explains. “It is also a human topic.”

Kleinberg is not alone. The emerging field of computational social science is attracting mathematically inclined scientists in ever-increasing numbers. This, in turn, is spurring the creation of academic departments and prompting companies such as the social-network giant Facebook, based in Menlo Park, California, to establish research teams to understand the structure of their networks and how information spreads across them.

“It's been really transformative,” says Michael Macy, a social scientist at Cornell and one of 15 co-authors of a 2009 manifesto4 seeking to raise the profile of the new discipline. “We were limited before to surveys, which are retrospective, and lab experiments, which are almost always done on small numbers of college sophomores.” Now, he says, the digital data-streams promise a portrait of individual and group behaviour at unprecedented scales and levels of detail. They also offer plenty of challenges — notably privacy issues, and the problem that the data sets may not truly be reflective of the population at large.

Nonetheless, says Macy, “I liken the opportunities to the changes in physics brought about by the particle accelerator, and in neuroscience by functional magnetic resonance imaging”.

Social calls

An early example of large-scale digital data being used on a social-science issue was a study in 2002 by Kleinberg and David Liben-Nowell, a computer scientist at Carleton College in Northfield, Minnesota. They looked at a mechanism that social scientists believed helped drive the formation of personal relationships: people tend to become friends with the friends of their friends. Although well established, the idea had never been tested on networks of more than a few tens or hundreds of people.

Kleinberg and Liben-Nowell studied the relationships formed in scientific collaborations. They looked at the thousands of physicists who uploaded papers to the arXiv preprint server during 1994–96. By writing software to automatically extract names from the papers, the pair built up a digital network several orders of magnitude larger than any that had been examined before, with each link representing two researchers who had collaborated. By following how the network changed over time, the researchers identified several measures of closeness among the researchers that could be used to forecast future collaborations5.


 
                                                                   
As expected, the results showed that new collaborations tended to spring from researchers whose spheres of existing collaborators overlapped — the research analogue of 'friends of friends'. But the mathematical sophistication of the predictions has allowed them to be used on even larger networks. Kleinberg's former PhD student, Lars Backstrom, also worked on the connection-prediction problem — experience that he has put to good use now that he works at Facebook, where he designed the social network's current friend-recommendation system.

Another long-standing social-science idea affirmed by computational researchers is the importance of 'weak ties' — relationships with distant acquaintances who are encountered relatively rarely. In 1973, Mark Granovetter, a social scientist now at Stanford University in Stanford, California, argued that weak ties form bridges between social cliques and so are important to the spread of information and to economic mobility6. In the pre-digital era it was almost impossible to verify his ideas at scale. But in 2007, a team led by Jukka-Pekka Onnela, a network scientist now at Harvard University in Cambridge, Massachusetts, used data on 4 million mobile-phone users to confirm that weak ties do indeed act as societal bridges7 (see 'The power of weak ties').

In 2010, a second group, which included Macy, showed that Granovetter was also right about the connection between economic mobility and weak ties. Using data from 65 million landlines and mobile phones in the United Kingdom, together with national census data, they revealed a powerful correlation between the diversity of individuals' relationships and economic development: the richer and more varied their connections, the richer their communities8 (see 'The economic link'). “We didn't imagine in the 1970s that we could work with data on this scale,” says Granovetter.

Infectious ideas

In some instances, big data have showed that long-standing ideas are wrong. This year, Kleinberg and his colleagues used data from the roughly 900 million users of Facebook to study contagion in social networks — a process that describes the spread of ideas such as fads, political opinions, new technologies and financial decisions. Almost all theories had assumed that the process mirrors viral contagion: the chance of a person adopting a new idea increases with the number of believers to which he or she is exposed.
 
Kleinberg's student Johan Ugander found that there is more to it than that: people's decision to join Facebook varies not with the total number of friends who are already using the site, but with the number of distinct social groups those friends occupy9. In other words, finding that Facebook is being used by people from, say, your work, your sports club and your close friends makes more of an impression than finding that friends from only one group use it. The conclusion — that the spread of ideas depends on the variety of people that hold them — could be important for marketing and public-health campaigns.

As computational social-science studies have proliferated, so have ideas about practical applications. At the Massachusetts Institute of Technology in Cambridge, computer scientist Alex Pentland's group uses smartphone apps and wearable recording devices to collect fine-grained data on subjects' daily movements and communications. By combining the data with surveys of emotional and physical health, the team has learned how to spot the emergence of health problems such as depression10. “We see groups that never call out,” says Pentland. “Being able to see isolation is really important when it comes to reaching people who need to be reached.” Ginger.io, a spin-off company in Cambridge, Massachusetts, led by Pentland's former student Anmol Madan, is now developing a smartphone app that notifies health-care providers when it spots a pattern in the data that may indicate a health problem.

Other companies are exploiting the more than 400 million messages that are sent every day on Twitter. Several research groups have developed software to analyse the sentiments expressed in tweets to predict real-world outcomes such as box-office revenues for films or election results11. Although the accuracy of such predictions is still a matter of debate12, Twitter began in August to post a daily political index for the US presidential election based on just such methods (election.twitter.com). At Indiana University in Bloomington, meanwhile, Johan Bollen and his colleagues have used similar software to search for correlations between public mood, as expressed on Twitter, and stock-market fluctuations13. Their results have been powerful enough for Derwent Capital, a London-based investment firm, to license Bollen's techniques.

Message received

When such Twitter-based polls began to appear around two years ago, critics wondered whether the service's relative popularity among specific demographic groups, such as young people, would skew the results. A similar debate revolves around all of the new data sets. Facebook, for example, now has close to a billion users, yet young people are still overrepresented among them. There are also differences between online and real-world communication, and it is not clear whether results from one sphere will apply in the other. “We often extrapolate from how one technology is used by one group to how humans in general interact,” notes Samuel Arbesman, a network scientist at Harvard University. But that, he says, “might not necessarily be reasonable”.

Proponents counter that these are not new problems. Almost all survey data contain some amount of demographic skew, and social scientists have developed a variety of weighting methods to redress the balance. If the bias in a particular data set, such as an excess of one group or another on Facebook, is understood, the results can be adjusted to account for it.

““We didn't imagine in the 1970s that we could work with data on this scale.””

Services such as Facebook and Twitter are also becoming increasingly widely used, reducing the bias. And even if the bias remains, it is arguably less severe than that in other data sets such as those for psychology and human behaviour, where most work is done on university students from Western, educated, industrialized, rich and democratic societies (often denoted WEIRD).

Granovetter has a more philosophical reservation about the influx of big data into his field. He says he is “very interested” in the new methods, but fears that the focus on data detracts from the need to get a better theoretical grasp on social systems. “Even the very best of these computational articles are largely focused on existing theories,” he says. “That's valuable, but it is only one piece of what needs to be done.” Granovetter's weak-ties paper6, for example, remains highly cited almost 40 years later. Yet it was “more or less data-free”, he says. “It didn't result from data analyses, it resulted from thinking about other studies. That is a separate activity and we need to have people doing that.”

The new breed of social scientists are also wrestling with the issue of data access. “Many of the emerging 'big data' come from private sources that are inaccessible to other researchers,” Bernardo Huberman, a computer scientist at HP Labs in Palo Alto, wrote in February14. “The data source may be hidden, compounding problems of verification, as well as concerns about the generality of the results.”

A prime example is Facebook's in-house research team, which routinely uses data about the interactions among the network's 900 million users for its own studies, including a re-evaluation of the famous claim that any two people on Earth are just six introductions apart. (It puts the figure at five15.) But the group publishes only the conclusions, not the raw data, in part because of privacy concerns. In July, Facebook announced that it was exploring a plan that would give external researchers the chance to check the in-house group's published conclusions against aggregated, anonymized data — but only for a limited time, and only if the outsiders first travelled to Facebook headquarters16.

In the short term, computational social scientists are more concerned about cultural problems in their discipline. Several institutions, including Harvard, have created programmes in the new field, but the power of academic boundaries is such that there is often little traffic between different departments. At Columbia University in New York, social scientist and network theorist Duncan Watts recalls a recent scheduling error that forced him to combine meetings with graduate students in computer science and sociology. “It was abundantly clear that these two groups could really use each other: the computer-science students had much better methodological chops than their sociology counterparts, but the sociologists had much more interesting questions,” he says. “And yet they'd never heard of each other, nor had it ever occurred to any of them to walk over to the other's department.”

Many researchers remain unaware of the power of the new data, agrees Harvard social scientist David Lazar, lead author on the 2009 manifesto. Little data-driven work is making it into top social-science journals. And computer-science conferences that focus on social issues, such as the Conference on Weblogs and Social Media, held in Dublin in June, attract few social scientists.

Nonetheless, says Lazar, with landmark papers appearing in leading journals and data sets on societal-wide behaviours available for the first time, those barriers are steadily breaking down. “The changes are more in front of us than behind us,” he says.

Certainly that is Kleinberg's perception. “I think of myself as a computer scientist who is interested in social questions,” he says. “But these boundaries are becoming hard to discern.”

Wednesday, August 22, 2012

Health IT's Next Big Challenge: Comparative Effectiveness Research


Innovative approach to medical data analysis can yield new treatment options at a lower cost.
 
Paul Cerrato,   InformationWeek August 21, 2012

Healthcare providers are being pushed to deliver more cost effective medical care and to improve the health of not just individual patients but large populations. One key to carrying out both mandates is finding more clinically effective treatment options.

Many academic medical thought leaders insist that the best way to find those treatment protocols is to test them in randomized controlled trials. Such RCTs require a large group of control subjects to receive either a placebo or conventional therapy and a large group to receive the experimental treatment in question. The problem is RCTs are outrageously expensive. In today's cost conscious healthcare system, that's a problem.

Enter comparative effectiveness research. CER compares two or more accepted treatments to determine which are most effective. Medical informatics comes into the picture because it's now possible to get these projects off the ground by analyzing huge patient databases. And much of that patient data can now be gleaned from electronic health record systems.

The American Recovery and Reinvestment Act of 2009 has earmarked $1.1 billion for CER. The Agency for Healthcare Research and Quality (AHRQ), the federal agency tasked with improving the quality, safety, efficiency, and effectiveness of health care, has been using part of that money to fund research on data infrastructure so that clinicians can figure how to take advantage of all the patient data in the Medicare system to compare treatment options. Other AHRQ-sponsored research has been looking at how to create an all-payer, all-claims database that clinicians can tap into for the same purpose.

Other CER-related projects include one led by David J. Magid, MD, director of research at the Colorado Permanente Group. His team searched through thousands of the group's EHRs to figure out which anti-hypertensive drugs are most effective when patients don't respond to first-line treatment with diuretics. The team managed to keep its research costs down to $200,000, a small fraction of what a randomized controlled trial would cost, and still came up with useful results, namely that beta blockers and ACE inhibitors work well.

Similarly a consortium of large healthcare systems, including Kaiser Permanente and Mayo Clinic, is capitalizing on the power of tens of millions of e-records to generate research. For example, they recently launched programs to mine their EHRs to compare treatment protocols for diabetes.

"With these large databases and detailed clinical information, we can conduct comparative, effective research in real world settings, with a full range of patients, not just those selected for clinical trials," Joe V. Selby, director of Kaiser's research division, states in a recent issue of Scientific American.

Boston's Beth Israel Deaconess Medical Center, one of the teaching hospitals affiliated with Harvard Medical School, recently entered the CER arena in a big way. Starting this month, the medical center launched Clinical Query, a searchable patient data repository that lets researchers and clinicians look for potential connections between diseases, treatment options, and risk factors, which in turn can become the jumping off point for a research project.

So if a Harvard researcher wants to compare the benefits of diuretics to ACE inhibitors among patients with hypertension, he can use Clinical Query to look at the records of more than 2 million patients and 200 million data points, including diagnoses, medications taken, lab values, and radiology images.

A comparison of data on the two classes of high blood pressure meds might reveal that one is more effective than the other. And while the results of that CER analysis may not carry the same weight as a randomized clinical trial in which groups of patients were actually given the drugs in real time to see which were more effective, the CER results can still guide clinicians on treatment options for their patients.

Given the fact that comparative effectiveness research will likely cost far less than a randomized clinical trial, it's time healthcare stakeholders take a closer look at this approach. The challenge for IT departments is going to be getting searchable patient data repositories up and running. Few hospitals have the resources to create their own version of Clinical Query. But at the very least, they need to start ramping up their data warehousing and data mining initiatives.

EHR systems are now collecting invaluable information that physicians can use to detect disease patterns, clusters of patients exposed to specific toxins, and groups of patients who respond well to various drug regimens. We can't waste this gold mine.

InformationWeek Healthcare brought together eight top IT execs to discuss BYOD, Meaningful Use, accountable care, and other contentious issues. Also in the new, all-digital CIO Roundtable issue: Why use IT systems to help cut medical costs if physicians ignore the cost of the care they provide? (Free with registration.)

Tuesday, August 21, 2012

Three kinds of big data

Looking ahead at big data's role in enterprise business intelligence, civil engineering, and customer relationship optimization.

Alistair Croll,  O'Reilly Radar,  August 21, 2012 
 
In the past couple of years, marketers and pundits have spent a lot of time labeling everything ”big data.” The reasoning goes something like this:

· Everything is on the Internet.

· The Internet has a lot of data.

· Therefore, everything is big data.

When you have a hammer, everything looks like a nail. When you have a Hadoop deployment, everything looks like big data. And if you’re trying to cloak your company in the mantle of a burgeoning industry, big data will do just fine. But seeing big data everywhere is a sure way to hasten the inevitable fall from the peak of high expectations to the trough of disillusionment.

We saw this with cloud computing. From early idealists saying everything would live in a magical, limitless, free data center to today’s pragmatism about virtualization and infrastructure, we soon took off our rose-colored glasses and put on welding goggles so we could actually build stuff.

So where will big data go to grow up?

Once we get over ourselves and start rolling up our sleeves, I think big data will fall into three major buckets: Enterprise BI, Civil Engineering, and Customer Relationship Optimization. This is where we’ll see most IT spending, most government oversight, and most early adoption in the next few years.

Enterprise BI 2.0

For decades, analysts have relied on business intelligence (BI) products like Hyperion, Microstrategy and Cognos to crunch large amounts of information and generate reports. Data warehouses and BI tools are great at answering the same question — such as “what were Mary’s sales this quarter?” — over and over again. But they’ve been less good at the exploratory, what-if, unpredictable questions that matter for planning and decision making because that kind of fast exploration of unstructured data is traditionally hard to do and therefore expensive.

Most “legacy” BI tools are constrained in two ways:

· First, they’ve been schema-then-capture tools in which the analyst decides what to collect, then later capture that data for analysis.

· Second, they’ve typically focused on reporting what Avinash Kaushik (channeling Donald Rumsfeld) refers to as “known unknowns” — things we know we don’t know, and generate reports for.

These tools are used for reporting and operational purposes, usually focused on controlling costs, executing against an existing plan, and reporting on how things are going.

As my Strata co-chair Edd Dumbill pointed out when I asked for thoughts on this piece:

“The predominant functional application of big data technologies today is in ETL (Extract, Transform, and Load). I’ve heard the figure that it’s about 80% of Hadoop applications. Just the real grunt work of log file or sensor processing before loading into an analytic database like Vertica.”

The availability of cheap, fast computers and storage, as well as open source tools, have made it okay to capture first and ask questions later. That changes how we use data because it makes it okay to speculate beyond the initial question that triggered the collection of data.

What’s more, the speed with which we can get results — sometimes as fast as a human can ask them — makes data easier to explore interactively. This combination of interactivity and speculation takes BI into the realm of “unknown unknowns,” the insights that can produce a competitive advantage or an out-of-the-box differentiator.

We saw this shift in cloud computing: first, big public clouds wooed green-field startups. Then, in a few years, incumbent IT vendors introduced their private cloud offerings. Private clouds included only a fraction of the benefits of public clouds, but were nevertheless a sufficient blend of smoke, mirrors, and features to delay the inevitable move to public resources by a few years and appease the business. For better or worse, that’s where most of IT budgets are being spent today according to IDC, Gartner, and others.

In the next few years, then, look for acquisitions and product introductions — and not a little vaporware — as BI vendors that enterprises trust bring them “big data lite”: enough to satisfy their CEO’s golf buddies, but not so much that their jobs are threatened. This, after all, is how change comes to big organizations.

Ultimately, we’ll see traditional “known unknowns” BI reporting living alongside big-data-powered data import and cleanup, and fast, exploratory data “unknown unknown” interactivity.

Civil Engineering

The second use of big data is in society and government. Already, data mining can be used to predict disease outbreaks, understand traffic patterns, and improve education.

Cities are facing budget crunches, infrastructure problems, and a crowding from rural citizens. Solving these problems is urgent, and cities are perfect labs for big data initiatives. Take a metropolis like New York: hackathons; open feeds of public data; and a population that generates a flood of information as it shops, commutes, gets sick, eats, and just goes about its daily life.



I think municipal data is one of the big three for several reasons: it’s a good tie breaker for partisanship, we have new interfaces everyone can understand, and we finally have a mostly-connected citizenry.

In an era of partisan bickering, hard numbers can settle the debate. So, they’re not just good government; they’re good politics. Expect to see big data applied to social issues, helping us to make funding more effective and scarce government resources more efficient (perhaps to the chagrin of some public servants and lobbyists). As this works in the world’s biggest cities, it’ll spread to smaller ones, to states, and to municipalities.

Making data accessible to citizens is possible, too: Siri and Google Now show the potential for personalized agents; Narrative Science takes complex data and turns it into words the masses can consume easily; Watson and Wolfram Alpha can give smart answers, either through curated reasoning or making smart guesses.

For the first time, we have a connected citizenry armed (for the most part) with smartphones. Nielsen estimated that smartphones would overtake feature phones in 2011, and that concentration is high in urban cores. The App Store is full of apps for bus schedules, commuters, local events, and other tools that can quickly become how governments connect with their citizens and manage their bureaucracies.

The consequence of all this, of course, is more data. Once governments go digital, their interactions with citizens can be easily instrumented and analyzed for waste or efficiency. That’s sure to provoke resistance from those who don’t like the scrutiny or accountability, but it’s a side effect of digitization: every industry that goes digital gets analyzed and optimized, whether it likes it or not.

Customer Relationship Optimization

The final home of applied big data is marketing. More specifically, it’s improving the relationship with consumers so companies can, as Sergio Zyman once said, sell them more stuff, more often, for more money, more efficiently.

The biggest data systems today are focused on web analytics, ad optimization, and the like. Many of today’s most popular architectures were weaned on ads and marketing, and have their ancestry in direct marketing plans. They’re just more focused than the comparatively blunt instruments with which direct marketers used to work.

The number of contact points in a company has multiplied significantly. Where once there was a phone number and a mailing address, today there are web pages, social media accounts, and more. Tracking users across all these channels — and turning every click, like, share, friend, or retweet into the start of a long funnel that leads, inexorably, to revenue is a big challenge. It’s also one that companies like Salesforce understand, with its investments in chat, social media monitoring, co-browsing, and more.

This is what’s lately been referred to as the “360-degree customer view” (though it’s not clear that companies will actually act on customer data if they have it, or whether doing so will become a compliance minefield). Big data is already intricately linked to online marketing, but it will branch out in two ways.

First, it’ll go from online to offline. Near-field-equipped smartphones with ambient check-in are a marketer’s wet dream, and they’re coming to pockets everywhere. It’ll be possible to track queue lengths, store traffic, and more, giving retailers fresh insights into their brick-and-mortar sales. Ultimately, companies will bring the optimization that online retail has enjoyed to an offline world as consumers become trackable.

Second, it’ll go from Wall Street (or maybe that’s Madison Avenue and Middlefield Road) to Main Street. Tools will get easier to use, and while small businesses might not have a BI platform, they’ll have a tablet or a smartphone that they can bring to their places of business. Mobile payment players like Square are already making them reconsider the checkout process. Adding portable customer intelligence to the tool suite of local companies will broaden how we use marketing tools.

Headlong into the trough

That’s my bet for the next three years, given the molasses of market confusion, vendor promises, and unrealistic expectations we’re about to contend with. Will big data change the world? Absolutely. Will it be able to defy the usual cycle of earnest adoption, crushing disappointment, and eventual rebirth all technologies must travel? Certainly not.


O'Reilly Radar (http://s.tt/1lhOc)

Saturday, August 18, 2012

Data and Analytics Key to Health Reform, but Challenges Stand in Way

Kate Ackerman, iHealthBeat, August 15, 2012
 
NATIONAL HARBOR, Md. -- At the eHealth Initiative's National Forum on Data and Analytics in Healthcare last week, stakeholders discussed the importance of data and analytics in implementing health reform, as well as the challenges associated with it.

Jennifer Covich Bordenick, CEO of the eHealth Initiative, said, "Our survey, CIOs and members kept telling us that they are concerned about analytics. They don't feel they have the tools necessary to meet the demands of accountable care and meaningful use." She noted that "93% of the CIOs believe it is very important, but 72% don't feel their organizations have what they need to meet the analytical needs."

By convening experts in health data and analytics, the forum aimed to highlight organizations that are leading the way and facilitate conversations around the need for improvement, she said.

Federal Government Touts Data and Analytics To Support Health Reform

Niall Brennan -- director of the Policy and Data Analysis Group at CMS -- told attendees that health data analytics is "absolutely central" to everything that has to do with health reform.

Brennan -- who stepped in for U.S. Chief Technology Officer Todd Park to give the afternoon's keynote speech -- highlighted the federal government's efforts "to make the data more helpful," while not compromising individuals' privacy.

He said that in the past CMS was "overly conservative" in terms of data release. However, in the last few years -- in large part because of the Affordable Care Act -- the agency has made great strides in liberating health data, he said.

"It might not look like it on the outside, but we are literally constantly pushing the envelope," Brennan said.

He cited the Blue Button Initiative and HealthData.gov as examples of the government's efforts to make privacy-protected health information available to the public.

He also noted that Section 10332 of the ACA authorizes the release of Medicare fee-for-service data to qualified entities if they agree to combine the CMS data with claims data from other sources to compile performance reports.

While Brennan touted the federal government's release of more health care data, he acknowledged, "You can have all the data in the world, but if you don't have the right analytics ... it's just a bunch of useless numbers."

Brennan said the federal government -- from CMS to HHS to the White House -- is "very, very committed" to data and analytics.

Survey Finds Industry Still Has Far To Go

At the forum, Jason Goldwater -- vice president of programs and research at eHI -- offered a sneak peek into the results of a survey eHI conducted with the College of Healthcare Information and Management Executives to get a picture of the types of data and analytics being used in health care.

The survey, which was conducted in July, focused on four areas:

· Types of data used;

· Types of analytic functions used;

· Types of functions needed; and

· Challenges to the use of data and analytics.

When asked what data their organizations actively exchange:

· 76.6% of respondents said lab results;

· 74.5% said demographics;

· 70.2% said discharge summaries;

· 46.8% said allergy information;

· 36.2% said continuity of care documents; and

· 36.2% said problem lists.

The survey found that most health care organizations are focusing their resources on retrospective analysis, with 58.3% citing that as the area in which they direct the majority of their analytical resources. According to the survey, 16.7% of respondents said their organizations direct the majority of their analytical resources toward real-time decision support, 13.9% said optimization and efficiency and 2.8% cited predictive analytics.

When asked what type of analytical functions their organizations primarily use:

· 87.5% of respondents said ad-hoc queries;

· 61.1% said data mining;

· 56.9% said data warehousing;

· 34.7% said exploratory data analysis;

· 30.6% said on-line analytical processing; and

· 23.6% said predictive modeling.

Goldwater said that despite health care organizations' interest in data and analytics, respondents cited several challenges, including:

· Lack of standardized data across systems;

· Lack of a system infrastructure to support analytics;

· Cost of analytical software;

· Concerns about privacy and security of the data; and

· Limited utility of the results to the organizations.

Goldwater said that eHI, CHIME and McKesson will host a webinar on Aug. 30 to discuss the survey results in more detail and that eHI will release an issue brief in the fall.

Speakers Highlight Challenges Associated With Health Data and Analytics

Several speakers offered real-life examples of the challenges cited by respondents to the eHI/CHIME survey.

Jason Williams -- vice president of business analytics at RelayHealth -- said that health care cost and quality reforms require inter-stakeholder transparency and that there needs to be "more emphasis" in that area. He also cited the demand for business and technology analysts and the need to find a balanced approach to privacy as areas for improvement.

Micky Tripathi -- president and CEO of the Massachusetts eHealth Collaborative and chair of eHI's Board of Directors -- said that variations in EHR systems can be problematic for analytics.

Brendan Mullen -- senior director of PINNACLE, the American College of Cardiology's outpatient registry -- noted that its system integration tool has to look at 41 locations in the NextGen EHR system to determine if a physician provided patient education on heart failure.

Tripathi said that as vendors are working to address the issues, new measures are coming down the pipeline.

For the Healthcare Information and Management Systems Society Conference in February, MAeHC compared the results of its certified Quality Data Center with the Office of the National Coordinator for Health IT-sponsored, open-source popHealth tool to evaluate meaningful use quality measures. The tools used the same exact data, but they did not produce the same results for any of the 44 measures, Tripathi said.

He explained that further investigation found several reasons for the discrepancies, including the definition of the continuity of care document, coding and mapping, and the interpretation of certain measures, such as age.

Tripathi said, "We have a ton of work to do," adding that the industry needs to keep pushing along.

Covich Bordenick told iHealthBeat, "It was clear by the end of that day that this was just the start of a conversation that is going to take years to explore," adding, "eHealth Initiative is going to help unpack this issue."

MORE ON THE WEB

· National Forum on Data and Analytics in Healthcare

· HealthData.gov

· "Data Analysis and the Future of Health Care" (Tibken, Wall Street Journal, 4/16
).