Showing posts with label Technological Advances. Show all posts
Showing posts with label Technological Advances. Show all posts

Thursday, September 20, 2012

Big Data for All


Omer Tene, Concurring Opinions, September 20, 2012

Much has been written over the past couple of years about “big data” (See, for example, here and here and here). In a new article, Big Data for All: Privacy and User Control in the Age of Analytics, which will be published in the Northwestern Journal of Technology and Intellectual Property, Jules Polonetsky and I try to reconcile the inherent tension between big data business models and individual privacy rights. We argue that going forward, organizations should provide individuals with practical, easy to use access to their information, so they can become active participants in the data economy. In addition, organizations should be required to be transparent about the decisional criteria underlying their data processing activities.

The term “big data” refers to advances in data mining and the massive increase in computing power and data storage capacity, which have expanded by orders of magnitude the scope of information available for organizations. Data are now available for analysis in raw form, escaping the confines of structured databases and enhancing researchers’ abilities to identify correlations and conceive of new, unanticipated uses for existing information. In addition, the increasing number of people, devices, and sensors that are now connected by digital networks has revolutionized the ability to generate, communicate, share, and access data.

Data creates enormous value for the world economy, driving innovation, productivity, efficiency and growth. In the article, we flesh out some compelling use cases for big data analysis. Consider, for example, a group of medical researchers who were able to parse out a harmful side effect of a combination of medications, which were used daily by millions of Americans, by analyzing massive amounts of online search queries. Or scientists who analyze mobile phone communications to better understand the needs of people who live in settlements or slums in developing countries.

At the same time, the “data deluge” presents formidable privacy concerns. Protecting privacy become harder as information is multiplied and shared ever more widely among multiple parties around the world. As more information regarding individuals’ health, financials, location, electricity use and online activity percolates, concerns arise about profiling, tracking, discrimination, exclusion, government surveillance and loss of control. From a more technical legal angle, big data challenges some of the most fundamental concepts of privacy law, including the definition of “personally identifiable information”, the role of individual control, and the principles of data minimization and purpose limitation.

In our article, we make the case for providing individuals with usable access to their data. The call for transparency is not new, of course. Rather the emphasis is on access to data in usable format, which can work to create value to individuals. Transparency and access alone have not emerged as potent tools because individuals do not care for, and cannot afford to indulge in transparency and access for their own sake (see one oft-cited counterexample here). The enabler of transparency and access is the ability to use the information and benefit from it in a tangible way. This will be achieved through “featurization” or “app-ification” of privacy. Organizations should build as many dials and levers as needed for individuals to engage with their data.

We expect that “featurization” of big data, harnessing its immense force for not only organizational but also individual benefit, will unleash a wave of innovation and create a market for personal data applications. The technological groundwork has already been completed with mash-ups and real-time APIs making it easier for organizations to combine information from different sources and services into a single user experience. Regardless of lingering questions concerning who – if anyone – “owns” the information, we think that fairness dictates that individuals enjoy beneficial use of the data about them.

Our second proposal would require organizations to disclose the decisional criteria underpinning their data analytics machinery. In a big data world, it is often not the data but rather the inferences drawn from them that give cause for concern. Inaccurate, manipulative or discriminatory conclusions may be drawn from perfectly innocuous, accurate data. Much like in quantum physics, the observer in big data analysis can affect the results of her research by defining the data set, proposing a hypothesis or writing an algorithm. At the end of the day, big data analysis is an interpretative process, in which one’s identity and perspective informs one’s results. Like any interpretative process, it is subject to error, inaccuracy and bias. Louis Brandeis, who together with Samuel Warren “invented” the legal right to privacy in 1890, has also written that “[s]unlight is said to be the best of disinfectants”. We trust if the existence and uses of databases were visible to the public, organizations would be more likely to avoid unethical or socially unacceptable uses of data.

Monday, September 10, 2012

Tech's New Wave, Driven by Data



Steve Lohr, The New York Times, September 8, 2012

From the article: "TECHNOLOGY tends to cascade into the marketplace in waves. Think of personal computers in the 1980s, the Internet in the 1990s and smartphones in the last five years.

Computing may be on the cusp of another such wave. This one, many researchers and entrepreneurs say, will be based on smarter machines and software that will automate more tasks and help people make better decisions in business, science and government. And the technological building blocks, both hardware and software, are falling into place, stirring optimism.

Michael R. Stonebraker, a pioneer in database research, is one of the optimists. Software used by companies and government agencies — in products sold by Oracle, I.B.M., Microsoft and others — descends from research done in the 1970s by Mr. Stonebraker and Eugene Wong, a colleague at the University of California, Berkeley, as well as a team of scientists at I.B.M.

Today, Mr. Stonebraker sees an opportunity for new kinds of ultrafast databases. The new software, he explains, takes advantage of rapid advances in computer hardware to help businesses and researchers find insights in the rising flood of data coming from so many sources, including Web-browsing trails, sensor data, genetic testing and stock trading.

So, at 68, Mr. Stonebraker is a co-founder and chief technology officer of two start-ups in the field of data-driven discovery, VoltDB and Paradigm4.

“Now is the time,” says Mr. Stonebraker, who is an adjunct professor at the Massachusetts Institute of Technology’s computer science and artificial intelligence laboratory. “The economics and the technology are ripe.”

The case for optimism is by no means unqualified. The march of these technologies raises social issues, including privacy concerns, and the timing is uncertain. All of the bold predictions in the 1990s that the Internet would disrupt traditional industries like media, advertising and retailing did come true — a decade later.

But a series of related technologies, scientists and entrepreneurs say, has reached a critical mass — come to a digital boiling point, so to speak — so that new products and capabilities become possible. The technical ingredients, they note, include powerful, low-cost computing and storage spread across thousands of computers. The digital engine rooms of Google and Amazon are prime examples.

Another fast-improving technology involves inexpensive and intelligent sensors, which are crucial to a new breed of automated machines like experimental driverless cars and battlefield drones. Clever software — notably machine-learning algorithms — animates much of the current wave of smarter technology. Two well-known examples are found in Watson, the “Jeopardy”-winning computer from I.B.M., and the movie recommendations on Netflix.

ADVANCES in such underlying technologies are fueling the current excitement in fields like artificial intelligence, robotics and data analysis and prediction. “All parts of the technology pipeline are gearing up at the same time, and that’s how you get this explosion of new applications and uses,” says Jon Kleinberg, a computer scientist at Cornell University.

Behind the seeming explosion, experts say, is a process of technology evolution. Paul Saffo, a technology forecaster, compares the process to the evolutionary biology concept known as “punctuated equilibria” formulated by the paleontologists Stephen Jay Gould and Niles Eldredge. The idea is that species often evolve in periodic spurts.

Yet, they say, there are typically years of progress before a commercial breakthrough in the technological realm.

“Even in Silicon Valley, it takes most technologies 20 years to become overnight successes,” says Mr. Saffo, a consulting professor at Stanford’s school of engineering.

The Internet provides a case study of both technology’s evolutionary progress and its exponential growth. In 1969, there were only four computers connected to the nascent Internet, compared with roughly a billion computing devices today, from laptops to cellphones, says Edward Lazowska, a computer scientist at the University of Washington.

The early increases in connected computers drew scant attention. “But at some point in the late 1990s,” Mr. Lazowska says, “you were going from 4 million to 8 million to 16 million to 32 million to 64 million, and people started to notice that something revolutionary was going on.”

Rocket Fuel is a four-year-old Silicon Valley start-up that uses artificial-intelligence software to place display advertisements for marketers on the Web. The company can not only tailor ads by demographic slices of viewers’ ages, gender and interests, but can also use its predictive algorithms to produce campaigns based on results, says George H. John, the company’s chief executive.

For example, a luxury carmaker might tell Rocket Fuel that it wants to place 100 million ads in the next month, and it will pay the company, say, $80 for generating a sales lead, as evidenced by a potential customer downloading a brochure or filling out an online form.

Rocket Fuel is growing fast, having nearly doubled its work force since the start of the year, to 240. So far in 2012, it has handled campaigns for more than 500 advertisers, including BMW, Duncan Hines, Allstate, Pizza Hut and Ace Hardware. It has raised $76 million in venture funding and debt, and its thousands of computers handle 19 billion bid requests a day on ad exchanges. Each online auction for ad space is typically completed in about 100 milliseconds, a tenth of a second.

Rocket Fuel, Mr. John says, is using some of the ideas he worked on in the 1990s as a doctoral student focusing on artificial intelligence at Stanford — research that was supported with government dollars from the National Science Foundation and other agencies, as is so often the case. In the last few years, building a business around those ideas has become achievable and affordable. “And a lot of it has to do with the underlying technology,” Mr. John says.

FOR Mr. Stonebraker, the hardware advance that opens the door to his start-ups is the striking improvement of solid-state memory, as performance climbs and prices plunge. Solid-state, or flash, memory is most widely known as the lightweight storage technology used in consumer devices like small music players and smartphones.

But increasingly, solid-state memory can be used in big computers, holding a hefty database in memory instead of sending data off to be stored on disk drives. According to Mr. Stonebraker, some data-handling tasks can now be completed 50 times faster than with conventional systems.

“Memory is the new disk,” he says. “The obvious thing to do is to exploit that technology.”

In the yin and yang of computing, it is software that exploits hardware, enabling a computer to do useful things. And machine-learning programs and other data-sifting software are advancing swiftly.

“There is no point in collecting and storing all this data if the algorithms are not able to find useful patterns and insights in the data,” says Mr. Kleinberg at Cornell. “But the software is scaling up to the task.”

A version of this article appeared in print on September 9, 2012, on page BU4 of the New York edition with the headline: Tech’s New Wave, Driven By Data.

Friday, September 7, 2012

One Woman's Data Trail Diary



Scott Shane, The New York Times, August 31, 2012

As part of The Agenda, The Times’s look at major issues facing the next administration, we have been examining the trade-offs, more than a decade after the Sept. 11 attacks, between security and privacy and civil liberties. Some readers have written in about the electronic data trail that all of us leave as we go about our lives, using the Internet and carrying smartphones.


Heidi Boghosian, a New Yorker and author of a book on surveillance scheduled for publication next year, “Spying on Democracy: A Short History of Government/Corporate Collusion in the Technology Age,” agreed to try to document her own data trail on one recent day. Her account, below, is nothing extraordinary – and that’s the point. It is impossible to live in urban, wired America without leaving clues about ourselves, our movements and our views everywhere. And it is all but impossible to be certain who is looking at the resulting data or video and how much of it is accessible to federal, state or local government.

Ms. Boghosian is executive director of the National Lawyers Guild, a group of self-described radical lawyers and law students founded in 1937, and between her day job and her book research, she thinks far more than most people about surveillance and privacy. But the exercise of documenting her day was nonetheless informative, she said.

“Definitely, for me, going through the process reinforced my sense of the role corporations play in our daily lives,” she said. “And I don’t think most people realize the extent to which corporations cooperate in turning over personal information to the government.”

Here is the record she made:

A Day of Surveillance:

(1) 8:30 a.m.: Closed Circuit Television (CCTV) in hallway permits private landlord to monitor departure of tenant from apartment building at 173 Avenue A, New York, N.Y. A sign is posted alerting tenants that their actions are being monitored.

(2) City-owned video surveillance camera, mounted atop a streetlight pole, records pedestrian and vehicular traffic on corner of Avenue A and 11th Street.

(3) 9:45 a.m.: Internet Protocol-based, closed-circuit television CCTV/video surveillance camera at Chase Bank A.T.M. on Second Avenue and 10th Street records clear image of person withdrawing money. I.P. video surveillance footage probably transmitted to a central monitoring room and digitally stored (allowing for advanced search techniques), or viewed over the Internet. Intelligent I.P. cameras with video analytics such as motion sensors, facial recognition and behavioral recognition are used to identify abnormal activity in and around banking locations.

(4) 10 a.m.: Customer Loyalty Card at East Village coffee shop Café Pick Me Up, Avenue A and 9th Street, likely allow the business to track and predict customer spending habits.

(5) 10:30 a.m.: iPhone (with G.P.S. tracking) in owner’s back pocket allows phone owner’s movement and location to be tracked (by government, if cellphone provider gives access) through day and evening, even if phone is turned off, as phone owner walks to Astor Place subway stop.

(6) 10:45 a.m.: Passed by “smart sign” (digital billboard with cameras that gauges demographics of passers-by) that delivers ads tailored to the demographics of the passer-by.

(7) 10:45 a.m.: CCTV in NY subway system monitor boarding #6 subway line to work.

(8) 11:11 a.m.: CCTV in elevator records ride to ninth-floor office in office building. Building security guard has four cameras behind front desk showing elevator, stairways and hallways.

(9) 11:20 a.m.: Facebook software tracks user activities on sites on Internet after logging in and reading a few comments. “Tag Suggestions” feature employs facial recognition technology and suggests name tags after uploading photos.

(10) 11:30 a.m.: Cookies on Internet monitor all Internet searches throughout day on range of subjects; ads appear on screen related to items purchased on line (athletic shoes, face cream).

(11) 1 p.m.: Credit card at Macy’s Department store, used for quick purchase, has embedded RFID (Radio Frequency ID) chip, tracking consumer spending habits and providing that information to big business.

(12) 1:15 p.m.: Shoes in Macy’s new shoe store all have RFID chips (unique identifier linked to database) in them.

(13) 1:30 p.m. Downloaded iTunes onto iPhone. Online music providers may share personal information with third parties.

(14) 2 p.m.: Video cameras and motion detectors in local supermarket track physical movements of customers (allegedly to aim for improved customer service) as customer drops in to pick up some fruit for lunch.

(15) 2:30 p.m.: Social security number and driver’s license information, required by Fulton Street Verizon cellular store winds up in Verizon’s digital database, as customer switches from AT&T. Allegedly needed Social Security number to conduct credit check even though customer has had a landline account with Verizon for many years.

(16) 3:30 p.m.: I.P. address may have been included on a bar code on a digital coupon while registering for at hotelcoupons.com to get a discount hotel deal.

(17) 4 p.m.: Continuous, systematic desktop monitoring surveillance of personal use of business computer to access site to order shoes could have been conducted had employer installed software to monitor my real time actions, purportedly to avoid discrimination and sexual harassment lawsuits that may result from inappropriate e-mails sent within company.

(18) 6 p.m.: Dropped by anti-fracking protest on West 14th Street. Unmarked police van with tinted windows probably had NYPD Technical Assistance Response Unit (TARU) personnel recording protest activities, especially because several Occupy protesters were present. TARU provides investigative technical equipment and tactical support to all N.Y.P.D. bureaus and also provides assistance to other city, state and federal agencies. The unit employs several forms of computer forensics.

(19) 7:30 p.m.: CCTV cameras inside East Village restaurant while meeting friend for dinner after the protest.

(20) 9:30 p.m.: Surveillance cameras on several buildings passed on way home is captured on tape.

(21) 10 p.m.: Final check of Gmail, and a few Google searches, allow Google to collect even more data on consumer habits and personal interests.

Wednesday, September 5, 2012

REINVENTING SOCIETY IN THE WAKE OF BIG DATA


A Conversation with Alex (Sandy) Pentland      Edge Video (24-Minutes)

ALEX 'SANDY' PENTLAND is a pioneer in big data, computational social science, mobile and health systems, and technology for developing countries. He is one of the most-cited computer scientists in the world and was named by Forbes as one of the world's seven most powerful data scientists. He currently directs MIT's Human Dynamics Laboratory and the MIT Media Lab Entrepreneurship Program, and advises the World Economic Forum, Nissan Motor Corporation, and a variety of start-up firms.

Recently I seem to have become MIT's Big Data guy, with people like Tim O'Reilly and "Forbes" calling me one of the seven most powerful data scientists in the world. I'm not sure what all of that means, but I have a distinctive view about Big Data, so maybe it is something that people want to hear.

I believe that the power of Big Data is that it is information about people's behavior instead of information about their beliefs. It's about the behavior of customers, employees, and prospects for your new business. It's not about the things you post on Facebook, and it's not about your searches on Google, which is what most people think about, and it's not data from internal company processes and RFIDs. This sort of Big Data comes from things like location data off of your cell phone or credit card, it's the little data breadcrumbs that you leave behind you as you move around in the world.

"What those breadcrumbs tell is the story of your life. It tells what you've chosen to do. That's very different than what you put on Facebook. What you put on Facebook is what you would like to tell people, edited according to the standards of the day. Who you actually are is determined by where you spend time, and which things you buy. Big data is increasingly about real behavior, and by analyzing this sort of data, scientists can tell an enormous amount about you. They can tell whether you are the sort of person who will pay back loans. They can tell you if you're likely to get diabetes."

Sandy Pentland's EDGE Profile page: http://edge.org/memberbio/alex_(sandy)_pentland

Permalink: http://www.edge.org/conversation/reinventing-society-in-the-wake-of-big-data

[ED. NOTE: Part of the ongoing series "COMPUTATIONAL SOCIAL SCIENCE @ Edge"

http://edge.org/event/special/computational-social-science]

Brookings Report: Big Data for Education: Data Mining, Data Analytics, and Web Dashboards



Darrell M. West, The Brookings Institution, September 4, 2012

Imagine this scenario: twelve-year-old Susan took a course designed to improve her reading skills. She read short stories and the teacher would give her and her fellow students a written test every other week measuring vocabulary and reading comprehension. A few days later, Susan’s instructor graded the paper and returned her exam. The test showed that she did well on vocabulary, but needed to work on retaining key concepts.

In the future, her younger brother Richard is likely to learn reading through a computerized software program. As he goes through each story, the computer will collect data on how long it takes him to master the material. After each assignment, a quiz will pop up on his screen and ask questions concerning vocabulary and reading comprehension. As he answers each item, Richard will get instant feedback showing whether his answer is correct and how his performance compares to classmates and students across the country. For items that are difficult, the computer will send him links to websites that explain words and concepts in greater detail. At the end of the session, his teacher will receive an automated readout on Richard and the other students in the class summarizing their reading time, vocabulary knowledge, reading comprehension, and use of supplemental electronic resources.

In comparing these two learning environments, it is apparent that current school evaluations suffer from several limitations. Many of the typical pedagogies provide little immediate feedback to students, require teachers to spend hours grading routine assignments, aren’t very proactive about showing students how to improve comprehension, and fail to take advantage of digital resources that can improve the learning process. This is unfortunate because data-driven approaches make it possible to study learning in real-time and offer systematic feedback to students and teachers.

In this report, I examine the potential for improved research, evaluation, and accountability through data mining, data analytics, and web dashboards. So-called “big data” make it possible to mine learning information for insights regarding student performance and learning approaches.[1] Rather than rely on periodic test performance, instructors can analyze what students know and what techniques are most effective for each pupil. By focusing on data analytics, teachers can study learning in far more nuanced ways.[2] Online tools enable evaluation of a much wider range of student actions, such as how long they devote to readings, where they get electronic resources, and how quickly they master key concepts.


[1] James Manyika, Michael Chui, Brad Brown, Jacques Bughin, Richard Dobbs, Charles Roxburgh, and Angela Byers, “Big Data: The Next Frontier for Innovation, Competition, and Productivity,” McKinsey Global Institute, May, 2011.
[2] Felix Castro, Alfredo Vellido, Angela Nebot, and Francisco Mugica, “Applying Data Mining Techniques to e-Learning Problems,” Studies in Computational Intelligence, Volume 62, 2007, pp. 183-221.

Tuesday, August 28, 2012

Gartner Says Big Data Makes Organizations Smarter, But Open Data Makes Them Richer

Open Data on the Agenda for Gartner Symposium/ITxpo, October 21-25, Orlando, Florida

STAMFORD, Conn., August 22, 2012—

Whereas "big data" will make organizations smarter, open data will be far more consequential for increasing revenue and business value in today's highly competitive environments, according to Gartner, Inc.

"Big data is a topic of growing interest for many business and IT leaders, and there is little doubt that it creates business value by enabling organizations to uncover previously unseen patterns and develop sharper insights about their businesses and environments," said David Newman, research vice president at Gartner. "However, for clients seeking competitive advantage through direct interactions with customers, partners and suppliers, open data is the solution. For example, more government agencies are now opening their data to the public Web to improve transparency, and more commercial organizations are using open data to get closer to customers, share costs with partners and generate revenue by monetizing information assets."

Gartner analysts believe an open data strategy should be a top priority for any organization that uses the Web as a channel for delivering goods and services. Open data strategies support outside-in business practices that generate growth and innovation. Enterprise architects help their organization connect independent open data projects by creating actionable deliverables and information-sharing practices that generate business-focused outcomes for achieving strategic customer growth and retention objectives.

Gartner analysts said that any business that has a data warehouse should consider how it can use data as a strategic asset and revenue generator. Maturing technologies for data quality and data anonymization can help mitigate regulatory restraints and risk factors. Open data APIs provide simple, Web-oriented means for data exchange, and linked data techniques are effective for generating big datasets. When considering the long-term benefits of an open data strategy, organizations should investigate the types of data exchange now emerging where information producers and consumers share data for profit.

Emerging data marketplaces are also places for organizations to open their data — potentially turning their "data into dollars." The challenge is to keep the barriers to entry low to enable participation by different types of business and streamlined processes for adding and vetting data sources. Monetizing data is a technological and operational challenge. If an enterprise's goal is to unlock its data's full revenue potential, it needs to be able to reach all possible data buyers efficiently.

"With tight budgets and continued economic uncertainty, organizations will need leaders who can craft breakthrough strategies that drive growth and innovation," said Mr. Newman. "As change agents, enterprise architects can help their organizations become richer through strategies such as open data."

Although openness is a pervasive and persistent issue in IT, there is very little agreement about exactly what "open" means. According to Gartner analysts, an informal definition of openness is a level playing field where everyone can play a game that can evolve. There is a positive relationship between the openness of information goods (for example, code, data, content and standards) and information services (for example, services that offer information goods, such as the Internet, Wikipedia, OpenStreetMap and GPS) and the size and diversity of the community sharing them. From the viewpoint of enterprise information architects, this is known as the information-sharing network effect: the business value of a data asset increases the more widely and easily it is shared.

Open data APIs are a lightweight approach to data exchange. Their use is now considered a best practice for opening data and functionality to developers and other businesses. Organizations use APIs to generate new sources of revenue, spur innovation, increase transparency and improve brand equity.

"The challenge for organizations is to determine how best to use APIs and how an open data strategy should align with business priorities," Mr. Newman said. "This is where enterprise architects can help. While some internal IT functions may be using APIs to fulfil local or specific application needs, the enterprise architecture process harvests and elevates good works as first-class strategic priorities that create business-focused outcomes. As a strategic enabler, APIs are a powerful means with which to build an ecosystem, and a first step toward monetizing data assets."

Additional information is available in the Gartner report "Open for Business: Learn to Profit by Open Data." The report is available on Gartner's website at http://www.gartner.com/resId=1947015.

About Gartner Symposium/ITxpo

Gartner Symposium/ITxpo is the world's most important gathering of CIOs and senior IT executives. This event delivers independent and objective content with the authority and weight of the world's leading IT research and advisory organization, and provides access to the latest solutions from key technology providers. Gartner's annual Symposium/ITxpo events are key components of attendees' annual planning efforts. IT executives rely on Gartner Symposium/ITxpo to gain insight into how their organizations can use IT to address business challenges and improve operational efficiency.

Additional information for Gartner Symposium/ITxpo in Orlando, Florida, October 21-25, is available at http://www.gartner.com/us/symposium. Members of the media can register for the event by contacting Janessa Rivera at janessa.rivera@gartner.com.

Additional information from the event will be shared on Twitter at http://twitter.com/Gartner_inc and using #GartnerSym.

Thursday, August 23, 2012

Computational social science: Making the links

From e-mails to social networks, the digital traces left by life in the modern world are transforming social science.
Jim Giles, Nature, August 22, 2012


Jon Kleinberg's early work was not for the mathematically faint of heart. His first publication1, in 1992, was a computer-science paper with contents as dense as its title: 'On dynamic Voronoi diagrams and the minimum Hausdorff distance for point sets under Euclidean motion in the plane'.

That was before the World-Wide Web exploded across the planet, driven by millions of individual users making independent decisions about who and what to link to. And it was before Kleinberg began to study the vast array of digital by-products generated by life in the modern world, from e-mails, mobile phone calls and credit-card purchases to Internet searches and social networks. Today, as a computer scientist at Cornell University in Ithaca, New York, Kleinberg uses these data to write papers such as 'How bad is forming your own opinion?'2 and 'You had me at hello: how phrasing affects memorability'3 — titles that would be at home in a social-science journal.


“I realized that computer science is not just about technology,” he explains. “It is also a human topic.”

Kleinberg is not alone. The emerging field of computational social science is attracting mathematically inclined scientists in ever-increasing numbers. This, in turn, is spurring the creation of academic departments and prompting companies such as the social-network giant Facebook, based in Menlo Park, California, to establish research teams to understand the structure of their networks and how information spreads across them.

“It's been really transformative,” says Michael Macy, a social scientist at Cornell and one of 15 co-authors of a 2009 manifesto4 seeking to raise the profile of the new discipline. “We were limited before to surveys, which are retrospective, and lab experiments, which are almost always done on small numbers of college sophomores.” Now, he says, the digital data-streams promise a portrait of individual and group behaviour at unprecedented scales and levels of detail. They also offer plenty of challenges — notably privacy issues, and the problem that the data sets may not truly be reflective of the population at large.

Nonetheless, says Macy, “I liken the opportunities to the changes in physics brought about by the particle accelerator, and in neuroscience by functional magnetic resonance imaging”.

Social calls

An early example of large-scale digital data being used on a social-science issue was a study in 2002 by Kleinberg and David Liben-Nowell, a computer scientist at Carleton College in Northfield, Minnesota. They looked at a mechanism that social scientists believed helped drive the formation of personal relationships: people tend to become friends with the friends of their friends. Although well established, the idea had never been tested on networks of more than a few tens or hundreds of people.

Kleinberg and Liben-Nowell studied the relationships formed in scientific collaborations. They looked at the thousands of physicists who uploaded papers to the arXiv preprint server during 1994–96. By writing software to automatically extract names from the papers, the pair built up a digital network several orders of magnitude larger than any that had been examined before, with each link representing two researchers who had collaborated. By following how the network changed over time, the researchers identified several measures of closeness among the researchers that could be used to forecast future collaborations5.


 
                                                                   
As expected, the results showed that new collaborations tended to spring from researchers whose spheres of existing collaborators overlapped — the research analogue of 'friends of friends'. But the mathematical sophistication of the predictions has allowed them to be used on even larger networks. Kleinberg's former PhD student, Lars Backstrom, also worked on the connection-prediction problem — experience that he has put to good use now that he works at Facebook, where he designed the social network's current friend-recommendation system.

Another long-standing social-science idea affirmed by computational researchers is the importance of 'weak ties' — relationships with distant acquaintances who are encountered relatively rarely. In 1973, Mark Granovetter, a social scientist now at Stanford University in Stanford, California, argued that weak ties form bridges between social cliques and so are important to the spread of information and to economic mobility6. In the pre-digital era it was almost impossible to verify his ideas at scale. But in 2007, a team led by Jukka-Pekka Onnela, a network scientist now at Harvard University in Cambridge, Massachusetts, used data on 4 million mobile-phone users to confirm that weak ties do indeed act as societal bridges7 (see 'The power of weak ties').

In 2010, a second group, which included Macy, showed that Granovetter was also right about the connection between economic mobility and weak ties. Using data from 65 million landlines and mobile phones in the United Kingdom, together with national census data, they revealed a powerful correlation between the diversity of individuals' relationships and economic development: the richer and more varied their connections, the richer their communities8 (see 'The economic link'). “We didn't imagine in the 1970s that we could work with data on this scale,” says Granovetter.

Infectious ideas

In some instances, big data have showed that long-standing ideas are wrong. This year, Kleinberg and his colleagues used data from the roughly 900 million users of Facebook to study contagion in social networks — a process that describes the spread of ideas such as fads, political opinions, new technologies and financial decisions. Almost all theories had assumed that the process mirrors viral contagion: the chance of a person adopting a new idea increases with the number of believers to which he or she is exposed.
 
Kleinberg's student Johan Ugander found that there is more to it than that: people's decision to join Facebook varies not with the total number of friends who are already using the site, but with the number of distinct social groups those friends occupy9. In other words, finding that Facebook is being used by people from, say, your work, your sports club and your close friends makes more of an impression than finding that friends from only one group use it. The conclusion — that the spread of ideas depends on the variety of people that hold them — could be important for marketing and public-health campaigns.

As computational social-science studies have proliferated, so have ideas about practical applications. At the Massachusetts Institute of Technology in Cambridge, computer scientist Alex Pentland's group uses smartphone apps and wearable recording devices to collect fine-grained data on subjects' daily movements and communications. By combining the data with surveys of emotional and physical health, the team has learned how to spot the emergence of health problems such as depression10. “We see groups that never call out,” says Pentland. “Being able to see isolation is really important when it comes to reaching people who need to be reached.” Ginger.io, a spin-off company in Cambridge, Massachusetts, led by Pentland's former student Anmol Madan, is now developing a smartphone app that notifies health-care providers when it spots a pattern in the data that may indicate a health problem.

Other companies are exploiting the more than 400 million messages that are sent every day on Twitter. Several research groups have developed software to analyse the sentiments expressed in tweets to predict real-world outcomes such as box-office revenues for films or election results11. Although the accuracy of such predictions is still a matter of debate12, Twitter began in August to post a daily political index for the US presidential election based on just such methods (election.twitter.com). At Indiana University in Bloomington, meanwhile, Johan Bollen and his colleagues have used similar software to search for correlations between public mood, as expressed on Twitter, and stock-market fluctuations13. Their results have been powerful enough for Derwent Capital, a London-based investment firm, to license Bollen's techniques.

Message received

When such Twitter-based polls began to appear around two years ago, critics wondered whether the service's relative popularity among specific demographic groups, such as young people, would skew the results. A similar debate revolves around all of the new data sets. Facebook, for example, now has close to a billion users, yet young people are still overrepresented among them. There are also differences between online and real-world communication, and it is not clear whether results from one sphere will apply in the other. “We often extrapolate from how one technology is used by one group to how humans in general interact,” notes Samuel Arbesman, a network scientist at Harvard University. But that, he says, “might not necessarily be reasonable”.

Proponents counter that these are not new problems. Almost all survey data contain some amount of demographic skew, and social scientists have developed a variety of weighting methods to redress the balance. If the bias in a particular data set, such as an excess of one group or another on Facebook, is understood, the results can be adjusted to account for it.

““We didn't imagine in the 1970s that we could work with data on this scale.””

Services such as Facebook and Twitter are also becoming increasingly widely used, reducing the bias. And even if the bias remains, it is arguably less severe than that in other data sets such as those for psychology and human behaviour, where most work is done on university students from Western, educated, industrialized, rich and democratic societies (often denoted WEIRD).

Granovetter has a more philosophical reservation about the influx of big data into his field. He says he is “very interested” in the new methods, but fears that the focus on data detracts from the need to get a better theoretical grasp on social systems. “Even the very best of these computational articles are largely focused on existing theories,” he says. “That's valuable, but it is only one piece of what needs to be done.” Granovetter's weak-ties paper6, for example, remains highly cited almost 40 years later. Yet it was “more or less data-free”, he says. “It didn't result from data analyses, it resulted from thinking about other studies. That is a separate activity and we need to have people doing that.”

The new breed of social scientists are also wrestling with the issue of data access. “Many of the emerging 'big data' come from private sources that are inaccessible to other researchers,” Bernardo Huberman, a computer scientist at HP Labs in Palo Alto, wrote in February14. “The data source may be hidden, compounding problems of verification, as well as concerns about the generality of the results.”

A prime example is Facebook's in-house research team, which routinely uses data about the interactions among the network's 900 million users for its own studies, including a re-evaluation of the famous claim that any two people on Earth are just six introductions apart. (It puts the figure at five15.) But the group publishes only the conclusions, not the raw data, in part because of privacy concerns. In July, Facebook announced that it was exploring a plan that would give external researchers the chance to check the in-house group's published conclusions against aggregated, anonymized data — but only for a limited time, and only if the outsiders first travelled to Facebook headquarters16.

In the short term, computational social scientists are more concerned about cultural problems in their discipline. Several institutions, including Harvard, have created programmes in the new field, but the power of academic boundaries is such that there is often little traffic between different departments. At Columbia University in New York, social scientist and network theorist Duncan Watts recalls a recent scheduling error that forced him to combine meetings with graduate students in computer science and sociology. “It was abundantly clear that these two groups could really use each other: the computer-science students had much better methodological chops than their sociology counterparts, but the sociologists had much more interesting questions,” he says. “And yet they'd never heard of each other, nor had it ever occurred to any of them to walk over to the other's department.”

Many researchers remain unaware of the power of the new data, agrees Harvard social scientist David Lazar, lead author on the 2009 manifesto. Little data-driven work is making it into top social-science journals. And computer-science conferences that focus on social issues, such as the Conference on Weblogs and Social Media, held in Dublin in June, attract few social scientists.

Nonetheless, says Lazar, with landmark papers appearing in leading journals and data sets on societal-wide behaviours available for the first time, those barriers are steadily breaking down. “The changes are more in front of us than behind us,” he says.

Certainly that is Kleinberg's perception. “I think of myself as a computer scientist who is interested in social questions,” he says. “But these boundaries are becoming hard to discern.”

Tuesday, August 14, 2012

Tim O'Reilly: Solving the Wanamaker problem for health care

Data science and technology give us the tools to revolutionize health care. Now we have to put them to use.

Tim O’Reilly, Julie Steele, Mike Lourdes and Colin Hill,  O'Reilly Radar, August 14, 2012

“The best minds of my generation are thinking about how to make people click ads.” — Jeff Hammerbacher, early Facebook employee
“Work on stuff that matters.” — Tim O’Reilly

In the early days of the 20th century, department store magnate John Wanamaker famously said, “I know that half of my advertising doesn’t work. The problem is that I don’t know which half.” 

The consumer Internet revolution was fueled by a search for the answer to Wanamaker’s question. Google AdWords and the pay-per-click model transformed a business in which advertisers paid for ad impressions into one in which they pay for results. “Cost per thousand impressions” (CPM) was replaced by “cost per click” (CPC), and a new industry was born. It’s important to understand why CPC replaced CPM, though. Superficially, it’s because Google was able to track when a user clicked on a link, and was therefore able to bill based on success. But billing based on success doesn’t fundamentally change anything unless you can also change the success rate, and that’s what Google was able to do. By using data to understand each user’s behavior, Google was able to place advertisements that an individual was likely to click. They knew “which half” of their advertising was more likely to be effective, and didn’t bother with the rest.

Since then, data and predictive analytics have driven ever deeper insight into user behavior such that companies like Google, Facebook, Twitter, Zynga, and LinkedIn are fundamentally data companies. And data isn’t just transforming the consumer Internet. It is transforming finance, design, and manufacturing — and perhaps most importantly, health care.

How is data science transforming health care? There are many ways in which health care is changing, and needs to change. We’re focusing on one particular issue: the problem Wanamaker described when talking about his advertising. How do you make sure you’re spending money effectively? Is it possible to know what will work in advance?

Too often, when doctors order a treatment, whether it’s surgery or an over-the-counter medication, they are applying a “standard of care” treatment or some variation that is based on their own intuition, effectively hoping for the best. The sad truth of medicine is that we don’t really understand the relationship between treatments and outcomes. We have studies to show that various treatments will work more often than placebos; but, like Wanamaker, we know that much of our medicine doesn’t work for half or our patients, we just don’t know which half. At least, not in advance. One of data science’s many promises is that, if we can collect data about medical treatments and use that data effectively, we’ll be able to predict more accurately which treatments will be effective for which patient, and which treatments won’t.

A better understanding of the relationship between treatments, outcomes, and patients will have a huge impact on the practice of medicine in the United States. Health care is expensive.

The U.S. spends over $2.6 trillion on health care every year, an amount that constitutes a serious fiscal burden for government, businesses, and our society as a whole. These costs include over $600 billion of unexplained variations in treatments: treatments that cause no differences in outcomes, or even make the patient’s condition worse. We have reached a point at which our need to understand treatment effectiveness has become vital — to the health care system and to the health and sustainability of the economy overall.

Why do we believe that data science has the potential to revolutionize health care? After all, the medical industry has had data for generations: clinical studies, insurance data, hospital records. But the health care industry is now awash in data in a way that it has never been before: from biological data such as gene expression, next-generation DNA sequence data, proteomics, and metabolomics, to clinical data and health outcomes data contained in ever more prevalent electronic health records (EHRs) and longitudinal drug and medical claims. 

We have entered a new era in which we can work on massive datasets effectively, combining data from clinical trials and direct observation by practicing physicians (the records generated by our $2.6 trillion of medical expense). When we combine data with the resources needed to work on the data, we can start asking the important questions, the Wanamaker questions, about what treatments work and for whom. 

The opportunities are huge: for entrepreneurs and data scientists looking to put their skills to work disrupting a large market, for researchers trying to make sense out of the flood of data they are now generating, and for existing companies (including health insurance companies, biotech, pharmaceutical, and medical device companies, hospitals and other care providers) that are looking to remake their businesses for the coming world of outcome-based payment models.

Making health care more effective
What, specifically, does data allow us to do that we couldn’t do before? For the past 60 or so years of medical history, we’ve treated patients as some sort of an average. A doctor would diagnose a condition and recommend a treatment based on what worked for most people, as reflected in large clinical studies. Over the years, we’ve become more sophisticated about what that average patient means, but that same statistical approach didn’t allow for differences between patients. A treatment was deemed effective or ineffective, safe or unsafe, based on double-blind studies that rarely took into account the differences between patients.

With the data that’s now available, we can go much further. The exceptions to this are relatively recent and have been dominated by cancer treatments, the first being Herceptin for breast cancer in women who over-express the Her2 receptor. With the data that’s now available, we can go much further for a broad range of diseases and interventions that are not just drugs but include surgery, disease management programs, medical devices, patient adherence, and care delivery. 

For a long time, we thought that Tamoxifen was roughly 80% effective for breast cancer patients. But now we know much more: we know that it’s 100% effective in 70 to 80% of the patients, and ineffective in the rest. That’s not word games, because we can now use genetic markers to tell whether it’s likely to be effective or ineffective for any given patient, and we can tell in advance whether to treat with Tamoxifen or to try something else. 

Two factors lie behind this new approach to medicine: a different way of using data, and the availability of new kinds of data. It’s not just stating that the drug is effective on most patients, based on trials (indeed, 80% is an enviable success rate); it’s using artificial intelligence techniques to divide the patients into groups and then determine the difference between those groups. We’re not asking whether the drug is effective; we’re asking a fundamentally different question: “for which patients is this drug effective?” We’re asking about the patients, not just the treatments. A drug that’s only effective on 1% of patients might be very valuable if we can tell who that 1% is, though it would certainly be rejected by any traditional clinical trial. 

More than that, asking questions about patients is only possible because we’re using data that wasn’t available until recently: DNA sequencing was only invented in the mid-1970s, and is only now coming into its own as a medical tool. What we’ve seen with Tamoxifen is as clear a solution to the Wanamaker problem as you could ask for: we now know when that treatment will be effective. If you can do the same thing with millions of cancer patients, you will both improve outcomes and save money. 

Dr. Lukas Wartman, a cancer researcher who was himself diagnosed with terminal leukemia, was successfully treated with sunitinib, a drug that was only approved for kidney cancer. Sequencing the genes of both the patient’s healthy cells and cancerous cells led to the discovery of a protein that was out of control and encouraging the spread of the cancer. The gene responsible for manufacturing this protein could potentially be inhibited by the kidney drug, although it had never been tested for this application. This unorthodox treatment was surprisingly effective: Wartman is now in remission. 

While this treatment was exotic and expensive, what’s important isn’t the expense but the potential for new kinds of diagnosis. The price of gene sequencing has been plummeting; it will be a common doctor’s office procedure in a few years. And through Amazon and Google, you can now “rent” a cloud-based supercomputing cluster that can solve huge analytic problems for a few hundred dollars per hour. What is now exotic inevitably becomes routine. 

But even more important: we’re looking at a completely different approach to treatment. Rather than a treatment that works 80% of the time, or even 100% of the time for 80% of the patients, a treatment might be effective for a small group. It might be entirely specific to the individual; the next cancer patient may have a different protein that’s out of control, an entirely different genetic cause for the disease. Treatments that are specific to one patient don’t exist in medicine as it’s currently practiced; how could you ever do an FDA trial for a medication that’s only going to be used once to treat a certain kind of cancer? 

Foundation Medicine is at the forefront of this new era in cancer treatment. They use next-generation DNA sequencing to discover DNA sequence mutations and deletions that are currently used in standard of care treatments, as well as many other actionable mutations that are tied to drugs for other types of cancer. They are creating a patient-outcomes repository that will be the fuel for discovering the relation between mutations and drugs. 

Foundation has identified DNA mutations in 50% of cancer cases for which drugs exist (information via a private communication), but are not currently used in the standard of care for the patient’s particular cancer.

The ability to do large-scale computing on genetic data gives us the ability to understand the origins of disease. If we can understand why an anti-cancer drug is effective (what specific proteins it affects), and if we can understand what genetic factors are causing the cancer to spread, then we’re able to use the tools at our disposal much more effectively. Rather than using imprecise treatments organized around symptoms, we’ll be able to target the actual causes of disease, and design treatments tuned to the biology of the specific patient. 

Eventually, we’ll be able to treat 100% of the patients 100% of the time, precisely because we realize that each patient presents a unique problem. 

Personalized treatment is just one area in which we can solve the Wanamaker problem with data. Hospital admissions are extremely expensive. Data can make hospital systems more efficient, and to avoid preventable complications such as blood clots and hospital re-admissions. It can also help address the challenge of hot-spotting (a term coined by Atul Gawande): finding people who use an inordinate amount of health care resources. By looking at data from hospital visits, Dr. Jeffrey Brenner of Camden, NJ, was able to determine that “just one per cent of the hundred thousand people who made use of Camden’s medical facilities accounted for thirty per cent of its costs.” Furthermore, many of these people came from two apartment buildings. Designing more effective medical care for these patients was difficult; it doesn’t fit our health insurance system, the patients are often dealing with many serious medical issues (addiction and obesity are frequent complications), and have trouble trusting doctors and social workers. It’s counter-intuitive, but spending more on some patients now results in spending less on them when they become really sick. While it’s a work in progress, it looks like building appropriate systems to target these high-risk patients and treat them before they’re hospitalized will bring significant savings. 

Many poor health outcomes are attributable to patients who don’t take their medications. Eliza, a Boston-based company started by Alexandra Drane, has pioneered approaches to improve compliance through interactive communication with patients. Eliza improves patient drug compliance by tracking which types of reminders work on which types of people; it’s similar to the way companies like Google target advertisements to individual consumers. By using data to analyze each patient’s behavior, Eliza can generate reminders that are more likely to be effective. The results aren’t surprising: if patients take their medicine as prescribed, they are more likely to get better. And if they get better, they are less likely to require further, more expensive treatment. Again, we’re using data to solve Wanamaker’s problem in medicine: we’re spending our resources on what’s effective, on appropriate reminders that are mostly to get patients to take their medications.

More data, more sources
The examples we’ve looked at so far have been limited to traditional sources of medical data: hospitals, research centers, doctor’s offices, insurers. The Internet has enabled the formation of patient networks aimed at sharing data. Health social networks now are some of the largest patient communities. As of November 2011, PatientsLikeMe has over 120,000 patients in 500 different condition groups; ACOR has over 100,000 patients in 127 cancer support groups; 23andMe has over 100,000 members in their genomic database; and diabetes health social network SugarStats has over 10,000 members. These are just the larger communities, thousands of small communities are created around rare diseases, or even uncommon experiences with common diseases. All of these communities are generating data that they voluntarily share with each other and the world. 

Increasingly, what they share is not just anecdotal, but includes an array of clinical data. For this reason, these groups are being recruited for large-scale crowdsourced clinical outcomes research.

Thanks to ubiquitous data networking through the mobile network, we can take several steps further. In the past two or three years, there’s been a flood of personal fitness devices (such as the Fitbit) for monitoring your personal activity. There are mobile apps for taking your pulse, and an iPhone attachment for measuring your glucose. There has been talk of mobile applications that would constantly listen to a patient’s speech and detect changes that might be the precursor for a stroke, or would use the accelerometer to report falls. Tanzeem Choudhury has developed an app called Be Well that is intended primarily for victims of depression, though it can be used by anyone. Be Well monitors the user’s sleep cycles, the amount of time they spend talking, and the amount of time they spend walking. The data is scored, and the app makes appropriate recommendations, based both on the individual patient and data collected across all the app’s users.

Continuous monitoring of critical patients in hospitals has been normal for years; but we now have the tools to monitor patients constantly, in their home, at work, wherever they happen to be. And if this sounds like big brother, at this point most of the patients are willing. We don’t want to transform our lives into hospital experiences; far from it! But we can collect and use the data we constantly emit, our “data smog,” to maintain our health, to become conscious of our behavior, and to detect oncoming conditions before they become serious. The most effective medical care is the medical care you avoid because you don’t need it.

Paying for results
Once we’re on the road toward more effective health care, we can look at other ways in which Wanamaker’s problem shows up in the medical industry. It’s clear that we don’t want to pay for treatments that are ineffective. Wanamaker wanted to know which part of his advertising was effective, not just to make better ads, but also so that he wouldn’t have to buy the advertisements that wouldn’t work. He wanted to pay for results, not for ad placements. Now that we’re starting to understand how to make treatment effective, now that we understand that it’s more than rolling the dice and hoping that a treatment that works for a typical patient will be effective for you, we can take the next step: Can we change the underlying incentives in the medical system? Can we make the system better by paying for results, rather than paying for procedures? 

It’s shocking just how badly the incentives in our current medical system are aligned with outcomes. If you see an orthopedist, you’re likely to get an MRI, most likely at a facility owned by the orthopedist’s practice. On one hand, it’s good medicine to know what you’re doing before you operate. But how often does that MRI result in a different treatment? How often is the MRI required just because it’s part of the protocol, when it’s perfectly obvious what the doctor needs to do? Many men have had PSA tests for prostate cancer; but in most cases, aggressive treatment of prostate cancer is a bigger risk than the disease itself. Yet the test itself is a significant profit center. Think again about Tamoxifen, and about the pharmaceutical company that makes it. In our current system, what does “100% effective in 80% of the patients” mean, except for a 20% loss in sales? That’s because the drug company is paid for the treatment, not for the result; it has no financial interest in whether any individual patient gets better. (Whether a statistically significant number of patients has side-effects is a different issue.) And at the same time, bringing a new drug to market is very expensive, and might not be worthwhile if it will only be used on the remaining 20% of the patients. And that’s assuming that one drug, not two, or 20, or 200 will be required to treat the unlucky 20% effectively.

It doesn’t have to be this way. 

In the U.K., Johnson & Johnson, faced with the possibility of losing reimbursements for their multiple myeloma drug Velcade, agreed to refund the money for patients who did not respond to the drug. Several other pay-for-performance drug deals have followed since, paving the way for the ultimate transition in pharmaceutical company business models in which their product is health outcomes instead of pills. Such a transition would rely more heavily on real-world outcome data (are patients actually getting better?), rather than controlled clinical trials, and would use molecular diagnostics to create personalized “treatment algorithms.” Pharmaceutical companies would also focus more on drug compliance to ensure health outcomes were being achieved. This would ultimately align the interests of drug makers with patients, their providers, and payors.

Similarly, rather than paying for treatments and procedures, can we pay hospitals and doctors for results? That’s what Accountable Care Organizations (ACOs) are about. ACOs are a leap forward in business model design, where the provider shoulders any financial risk. 
ACOs represent a new framing of the much maligned HMO approaches from the ’90s, which did not work. HMOs tried to use statistics to predict and prevent unneeded care. The ACO model, rather than controlling doctors with what the data says they “should” do, uses data to measure how each doctor performs. Doctors are paid for successes, not for the procedures they administer. The main advantage that the ACO model has over the HMO model is how good the data is, and how that data is leveraged. The ACO model aligns incentives with outcomes: a practice that owns an MRI facility isn’t incentivized to order MRIs when they’re not necessary. It is incentivized to use all the data at its disposal to determine the most effective treatment for the patient, and to follow through on that treatment with a minimum of unnecessary testing.

When we know which procedures are likely to be successful, we’ll be in a position where we can pay only for the health care that works. When we can do that, we’ve solved Wanamaker’s problem for health care.

Enabling data
Data science is not optional in health care reform; it is the linchpin of the whole process. All of the examples we’ve seen, ranging from cancer treatment to detecting hot spots where additional intervention will make hospital admission unnecessary, depend on using data effectively: taking advantage of new data sources and new analytics techniques, in addition to the data the medical profession has had all along. 

But it’s too simple just to say “we need data.” We’ve had data all along: handwritten records in manila folders on acres and acres of shelving. Insurance company records. But it’s all been locked up in silos: insurance silos, hospital silos, and many, many doctor’s office silos. Data doesn’t help if it can’t be moved, if data sources can’t be combined. 

There are two big issues here. First, a surprising amount of medical records are still either hand-written, or in digital formats that are scarcely better than hand-written (for example, scanned images of hand-written records). Getting medical records into a format that’s computable is a prerequisite for almost any kind of progress. Second, we need to break down those silos. 

Anyone who has worked with data knows that, in any problem, 90% of the work is getting the data in a form in which it can be used; the analysis itself is often simple. We need electronic health records: patient data in a more-or-less standard form that can be shared efficiently, data that can be moved from one location to another at the speed of the Internet. Not all data formats are created equal, and some are certainly better than others: but at this point, any machine-readable format, even simple text files, is better than nothing. While there are currently hundreds of different formats for electronic health records, the fact that they’re electronic means that they can be converted from one form into another. Standardizing on a single format would make things much easier, but just getting the data into some electronic form, any, is the first step.

Once we have electronic health records, we can link doctor’s offices, labs, hospitals, and insurers into a data network, so that all patient data is immediately stored in a data center: every prescription, every procedure, and whether that treatment was effective or not. This isn’t some futuristic dream; it’s technology we have now. Building this network would be substantially simpler and cheaper than building the networks and data centers now operated by Google, Facebook, Amazon, Apple, and many other large technology companies. It’s not even close to pushing the limits. 

Electronic health records enable us to go far beyond the current mechanism of clinical trials. In the past, once a drug has been approved in trials, that’s effectively the end of the story: running more tests to determine whether it’s effective in practice would be a huge expense. A physician might get a sense for whether any treatment worked, but that evidence is essentially anecdotal: it’s easy to believe that something is effective because that’s what you want to see. And if it’s shared with other doctors, it’s shared while chatting at a medical convention. But with electronic health records, it’s possible (and not even terribly expensive) to collect documentation from thousands of physicians treating millions of patients. We can find out when and where a drug was prescribed, why, and whether there was a good outcome. We can ask questions that are never part of clinical trials: is the medication used in combination with anything else? What other conditions is the patient being treated for? We can use machine learning techniques to discover unexpected combinations of drugs that work well together, or to predict adverse reactions. We’re no longer limited by clinical trials; every patient can be part of an ongoing evaluation of whether his treatment is effective, and under what conditions. Technically, this isn’t hard. The only difficult part is getting the data to move, getting data in a form where it’s easily transferred from the doctor’s office to analytics centers.

To solve problems of hot-spotting (individual patients or groups of patients consuming inordinate medical resources) requires a different combination of information. You can’t locate hot spots if you don’t have physical addresses. Physical addresses can be geocoded (converted from addresses to longitude and latitude, which is more useful for mapping problems) easily enough, once you have them, but you need access to patient records from all the hospitals operating in the area under study. And you need access to insurance records to determine how much health care patients are requiring, and to evaluate whether special interventions for these patients are effective. Not only does this require electronic records, it requires cooperation across different organizations (breaking down silos), and assurance that the data won’t be misused (patient privacy). Again, the enabling factor is our ability to combine data from different sources; once you have the data, the solutions come easily. 

Breaking down silos has a lot to do with aligning incentives. Currently, hospitals are trying to optimize their income from medical treatments, while insurance companies are trying to optimize their income by minimizing payments, and doctors are just trying to keep their heads above water. There’s little incentive to cooperate. But as financial pressures rise, it will become critically important for everyone in the health care system, from the patient to the insurance executive, to assume that they are getting the most for their money. While there’s intense cultural resistance to be overcome (through our experience in data science, we’ve learned that it’s often difficult to break down silos within an organization, let alone between organizations), the pressure of delivering more effective health care for less money will eventually break the silos down. The old zero-sum game of winners and losers must end if we’re going to have a medical system that’s effective over the coming decades.

Data becomes infinitely more powerful when you can mix data from different sources: many doctor’s offices, hospital admission records, address databases, and even the rapidly increasing stream of data coming from personal fitness devices. The challenge isn’t employing our statistics more carefully, precisely, or guardedly. It’s about letting go of an old paradigm that starts by assuming only certain variables are key and ends by correlating only these variables. This paradigm worked well when data was scarce, but if you think about, these assumptions arise precisely because data is scarce. We didn’t study the relationship between leukemia and kidney cancers because that would require asking a huge set of questions that would require collecting a lot of data; and a connection between leukemia and kidney cancer is no more likely than a connection between leukemia and flu. But the existence of data is no longer a problem: we’re collecting the data all the time. Electronic health records let us move the data around so that we can assemble a collection of cases that goes far beyond a particular practice, a particular hospital, a particular study. So now, we can use machine learning techniques to identify and test all possible hypotheses, rather than just the small set that intuition might suggest. And finally, with enough data, we can get beyond correlation to causation: rather than saying “A and B are correlated,” we’ll be able to say “A causes B,” and know what to do about it. 

Building the health care system we want
The U.S. ranks 37th out of developed economies in life expectancy and other measures of health, while by far outspending other countries on per-capita health care costs. We spend 18% of GDP on health care, while other countries on average spend on the order of 10% of GDP. We spend a lot of money on treatments that don’t work, because we have a poor understanding at best of what will and won’t work. 

Part of the problem is cultural. In a country where even pets can have hip replacement surgery, it’s hard to imagine not spending every penny you have to prolong Grandma’s life — or your own. The U.S. is a wealthy nation, and health care is something we choose to spend our money on. But wealthy or not, nobody wants ineffective treatments. Nobody wants to roll the dice and hope that their biology is similar enough to a hypothetical “average” patient. No one wants a “winner take all” payment system in which the patient is always the loser, paying for procedures whether or not they are helpful or necessary. Like Wanamaker with his advertisements, we want to know what works, and we want to pay for what works. We want a smarter system where treatments are designed to be effective on our individual biologies; where treatments are administered effectively; where our hospitals our used effectively; and where we pay for outcomes, not for procedures.

We’re on the verge of that new system now. We don’t have it yet, but we can see it around the corner. Ultra-cheap DNA sequencing in the doctor’s office, massive inexpensive computing power, the availability of EHRs to study whether treatments are effective even after the FDA trials are over, and improved techniques for analyzing data are the tools that will bring this new system about. The tools are here now; it’s up to us to put them into use.

Recommended reading:

We recommend the following books regarding technology, data, and health care reform:
·       Ahier, Brian. “Big data is the next big thing in health IT,” O’Reilly Radar. February 27, 2012.
·       Bigelow, Bruce. “Big Data, Big Biology, and the ‘Tipping Point’ in Quantified Health,” Xconomy. April 26, 2012.
·       Brawley, Otis Webb. How We Do Harm: A Doctor Breaks Ranks About Being Sick in America. St. Marten’s Press, 2012.
·       Christensen, Clayton M. et al. The Innovator’s Prescription: A Disruptive Solution for Health Care. McGraw Hill, 2008.
·       Howard, Alex. “Data for the Public Good,” O’Reilly Radar. February 22, 2012.
·       Manyika, James et al. “Big data: The next frontier for innovation, competition, and productivity,” McKinsey Global Institute. May, 2011.
·       Oram, Andy. “Five tough lessons I had to learn about health care,” O’Reilly Radar. March 26, 2012.
·       Shah, Nigam H and Jessica D Tenenbaum. “The coming age of data-driven medicine: translational bioinformatics’ next frontier,” Journal of the American Medical Informatics Association (JAMIA). March 26, 2012.
·       Trotter, Fred and David Uhlman. Meaningful Use and Beyond. O’Reilly Media, 2011.
·       Wilbanks, John. “Valuing Health Care: Improving Productivity and Quality” [PDF], Ewing Marion Kauffman Foundation. April, 2012.