Tuesday, October 4, 2011

Rob Atkinson: Government Opportunities to Harness “Big Data”

5150336351_ae2a64336a
Recently more attention has been drawn to the emergence of "Big Data"—large scale data sets that businesses and government are using to unlock new value using today's computing and communications power. As a McKinsey Global Institute (MGI) study recently showed, Big Data offers a wide range of commercial opportunities in virtually every sector of the economy for the United States. To take one example, the MGI estimates that better use of big data in health care could generate an additional $300 billion, with approximately two-thirds of that values coming from more efficient delivery of health care.

The use of Big Data should not be confined to just the private sector; data offers incredible new opportunities to the public sector as well.Policymakers have the opportunity to use Big Data to improve government in areas such as public safety, public health, public utilities and public transportation.

ITIF has discussed many of these opportunities before.
Consider the following:
Better use of data can help government agencies, from city agencies to federal bureaucracies, operate more efficiently, create more transparency, and make more informed decisions. And government can use cloud computing to more efficiently develop online systems that provide anytime, anywhere access to information. However, government officials should do more to spur uses of data. Taking advantage of these opportunities will require federal government leadership, such as the Department of Commerce creating a data policy office to spur data innovation and overcome obstacles to adoption, all the while protecting privacy. And going forward, government agencies will increasingly have to deal with issues such as data security and identity management, so these issues do not become impediments to successful utilization of data analytics. Local governments can help pioneer the use of data as well.  For example, the city of Boston city sponsored the development of a mobile app "Street Bump" to automatically determine where potholes are based on data collected using citizen's smart phones equipped with GPS and accelerometers. Tools like these are helping create "smart cities" and build a world that is alive with information.

Although there have been many successes in this area, much more can be done. For example, in homeland security, law enforcement must deal with a changing threat landscape. While corporations and individuals can increasingly use better technology to communicate and store data security, criminals can also use these same tools.  As a result, law enforcement is increasingly confronting the "
Going Dark" problem where they have less access to investigative data, not because of a lack of legal authority, but because of technological hurdles. Yet while law enforcement may have a reduced ability to intercept some types of communication, they now have many more sources of data, such as transactional data, to use to detect threats. As ITIF discussed at an event in 2010 following the Christmas Day terrorist attempt, the intelligence community still needs to develop better analytical tools to "connect the dots" and allow intelligence officers to do a better job. Similarly in many other sectors, Big Data offers government opportunities to reinvent how to operate effectively. 

Overall, more investment in data infrastructure and analytics will enable government to better provide and efficiently deliver values and services to its citizens.

Silicon Valley reborn as Smartphone Valley (FT)

By Chris Nuttall in Palo Alto    October 3, 2011 4:55 pm   Finncial Times

 Apple’s unveiling of the iPhone 5 on Tuesday at its Cupertino headquarters is just the latest sign that Silicon Valley is taking on a fresh mantle of Smartphone Valley, with its growing reputation making it a magnet for mobile operators around the world.

AT&T, Verizon and Vodafone have all just opened research, testing and incubation centres in San Francisco and the Valley only weeks apart.

 “The reason we are here is this is the centre of the Earth right now when it comes to innovation,” said Fay Arjomandi, head of research and development for Vodafone in the US, at the opening of its Xone Lab in Redwood City last month.

Apple’s iPhone and Google’s Android operating system have set new standards for hardware and software in the mobile industry.

The app culture that both have engendered means there is a race on among operators to feature the newest trends and advances in software first, amid fierce competition for the attention of Silicon Valley developers.
Vodafone’s lab allows developers to incubate their ideas and test how apps and services would perform on the worldwide networks it replicates at its facility.

“We want to bring to market these ideas and accelerate the process – the hope is that six to nine months after we prove something here, it can be in the hands of a user in one of our operating countries,” says Ms Arjomandi.

A week after the Vodafone opening, AT&T launched its Foundry innovation lab in an old French steam laundry in Palo Alto.

It too emphasised the need to move “at web developer rather than carrier speed” – to the extent that much of the office furniture is on wheels and is easily reconfigurable.

“This is part of our transformation – to move from being a telecoms company to a technology company,” said John Donovan, AT&T’s chief technology officer.

The operators have felt the need to change as technology companies such as Apple and Google, which is buying Motorola, have become more influential and handset makers have aligned themselves behind different operating systems.

“They have to show they are a part of the ecosystem, there’s been a big shift in the last year or so, when there has been the threat that carriers just become a dumb pipe [for data],” says Chris Jones, a telecoms analyst based in the Valley for the Canalys research firm.

Trevor Healy, chief executive of the mobile advertising platform Amobee, whose previous Valley company Jajah was bought by Spain’s Telefónica, has also noticed a change in attitude forced upon the operators.
“It’s not long ago that these companies were myopic and closed, but now there’s a carrot and stick for them in the shape of new revenue streams when their core revenues are stagnant and [they face] the stick of other big players like Google and Apple making ground,” he says.

The operators are arriving late to the Valley compared with infrastructure and handset players such as Ericsson and Nokia, who have had research centres here for years and base their chief technology officers in the Bay Area.

“The operators may have abdicated their positions with the operating systems and handsets, but now they’re putting a stake in the ground and focusing on app development, advertising and payments and billing,” says Mr Healy.
The consumer smartphone market may appear to have been carved up between Apple’s App Store and the Android Market, but operators feel they can still offer premium apps and services based on their network’s ability to know location, process payments, serve advertising, communicate with a range of devices and combine disparate services.

Mr Donovan boasted that the Foundry had 106 projects under way and that AT&T was now measuring itself on the openness of its network: 82 APIs (or access points into the network) had been launched over the past year.

Start-up demonstrations at both the AT&T and Vodafone opening events hinted at the “mash-ups” that were possible through mobile networks – for example, how an app recording running performances could link up with ones relaying sensor information on a person’s health.

Augmented reality applications, improved graphics for gaming and the ability for machines to talk to other machines – “the internet of things” – could all benefit from being hard-wired into the operators’ networks, the demonstrations proved.

It remains to be seen whether developers will be attracted in significant numbers to working with the operators, but Ms Arjomandi argues all parties can benefit.

“We would like to have the advantage of [perhaps one-year exclusivity] lead time to the market – to have a differentiation for us, we need to get that lead time,” she said.

“In exchange, we give them access to 340m customers around the world in 30-plus countries – so it’s definitely a win-win situation.”

Thursday, September 29, 2011

Big Data Take Center Stage at Health 2.0 Conference

George Lauer, iHealthBeat Contributing Editor September 29, 2011


SAN FRANCISCO -- If the Health 2.0 movement is perceived as an ongoing conversation, the first four years could be seen as determining who was going to be talking and what the means of communication was going to be. Now, in year five, the focus is turning to what, exactly, everybody's going to be talking about.

"Data is everything," said Health 2.0 cofounder Mathew Holt at the Fifth Annual Health 2.0 Conference this week.


Raw information -- the gathering of it, collating, crunching and ultimately putting it to work -- was a major theme of this year's event.


"If we see data in health care as this huge glacier, we really are just now coming to see the tip of the iceberg," said Indu Subaiya, cofounder with Holt. Together, in 2007, they launched the first Health 2.0 event aimed at sparking innovative ways to get consumers involved in actively managing their own health care. Between 400 and 500 people ("mostly techies and nerds," according to Holt) showed up for the first conference. This week, more than 1,500 people from several countries attended.


"We're getting to the stage that is kind of the holy grail of Health 2.0, where the rubber meets the road," Subaiya said. "This year at Health 2.0 we're really beginning to see the intelligent mining of big data that can truly change how the system works. It's all around us, this data, but we have so much more work to do to link it and make use of it."


The kind of data examined during the conference ranged from tiny sensors and devices to monitor heart rate and glucose levels to "Big Data" gathered by the government and huge corporations to analyze big-picture situations ranging from flu epidemics to payer and provider databases.


"Data at every level is going to be what drives things forward," Holt said. 


"Almost everything we talk about here is being generated and displayed on the Internet, and we're only talking about a small, tiny fraction of what's out there," Holt added.


Government Embracing Health 2.0
What started as a grassroots effort to get patients involved with health IT has landed the federal government as a partner. Two HHS representatives participated in the Health 2.0 event this week -- Farzad Mostashari, national coordinator for health IT, and Lygeia Ricciardi, senior adviser for consumer eHealth at the Office of the National Coordinator for Health IT.


"We are embracing this with the passion of a recent convert," Mostashari said during a patient-centered panel presentation. "We come to this community with humility. This is not something that I've been at the forefront fighting for, but I really do believe we're on the right track."


Ricciardi urged attendees to check out the government's new offerings at HealthIT.gov, including a new educational campaign, "Putting the I in Health IT."


"We had this big formal launch in Washington two weeks ago as part of National Health IT Week," Ricciardi said. "One side (of the site's home page) is set up for providers and professionals and the other for patients and families. There's a whole body of information on health IT for the general public to understand it better," Ricciardi added.


Health 2.0 and ONC announced the launch of another in its series of challenges to health care technology innovators -- this one an invitation to multidisciplinary teams of developers to create a user-friendly way for patients to manage cardiovascular health using health IT. The project is part of the larger Health 2.0 Developer Challenge, a partnership between ONC and Health 2.0 designed to spur innovations in the use of technology to improve health care outcomes.


Earlier this month, ONC and Health 2.0 launched two competitions to encourage development of health IT applications to improve outcomes in patients transitioning from hospital to home and to facilitate reporting of adverse events related to medical devices.


'About To Change the World'
Although health care faces many obstacles in trying to bring the industry into the 21st century, the people and ideas involved in Health 2.0 are "on the verge of doing all the things that are about to change the world," according to Mark Smith, president and CEO of the California HealthCare Foundation and the conference's keynote speaker. CHCF publishes California Healthline.


Compared to other parts of modern life, such as banking and research, "in the trenches of American medical practice, we are still about 20 years behind," Smith said.


The good news, Smith said, is that technology is maturing and policy is evolving. The bad news is that time is running out in the race to fix an "enormously complex" health care system with "perverse incentives."
Smith pointed to Bear Stearns and Lehman Brothers, investment banks that folded in the credit crisis of 2008, as examples of "what happens when you run out of time."


"The fact of the matter is we're seeing lots of niceties in health policy swamped by fiscal realities," Smith said. "The amount of time we have is dwindling."


In November 2010, CHCF launched a $10 million fund to support health care innovation.
"We are interested in helping to stimulate innovations that can get some kind of market traction," Smith told a room full of entrepreneurs.


To illustrate the kinds of ideas that don't get traction, Smith showed a slide of the cartoon South Park's Underpants Business Model:
·       Phase One -- Collect underpants
·       Phase Two -- ?????
·       Phase Three -- Profits!!!


"For a lot of health care ideas, that's really what the business model amounts to," Smith said. CHCF wants "to fund successful companies because we're tired of seeing good ideas dying on the vine."


The CHCF Health Innovation Fund invests in for-profit and not-for-profit companies and organizations, with a focus on start-ups. Investments range from seed grants of $50,000 to as much as $3 million over the life of a company or service.


Sampling of Launches, Rollouts
Dozens of new products, projects, partnerships and campaigns were announced during the two-day conference, including:

·       Health eVillages, a partnership between the Robert F. Kennedy Center for Justice and Human Rights and Massachusetts-based Physicians Interactive Holdings. The project aims to get mobile phones and handheld devices to medical professionals in poor, remote and underserved parts of the world. So far, the organization has mounted pilot projects in Haiti, Uganda, Kenya and the U.S. Gulf Coast;

·       Society for Participatory Medicine, which encourages patients to take an active role in their health care and encourages physicians to make that happen, announced a campaign to get 10,000 physicians enrolled in a program to recieve the Society Seal, proclaiming their compliance with three basic tenets of the organization -- providing patients with all medical data, convening a patient advisory group, soliciting input from patients and providing resources to patients; and

·       Clarimed, billed as a "first-of-its-kind" health care rating organization, launched its online rating service. Calling itself the first independent rating agency for medical devices, diseases and health care products and services "in the gap that exists between scientific portals and medical websites for consumers," Clarimed was formed in April, launched a beta site in June and went live this week.


MORE ON THE WEB
·       Health 2.0 San Francisco 2011
·       HealthIT.gov


Read more: http://www.ihealthbeat.org/features/2011/data-take-center-stage-at-health-2-0-conference.aspx#ixzz1ZMzOTLWZ

Information explosion: how rapidly expanding storage spurs innovation

Lee Hutchinson   Ars Technica  | Published a day ago

Moore's Law gets all the press. It's easy to present even to non-technical readers, and the way it's most often expressed is something like, "computers double in speed every year," though that's a bastardization of the axiom, which actually states that the transistor count of integrated circuits tends to double every eighteen months or so. This formulation does succinctly capture how fast computers have gotten in so short a time.

But integrated circuit density hasn't been the only computing tech which has shown extremely rapid progress over the past thirty years. Consider magnetic storage. Modern hard drives are precisely manufactured miracles, products of billions of dollars and decades of research into magnetism and quantum mechanics, squeezing ludicrously large amounts of data into ludicrously tiny spaces. A hard drive with about three terabytes of capacity can be had for less than $150 today; a PC equipped with two or three of these would have more on-board storage than most large enterprises had in aggregate even a decade ago.


That kind of inexpensive capacity has revolutionized the way people keep and use data, both at home and at work. From complex storage- and compute-intensive tasks like oil and gas upstream processing all the way down to editing a home vacation video, the ability to store and manipulate increasingly voluminous data actually drives serious innovation.

A PCM .wav file of a three-minute song might be 30 or 40 MB, while an mp3 of the same song might be 3 MB. If all you've got is a 100 MB hard disk to work with, even the ten-fold decrease in file size doesn't make keeping an mp3 collection practical. But in 1995-6, when mp3s began to flow back and forth on the USENET and FTP and websites, the format suddenly became an extremely attractive way to collect music. With most folks having a handful of gigabytes, having an entire music collection available on a computer suddenly became possible. Copious local storage paved the way for entire new industries—think of early iPods and the iTunes store.

(Despite the sea change currently underway to solid state disks, there's every indication that spinning magnetic storage will continue to be with us for years to come—Seagate already sells a 4 TB external 3.5-inch hard drive.)

But data in and of itself doesn't do much. In this era of YouTube and Facebook, data is useless without a means to transport it from one place to another, a way to share it with others. Unfortunately, data transport has not been an area where the brain-smashing innovation race in integrated circuits and magnetic storage has been mirrored. Just as the density explosion in storage has opened the door to people and companies doing tremendous new things with computers, the lack of commensurate scaling in commonly available telecommunications has greatly hamstrung people and industries.

This has led to one of the most annoying paradoxes of the Internet age. On one hand, the massive increase in storage density encourages creation at all layers, from the biggest of businesses leveraging the ability to analyze and mine targeted data out of petabytes of raw information down to the adorable old couple trying to make a video on their home computer (both of these things can eat up a lot of hard drive space!).

On the other hand, the asymmetrical nature of most broadband solutions available to consumers in the US and Europe and a stagnation in their speed encourages only consumption at the "lower" levels of that stack. Companies that need both the ability to transmit and receive data over distance can usually afford to pay for symmetrical high-speed network links, while consumers (at least in the US) typically can pick from two choices for Internet access—DSL or cable. Both access methods typically provide plenty of download bandwidth for Netflixing and iTunesing and YouTubeing, but comparatively tiny upload bandwidth for sending data (most DSL and cable Internet plans have upload speeds that are less than 25 percent of the download speeds).

This asymmetry of access leads us to a strange place, where most folks have the ability to store and create more amazing things than ever before, while at the same time they lack the ability to quickly and easily share any of those things with each other.

Storage through the decades


Let's take a look back at a typical consumer's (i.e., my) hard disk drive at five-year intervals over the past three decades:

The chart on the top shows what appears to be a roughly linear progression of capacity in storage, but the vertical axes actually increase in powers of ten, showing almost a true exponential growth curve. To better appreciate the growth, consult the chart on the bottom, which uses fixed vertical axes. Starting at 5 MB in 1981, we jump quickly to 20 MB in five years, then to 120 MB in five more years, then to 2 GB, to 40 GB, to 500 GB, and finally to 3 TB.
The capacity increases get truly staggering near the end. This alone doesn't tell us much, since it's an accepted bit of conventional wisdom that hard drives will always grow larger over time. The important thing about this increase in storage capacity isn't that it has let us store more data, but that having fast random access to more data lets us do more with that data.

No one keeps files for the sake of keeping files. Even digital hoarders who have every single episode of Doctor Who ever created don't simply have those files to have files—they have those files because they comprise a collection of data with value as a collection. An Excel or Word document has no value in and of itself; rather, documents gain value when they are consulted or shared, when the data that they contain is compared against or integrated with other data and transformed into new data. With this in mind, the exponential increase in storage density coupled with a decrease in how much it costs per-gigabyte to purchase that storage has driven a huge shift in what it means to create and innovate.

In the early '80s, microcomputers like the IBM PC and the Apple II were used largely for traditional computing tasks—crunching numbers or programming applications that would be used for crunching numbers. (Well, that and games.) A 5MB ProFile hard drive for an Apple Lisa cost about $3,500 when it first became available for that system in 1981, but the price could be justified by the fast and random access that ProFile offered to what at the time was a significant amount of data (equivalent to dozens of double-sided floppy disks). Because of the cost, though, home computers didn't typically have internal hard disk drives until the late 1980s. As their storage capacities increased, so did the things that computers were typically used to do.

Serious music production at home, a trend that arguably got started with module tracking on the Amiga platform and spread rapidly in the early 1990s to the PC, relied on bigger hard drives to hold the digital samples that made up the music files. Accelerating storage density pushed this kind of music production out of the basement and into the home studio, today letting artists like Pomplamoose record and produce entire albums on personal computers without relying on time in hugely expensive recording studios. Entry barriers into a recording career are being quickly demolished and a tidal wave of skilled and talented artists are appearing, doing things that would have been impossible even a few years ago.

Terabytes of fast disks available in most home computers also enable software like iMovie and Windows Movie Maker to exist and thrive, bringing movie-creation tools to the average home user. Not everyone is a talented auteur with a Citizen Kane waiting inside —in fact, most folks tend to create videos of their babies farting or themselves farting or animals farting—but folks like Neill Blomkamp or Bruce Branit have also risen above the noise to create some amazing things.

This sort of storage-driven innovation holds true for the enterprise just as much as it does for the home—the fact that I've been linking to YouTube videos above is telling. Companies like Google (which owns YouTube) wouldn't exist without huge amounts of storage. Google began out of a need to locate and organize information, and it manages and stores petabytes and petabytes of data. Even more to the point, it does so in a way that underlines the commoditzation of information technology in general—most big enterprises these days store data in large monolithic storage arrays, but Google keeps all of its data on throwaway servers with consumer-grade hard disks, relying on redundancy and smart design to route around failed components. As storage densities continue to increase and the transformational things that people can do with that storage scales, more and more enterprises will shift in Google's direction.

Advanced trending and data analytics can also be done now with more data available—for example, T-Mobile was able to take the huge amount of network and subscriber data they store, sift it, and come up with insights into why customers leave (disclaimer: the linked article mentions EMC, and I'm an EMC employee). Having more data at hand lets us make connections or create new things that simply wouldn't have been possible without massive storage.
There's a darker side to all of this shiny creativity, though. It can still be too hard to share it with the world.

Storage enough at last?
There's a famous episode of The Twilight Zone named "Time Enough at Last," wherein a bookish fellow played by Burgess Meredith laments that there's not enough time to indulge in reading. He decides to take his lunch break in the vault of the bank where he works, and while he's down there, the Commies nuke everything. He emerges from the vault stunned at the ruins around him and, finding no one alive but himself, eventually decides to take his own life. Instead, though, he spies a library in the distance; he approaches it and, finding that it's stocked full of undamaged books, realizes that his prayers have been answered—he finally has time, time enough at last to read all he wants! The hook, though—spoiler alert from 1959!—is that as he sits down to read his first book, he breaks his glasses. The episode closes with Burgess howling about the unfairness of it all, amidst stacks of books that he will now be forever unable to read.

This parallels the situation most consumers (at least within the US and much of Europe) find themselves today. We have, finally, storage enough at last to do just about anything we want. We can create and keep terabytes of digital stuff, but the value of that stuff is in what you can do with it, and a home movie doesn't do you any good sitting on your hard drive—its value comes from sharing it. Increasingly, that sharing takes place over the Internet.

And sharing is a problem. Let's take a look at the typical amount of inbound and outbound network bandwidth a US home might have had for the same years in which we look at hard drive capacity. Internet connections to the home weren't terribly common before about 1994, but those of us old enough to experience the days before USENET and the Web will remember local BBS systems and fledgling "information services" like Prodigy and Compuserve, all of which served many of the same purposes then that the Internet does now—connecting with others and exchanging files and messages.

From 300bps in 1981, we travel through 1200bps and the halcyon days of local BBSs at 14.4Kbps, through the brief 56K era and into broadband, going from 1Mb to 3Mb to the 10Mb (and beyond) connections common today. The same kind of exponential-ish increase in speeds is evident, especially in the break between 1996 and 2001 when the widespread shift to broadband connectivity began to pull folks out of needing to use modems over POTS and onto DSL and cable Internet access. However, with that came huge gulfs between the speed at which a person could download things to his computer versus the speed at which he could upload things from it to others.

To be sure, upload/download asymmetry has been around for a long time. Back in my own BBS days, I used to look with envy on the guys with USRobotics Courier HST Dual Standard modems, which could talk to other Courier HST modems at a ridiculously fast (for the time) 16.8Kbps and at the same time have a 450bps back-channel available for control traffic. Similarly, most folks who lived through the 56K days remember that 56K was really only one way—traffic going out from your modem was still limited to 33.8Kbps (and in reality, most folks only saw 22-26 Kbps). But nowhere is the asymmetry as great as with modern cable and DSL connections. The practical reason behind the split between up and down speeds is one of frequency utilization—there's only so many hertz available to carry a signal, and so more of the available frequency space is allocated to carry information to the consumer than from the consumer.

Unfortunately, this common asymmetric split now affects a good chunk of the world's Internet users.

You can't take it with you
Most folks carry what feels like significant chunks of our lives on our hard drives—we express ourselves through the things we create, and we keep things that have powerful meaning for us. A hard drive crash where data is lost can lead to not just technical or financial problems, but actual personal anguish—our drives are surrogates for our memories, and if they die we can lose parts of ourselves.

And yet, with all the innovation that has arrived as a consequence of more and faster data storage, most lack the ability to easily move that data around. We can download multi-gigabyte high definition movies from iTunes or NetFlix in astonishingly short amounts of time, but how many hours does it take to upload a multi-gigabyte file?
Amount of time it would take to upload an entire hard drive at each point in time


Companies like Microsoft and Apple and others continue to roll out cloud sync services, which among other things suggest you put your important files "in the cloud" so they can be accessible anywhere—but why should that be necessary when they could just as easily be accessed in situ on your own computer? I have hundreds of movies that I've ripped from DVDs that I own just sitting on my NAS; why shouldn't I be able to access them wherever I am and stream them to myself at a friend's house as fast as NetFlix can? Why shouldn't I be able to get to any part of my digital identity from anywhere in the world?

Cloud backup services like Backblaze or Mozy have their place, too, providing off-site storage for important data in case of a disaster, but it can take weeks—literally weeks—to upload a few terabytes of data. Most people today don't bother with this kind of backup service, precisely because it takes so long to populate initially. Worse still, for people with metered Internet connections, that initial data copy will almost certainly count against your 50 or 150 or 250 GB monthly cap, which might lead to your Internet access being terminated without explanation by your ISP.
The lack of upstream bandwidth in a typical consumer broadband connection holds back some of the innovation that could come from huge amounts of storage. For instance, with a 10 Mb upload connection to go along with your 10 Mb download connection, content sharing sites like YouTube suddenly lose some of their relevance. "Digital locker" services like RapidShare and MegaUpload serve fewer purposes. An entire infrastructure of sites and services that have evolved as buffers and workarounds for tiny upload speeds are made redundant and can be done away with.
Do more with more
Even if our wings remain clipped by upload restrictions, we're still innovating through storage. We have more and faster storage available to us at home and at work, and the things we can do with that capacity continue to evolve. Parkinson's Law states that "work expands so as to fill the time available for its completion." A common corollary to that law is that "data expands to fill the space available for storage." Put even more simply, as the ghostly voice said to Kevin Costner in Field of Dreams, "If you build it, they will come."
Nothing exists in a vacuum, of course, and it's disingenuous to suggest that increases in storage capacity can alone be responsible for anything, much in the way that it's wrong to say an increase in CPU clock speeds alone is responsible for anything, or that the change from IDE to PCI buses in a PC was by itself responsible for anything. However, the rapid increase in storage capacities to which the average consumer has access has removed barriers and created big new opportunities.


Wednesday, September 14, 2011

Six Provocations for Big Data

Six Provocations for Big Data
danah boyd, Microsoft Research; University of New South Wales (UNSW); Harvard University - Berkman Center for Internet & Society
Kate Crawford, University of New South Wales (UNSW)

September 21, 2011

Abstract:
   
The era of Big Data has begun. Computer scientists, physicists, economists, mathematicians, political scientists, bio-informaticists, sociologists, and many others are clamoring for access to the massive quantities of information produced by and about people, things, and their interactions. Diverse groups argue about the potential benefits and costs of analyzing information from Twitter, Google, Verizon, 23andMe, Facebook, Wikipedia, and every space where large groups of people leave digital traces and deposit data. Significant questions emerge.



Will large-scale analysis of DNA help cure diseases? Or will it usher in a new wave of medical inequality? Will data analytics help make people’s access to information more efficient and effective? Or will it be used to track protesters in the streets of major cities? Will it transform how we study human communication and culture, or narrow the palette of research options and alter what ‘research’ means? Some or all of the above?

This essay offers six provocations that we hope can spark conversations about the issues of Big Data. Given the rise of Big Data as both a phenomenon and a methodological persuasion, we believe that it is time to start critically interrogating this phenomenon, its assumptions, and its biases.

(This paper was presented at Oxford Internet Institute’s “A Decade in Internet Time: Symposium on the Dynamics of theInternet and Society” on September 21, 2011.)

Tuesday, September 13, 2011

Big Data: Spy Agency Seeks Digital Mosaic to Divine Future

The U.S. intelligence community wants to mine lots and lots of the tidbits bopping around on the Internet to suss out trends before they make the news.

Emily Badger     Miller McCune   September 9, 2011

U.S. intelligence agencies hope those tidbits bopping around on the Internet will help discover all kinds of trends before they make the news. (John Foxx/Stockbyte)

Governments have been caught off guard a lot lately: by revolutions, by riots, even by unemployment rates (or, to go back even further, by events like 9/11).

In the information age — where there's no limit to publicly available data on everything from political chatter to gas prices — it seems policymakers should be better at predicting major societal shifts and events than at any point in history. Shouldn't all these little pieces of information be telling us something big? Shouldn't they be telling us about where the next mass migration will come from or where the next riot will be?

The government's Office of the Director of National Intelligence is betting this is the case. Its Intelligence Advanced Research Projects Activity ( or IARPA — which sounds like a less menacing cousin to DARPA) is rolling out a new R&D project to test tools that would mine publicly available data to predict political and humanitarian crises, disease outbreaks, mass violence and instability. The project, the Open Source Indicators Program, is premised on the idea that big events are preceded by population-level changes, and that those population-level changes should be identifiable if we just look in the right places.

Idea Lobby
THE IDEA LOBBY
Miller-McCune's Washington correspondent Emily Badger follows the ideas informing, explaining and influencing government, from the local think tank circuit to academic research that shapes D.C. policy from afar.

The concept isn't new. Allied intelligence officers did something similar during World War II, for example, mining letters to the editors of local newspapers and radio transmissions for clues as to what was going on inside Nazi Germany.

"The difference, the big difference — and the thing this initiative is picking up on — is that whereas we used to have to do this with a relatively small sample of radio shows and newspapers and stuff like that, we're now drinking from the fire hose of the Internet," said Philip Schrodt, a political scientist at Penn State.

An IARPA public affairs officer said officials could not discuss the program while the government is still soliciting proposals. But Schrodt and several other researchers affiliated with academic teams that may eventually wind up working on the project, alongsideprivate contractors, offered a look into an expanding, multidisciplinary field of real-time data analysis that the government hopes may allow it to "beat the news."

This experiment — which will test methods in Latin America, not the U.S. — dovetails with broader trends in computer and social science toward mining large-scale sets of social data. Twitter, for instance, is a jackpot: The network can be scoured both for the content of messages and the connections that are revealed between people writing and reading them.


About half of the populations in most major Western countries are now on Facebook (and a surprising amount of data on Facebook remains accessible to the public). Economic data can be collected from e-commerce sites like Amazon. There are also open-source indicators embedded in Internet news sites, message boards, unemployment data, Web search queries, traffic Web cams and financial markets.

Think of what researchers could learn about the economic mood of consumers if they suddenly discovered hundreds of people (each anonymously identified) all trying to sell their best jewelry on Craigslist tomorrow. Other open-source indicators could provide a trove of information on a topic social scientists have been weighing a lot lately: the effect of price changes — whether for housing, gas or food — on social stability.

And because all this data can be collected in real time and by automated systems, it's more up-to-date, it's larger in sheer quantity, and it's becoming more cost-effective and realistic than ever to analyze.
• • • • • • • • • • • • • • •
The easiest place to understand the potential of all this data is in public health, where live analysis is already helping to identify and track disease epidemics at a speed that was never possible before the Internet. Google Flu Trends pioneered the technique. The tool measures the frequency of certain search terms commonly associated with the flu to identify outbreaks down to the city level. Now it's doing the same with dengue fever.

Researchers at Harvard developed a similar tool five years ago calledHealthMap, which offers a prototype of the model IARPA envisions testing beyond public health. HealthMap continuously scrapes the Web for health-related keywords and deletes noise, then classifies the data, geocodes it, filters it into different categories and maps the results. The project is particularly useful in identifying early disease indicators in countries that have no public health capacity and weak or opaque data reporting (because these are often the same countries where sick people are unlikely to Google their symptoms on a home computer, HealthMap also taps into SMS messages and smartphone data).

"These discussions are taking place by the minute, and we're caching those in real time," said John Brownstein, the director and co-founder of HealthMap. "Our goal is as soon as anybody's talking about an outbreak on the Web, within the hour it ends up on HealthMap."


In theory, similar methods might track contagious ideas, economic unrest or resource shortages — and in a way that would actually move faster than news reports.

Automated systems could certainly be more comprehensive than traditional media. The BBC and Associated Press don't have correspondents in every small town in Mexico, but an automated computer program could analyze in real time reports from every local news site in the country. Or, better yet, it could analyze social network chatter before it even gets to the local newsroom.

"If everybody is suddenly upset about the drought in Texas or something like that," said Schrodt, "we should be able to pick that up immediately without having some reporter go out and talk to some farmer whose cows have died."

This type of analysis could also help policymakers be less surprised by news when it happens, and to monitor it literally in real time.

The bigger question, though, is how researchers take the step from watching trends develop in a live time series to anticipating them before they happen.
Some trends lend themselves more easily to this challenge — disease epidemics are one. HealthMap can start to identify outbreaks before officials even know what disease they're looking at because the tool is designed to scan for basic symptoms like runny noses or a run on aspirin. HealthMap isn't merely crawling the Web for Google searches of "Do I have SARS?"

But how do we anticipate "riots in London" when we don't even know "riots in London" is what we're looking for? Could open-source indicators have identified the rise of the Tea Party before it even started calling itself that?

• • • • • • • • • • • • • • •
"That's the really, really big question, that's the single biggest challenge in this IARPA initiative," Schrodt said. He has worked on other government research projects, including one called the Political Instability Task Force. It's clear, though, in that project, what researchers are looking for: political instability, which is suggested by fewer than five indicators. Here, though, the goal is to identify everything that's publicly available, for anything that might be interesting.

"I can't take a book off my shelf and open it up and say, 'Oh, here's what you do if you're monitoring 100 different indicators using 200 different sources,'" Schrodt said. "That's the new science on this. I don't know if it will work or not." (The political instability project, as an example, has gotten to about 80 percent accuracy forecasting probabilities that countries will collapse.)
Peter Gloor's research at MIT's Center for Collective Intelligence has been trying to solve this unknown prediction problem by identifying the most creative people – both those who are destructively creative, like terrorist plotters, and those who are constructively creative, like Justin Bieber. He's trying to find the people behind trends — information producers, not the information itself — before trends take off.

"We're looking for the trendsetters while they are being born and made," Gloor said. "In Twitter, once you are Ashton Kutcher, everyone knows you; it's clear you are a trendsetter. But Justin Bieber, whenever he was posting his first video on YouTube, it was not as clear."

Gloor has done a lot of this work in Wikipedia, where pages are written and edited according to a revealing pattern: Generally, about 1 percent of contributors write 90 percent of the content; 9 percent of people write another 9 percent of the content; and 90 percent of contributors write just 1 percent of the crowdsourced encyclopedia. If you find the 1 percent of people who are doing all the heavy lifting, those are your future trendsetters (or "coolhunters").

Charles Elkan, a computer scientist at the University of California at San Diego, suspects the biggest value of all these tools will come from quantifying what social scientists already know.

"Social scientists have said for 100 years that revolutions happen at times of rising expectations," he said. "If, for example, society is becoming more prosperous, people start have rising economic expectations, that spills over into rising political expectations, and that can make a revolution more likely. Social scientists have said this for a long time, and now it's beginning to be possible to quantify it."

Finland, for example, is obviously a more stable country than Sudan. But attaching a number to that statement — say, Sudan is eight times more likely to experience a coup or revolution in the next two years — is much trickier.
• • • • • • • • • • • • • • •
One thing open-source indicators probably can't do is write the news before it happens.
"Predicting events is even harder than predicting trends," Elkan said. "If you think of a trend as creating the probability for something, that probability combined with some spark makes the actual event."

And then you have to predict the spark, too.

"We can say unemployment is trending upwards, and that layoffs are increasing," Elkan went on, "but it's still very difficult to predict which manager of which company will wake up in the morning and decide, 'I can't wait any longer, I need to lay off 10 people today.'"

It's possible to predict, for instance, that riots are more likely in London than they are in Orlando. But no one would have predicted a week ahead of time that a specific police shooting would catalyze riots (as plenty of other police shootings — even in other seething areas — don't set off riots). Similarly, no one would have predicted that the suicide of a Tunisian fruit vendor would spark a revolution that would sweep out of North Africa and into the Arabian Peninsula and Mideast.

In this sense, there may be a sizable gap between our imagination of what's possible with trend prediction and the reality of what it can do for policymakers.

The primary limitation isn't the data collection; it's the data analysis. And there's a point in the process where human judgment must take over from automated computer systems. That is inevitably the moment when controversial or expensive action looms: Should we double the FEMA budget, send support to Syrian revolutionaries or shift course on jobs policy to quell a domestic uprising?

The attention span of decision-makers, Elkan cautions, is a limited resource, too.

There's another factor: Governments aren't the only ones who might like to leverage this data. Its business application is obvious. Movie theaters could use trend prediction to book films. Department stores could use it to stock products. Corporations could mine Twitter chatter about them to craft corporate social responsibility policies.

"That's one reason to be skeptical about the extent of the possibility," Elkan said. "If it was really possible, for example, to predict unemployment much better than existing methods do, then people on Wall Street would be doing it and taking advantage of it."

Like the other researchers, he's cautious about where all of this is headed, even as he tries to design computer models to more accurately predict the future.

"We can easily specify, 'Let's monitor every newspaper and every radio station everywhere in the world, then as soon as there's one school that has a local outbreak of some disease, then we can put that into our computer models and make predictions of when it will come to the United States,'" Elkan said. "That's a realistic dream, but it's still a dream at this point."

Monday, September 12, 2011

Big Data : Data-Driven Decisions in Asthma Care

Asthmapolis: Tools to Track, Manage, and Research Asthma

GPS Sensors and Mobile App Track Asthma Symptoms, Triggers, and Inhaler Use

California Healthcare Foundation September 2011

Despite dramatic advances in medication effectiveness, many people with asthma experience episodes of uncontrolled disease, causing them to be at greater risk for acute exacerbations. Historically health care providers have had inadequate tools to understand the severity of a patient's asthma, limiting their ability to predict and prevent serious and costly asthma attacks. To address this problem Asthmapolis developed GPS sensors that record exactly when and where patients use their inhalers and a digital interface to display the information captured. The patient-level data may then help physicians make more informed clinical interventions and help patients track and manage their condition. 

With the support of the California HealthCare Foundation, Asthmapolis has begun a pilot with Woodland HealthCare, a member of Catholic HealthCare West, to determine if the data collected by sensors and presented to physicians and patients results in actions that improve asthma control. This study is being conducted with English- and Spanish-speaking subjects and will include underserved patients in the Sacramento area.

CHCF invested in Asthmapolis because we see the potential for patients to benefit from technology that provides tailored information, allowing them to better understand and manage asthma, which will hopefully lead to fewer serious and costly episodes.

Read more: http://www.chcf.org/projects/2011/asthmapolis#ixzz1Xm0uua3l

Monday, August 22, 2011

Big Data: Why HP Wants Autonomy: Math Skills

Autonomy excels at analyzing the vast amounts of "unstructured data" being produced every day.

Tom Simonite    Technology Review (by MIT) Thursday, August 18, 2011

News broke today that HP, the world's biggest manufacturer of personal computers, had offered to acquire the British software company Autonomy. While the latter is hardly a household name, it gets close to $1 billion in revenue each year from software that can turn huge volumes of images, text, and video into useful statistics and insights for businesses.

Acquiring that technology will enable HP to expand its business software products, and put it in a good position to exploit a trend dubbed "big data." Businesses are increasingly interested in finding ways to distill meaning from the growing piles of digital information, from tweets to video, flowing through our lives at work and at home.

Whit Andrews, a vice president and analyst with Gartner who specializes in technology that processes and organizes information, says that Autonomy was years ahead of other companies in making such analysis possible. "They have had this vision for over a decade that there was immense value in being able to do statistical analysis for data like audio and video that conventional technology cannot handle."

Autonomy's products enable companies to do things like analyze transcripts from call centers; discover which sales strategies work best; and process troves of e-mails and other documents to match whether what is being said and done comports with a company's legal responsibilities. Such analysis can be automated using software that looks for certain things automatically, or performed manually by a person entering queries, and then sifting through the results themselves.

Andrews says business and technology companies are beginning to realize that both types of analysis could offer much more than conventional approaches, which rely on so-called "structured" data, such as a spreadsheet organized into labelled columns. "Business analytics is about structured data, like spreadsheets," says Andrews. "Autonomy does an exceptional job at analyzing unstructured data, which may prove even more valuable."
In an interview with Technology Review published last year, Autonomy's founder and CEO, Mike Lynch, estimated that about 85 percent of the information inside a business is unstructured. "[W]e are human beings, and unstructured information is at the core of everything we do," he said. "Most business is done using this kind of human-friendly information."

Lynch founded the company to commercialize statistical techniques developed at Cambridge University based on Bayesian inference, a mathematical technique that can estimate the probability of potential outcomes based on previous evidence.

Companies like IBM are working hard on their own approaches to analyzing unstructured data, but Autonomy has been at it for longer, says Andrews. Acquiring the company could enable HP to take a much more dominant position in the growing market for what Autonomy's Lynch dubs "meaning-based computing."

4G a boon to U.S. economy and jobs, study says

Roger Cheng    CNet News    August 21, 2011 9:01 PM PDT

The wireless carriers' investment in 4G networks could be the salve that the ailing U.S. economy is looking for.

The carriers could invest between $25 billion and $53 billion in building out their 4G network through 2016, according to a study from Deloitte. That in turn could lead to the creation of 371,000 to 771,000 jobs, and gross domestic product growth of $73 billion to $151 billion.
"Investment in such a powerful form of communication contributes to the economic recovery and provides a job-creating engine for the future," said Phil Asmundson a consultant for Deloitte.

The rise of 4G networks could provide some support for an economy still struggling to recover, and which some believe is slipping back into a recession. The recent bitter political struggle to raise the national debt ceiling, the sharp declines in the stock market, and the continued high rate of unemployment have many still concerned.
The wide-ranging estimate assumes two different scenarios. The baseline scenario has the carriers deploying 4G technology at a moderate pace with a slow transition from 3G to 4G extending into the middle of the decade. Deloitte warns that under these conditions, the U.S. firms will be vulnerable to foreign competitors looking to jump ahead in the 4G race.

The second scenario predicts a more rapid investment in 4G networks and the creation of 4G-based services before global competitors gain momentum. Deloitte said the demand stimulated by the new services would propel more network investment, "setting off a virtuous cycle of investment and market response." Cloud services are among the primary catalysts for investment, according to Deloitte.
Telecommunications investment has remained strong over the past few years, with large companies such as AT&T and Verizon pouring money into network upgrades. AT&T has argued that its acquisition of T-Mobile would lead to new jobs and investment because it is able to expand its future 4G network across a wider stretch of the country. Chief Executive Randall Stephenson has long argued that investment in telecommunications technology is crucial to staying ahead in the world.

AT&T plans to launch its first 4G LTE markets later this summer.
Verizon, meanwhile, has spent billions of dollars upgrading its fixed-line infrastructure with fiber-optic lines to deliver television content and faster Internet service. On the wireless side, the company has been racing to build out its 4G LTE network, and recently said that it covers more than half the country.

Clearwire is also switching gears from its current 4G WiMax standard and moving toward LTE, but the company said it plans to spend only $600 million for the upgrade.

Clearwire's largest customer and shareholder, Sprint Nextel, is planning to unveil its 4G plans in October. The company has already signed a network-hosting deal with LightSquared, which plans to start testing its 4G network with customers next year.

The widespread adoption of 4G services should also help certain segments such as disadvantaged minority groups; rural communities and areas with limited full broadband access; and small businesses, the firm said, adding the advent of 4G could work to bring those groups further into the economic mainstream.

Read more: http://news.cnet.com/8301-1035_3-20094588-94/4g-a-boon-to-u.s-economy-and-jobs-study-says/#ixzz1VleV8J20

Friday, August 19, 2011

Computer Analysis Could Find New Uses for Existing Drugs

I HealthBeat     August 18, 011

A computer program that analyzes drug data and genetic information could help discover new uses for medicines already on the market, according to two new studies published in the journal Science Translational Medicine, United Press International reports.

Methodology

For the NIH-funded research, Stanford University scientists extracted data from NIH's Gene Expression Omnibus, a public database containing findings from thousands of genomic studies conducted throughout the world.

Researchers focused on 100 diseases and 164 drugs. They analyzed thousands of possible drug-disease combinations to identify which medications and medical conditions had gene expression patterns that could cancel each other out. Such matches indicate that the drug potentially could mitigate the effects of the disease (United Press International, 8/17).

Research Findings

Researchers identified possible drug-disease matches for 53 of the 100 diseases analyzed (Renick, Bloomberg, 8/17).

For example, they found that an epilepsy treatment potentially could treat inflammatory bowel disease and that an ulcer drug might be an effective lung cancer medication (Dockser Marcus, Wall Street Journal, 8/18).

Possible Implications

Researchers noted that re-purposing existing medications to treat different diseases could reduce some of the costs and requirements involved in the drug development process. They noted that it takes an average of 15 years and about $1 billion to bring a single new drug to market (Bloomberg, 8/17).

In an accompanying commentary on the research, Yves Lussier -- a professor of medicine and engineering at the University of Illinois in Chicago -- wrote that the findings should not prompt physicians to prescribe an ulcer drug to treat lung cancer. However, he added that the findings are "impressive enough to be improved upon and studied further."

Lussier also noted that if the computer program is effective at detecting possible off-label uses for existing drugs, it "opens the door to very low-cost, individualized personal therapies" (Wall Street Journal, 8/18).


Read more: http://www.ihealthbeat.org/articles/2011/8/18/computer-analysis-could-find-new-uses-for-existing-drugs.aspx#ixzz1VTsrgrjJ