Lee Hutchinson Ars Technica | Published a day ago
Moore's Law gets all the press. It's easy to present even to non-technical readers, and the way it's most often expressed is something like, "computers double in speed every year," though that's a bastardization of the axiom, which actually states that the transistor count of integrated circuits tends to double every eighteen months or so. This formulation does succinctly capture how fast computers have gotten in so short a time.
But integrated circuit density hasn't been the only computing tech which has shown extremely rapid progress over the past thirty years. Consider magnetic storage. Modern hard drives are precisely manufactured miracles, products of billions of dollars and decades of research into magnetism and quantum mechanics, squeezing ludicrously large amounts of data into ludicrously tiny spaces. A hard drive with about three terabytes of capacity can be had for less than $150 today; a PC equipped with two or three of these would have more on-board storage than most large enterprises had in aggregate even a decade ago.
That kind of inexpensive capacity has revolutionized the way people keep and use data, both at home and at work. From complex storage- and compute-intensive tasks like oil and gas upstream processing all the way down to editing a home vacation video, the ability to store and manipulate increasingly voluminous data actually drives serious innovation.
A PCM .wav file of a three-minute song might be 30 or 40 MB, while an mp3 of the same song might be 3 MB. If all you've got is a 100 MB hard disk to work with, even the ten-fold decrease in file size doesn't make keeping an mp3 collection practical. But in 1995-6, when mp3s began to flow back and forth on the USENET and FTP and websites, the format suddenly became an extremely attractive way to collect music. With most folks having a handful of gigabytes, having an entire music collection available on a computer suddenly became possible. Copious local storage paved the way for entire new industries—think of early iPods and the iTunes store.
(Despite the sea change currently underway to solid state disks, there's every indication that spinning magnetic storage will continue to be with us for years to come—Seagate already sells a 4 TB external 3.5-inch hard drive.)
But data in and of itself doesn't do much. In this era of YouTube and Facebook, data is useless without a means to transport it from one place to another, a way to share it with others. Unfortunately, data transport has not been an area where the brain-smashing innovation race in integrated circuits and magnetic storage has been mirrored. Just as the density explosion in storage has opened the door to people and companies doing tremendous new things with computers, the lack of commensurate scaling in commonly available telecommunications has greatly hamstrung people and industries.
This has led to one of the most annoying paradoxes of the Internet age. On one hand, the massive increase in storage density encourages creation at all layers, from the biggest of businesses leveraging the ability to analyze and mine targeted data out of petabytes of raw information down to the adorable old couple trying to make a video on their home computer (both of these things can eat up a lot of hard drive space!).
On the other hand, the asymmetrical nature of most broadband solutions available to consumers in the US and Europe and a stagnation in their speed encourages only consumption at the "lower" levels of that stack. Companies that need both the ability to transmit and receive data over distance can usually afford to pay for symmetrical high-speed network links, while consumers (at least in the US) typically can pick from two choices for Internet access—DSL or cable. Both access methods typically provide plenty of download bandwidth for Netflixing and iTunesing and YouTubeing, but comparatively tiny upload bandwidth for sending data (most DSL and cable Internet plans have upload speeds that are less than 25 percent of the download speeds).
This asymmetry of access leads us to a strange place, where most folks have the ability to store and create more amazing things than ever before, while at the same time they lack the ability to quickly and easily share any of those things with each other.
Storage through the decades
Let's take a look back at a typical consumer's (i.e., my) hard disk drive at five-year intervals over the past three decades:
The chart on the top shows what appears to be a roughly linear progression of capacity in storage, but the vertical axes actually increase in powers of ten, showing almost a true exponential growth curve. To better appreciate the growth, consult the chart on the bottom, which uses fixed vertical axes. Starting at 5 MB in 1981, we jump quickly to 20 MB in five years, then to 120 MB in five more years, then to 2 GB, to 40 GB, to 500 GB, and finally to 3 TB.
The capacity increases get truly staggering near the end. This alone doesn't tell us much, since it's an accepted bit of conventional wisdom that hard drives will always grow larger over time. The important thing about this increase in storage capacity isn't that it has let us store more data, but that having fast random access to more data lets us do more with that data.
No one keeps files for the sake of keeping files. Even digital hoarders who have every single episode of Doctor Who ever created don't simply have those files to have files—they have those files because they comprise a collection of data with value as a collection. An Excel or Word document has no value in and of itself; rather, documents gain value when they are consulted or shared, when the data that they contain is compared against or integrated with other data and transformed into new data. With this in mind, the exponential increase in storage density coupled with a decrease in how much it costs per-gigabyte to purchase that storage has driven a huge shift in what it means to create and innovate.
In the early '80s, microcomputers like the IBM PC and the Apple II were used largely for traditional computing tasks—crunching numbers or programming applications that would be used for crunching numbers. (Well, that and games.) A 5MB ProFile hard drive for an Apple Lisa cost about $3,500 when it first became available for that system in 1981, but the price could be justified by the fast and random access that ProFile offered to what at the time was a significant amount of data (equivalent to dozens of double-sided floppy disks). Because of the cost, though, home computers didn't typically have internal hard disk drives until the late 1980s. As their storage capacities increased, so did the things that computers were typically used to do.
Serious music production at home, a trend that arguably got started with module tracking on the Amiga platform and spread rapidly in the early 1990s to the PC, relied on bigger hard drives to hold the digital samples that made up the music files. Accelerating storage density pushed this kind of music production out of the basement and into the home studio, today letting artists like Pomplamoose record and produce entire albums on personal computers without relying on time in hugely expensive recording studios. Entry barriers into a recording career are being quickly demolished and a tidal wave of skilled and talented artists are appearing, doing things that would have been impossible even a few years ago.
Terabytes of fast disks available in most home computers also enable software like iMovie and Windows Movie Maker to exist and thrive, bringing movie-creation tools to the average home user. Not everyone is a talented auteur with a Citizen Kane waiting inside —in fact, most folks tend to create videos of their babies farting or themselves farting or animals farting—but folks like Neill Blomkamp or Bruce Branit have also risen above the noise to create some amazing things.
This sort of storage-driven innovation holds true for the enterprise just as much as it does for the home—the fact that I've been linking to YouTube videos above is telling. Companies like Google (which owns YouTube) wouldn't exist without huge amounts of storage. Google began out of a need to locate and organize information, and it manages and stores petabytes and petabytes of data. Even more to the point, it does so in a way that underlines the commoditzation of information technology in general—most big enterprises these days store data in large monolithic storage arrays, but Google keeps all of its data on throwaway servers with consumer-grade hard disks, relying on redundancy and smart design to route around failed components. As storage densities continue to increase and the transformational things that people can do with that storage scales, more and more enterprises will shift in Google's direction.
Advanced trending and data analytics can also be done now with more data available—for example, T-Mobile was able to take the huge amount of network and subscriber data they store, sift it, and come up with insights into why customers leave (disclaimer: the linked article mentions EMC, and I'm an EMC employee). Having more data at hand lets us make connections or create new things that simply wouldn't have been possible without massive storage.
There's a darker side to all of this shiny creativity, though. It can still be too hard to share it with the world.
Storage enough at last?
There's a famous episode of The Twilight Zone named "Time Enough at Last," wherein a bookish fellow played by Burgess Meredith laments that there's not enough time to indulge in reading. He decides to take his lunch break in the vault of the bank where he works, and while he's down there, the Commies nuke everything. He emerges from the vault stunned at the ruins around him and, finding no one alive but himself, eventually decides to take his own life. Instead, though, he spies a library in the distance; he approaches it and, finding that it's stocked full of undamaged books, realizes that his prayers have been answered—he finally has time, time enough at last to read all he wants! The hook, though—spoiler alert from 1959!—is that as he sits down to read his first book, he breaks his glasses. The episode closes with Burgess howling about the unfairness of it all, amidst stacks of books that he will now be forever unable to read.
This parallels the situation most consumers (at least within the US and much of Europe) find themselves today. We have, finally, storage enough at last to do just about anything we want. We can create and keep terabytes of digital stuff, but the value of that stuff is in what you can do with it, and a home movie doesn't do you any good sitting on your hard drive—its value comes from sharing it. Increasingly, that sharing takes place over the Internet.
And sharing is a problem. Let's take a look at the typical amount of inbound and outbound network bandwidth a US home might have had for the same years in which we look at hard drive capacity. Internet connections to the home weren't terribly common before about 1994, but those of us old enough to experience the days before USENET and the Web will remember local BBS systems and fledgling "information services" like Prodigy and Compuserve, all of which served many of the same purposes then that the Internet does now—connecting with others and exchanging files and messages.
From 300bps in 1981, we travel through 1200bps and the halcyon days of local BBSs at 14.4Kbps, through the brief 56K era and into broadband, going from 1Mb to 3Mb to the 10Mb (and beyond) connections common today. The same kind of exponential-ish increase in speeds is evident, especially in the break between 1996 and 2001 when the widespread shift to broadband connectivity began to pull folks out of needing to use modems over POTS and onto DSL and cable Internet access. However, with that came huge gulfs between the speed at which a person could download things to his computer versus the speed at which he could upload things from it to others.
To be sure, upload/download asymmetry has been around for a long time. Back in my own BBS days, I used to look with envy on the guys with USRobotics Courier HST Dual Standard modems, which could talk to other Courier HST modems at a ridiculously fast (for the time) 16.8Kbps and at the same time have a 450bps back-channel available for control traffic. Similarly, most folks who lived through the 56K days remember that 56K was really only one way—traffic going out from your modem was still limited to 33.8Kbps (and in reality, most folks only saw 22-26 Kbps). But nowhere is the asymmetry as great as with modern cable and DSL connections. The practical reason behind the split between up and down speeds is one of frequency utilization—there's only so many hertz available to carry a signal, and so more of the available frequency space is allocated to carry information to the consumer than from the consumer.
Unfortunately, this common asymmetric split now affects a good chunk of the world's Internet users.
You can't take it with you
Most folks carry what feels like significant chunks of our lives on our hard drives—we express ourselves through the things we create, and we keep things that have powerful meaning for us. A hard drive crash where data is lost can lead to not just technical or financial problems, but actual personal anguish—our drives are surrogates for our memories, and if they die we can lose parts of ourselves.
And yet, with all the innovation that has arrived as a consequence of more and faster data storage, most lack the ability to easily move that data around. We can download multi-gigabyte high definition movies from iTunes or NetFlix in astonishingly short amounts of time, but how many hours does it take to upload a multi-gigabyte file?
Amount of time it would take to upload an entire hard drive at each point in time
Companies like Microsoft and Apple and others continue to roll out cloud sync services, which among other things suggest you put your important files "in the cloud" so they can be accessible anywhere—but why should that be necessary when they could just as easily be accessed in situ on your own computer? I have hundreds of movies that I've ripped from DVDs that I own just sitting on my NAS; why shouldn't I be able to access them wherever I am and stream them to myself at a friend's house as fast as NetFlix can? Why shouldn't I be able to get to any part of my digital identity from anywhere in the world?
Cloud backup services like Backblaze or Mozy have their place, too, providing off-site storage for important data in case of a disaster, but it can take weeks—literally weeks—to upload a few terabytes of data. Most people today don't bother with this kind of backup service, precisely because it takes so long to populate initially. Worse still, for people with metered Internet connections, that initial data copy will almost certainly count against your 50 or 150 or 250 GB monthly cap, which might lead to your Internet access being terminated without explanation by your ISP.
The lack of upstream bandwidth in a typical consumer broadband connection holds back some of the innovation that could come from huge amounts of storage. For instance, with a 10 Mb upload connection to go along with your 10 Mb download connection, content sharing sites like YouTube suddenly lose some of their relevance. "Digital locker" services like RapidShare and MegaUpload serve fewer purposes. An entire infrastructure of sites and services that have evolved as buffers and workarounds for tiny upload speeds are made redundant and can be done away with.
Do more with more
Even if our wings remain clipped by upload restrictions, we're still innovating through storage. We have more and faster storage available to us at home and at work, and the things we can do with that capacity continue to evolve. Parkinson's Law states that "work expands so as to fill the time available for its completion." A common corollary to that law is that "data expands to fill the space available for storage." Put even more simply, as the ghostly voice said to Kevin Costner in Field of Dreams, "If you build it, they will come."
Nothing exists in a vacuum, of course, and it's disingenuous to suggest that increases in storage capacity can alone be responsible for anything, much in the way that it's wrong to say an increase in CPU clock speeds alone is responsible for anything, or that the change from IDE to PCI buses in a PC was by itself responsible for anything. However, the rapid increase in storage capacities to which the average consumer has access has removed barriers and created big new opportunities.
Thursday, September 29, 2011
Wednesday, September 14, 2011
Six Provocations for Big Data
Six Provocations for Big Data
danah boyd, Microsoft Research; University of New South Wales (UNSW); Harvard University - Berkman Center for Internet & Society
Kate Crawford, University of New South Wales (UNSW)
September 21, 2011
Abstract:
The era of Big Data has begun. Computer scientists, physicists, economists, mathematicians, political scientists, bio-informaticists, sociologists, and many others are clamoring for access to the massive quantities of information produced by and about people, things, and their interactions. Diverse groups argue about the potential benefits and costs of analyzing information from Twitter, Google, Verizon, 23andMe, Facebook, Wikipedia, and every space where large groups of people leave digital traces and deposit data. Significant questions emerge.
Will large-scale analysis of DNA help cure diseases? Or will it usher in a new wave of medical inequality? Will data analytics help make people’s access to information more efficient and effective? Or will it be used to track protesters in the streets of major cities? Will it transform how we study human communication and culture, or narrow the palette of research options and alter what ‘research’ means? Some or all of the above?
This essay offers six provocations that we hope can spark conversations about the issues of Big Data. Given the rise of Big Data as both a phenomenon and a methodological persuasion, we believe that it is time to start critically interrogating this phenomenon, its assumptions, and its biases.
(This paper was presented at Oxford Internet Institute’s “A Decade in Internet Time: Symposium on the Dynamics of theInternet and Society” on September 21, 2011.)
danah boyd, Microsoft Research; University of New South Wales (UNSW); Harvard University - Berkman Center for Internet & Society
Kate Crawford, University of New South Wales (UNSW)
September 21, 2011
Abstract:
The era of Big Data has begun. Computer scientists, physicists, economists, mathematicians, political scientists, bio-informaticists, sociologists, and many others are clamoring for access to the massive quantities of information produced by and about people, things, and their interactions. Diverse groups argue about the potential benefits and costs of analyzing information from Twitter, Google, Verizon, 23andMe, Facebook, Wikipedia, and every space where large groups of people leave digital traces and deposit data. Significant questions emerge.
Will large-scale analysis of DNA help cure diseases? Or will it usher in a new wave of medical inequality? Will data analytics help make people’s access to information more efficient and effective? Or will it be used to track protesters in the streets of major cities? Will it transform how we study human communication and culture, or narrow the palette of research options and alter what ‘research’ means? Some or all of the above?
This essay offers six provocations that we hope can spark conversations about the issues of Big Data. Given the rise of Big Data as both a phenomenon and a methodological persuasion, we believe that it is time to start critically interrogating this phenomenon, its assumptions, and its biases.
(This paper was presented at Oxford Internet Institute’s “A Decade in Internet Time: Symposium on the Dynamics of theInternet and Society” on September 21, 2011.)
Tuesday, September 13, 2011
Big Data: Spy Agency Seeks Digital Mosaic to Divine Future
The U.S. intelligence community wants to mine lots and lots of the tidbits bopping around on the Internet to suss out trends before they make the news.
Emily Badger Miller McCune September 9, 2011
U.S. intelligence agencies hope those tidbits bopping around on the Internet will help discover all kinds of trends before they make the news. (John Foxx/Stockbyte)
Governments have been caught off guard a lot lately: by revolutions, by riots, even by unemployment rates (or, to go back even further, by events like 9/11).
In the information age — where there's no limit to publicly available data on everything from political chatter to gas prices — it seems policymakers should be better at predicting major societal shifts and events than at any point in history. Shouldn't all these little pieces of information be telling us something big? Shouldn't they be telling us about where the next mass migration will come from or where the next riot will be?
The government's Office of the Director of National Intelligence is betting this is the case. Its Intelligence Advanced Research Projects Activity ( or IARPA — which sounds like a less menacing cousin to DARPA) is rolling out a new R&D project to test tools that would mine publicly available data to predict political and humanitarian crises, disease outbreaks, mass violence and instability. The project, the Open Source Indicators Program, is premised on the idea that big events are preceded by population-level changes, and that those population-level changes should be identifiable if we just look in the right places.
Idea Lobby
THE IDEA LOBBY
Miller-McCune's Washington correspondent Emily Badger follows the ideas informing, explaining and influencing government, from the local think tank circuit to academic research that shapes D.C. policy from afar.
The concept isn't new. Allied intelligence officers did something similar during World War II, for example, mining letters to the editors of local newspapers and radio transmissions for clues as to what was going on inside Nazi Germany.
"The difference, the big difference — and the thing this initiative is picking up on — is that whereas we used to have to do this with a relatively small sample of radio shows and newspapers and stuff like that, we're now drinking from the fire hose of the Internet," said Philip Schrodt, a political scientist at Penn State.
An IARPA public affairs officer said officials could not discuss the program while the government is still soliciting proposals. But Schrodt and several other researchers affiliated with academic teams that may eventually wind up working on the project, alongsideprivate contractors, offered a look into an expanding, multidisciplinary field of real-time data analysis that the government hopes may allow it to "beat the news."
This experiment — which will test methods in Latin America, not the U.S. — dovetails with broader trends in computer and social science toward mining large-scale sets of social data. Twitter, for instance, is a jackpot: The network can be scoured both for the content of messages and the connections that are revealed between people writing and reading them.
About half of the populations in most major Western countries are now on Facebook (and a surprising amount of data on Facebook remains accessible to the public). Economic data can be collected from e-commerce sites like Amazon. There are also open-source indicators embedded in Internet news sites, message boards, unemployment data, Web search queries, traffic Web cams and financial markets.
Think of what researchers could learn about the economic mood of consumers if they suddenly discovered hundreds of people (each anonymously identified) all trying to sell their best jewelry on Craigslist tomorrow. Other open-source indicators could provide a trove of information on a topic social scientists have been weighing a lot lately: the effect of price changes — whether for housing, gas or food — on social stability.
And because all this data can be collected in real time and by automated systems, it's more up-to-date, it's larger in sheer quantity, and it's becoming more cost-effective and realistic than ever to analyze.
• • • • • • • • • • • • • • •
The easiest place to understand the potential of all this data is in public health, where live analysis is already helping to identify and track disease epidemics at a speed that was never possible before the Internet. Google Flu Trends pioneered the technique. The tool measures the frequency of certain search terms commonly associated with the flu to identify outbreaks down to the city level. Now it's doing the same with dengue fever.
Researchers at Harvard developed a similar tool five years ago calledHealthMap, which offers a prototype of the model IARPA envisions testing beyond public health. HealthMap continuously scrapes the Web for health-related keywords and deletes noise, then classifies the data, geocodes it, filters it into different categories and maps the results. The project is particularly useful in identifying early disease indicators in countries that have no public health capacity and weak or opaque data reporting (because these are often the same countries where sick people are unlikely to Google their symptoms on a home computer, HealthMap also taps into SMS messages and smartphone data).
"These discussions are taking place by the minute, and we're caching those in real time," said John Brownstein, the director and co-founder of HealthMap. "Our goal is as soon as anybody's talking about an outbreak on the Web, within the hour it ends up on HealthMap."
In theory, similar methods might track contagious ideas, economic unrest or resource shortages — and in a way that would actually move faster than news reports.
Automated systems could certainly be more comprehensive than traditional media. The BBC and Associated Press don't have correspondents in every small town in Mexico, but an automated computer program could analyze in real time reports from every local news site in the country. Or, better yet, it could analyze social network chatter before it even gets to the local newsroom.
"If everybody is suddenly upset about the drought in Texas or something like that," said Schrodt, "we should be able to pick that up immediately without having some reporter go out and talk to some farmer whose cows have died."
This type of analysis could also help policymakers be less surprised by news when it happens, and to monitor it literally in real time.
The bigger question, though, is how researchers take the step from watching trends develop in a live time series to anticipating them before they happen.
Some trends lend themselves more easily to this challenge — disease epidemics are one. HealthMap can start to identify outbreaks before officials even know what disease they're looking at because the tool is designed to scan for basic symptoms like runny noses or a run on aspirin. HealthMap isn't merely crawling the Web for Google searches of "Do I have SARS?"
But how do we anticipate "riots in London" when we don't even know "riots in London" is what we're looking for? Could open-source indicators have identified the rise of the Tea Party before it even started calling itself that?
• • • • • • • • • • • • • • •
"That's the really, really big question, that's the single biggest challenge in this IARPA initiative," Schrodt said. He has worked on other government research projects, including one called the Political Instability Task Force. It's clear, though, in that project, what researchers are looking for: political instability, which is suggested by fewer than five indicators. Here, though, the goal is to identify everything that's publicly available, for anything that might be interesting.
"I can't take a book off my shelf and open it up and say, 'Oh, here's what you do if you're monitoring 100 different indicators using 200 different sources,'" Schrodt said. "That's the new science on this. I don't know if it will work or not." (The political instability project, as an example, has gotten to about 80 percent accuracy forecasting probabilities that countries will collapse.)
Peter Gloor's research at MIT's Center for Collective Intelligence has been trying to solve this unknown prediction problem by identifying the most creative people – both those who are destructively creative, like terrorist plotters, and those who are constructively creative, like Justin Bieber. He's trying to find the people behind trends — information producers, not the information itself — before trends take off.
"We're looking for the trendsetters while they are being born and made," Gloor said. "In Twitter, once you are Ashton Kutcher, everyone knows you; it's clear you are a trendsetter. But Justin Bieber, whenever he was posting his first video on YouTube, it was not as clear."
Gloor has done a lot of this work in Wikipedia, where pages are written and edited according to a revealing pattern: Generally, about 1 percent of contributors write 90 percent of the content; 9 percent of people write another 9 percent of the content; and 90 percent of contributors write just 1 percent of the crowdsourced encyclopedia. If you find the 1 percent of people who are doing all the heavy lifting, those are your future trendsetters (or "coolhunters").
Charles Elkan, a computer scientist at the University of California at San Diego, suspects the biggest value of all these tools will come from quantifying what social scientists already know.
"Social scientists have said for 100 years that revolutions happen at times of rising expectations," he said. "If, for example, society is becoming more prosperous, people start have rising economic expectations, that spills over into rising political expectations, and that can make a revolution more likely. Social scientists have said this for a long time, and now it's beginning to be possible to quantify it."
Finland, for example, is obviously a more stable country than Sudan. But attaching a number to that statement — say, Sudan is eight times more likely to experience a coup or revolution in the next two years — is much trickier.
• • • • • • • • • • • • • • •
One thing open-source indicators probably can't do is write the news before it happens.
"Predicting events is even harder than predicting trends," Elkan said. "If you think of a trend as creating the probability for something, that probability combined with some spark makes the actual event."
And then you have to predict the spark, too.
"We can say unemployment is trending upwards, and that layoffs are increasing," Elkan went on, "but it's still very difficult to predict which manager of which company will wake up in the morning and decide, 'I can't wait any longer, I need to lay off 10 people today.'"
It's possible to predict, for instance, that riots are more likely in London than they are in Orlando. But no one would have predicted a week ahead of time that a specific police shooting would catalyze riots (as plenty of other police shootings — even in other seething areas — don't set off riots). Similarly, no one would have predicted that the suicide of a Tunisian fruit vendor would spark a revolution that would sweep out of North Africa and into the Arabian Peninsula and Mideast.
In this sense, there may be a sizable gap between our imagination of what's possible with trend prediction and the reality of what it can do for policymakers.
The primary limitation isn't the data collection; it's the data analysis. And there's a point in the process where human judgment must take over from automated computer systems. That is inevitably the moment when controversial or expensive action looms: Should we double the FEMA budget, send support to Syrian revolutionaries or shift course on jobs policy to quell a domestic uprising?
The attention span of decision-makers, Elkan cautions, is a limited resource, too.
There's another factor: Governments aren't the only ones who might like to leverage this data. Its business application is obvious. Movie theaters could use trend prediction to book films. Department stores could use it to stock products. Corporations could mine Twitter chatter about them to craft corporate social responsibility policies.
"That's one reason to be skeptical about the extent of the possibility," Elkan said. "If it was really possible, for example, to predict unemployment much better than existing methods do, then people on Wall Street would be doing it and taking advantage of it."
Like the other researchers, he's cautious about where all of this is headed, even as he tries to design computer models to more accurately predict the future.
"We can easily specify, 'Let's monitor every newspaper and every radio station everywhere in the world, then as soon as there's one school that has a local outbreak of some disease, then we can put that into our computer models and make predictions of when it will come to the United States,'" Elkan said. "That's a realistic dream, but it's still a dream at this point."
Emily Badger Miller McCune September 9, 2011
U.S. intelligence agencies hope those tidbits bopping around on the Internet will help discover all kinds of trends before they make the news. (John Foxx/Stockbyte)
Governments have been caught off guard a lot lately: by revolutions, by riots, even by unemployment rates (or, to go back even further, by events like 9/11).
In the information age — where there's no limit to publicly available data on everything from political chatter to gas prices — it seems policymakers should be better at predicting major societal shifts and events than at any point in history. Shouldn't all these little pieces of information be telling us something big? Shouldn't they be telling us about where the next mass migration will come from or where the next riot will be?
The government's Office of the Director of National Intelligence is betting this is the case. Its Intelligence Advanced Research Projects Activity ( or IARPA — which sounds like a less menacing cousin to DARPA) is rolling out a new R&D project to test tools that would mine publicly available data to predict political and humanitarian crises, disease outbreaks, mass violence and instability. The project, the Open Source Indicators Program, is premised on the idea that big events are preceded by population-level changes, and that those population-level changes should be identifiable if we just look in the right places.
Idea Lobby
THE IDEA LOBBY
Miller-McCune's Washington correspondent Emily Badger follows the ideas informing, explaining and influencing government, from the local think tank circuit to academic research that shapes D.C. policy from afar.
The concept isn't new. Allied intelligence officers did something similar during World War II, for example, mining letters to the editors of local newspapers and radio transmissions for clues as to what was going on inside Nazi Germany.
"The difference, the big difference — and the thing this initiative is picking up on — is that whereas we used to have to do this with a relatively small sample of radio shows and newspapers and stuff like that, we're now drinking from the fire hose of the Internet," said Philip Schrodt, a political scientist at Penn State.
An IARPA public affairs officer said officials could not discuss the program while the government is still soliciting proposals. But Schrodt and several other researchers affiliated with academic teams that may eventually wind up working on the project, alongsideprivate contractors, offered a look into an expanding, multidisciplinary field of real-time data analysis that the government hopes may allow it to "beat the news."
This experiment — which will test methods in Latin America, not the U.S. — dovetails with broader trends in computer and social science toward mining large-scale sets of social data. Twitter, for instance, is a jackpot: The network can be scoured both for the content of messages and the connections that are revealed between people writing and reading them.
About half of the populations in most major Western countries are now on Facebook (and a surprising amount of data on Facebook remains accessible to the public). Economic data can be collected from e-commerce sites like Amazon. There are also open-source indicators embedded in Internet news sites, message boards, unemployment data, Web search queries, traffic Web cams and financial markets.
Think of what researchers could learn about the economic mood of consumers if they suddenly discovered hundreds of people (each anonymously identified) all trying to sell their best jewelry on Craigslist tomorrow. Other open-source indicators could provide a trove of information on a topic social scientists have been weighing a lot lately: the effect of price changes — whether for housing, gas or food — on social stability.
And because all this data can be collected in real time and by automated systems, it's more up-to-date, it's larger in sheer quantity, and it's becoming more cost-effective and realistic than ever to analyze.
• • • • • • • • • • • • • • •
The easiest place to understand the potential of all this data is in public health, where live analysis is already helping to identify and track disease epidemics at a speed that was never possible before the Internet. Google Flu Trends pioneered the technique. The tool measures the frequency of certain search terms commonly associated with the flu to identify outbreaks down to the city level. Now it's doing the same with dengue fever.
Researchers at Harvard developed a similar tool five years ago calledHealthMap, which offers a prototype of the model IARPA envisions testing beyond public health. HealthMap continuously scrapes the Web for health-related keywords and deletes noise, then classifies the data, geocodes it, filters it into different categories and maps the results. The project is particularly useful in identifying early disease indicators in countries that have no public health capacity and weak or opaque data reporting (because these are often the same countries where sick people are unlikely to Google their symptoms on a home computer, HealthMap also taps into SMS messages and smartphone data).
"These discussions are taking place by the minute, and we're caching those in real time," said John Brownstein, the director and co-founder of HealthMap. "Our goal is as soon as anybody's talking about an outbreak on the Web, within the hour it ends up on HealthMap."
In theory, similar methods might track contagious ideas, economic unrest or resource shortages — and in a way that would actually move faster than news reports.
Automated systems could certainly be more comprehensive than traditional media. The BBC and Associated Press don't have correspondents in every small town in Mexico, but an automated computer program could analyze in real time reports from every local news site in the country. Or, better yet, it could analyze social network chatter before it even gets to the local newsroom.
"If everybody is suddenly upset about the drought in Texas or something like that," said Schrodt, "we should be able to pick that up immediately without having some reporter go out and talk to some farmer whose cows have died."
This type of analysis could also help policymakers be less surprised by news when it happens, and to monitor it literally in real time.
The bigger question, though, is how researchers take the step from watching trends develop in a live time series to anticipating them before they happen.
Some trends lend themselves more easily to this challenge — disease epidemics are one. HealthMap can start to identify outbreaks before officials even know what disease they're looking at because the tool is designed to scan for basic symptoms like runny noses or a run on aspirin. HealthMap isn't merely crawling the Web for Google searches of "Do I have SARS?"
But how do we anticipate "riots in London" when we don't even know "riots in London" is what we're looking for? Could open-source indicators have identified the rise of the Tea Party before it even started calling itself that?
• • • • • • • • • • • • • • •
"That's the really, really big question, that's the single biggest challenge in this IARPA initiative," Schrodt said. He has worked on other government research projects, including one called the Political Instability Task Force. It's clear, though, in that project, what researchers are looking for: political instability, which is suggested by fewer than five indicators. Here, though, the goal is to identify everything that's publicly available, for anything that might be interesting.
"I can't take a book off my shelf and open it up and say, 'Oh, here's what you do if you're monitoring 100 different indicators using 200 different sources,'" Schrodt said. "That's the new science on this. I don't know if it will work or not." (The political instability project, as an example, has gotten to about 80 percent accuracy forecasting probabilities that countries will collapse.)
Peter Gloor's research at MIT's Center for Collective Intelligence has been trying to solve this unknown prediction problem by identifying the most creative people – both those who are destructively creative, like terrorist plotters, and those who are constructively creative, like Justin Bieber. He's trying to find the people behind trends — information producers, not the information itself — before trends take off.
"We're looking for the trendsetters while they are being born and made," Gloor said. "In Twitter, once you are Ashton Kutcher, everyone knows you; it's clear you are a trendsetter. But Justin Bieber, whenever he was posting his first video on YouTube, it was not as clear."
Gloor has done a lot of this work in Wikipedia, where pages are written and edited according to a revealing pattern: Generally, about 1 percent of contributors write 90 percent of the content; 9 percent of people write another 9 percent of the content; and 90 percent of contributors write just 1 percent of the crowdsourced encyclopedia. If you find the 1 percent of people who are doing all the heavy lifting, those are your future trendsetters (or "coolhunters").
Charles Elkan, a computer scientist at the University of California at San Diego, suspects the biggest value of all these tools will come from quantifying what social scientists already know.
"Social scientists have said for 100 years that revolutions happen at times of rising expectations," he said. "If, for example, society is becoming more prosperous, people start have rising economic expectations, that spills over into rising political expectations, and that can make a revolution more likely. Social scientists have said this for a long time, and now it's beginning to be possible to quantify it."
Finland, for example, is obviously a more stable country than Sudan. But attaching a number to that statement — say, Sudan is eight times more likely to experience a coup or revolution in the next two years — is much trickier.
• • • • • • • • • • • • • • •
One thing open-source indicators probably can't do is write the news before it happens.
"Predicting events is even harder than predicting trends," Elkan said. "If you think of a trend as creating the probability for something, that probability combined with some spark makes the actual event."
And then you have to predict the spark, too.
"We can say unemployment is trending upwards, and that layoffs are increasing," Elkan went on, "but it's still very difficult to predict which manager of which company will wake up in the morning and decide, 'I can't wait any longer, I need to lay off 10 people today.'"
It's possible to predict, for instance, that riots are more likely in London than they are in Orlando. But no one would have predicted a week ahead of time that a specific police shooting would catalyze riots (as plenty of other police shootings — even in other seething areas — don't set off riots). Similarly, no one would have predicted that the suicide of a Tunisian fruit vendor would spark a revolution that would sweep out of North Africa and into the Arabian Peninsula and Mideast.
In this sense, there may be a sizable gap between our imagination of what's possible with trend prediction and the reality of what it can do for policymakers.
The primary limitation isn't the data collection; it's the data analysis. And there's a point in the process where human judgment must take over from automated computer systems. That is inevitably the moment when controversial or expensive action looms: Should we double the FEMA budget, send support to Syrian revolutionaries or shift course on jobs policy to quell a domestic uprising?
The attention span of decision-makers, Elkan cautions, is a limited resource, too.
There's another factor: Governments aren't the only ones who might like to leverage this data. Its business application is obvious. Movie theaters could use trend prediction to book films. Department stores could use it to stock products. Corporations could mine Twitter chatter about them to craft corporate social responsibility policies.
"That's one reason to be skeptical about the extent of the possibility," Elkan said. "If it was really possible, for example, to predict unemployment much better than existing methods do, then people on Wall Street would be doing it and taking advantage of it."
Like the other researchers, he's cautious about where all of this is headed, even as he tries to design computer models to more accurately predict the future.
"We can easily specify, 'Let's monitor every newspaper and every radio station everywhere in the world, then as soon as there's one school that has a local outbreak of some disease, then we can put that into our computer models and make predictions of when it will come to the United States,'" Elkan said. "That's a realistic dream, but it's still a dream at this point."
Monday, September 12, 2011
Big Data : Data-Driven Decisions in Asthma Care
Asthmapolis: Tools to Track, Manage, and Research Asthma
GPS Sensors and Mobile App Track Asthma Symptoms, Triggers, and Inhaler Use
California Healthcare Foundation September 2011
Despite dramatic advances in medication effectiveness, many people with asthma experience episodes of uncontrolled disease, causing them to be at greater risk for acute exacerbations. Historically health care providers have had inadequate tools to understand the severity of a patient's asthma, limiting their ability to predict and prevent serious and costly asthma attacks. To address this problem Asthmapolis developed GPS sensors that record exactly when and where patients use their inhalers and a digital interface to display the information captured. The patient-level data may then help physicians make more informed clinical interventions and help patients track and manage their condition.
With the support of the California HealthCare Foundation, Asthmapolis has begun a pilot with Woodland HealthCare, a member of Catholic HealthCare West, to determine if the data collected by sensors and presented to physicians and patients results in actions that improve asthma control. This study is being conducted with English- and Spanish-speaking subjects and will include underserved patients in the Sacramento area.
CHCF invested in Asthmapolis because we see the potential for patients to benefit from technology that provides tailored information, allowing them to better understand and manage asthma, which will hopefully lead to fewer serious and costly episodes.
Read more: http://www.chcf.org/projects/2011/asthmapolis#ixzz1Xm0uua3l
GPS Sensors and Mobile App Track Asthma Symptoms, Triggers, and Inhaler Use
California Healthcare Foundation September 2011
Despite dramatic advances in medication effectiveness, many people with asthma experience episodes of uncontrolled disease, causing them to be at greater risk for acute exacerbations. Historically health care providers have had inadequate tools to understand the severity of a patient's asthma, limiting their ability to predict and prevent serious and costly asthma attacks. To address this problem Asthmapolis developed GPS sensors that record exactly when and where patients use their inhalers and a digital interface to display the information captured. The patient-level data may then help physicians make more informed clinical interventions and help patients track and manage their condition.
With the support of the California HealthCare Foundation, Asthmapolis has begun a pilot with Woodland HealthCare, a member of Catholic HealthCare West, to determine if the data collected by sensors and presented to physicians and patients results in actions that improve asthma control. This study is being conducted with English- and Spanish-speaking subjects and will include underserved patients in the Sacramento area.
CHCF invested in Asthmapolis because we see the potential for patients to benefit from technology that provides tailored information, allowing them to better understand and manage asthma, which will hopefully lead to fewer serious and costly episodes.
Read more: http://www.chcf.org/projects/2011/asthmapolis#ixzz1Xm0uua3l
Monday, August 22, 2011
Big Data: Why HP Wants Autonomy: Math Skills
Autonomy excels at analyzing the vast amounts of "unstructured data" being produced every day.
Tom Simonite Technology Review (by MIT) Thursday, August 18, 2011
News broke today that HP, the world's biggest manufacturer of personal computers, had offered to acquire the British software company Autonomy. While the latter is hardly a household name, it gets close to $1 billion in revenue each year from software that can turn huge volumes of images, text, and video into useful statistics and insights for businesses.
Acquiring that technology will enable HP to expand its business software products, and put it in a good position to exploit a trend dubbed "big data." Businesses are increasingly interested in finding ways to distill meaning from the growing piles of digital information, from tweets to video, flowing through our lives at work and at home.
Whit Andrews, a vice president and analyst with Gartner who specializes in technology that processes and organizes information, says that Autonomy was years ahead of other companies in making such analysis possible. "They have had this vision for over a decade that there was immense value in being able to do statistical analysis for data like audio and video that conventional technology cannot handle."
Autonomy's products enable companies to do things like analyze transcripts from call centers; discover which sales strategies work best; and process troves of e-mails and other documents to match whether what is being said and done comports with a company's legal responsibilities. Such analysis can be automated using software that looks for certain things automatically, or performed manually by a person entering queries, and then sifting through the results themselves.
Andrews says business and technology companies are beginning to realize that both types of analysis could offer much more than conventional approaches, which rely on so-called "structured" data, such as a spreadsheet organized into labelled columns. "Business analytics is about structured data, like spreadsheets," says Andrews. "Autonomy does an exceptional job at analyzing unstructured data, which may prove even more valuable."
In an interview with Technology Review published last year, Autonomy's founder and CEO, Mike Lynch, estimated that about 85 percent of the information inside a business is unstructured. "[W]e are human beings, and unstructured information is at the core of everything we do," he said. "Most business is done using this kind of human-friendly information."
Lynch founded the company to commercialize statistical techniques developed at Cambridge University based on Bayesian inference, a mathematical technique that can estimate the probability of potential outcomes based on previous evidence.
Companies like IBM are working hard on their own approaches to analyzing unstructured data, but Autonomy has been at it for longer, says Andrews. Acquiring the company could enable HP to take a much more dominant position in the growing market for what Autonomy's Lynch dubs "meaning-based computing."
Tom Simonite Technology Review (by MIT) Thursday, August 18, 2011
News broke today that HP, the world's biggest manufacturer of personal computers, had offered to acquire the British software company Autonomy. While the latter is hardly a household name, it gets close to $1 billion in revenue each year from software that can turn huge volumes of images, text, and video into useful statistics and insights for businesses.
Acquiring that technology will enable HP to expand its business software products, and put it in a good position to exploit a trend dubbed "big data." Businesses are increasingly interested in finding ways to distill meaning from the growing piles of digital information, from tweets to video, flowing through our lives at work and at home.
Whit Andrews, a vice president and analyst with Gartner who specializes in technology that processes and organizes information, says that Autonomy was years ahead of other companies in making such analysis possible. "They have had this vision for over a decade that there was immense value in being able to do statistical analysis for data like audio and video that conventional technology cannot handle."
Autonomy's products enable companies to do things like analyze transcripts from call centers; discover which sales strategies work best; and process troves of e-mails and other documents to match whether what is being said and done comports with a company's legal responsibilities. Such analysis can be automated using software that looks for certain things automatically, or performed manually by a person entering queries, and then sifting through the results themselves.
Andrews says business and technology companies are beginning to realize that both types of analysis could offer much more than conventional approaches, which rely on so-called "structured" data, such as a spreadsheet organized into labelled columns. "Business analytics is about structured data, like spreadsheets," says Andrews. "Autonomy does an exceptional job at analyzing unstructured data, which may prove even more valuable."
In an interview with Technology Review published last year, Autonomy's founder and CEO, Mike Lynch, estimated that about 85 percent of the information inside a business is unstructured. "[W]e are human beings, and unstructured information is at the core of everything we do," he said. "Most business is done using this kind of human-friendly information."
Lynch founded the company to commercialize statistical techniques developed at Cambridge University based on Bayesian inference, a mathematical technique that can estimate the probability of potential outcomes based on previous evidence.
Companies like IBM are working hard on their own approaches to analyzing unstructured data, but Autonomy has been at it for longer, says Andrews. Acquiring the company could enable HP to take a much more dominant position in the growing market for what Autonomy's Lynch dubs "meaning-based computing."
4G a boon to U.S. economy and jobs, study says
Roger Cheng CNet News August 21, 2011 9:01 PM PDT
The wireless carriers' investment in 4G networks could be the salve that the ailing U.S. economy is looking for.
The carriers could invest between $25 billion and $53 billion in building out their 4G network through 2016, according to a study from Deloitte. That in turn could lead to the creation of 371,000 to 771,000 jobs, and gross domestic product growth of $73 billion to $151 billion.
"Investment in such a powerful form of communication contributes to the economic recovery and provides a job-creating engine for the future," said Phil Asmundson a consultant for Deloitte.
The rise of 4G networks could provide some support for an economy still struggling to recover, and which some believe is slipping back into a recession. The recent bitter political struggle to raise the national debt ceiling, the sharp declines in the stock market, and the continued high rate of unemployment have many still concerned.
The wide-ranging estimate assumes two different scenarios. The baseline scenario has the carriers deploying 4G technology at a moderate pace with a slow transition from 3G to 4G extending into the middle of the decade. Deloitte warns that under these conditions, the U.S. firms will be vulnerable to foreign competitors looking to jump ahead in the 4G race.
The second scenario predicts a more rapid investment in 4G networks and the creation of 4G-based services before global competitors gain momentum. Deloitte said the demand stimulated by the new services would propel more network investment, "setting off a virtuous cycle of investment and market response." Cloud services are among the primary catalysts for investment, according to Deloitte.
Telecommunications investment has remained strong over the past few years, with large companies such as AT&T and Verizon pouring money into network upgrades. AT&T has argued that its acquisition of T-Mobile would lead to new jobs and investment because it is able to expand its future 4G network across a wider stretch of the country. Chief Executive Randall Stephenson has long argued that investment in telecommunications technology is crucial to staying ahead in the world.
AT&T plans to launch its first 4G LTE markets later this summer.
Verizon, meanwhile, has spent billions of dollars upgrading its fixed-line infrastructure with fiber-optic lines to deliver television content and faster Internet service. On the wireless side, the company has been racing to build out its 4G LTE network, and recently said that it covers more than half the country.
Clearwire is also switching gears from its current 4G WiMax standard and moving toward LTE, but the company said it plans to spend only $600 million for the upgrade.
Clearwire's largest customer and shareholder, Sprint Nextel, is planning to unveil its 4G plans in October. The company has already signed a network-hosting deal with LightSquared, which plans to start testing its 4G network with customers next year.
The widespread adoption of 4G services should also help certain segments such as disadvantaged minority groups; rural communities and areas with limited full broadband access; and small businesses, the firm said, adding the advent of 4G could work to bring those groups further into the economic mainstream.
Read more: http://news.cnet.com/8301-1035_3-20094588-94/4g-a-boon-to-u.s-economy-and-jobs-study-says/#ixzz1VleV8J20
The wireless carriers' investment in 4G networks could be the salve that the ailing U.S. economy is looking for.
The carriers could invest between $25 billion and $53 billion in building out their 4G network through 2016, according to a study from Deloitte. That in turn could lead to the creation of 371,000 to 771,000 jobs, and gross domestic product growth of $73 billion to $151 billion.
"Investment in such a powerful form of communication contributes to the economic recovery and provides a job-creating engine for the future," said Phil Asmundson a consultant for Deloitte.
The rise of 4G networks could provide some support for an economy still struggling to recover, and which some believe is slipping back into a recession. The recent bitter political struggle to raise the national debt ceiling, the sharp declines in the stock market, and the continued high rate of unemployment have many still concerned.
The wide-ranging estimate assumes two different scenarios. The baseline scenario has the carriers deploying 4G technology at a moderate pace with a slow transition from 3G to 4G extending into the middle of the decade. Deloitte warns that under these conditions, the U.S. firms will be vulnerable to foreign competitors looking to jump ahead in the 4G race.
The second scenario predicts a more rapid investment in 4G networks and the creation of 4G-based services before global competitors gain momentum. Deloitte said the demand stimulated by the new services would propel more network investment, "setting off a virtuous cycle of investment and market response." Cloud services are among the primary catalysts for investment, according to Deloitte.
Telecommunications investment has remained strong over the past few years, with large companies such as AT&T and Verizon pouring money into network upgrades. AT&T has argued that its acquisition of T-Mobile would lead to new jobs and investment because it is able to expand its future 4G network across a wider stretch of the country. Chief Executive Randall Stephenson has long argued that investment in telecommunications technology is crucial to staying ahead in the world.
AT&T plans to launch its first 4G LTE markets later this summer.
Verizon, meanwhile, has spent billions of dollars upgrading its fixed-line infrastructure with fiber-optic lines to deliver television content and faster Internet service. On the wireless side, the company has been racing to build out its 4G LTE network, and recently said that it covers more than half the country.
Clearwire is also switching gears from its current 4G WiMax standard and moving toward LTE, but the company said it plans to spend only $600 million for the upgrade.
Clearwire's largest customer and shareholder, Sprint Nextel, is planning to unveil its 4G plans in October. The company has already signed a network-hosting deal with LightSquared, which plans to start testing its 4G network with customers next year.
The widespread adoption of 4G services should also help certain segments such as disadvantaged minority groups; rural communities and areas with limited full broadband access; and small businesses, the firm said, adding the advent of 4G could work to bring those groups further into the economic mainstream.
Read more: http://news.cnet.com/8301-1035_3-20094588-94/4g-a-boon-to-u.s-economy-and-jobs-study-says/#ixzz1VleV8J20
Friday, August 19, 2011
Computer Analysis Could Find New Uses for Existing Drugs
I HealthBeat August 18, 011
A computer program that analyzes drug data and genetic information could help discover new uses for medicines already on the market, according to two new studies published in the journal Science Translational Medicine, United Press International reports.
Methodology
For the NIH-funded research, Stanford University scientists extracted data from NIH's Gene Expression Omnibus, a public database containing findings from thousands of genomic studies conducted throughout the world.
Researchers focused on 100 diseases and 164 drugs. They analyzed thousands of possible drug-disease combinations to identify which medications and medical conditions had gene expression patterns that could cancel each other out. Such matches indicate that the drug potentially could mitigate the effects of the disease (United Press International, 8/17).
Research Findings
Researchers identified possible drug-disease matches for 53 of the 100 diseases analyzed (Renick, Bloomberg, 8/17).
For example, they found that an epilepsy treatment potentially could treat inflammatory bowel disease and that an ulcer drug might be an effective lung cancer medication (Dockser Marcus, Wall Street Journal, 8/18).
Possible Implications
Researchers noted that re-purposing existing medications to treat different diseases could reduce some of the costs and requirements involved in the drug development process. They noted that it takes an average of 15 years and about $1 billion to bring a single new drug to market (Bloomberg, 8/17).
In an accompanying commentary on the research, Yves Lussier -- a professor of medicine and engineering at the University of Illinois in Chicago -- wrote that the findings should not prompt physicians to prescribe an ulcer drug to treat lung cancer. However, he added that the findings are "impressive enough to be improved upon and studied further."
Lussier also noted that if the computer program is effective at detecting possible off-label uses for existing drugs, it "opens the door to very low-cost, individualized personal therapies" (Wall Street Journal, 8/18).
Read more: http://www.ihealthbeat.org/articles/2011/8/18/computer-analysis-could-find-new-uses-for-existing-drugs.aspx#ixzz1VTsrgrjJ
A computer program that analyzes drug data and genetic information could help discover new uses for medicines already on the market, according to two new studies published in the journal Science Translational Medicine, United Press International reports.
Methodology
For the NIH-funded research, Stanford University scientists extracted data from NIH's Gene Expression Omnibus, a public database containing findings from thousands of genomic studies conducted throughout the world.
Researchers focused on 100 diseases and 164 drugs. They analyzed thousands of possible drug-disease combinations to identify which medications and medical conditions had gene expression patterns that could cancel each other out. Such matches indicate that the drug potentially could mitigate the effects of the disease (United Press International, 8/17).
Research Findings
Researchers identified possible drug-disease matches for 53 of the 100 diseases analyzed (Renick, Bloomberg, 8/17).
For example, they found that an epilepsy treatment potentially could treat inflammatory bowel disease and that an ulcer drug might be an effective lung cancer medication (Dockser Marcus, Wall Street Journal, 8/18).
Possible Implications
Researchers noted that re-purposing existing medications to treat different diseases could reduce some of the costs and requirements involved in the drug development process. They noted that it takes an average of 15 years and about $1 billion to bring a single new drug to market (Bloomberg, 8/17).
In an accompanying commentary on the research, Yves Lussier -- a professor of medicine and engineering at the University of Illinois in Chicago -- wrote that the findings should not prompt physicians to prescribe an ulcer drug to treat lung cancer. However, he added that the findings are "impressive enough to be improved upon and studied further."
Lussier also noted that if the computer program is effective at detecting possible off-label uses for existing drugs, it "opens the door to very low-cost, individualized personal therapies" (Wall Street Journal, 8/18).
Read more: http://www.ihealthbeat.org/articles/2011/8/18/computer-analysis-could-find-new-uses-for-existing-drugs.aspx#ixzz1VTsrgrjJ
Tuesday, August 16, 2011
The web turns twenty
Difference Engine: Happy anniversary?
Aug 12th 2011, 10:37 by N.V. | LOS ANGELES
IT IS always a little disconcerting to realise a generation has grown up never knowing what it was like to manage without something that is taken for granted today. A case in point: the World Wide Web (WWW), which celebrated the 20th anniversary of its introduction last Saturday. It is no exaggeration to say that not since the invention of the printing press has a new media technology altered the way people think, work and play quite so extensively. With the web having been so thoroughly embraced socially, politically and economically, the world has become an entirely different place from what it was just two decades ago. Whether the web has made it a better place or a worse one is for readers to decide.
It was on August 6th, 1991, that Tim Berners-Lee, a British physicist at the European Organisation for Nuclear Research (CERN), in Geneva, created the first-ever web page—a summary of his WWW project along with explanations to help visitors build websites of their own and to search the web for information. No screen-shots survive of the original web page; its original address simply redirects visitors to a contemporary site providing details of the project’s early days at CERN.
First, however, a few things to get straight. The web is not to be confused with the internet—a global system of interconnected networks developed in the 1960s, originally for academic and government researchers in America. The internet sends information as discrete packets of data using a suite of protocols known as TCP/IP. The genius of the system is that the data tell the network where they want to go, instead of the network telling the data where they are being sent. All networks adopting this procedure—no matter where they are or how they actually function—are then reduced effectively to the same bare essentials, allowing them to interconnect and exchange data seamlessly.
The web, by contrast, is simply a way of organising information on a computer network by means of “hyperlinks”—ie, references to other resources on the network that users can visit directly from the document they are reading. As conceived, the web is simply another service—albeit a very important one—running on top of the internet.
Apart from coming up with the idea for sharing information embedded with hypertext links over the internet, to make it happen Mr Berners-Lee (subsequently knighted for his efforts) had to create the first web browser-editor, the first web server, and the first version of the hypertext mark-up language (HTML), which would become the primary means for publishing information on the web. Within a year or two of the web’s introduction, software packages such as Viola, Cello and Mosaic had made it possible for users to browse the web graphically—by clicking on highlighted hyperlinks in web pages and being redirected to yet other web pages, and so on.
It is fair to say that, without the internet, the web would not have existed—at least, not in the form we know it today. And without the web, the internet would have remained essentially a tool for geeks and professionals. No doubt, e-mail would have continued to flourish without the web: it was one of the internet’s earliest applications. So would news groups, bulletin boards, instant messaging and listservs. In due course, internet telephony applications like Skype and even streaming video services similar to Hulu or YouTube would have emerged as well. But users would have had to master the vagaries of Archie, Finger, Gopher, Telnet, Veronica and WAIS (don’t even ask). Thanks to the web’s ease of navigation and the richness of its HTML formatting language, most of these arcane internet tools have gone the way of the dodo.
No question that, over the past 20 years, the web has brought numerous benefits. But it has had its dark side, too. Cybercrime has become prevalent as thieves, hucksters, predators, child pornographers, terrorists, drug cartels and even foreign powers have used the anonymity of the so-called “deep web” to perpetrate crimes. In his pioneering study in 2001, Michael Bergman, a semantics-search-engine whiz based in Iowa, reckoned there was 400 to 550 times more information lurking underground in the deep web than on the surface in the public web. Information in the deep web lay hidden from Google’s crawlers by residing behind password-protected firewalls or requiring admission forms to be completed manually to gain access. By Mr Bergman’s estimate, the deep web contained some 7,500 terabytes of information, compared with a mere 19 terabytes in the public web at the time. Put another way, search engines were indexing less than 0.25% of the web pages available.
Things are probably no different today. By and large, though, the bulk of information in such hidden repositories is legitimate, stashed there by private companies, research institutions and government agencies for security reasons. “There’s a lot or legitimate and valuable content in the deep web,” says Juliana Freire, the former leader of a University of Utah project called DeepPeep. Even so, the fact that there is vastly more information on the web that is inaccessible, compared with what is open to public view, gives one pause for thought.
On balance, the world is grateful for what the web has wrought. Despite their cavalier attitudes to privacy, websites like Facebook, Twitter, Tumblr and Foursquare have changed the way a whole generation of people communicates—creating new ways to make friends, find old acquaintances, socialise online and pursue common interests. Business sites like LinkedIn help them further their careers. YouTube and Flickr let enthusiasts share their home videos and snap shots with millions of others. Online dating sites such as Match, with its algorithms for compatibility, have fostered meaningful relationships for many a lonely heart.
From Amazon to Zappos, online retailing sites have taken the drudgery out of shopping, allowing goods to be bought with the click of a mouse at home. E-Bay lets people sell those they no longer want. Meanwhile, music-streaming sites like Spotify have opened millions of ears to melodies they might never otherwise have heard.
At a keystroke, it has become possible to find all sorts of obscure information, thanks to Google, Bing, Ask and other search engines. Wikipedia may not be the most reliable of sources, but at least it provides a quick run-down on practically anything you need to know in a hurry. Compared with printed encyclopaedias and public libraries, the web has democratised the collected wisdom of ages, and redistributed it in a way unimaginable a few decades ago. Meanwhile, people no longer have to wait for newspapers to be delivered in the morning, or for broadcasters to assemble their news shows. Web pages, tweets and blogs deliver the news as it happens.
Few would deny that such services have made the world a smarter, livelier, more interesting place. But while the news travels faster than ever courtesy of the web, so do lies, hyperbole and distortions. All those with access to the web now have a voice to air their grievances, vent their anger, parade their biases, push the boundaries of decency, spill the beans. The gatekeepers have gone.
When WikiLeaks dumps massive volumes of diplomatic correspondence stolen from government computers on its website, it is not engaging in some heroic act of free speech, nor bringing specific cases of wrongdoing to the public’s attention. In a deliberate and calculated manner, it is making the world a more dangerous place. In dealing with issues of privacy, public safety and national security, governments have every right to discuss such matters behind closed doors—indeed, we insist they do. It is dangerously naïve to argue otherwise.
Meanwhile, for every online job the web has created, several others have been lost in the bricks-and-mortar world. And unlike the latter, many of the new online jobs lie beyond a country’s shores. Likewise, for all the new freedoms and certainties the web has created, numerous old ones have disappeared. Consider copyright. Once it provided authors, artists and musicians with a living, and ensured that the fourth estate could do its job of rooting out injustice and corruption. Illegal downloading from the web, and the widespread erosion of copyright protection generally, has put paid to much of that.
You have to wonder whether something is wrong when so many people spend so much of their time these days in front of a computer screen tapping away on a keyboard, instead of going out into the real world to experience life’s actual (as opposed to virtual) adventures. Ironically, for all the labour-saving tools the web has given us, and all the personal connections it has allowed us to make, we seem to have become lonelier and more isolated than ever. That is a rather sorry state of affairs.
Aug 12th 2011, 10:37 by N.V. | LOS ANGELES
IT IS always a little disconcerting to realise a generation has grown up never knowing what it was like to manage without something that is taken for granted today. A case in point: the World Wide Web (WWW), which celebrated the 20th anniversary of its introduction last Saturday. It is no exaggeration to say that not since the invention of the printing press has a new media technology altered the way people think, work and play quite so extensively. With the web having been so thoroughly embraced socially, politically and economically, the world has become an entirely different place from what it was just two decades ago. Whether the web has made it a better place or a worse one is for readers to decide.
It was on August 6th, 1991, that Tim Berners-Lee, a British physicist at the European Organisation for Nuclear Research (CERN), in Geneva, created the first-ever web page—a summary of his WWW project along with explanations to help visitors build websites of their own and to search the web for information. No screen-shots survive of the original web page; its original address simply redirects visitors to a contemporary site providing details of the project’s early days at CERN.
First, however, a few things to get straight. The web is not to be confused with the internet—a global system of interconnected networks developed in the 1960s, originally for academic and government researchers in America. The internet sends information as discrete packets of data using a suite of protocols known as TCP/IP. The genius of the system is that the data tell the network where they want to go, instead of the network telling the data where they are being sent. All networks adopting this procedure—no matter where they are or how they actually function—are then reduced effectively to the same bare essentials, allowing them to interconnect and exchange data seamlessly.
The web, by contrast, is simply a way of organising information on a computer network by means of “hyperlinks”—ie, references to other resources on the network that users can visit directly from the document they are reading. As conceived, the web is simply another service—albeit a very important one—running on top of the internet.
Apart from coming up with the idea for sharing information embedded with hypertext links over the internet, to make it happen Mr Berners-Lee (subsequently knighted for his efforts) had to create the first web browser-editor, the first web server, and the first version of the hypertext mark-up language (HTML), which would become the primary means for publishing information on the web. Within a year or two of the web’s introduction, software packages such as Viola, Cello and Mosaic had made it possible for users to browse the web graphically—by clicking on highlighted hyperlinks in web pages and being redirected to yet other web pages, and so on.
It is fair to say that, without the internet, the web would not have existed—at least, not in the form we know it today. And without the web, the internet would have remained essentially a tool for geeks and professionals. No doubt, e-mail would have continued to flourish without the web: it was one of the internet’s earliest applications. So would news groups, bulletin boards, instant messaging and listservs. In due course, internet telephony applications like Skype and even streaming video services similar to Hulu or YouTube would have emerged as well. But users would have had to master the vagaries of Archie, Finger, Gopher, Telnet, Veronica and WAIS (don’t even ask). Thanks to the web’s ease of navigation and the richness of its HTML formatting language, most of these arcane internet tools have gone the way of the dodo.
No question that, over the past 20 years, the web has brought numerous benefits. But it has had its dark side, too. Cybercrime has become prevalent as thieves, hucksters, predators, child pornographers, terrorists, drug cartels and even foreign powers have used the anonymity of the so-called “deep web” to perpetrate crimes. In his pioneering study in 2001, Michael Bergman, a semantics-search-engine whiz based in Iowa, reckoned there was 400 to 550 times more information lurking underground in the deep web than on the surface in the public web. Information in the deep web lay hidden from Google’s crawlers by residing behind password-protected firewalls or requiring admission forms to be completed manually to gain access. By Mr Bergman’s estimate, the deep web contained some 7,500 terabytes of information, compared with a mere 19 terabytes in the public web at the time. Put another way, search engines were indexing less than 0.25% of the web pages available.
Things are probably no different today. By and large, though, the bulk of information in such hidden repositories is legitimate, stashed there by private companies, research institutions and government agencies for security reasons. “There’s a lot or legitimate and valuable content in the deep web,” says Juliana Freire, the former leader of a University of Utah project called DeepPeep. Even so, the fact that there is vastly more information on the web that is inaccessible, compared with what is open to public view, gives one pause for thought.
On balance, the world is grateful for what the web has wrought. Despite their cavalier attitudes to privacy, websites like Facebook, Twitter, Tumblr and Foursquare have changed the way a whole generation of people communicates—creating new ways to make friends, find old acquaintances, socialise online and pursue common interests. Business sites like LinkedIn help them further their careers. YouTube and Flickr let enthusiasts share their home videos and snap shots with millions of others. Online dating sites such as Match, with its algorithms for compatibility, have fostered meaningful relationships for many a lonely heart.
From Amazon to Zappos, online retailing sites have taken the drudgery out of shopping, allowing goods to be bought with the click of a mouse at home. E-Bay lets people sell those they no longer want. Meanwhile, music-streaming sites like Spotify have opened millions of ears to melodies they might never otherwise have heard.
At a keystroke, it has become possible to find all sorts of obscure information, thanks to Google, Bing, Ask and other search engines. Wikipedia may not be the most reliable of sources, but at least it provides a quick run-down on practically anything you need to know in a hurry. Compared with printed encyclopaedias and public libraries, the web has democratised the collected wisdom of ages, and redistributed it in a way unimaginable a few decades ago. Meanwhile, people no longer have to wait for newspapers to be delivered in the morning, or for broadcasters to assemble their news shows. Web pages, tweets and blogs deliver the news as it happens.
Few would deny that such services have made the world a smarter, livelier, more interesting place. But while the news travels faster than ever courtesy of the web, so do lies, hyperbole and distortions. All those with access to the web now have a voice to air their grievances, vent their anger, parade their biases, push the boundaries of decency, spill the beans. The gatekeepers have gone.
When WikiLeaks dumps massive volumes of diplomatic correspondence stolen from government computers on its website, it is not engaging in some heroic act of free speech, nor bringing specific cases of wrongdoing to the public’s attention. In a deliberate and calculated manner, it is making the world a more dangerous place. In dealing with issues of privacy, public safety and national security, governments have every right to discuss such matters behind closed doors—indeed, we insist they do. It is dangerously naïve to argue otherwise.
Meanwhile, for every online job the web has created, several others have been lost in the bricks-and-mortar world. And unlike the latter, many of the new online jobs lie beyond a country’s shores. Likewise, for all the new freedoms and certainties the web has created, numerous old ones have disappeared. Consider copyright. Once it provided authors, artists and musicians with a living, and ensured that the fourth estate could do its job of rooting out injustice and corruption. Illegal downloading from the web, and the widespread erosion of copyright protection generally, has put paid to much of that.
You have to wonder whether something is wrong when so many people spend so much of their time these days in front of a computer screen tapping away on a keyboard, instead of going out into the real world to experience life’s actual (as opposed to virtual) adventures. Ironically, for all the labour-saving tools the web has given us, and all the personal connections it has allowed us to make, we seem to have become lonelier and more isolated than ever. That is a rather sorry state of affairs.
Wednesday, August 3, 2011
Medicines bright future
by Vivek Wadhwa • July 28, 2011
Internet and social media are capturing the public’s attention, but some of the most significant advances today are happening in medicine. Technology and medicine are converging in new ways to make possible the types of innovations that could be seen on “Star Trek.” Consider this: We spend the majority of our health-care dollars on treating chronic diseases. Technological advances will enable us to shift those investments into improving our health and preventing disease.
My colleague Daniel Kraft is a physician who chairs the medicine track and heads the FutureMed Program for Singularity University. The Silicon Valley-based university teaches business executives, technologists and government leaders about “exponential technologies.” These are inventions in fields that experience faster growth than average — such as robotics, nanotechnology and artificial intelligence. Singularity University’s founders believe that these technologies, when combined in new ways, could solve some of the world’s major problems, such as poverty, hunger, energy shortages and disease.
Here are the three major trends Kraft sees in health and medicine:
Medicine goes mobile and goes home.
Many aspects of health care and disease management will become cheaper and more effective as our mobile phones and other, similar technology platforms become smaller, Web-enabled and interconnected. In essence, these smartphones will become health platforms. They already contain a wide array of sensors, including an accelerometer that can serve as a pedometer, a camera that can photograph external ailments and transmit them for analysis, and a global positioning system (GPS) that can track our locations.
Developers are also looking beyond the smartphone when it comes to developing these new technologies.
For example, Fitbit is a clip-on, Web-integrated device that helps track how many calories you burn during the day; Zeo is a wireless headband that helps you track the quality and duration of your sleep. Other devices, such as the Basis monitor, which is still in development, can keep track of heart rates and movement. Meanwhile, an array of devices, including scales, blood pressure monitors and blood glucose monitors, are becoming Wi-Fi-enabled. These technologies, when they are connected to electronic and personal health records and to social networks, can create powerful feedback loops with friends, and provide clinicians with better information for helping their patients.
Expect to see products that keep track of your health by connecting to global health-care systems similar to the in-car assistance program OnStar. These will incorporate ubiquitous sensors embedded in toothbrushes and clothes, for example. They may even analyze our bathroom visits and food intake. They will likely use artificial-intelligence to constantly monitor our health data, predict disease and summon help in the event we fall ill.
Personalization: From genomics to proteomics
We learned how to sequence the genome a decade ago, and doing it cost billions of dollars. Companies like 23andMe are now offering partial DNA genotyping for as low as $99 (with a one-year subscription to their information service). Expect prices to continue to rapidly decline to that of a regular blood test.
This means that it is now becoming more affordable to compare one person’s DNA with another’s, learn what diseases those with similar genetics have had and discover how effective different medications or other interventions were in treating them. Imagine doing a Google search on specific genes to find others like you and learn their abilities, allergies, likes and dislikes and what diseases they are predisposed to. That future is closer than you may think.
This opens up an era of crowd-sourced, data-driven, participatory, genomics-based medicine. Today, medicines are prescribed on a “one size fits all” basis. When a particular medication causes a significant negative reaction with a small part of the population, it is prevented from being available to anyone. In the future, expect to see doctors prescribing and selecting the most patient-appropriate medicines based on a person’s DNA (the field of “pharmacogenomics”).
Regenerative medicine
Physicians have been conducting adult stem-cell therapy for more than 40 years in the field of bone-marrow transplantation, which involves transplanting stem cells that become red blood cells. Adult stem cells are now being applied in a variety of arenas, from orthopedics to cardiovascular therapy. The first trials using cells derived from embryonic stem cells were for acute spinal cord injury and started within the last year. But embryonic stem cells have raised ethical and moral controversy even though the research remains critical for future progress.
The good news is a new type of cell, induced pluripotent stem cells, which will enable the generation of personalized stem cell lines for use in diagnostics, prognosis or potentially for therapy in the same patient. IPS cells can replace embryonic stem cells for some applications and are, for example, being used to develop neurons from patients with ALS/Lou Gehrig’s disease in order to better understand the disease and develop new therapies.
Tissue engineering and 3-D printing technologies are also beginning to merge. The combination of the two technologies could lead to an era of personalized organ generation. Indeed, earlier this month, surgeons in Sweden carried out the world’s first synthetic organ transplant— a synthetic trachea/windpipe structure created and seeded with the patient’s own progenitor cells.
These developments are just the beginning. There will undoubtedly be regulatory, reimbursement and other challenges. And there will be heated debates about ethics and morals. But it won’t be long before we are using devices similar to the “Star Trek” tricorder and synthesizing our medications.
Internet and social media are capturing the public’s attention, but some of the most significant advances today are happening in medicine. Technology and medicine are converging in new ways to make possible the types of innovations that could be seen on “Star Trek.” Consider this: We spend the majority of our health-care dollars on treating chronic diseases. Technological advances will enable us to shift those investments into improving our health and preventing disease.
My colleague Daniel Kraft is a physician who chairs the medicine track and heads the FutureMed Program for Singularity University. The Silicon Valley-based university teaches business executives, technologists and government leaders about “exponential technologies.” These are inventions in fields that experience faster growth than average — such as robotics, nanotechnology and artificial intelligence. Singularity University’s founders believe that these technologies, when combined in new ways, could solve some of the world’s major problems, such as poverty, hunger, energy shortages and disease.
Here are the three major trends Kraft sees in health and medicine:
Medicine goes mobile and goes home.
Many aspects of health care and disease management will become cheaper and more effective as our mobile phones and other, similar technology platforms become smaller, Web-enabled and interconnected. In essence, these smartphones will become health platforms. They already contain a wide array of sensors, including an accelerometer that can serve as a pedometer, a camera that can photograph external ailments and transmit them for analysis, and a global positioning system (GPS) that can track our locations.
Developers are also looking beyond the smartphone when it comes to developing these new technologies.
For example, Fitbit is a clip-on, Web-integrated device that helps track how many calories you burn during the day; Zeo is a wireless headband that helps you track the quality and duration of your sleep. Other devices, such as the Basis monitor, which is still in development, can keep track of heart rates and movement. Meanwhile, an array of devices, including scales, blood pressure monitors and blood glucose monitors, are becoming Wi-Fi-enabled. These technologies, when they are connected to electronic and personal health records and to social networks, can create powerful feedback loops with friends, and provide clinicians with better information for helping their patients.
Expect to see products that keep track of your health by connecting to global health-care systems similar to the in-car assistance program OnStar. These will incorporate ubiquitous sensors embedded in toothbrushes and clothes, for example. They may even analyze our bathroom visits and food intake. They will likely use artificial-intelligence to constantly monitor our health data, predict disease and summon help in the event we fall ill.
Personalization: From genomics to proteomics
We learned how to sequence the genome a decade ago, and doing it cost billions of dollars. Companies like 23andMe are now offering partial DNA genotyping for as low as $99 (with a one-year subscription to their information service). Expect prices to continue to rapidly decline to that of a regular blood test.
This means that it is now becoming more affordable to compare one person’s DNA with another’s, learn what diseases those with similar genetics have had and discover how effective different medications or other interventions were in treating them. Imagine doing a Google search on specific genes to find others like you and learn their abilities, allergies, likes and dislikes and what diseases they are predisposed to. That future is closer than you may think.
This opens up an era of crowd-sourced, data-driven, participatory, genomics-based medicine. Today, medicines are prescribed on a “one size fits all” basis. When a particular medication causes a significant negative reaction with a small part of the population, it is prevented from being available to anyone. In the future, expect to see doctors prescribing and selecting the most patient-appropriate medicines based on a person’s DNA (the field of “pharmacogenomics”).
Regenerative medicine
Physicians have been conducting adult stem-cell therapy for more than 40 years in the field of bone-marrow transplantation, which involves transplanting stem cells that become red blood cells. Adult stem cells are now being applied in a variety of arenas, from orthopedics to cardiovascular therapy. The first trials using cells derived from embryonic stem cells were for acute spinal cord injury and started within the last year. But embryonic stem cells have raised ethical and moral controversy even though the research remains critical for future progress.
The good news is a new type of cell, induced pluripotent stem cells, which will enable the generation of personalized stem cell lines for use in diagnostics, prognosis or potentially for therapy in the same patient. IPS cells can replace embryonic stem cells for some applications and are, for example, being used to develop neurons from patients with ALS/Lou Gehrig’s disease in order to better understand the disease and develop new therapies.
Tissue engineering and 3-D printing technologies are also beginning to merge. The combination of the two technologies could lead to an era of personalized organ generation. Indeed, earlier this month, surgeons in Sweden carried out the world’s first synthetic organ transplant— a synthetic trachea/windpipe structure created and seeded with the patient’s own progenitor cells.
These developments are just the beginning. There will undoubtedly be regulatory, reimbursement and other challenges. And there will be heated debates about ethics and morals. But it won’t be long before we are using devices similar to the “Star Trek” tricorder and synthesizing our medications.
Thursday, July 28, 2011
5 real-world uses of big data
5 real-world uses of big data
By David Smith Jul. 17, 2011,
http://gigaom.com/cloud/5-real-world-uses-of-big-data/
In the past year, big data has emerged as one of the most closely watched trends in IT. Organizations today are generating more data in a single day than that the entire Internet was generated as recently as 2000. The explosion of “big data”–much of it in complex and unstructured formats–has presented companies with a tremendous opportunity to leverage their data for better business insights through analytics.
Wal-Mart was one of the early pioneers in this field, using predictive analytics to better identify customer preferences on a regional basis and stock their branch locations accordingly. It was an incredibly effective tactic that yielded strong ROI and allowed them to separate themselves from the retail pack. Other industries took notice of Wal-Mart’s tactics — and the success they gleaned from processing and analyzing their data — and began to employ the same tactics.
While data analytics was once considered a competitive advantage, it’s increasingly being seen as a necessity for enterprises–to the point that those that aren’t employing some kind of analytics are seen to be at a competitive disadvantage. Driven by the rise of modern statistical languages like R, there’s been a surge in enterprises hiring data analysts–which has in turn given rise to the larger data science movement. Data is a huge asset for enterprises, and they’re beginning to treat it accordingly.
For all the talk about the need to effectively analyze your data, though, there’s been relatively little written about how organizations are using data to achieve actionable results. With that in mind, here are five use cases involving analyses of large data sets that brought about valuable new insight:
· NYU Ph.D. student conducts comprehensive analysis of Wikileaks data for greater insight into the Afghanistan conflict: Drew Conway is a Ph.D. student at New York University who also runs the popular, data-centric Zero Intelligence Agents blog. Last year, he analyzed several terabytes worth of Wikileaks data to determine key trends around U.S. and coalition troop activity in Afghanistan. Conway used the R statistics language first to sort the overall flow of information in the five Afghanistan regions, categorized by type of activity (enemy, neutral, ally), and then to identify key patterns from the data. His findings gave credence to a number of popular theories on troop activity there–that there were seasonal spikes in conflict with the Taliban and most coalition activity stemmed from the “Ring Road” that surrounds the capitol, Kabul, to name a few. Through this work, Conway helped the public glean additional insight into the state of affairs for American troops in Afghanistan and the high degree of combat they experienced there.
· International non-profit organization uses data science to confirm Guatemalan genocide: Benetech is a non-profit organization that has been contracted by the likes of Amnesty International and Human Rights Watch to address controversial geopolitical issues through data science. Several years ago, they were contracted to analyze a massive trove of secret files from Guatemala’s National Police that were discovered in an abandoned munitions depot. The documents, of which there were over 80 million, detailed state-sanctioned arrests and disappearances that occurred during the country’s decades-long civil conflict that occurred between 1960 and 1996. There had long been whispers of a genocide against the country’s Mayan population during that period, but no hard evidence had previously emerged to verify these claims. Benetech’s scientists set up a random sample of the data to analyze its content for details on missing victims from the decades-long conflict. After exhaustive analysis, Benetech was able to come to the grim conclusion that genocide had in fact occurred in Guatemala. In the process, they were able to give closure to grieving relatives that had wondered about the fate of their loved ones for decades.
· Statistician develops innovative metrics tracking for baseball players, gains widespread recognition and a job with the Boston Red Sox: Bill James (he of Moneyball fame) is a well-known figure in the world of both baseball and statistics at this point, but that has not always been the case. James, a classically trained statistician and avid baseball fan, began publishing research in the early 1970s that took a more quantitative approach to analyzing the performance of baseball players. His work focused on providing specific metrics that could empirically support or refute specific claims about players, be it the amount of runs they contributed to in a given season or how their defensive abilities contributed to or detracted from a team’s success. James’ approach became known as sabermetrics and has since expanded to incorporate a wide range of quantitative analyses for measuring baseball metrics. Over time, sabermetrics has gained wide recognition in baseball to the point that it’s now employed by all 30 Major League Baseball teams for tracking player metrics. In 2003, James was named Senior Advisor of Baseball Operations by the Boston Red Sox, a position he holds to this day.
· U.S. government uses R to coordinate disaster response to BP oil spill: In the early days of last year’s Deepwater Horizon disaster, the flow of oil rate from the spill was of primary concern; estimating it accurately was key to coordinating the scale and scope of the U.S. government’s response to the emergency. The National Institute of Science and Technology (NIST) was charged with making sense of the varying estimates that existed from both BP and independent third-parties. To do so, NIST used the open source R language to run an uncertainty analysis that harmonized the estimates from various sources to come up with actionable intelligence around which disaster response efforts could be coordinated.
· Medical diagnostics company analyzes millions of lines of data to develop first non-intrusive test for predicting coronary artery disease: CardioDX is a relatively small, Palo Alto, Calif.-based company that performs genomic research. One of their major initiatives over the past several years was developing a predictive test that could identify coronary artery disease in its most nascent stages. To do so, researchers at the company analyzed over 100 million gene samples to ultimately identify the 23 primary predictive genes for coronary artery disease. The resulting test, known as the “Corus CAD Test,” was recognized as on of the “Top Ten Medical Breakthroughs of 2010” by TIME Magazine.
These are but a few brief examples of the exciting work that’s being undertaken in the rapidly growing discipline of data science. More and more, data analysis is being relied on to provide context for critical business decisions, a trend that promises to increase as data sets grow larger and more complex and scientists continue to push the limits of statistical innovation.
David Smith is vice president of community at Revolution Analytics, a company founded in 2007 to foster R analytics by creating programs to make it easier for data scientists to analyze large amounts of data.
By David Smith Jul. 17, 2011,
http://gigaom.com/cloud/5-real-world-uses-of-big-data/
In the past year, big data has emerged as one of the most closely watched trends in IT. Organizations today are generating more data in a single day than that the entire Internet was generated as recently as 2000. The explosion of “big data”–much of it in complex and unstructured formats–has presented companies with a tremendous opportunity to leverage their data for better business insights through analytics.
Wal-Mart was one of the early pioneers in this field, using predictive analytics to better identify customer preferences on a regional basis and stock their branch locations accordingly. It was an incredibly effective tactic that yielded strong ROI and allowed them to separate themselves from the retail pack. Other industries took notice of Wal-Mart’s tactics — and the success they gleaned from processing and analyzing their data — and began to employ the same tactics.
While data analytics was once considered a competitive advantage, it’s increasingly being seen as a necessity for enterprises–to the point that those that aren’t employing some kind of analytics are seen to be at a competitive disadvantage. Driven by the rise of modern statistical languages like R, there’s been a surge in enterprises hiring data analysts–which has in turn given rise to the larger data science movement. Data is a huge asset for enterprises, and they’re beginning to treat it accordingly.
For all the talk about the need to effectively analyze your data, though, there’s been relatively little written about how organizations are using data to achieve actionable results. With that in mind, here are five use cases involving analyses of large data sets that brought about valuable new insight:
· NYU Ph.D. student conducts comprehensive analysis of Wikileaks data for greater insight into the Afghanistan conflict: Drew Conway is a Ph.D. student at New York University who also runs the popular, data-centric Zero Intelligence Agents blog. Last year, he analyzed several terabytes worth of Wikileaks data to determine key trends around U.S. and coalition troop activity in Afghanistan. Conway used the R statistics language first to sort the overall flow of information in the five Afghanistan regions, categorized by type of activity (enemy, neutral, ally), and then to identify key patterns from the data. His findings gave credence to a number of popular theories on troop activity there–that there were seasonal spikes in conflict with the Taliban and most coalition activity stemmed from the “Ring Road” that surrounds the capitol, Kabul, to name a few. Through this work, Conway helped the public glean additional insight into the state of affairs for American troops in Afghanistan and the high degree of combat they experienced there.
· International non-profit organization uses data science to confirm Guatemalan genocide: Benetech is a non-profit organization that has been contracted by the likes of Amnesty International and Human Rights Watch to address controversial geopolitical issues through data science. Several years ago, they were contracted to analyze a massive trove of secret files from Guatemala’s National Police that were discovered in an abandoned munitions depot. The documents, of which there were over 80 million, detailed state-sanctioned arrests and disappearances that occurred during the country’s decades-long civil conflict that occurred between 1960 and 1996. There had long been whispers of a genocide against the country’s Mayan population during that period, but no hard evidence had previously emerged to verify these claims. Benetech’s scientists set up a random sample of the data to analyze its content for details on missing victims from the decades-long conflict. After exhaustive analysis, Benetech was able to come to the grim conclusion that genocide had in fact occurred in Guatemala. In the process, they were able to give closure to grieving relatives that had wondered about the fate of their loved ones for decades.
· Statistician develops innovative metrics tracking for baseball players, gains widespread recognition and a job with the Boston Red Sox: Bill James (he of Moneyball fame) is a well-known figure in the world of both baseball and statistics at this point, but that has not always been the case. James, a classically trained statistician and avid baseball fan, began publishing research in the early 1970s that took a more quantitative approach to analyzing the performance of baseball players. His work focused on providing specific metrics that could empirically support or refute specific claims about players, be it the amount of runs they contributed to in a given season or how their defensive abilities contributed to or detracted from a team’s success. James’ approach became known as sabermetrics and has since expanded to incorporate a wide range of quantitative analyses for measuring baseball metrics. Over time, sabermetrics has gained wide recognition in baseball to the point that it’s now employed by all 30 Major League Baseball teams for tracking player metrics. In 2003, James was named Senior Advisor of Baseball Operations by the Boston Red Sox, a position he holds to this day.
· U.S. government uses R to coordinate disaster response to BP oil spill: In the early days of last year’s Deepwater Horizon disaster, the flow of oil rate from the spill was of primary concern; estimating it accurately was key to coordinating the scale and scope of the U.S. government’s response to the emergency. The National Institute of Science and Technology (NIST) was charged with making sense of the varying estimates that existed from both BP and independent third-parties. To do so, NIST used the open source R language to run an uncertainty analysis that harmonized the estimates from various sources to come up with actionable intelligence around which disaster response efforts could be coordinated.
· Medical diagnostics company analyzes millions of lines of data to develop first non-intrusive test for predicting coronary artery disease: CardioDX is a relatively small, Palo Alto, Calif.-based company that performs genomic research. One of their major initiatives over the past several years was developing a predictive test that could identify coronary artery disease in its most nascent stages. To do so, researchers at the company analyzed over 100 million gene samples to ultimately identify the 23 primary predictive genes for coronary artery disease. The resulting test, known as the “Corus CAD Test,” was recognized as on of the “Top Ten Medical Breakthroughs of 2010” by TIME Magazine.
These are but a few brief examples of the exciting work that’s being undertaken in the rapidly growing discipline of data science. More and more, data analysis is being relied on to provide context for critical business decisions, a trend that promises to increase as data sets grow larger and more complex and scientists continue to push the limits of statistical innovation.
David Smith is vice president of community at Revolution Analytics, a company founded in 2007 to foster R analytics by creating programs to make it easier for data scientists to analyze large amounts of data.
Tuesday, July 19, 2011
PM-ISE Releases the 2011 ISE Annual Report to the Congress
The PM-ISE has officially released its 2011 ISE Annual Report to the Congress and we are proud of the information sharing success stories featured in the Report – stories that describe the outstanding accomplishments of our mission partners across the federal, state, local, and tribal governments, the private sector, and foreign allies.
The Annual Report is required by law to provide the Congress “a progress report on the extent to which the ISE has been implemented.”[1] The Report highlights major ISE activities since July 2010 and is organized around five themes:
Strengthening Management and Oversight - The Annual Report describes the work of the Information Sharing and Access Interagency Policy Committee (ISA IPC) and its sub-committees and working groups; of particular note, the Report highlights how these bodies expanded to include representatives of non-federal organizations and are reaching out to engage the private sector in developing the ISE, as well.
Improving Information Sharing Activities - Among the many activities presented, the Report describes how the Nationwide Suspicious Activity Reporting Initiative has made substantial progress toward streamlining reporting and analysis within fusion centers by implementing new standards, policies, and processes. Another notable interagency effort involved the Baseline Capabilities Assessment, during which federal, state, and local officials completed the first nationwide, in-depth assessment of fusion centers to baseline their capabilities.
Establishing Standards for Responsible Information Sharing and Protection - Standards are critical to powering the ISE, and so the Report describes the efforts by the PM-ISE, its mission partners, and standards organizations to identify the best existing standards for reuse and implementation across the ISE.
Enabling Assured Interoperability across Networks - The Report details the tremendous progress made toward implementing a Simplified Sign On that will enable federal, state, local, and tribal law enforcement officers and analysts to more easily access a rich variety of data services provided by Assured Sensitive but Unclassified (SBU) networks. The Report also describes similar efforts for classified information sharing.
Enhancing Privacy, Civil Rights, and Civil Liberties Protections - Balancing the need for national security with the need to protect privacy and civil liberties, the Report provides information on policies and training activities designed to enhance these protections.
These are only a few of the activities that are helping the nation build a robust information sharing environment. And, while the Annual Report is primarily focused on terrorism-related initiatives, it also describes mission partner accomplishments that may not have been developed explicitly to support CT, but which may ultimately become best practices for information sharing and collaboration government-wide.
The Annual Report is required by law to provide the Congress “a progress report on the extent to which the ISE has been implemented.”[1] The Report highlights major ISE activities since July 2010 and is organized around five themes:
Strengthening Management and Oversight - The Annual Report describes the work of the Information Sharing and Access Interagency Policy Committee (ISA IPC) and its sub-committees and working groups; of particular note, the Report highlights how these bodies expanded to include representatives of non-federal organizations and are reaching out to engage the private sector in developing the ISE, as well.
Improving Information Sharing Activities - Among the many activities presented, the Report describes how the Nationwide Suspicious Activity Reporting Initiative has made substantial progress toward streamlining reporting and analysis within fusion centers by implementing new standards, policies, and processes. Another notable interagency effort involved the Baseline Capabilities Assessment, during which federal, state, and local officials completed the first nationwide, in-depth assessment of fusion centers to baseline their capabilities.
Establishing Standards for Responsible Information Sharing and Protection - Standards are critical to powering the ISE, and so the Report describes the efforts by the PM-ISE, its mission partners, and standards organizations to identify the best existing standards for reuse and implementation across the ISE.
Enabling Assured Interoperability across Networks - The Report details the tremendous progress made toward implementing a Simplified Sign On that will enable federal, state, local, and tribal law enforcement officers and analysts to more easily access a rich variety of data services provided by Assured Sensitive but Unclassified (SBU) networks. The Report also describes similar efforts for classified information sharing.
Enhancing Privacy, Civil Rights, and Civil Liberties Protections - Balancing the need for national security with the need to protect privacy and civil liberties, the Report provides information on policies and training activities designed to enhance these protections.
These are only a few of the activities that are helping the nation build a robust information sharing environment. And, while the Annual Report is primarily focused on terrorism-related initiatives, it also describes mission partner accomplishments that may not have been developed explicitly to support CT, but which may ultimately become best practices for information sharing and collaboration government-wide.
FT: A brave new networked world
July 18, 2011 11:15 pm
A brave new networked world
By Philip Delves Broughton
http://www.ft.com/intl/cms/s/0/2fe38490-b176-11e0-9444-00144feab49a.html#axzz1SSrsjjWo
We are under siege by social networks, from Facebook to LinkedIn to school, university and even corporate alumni organisations, which are ever more aggressive and sophisticated in their networking efforts.
Facebook has more than 750m users, LinkedIn 100m, and Twitter is handling 1bn tweets a week. Technology-driven social networks have been credited with propelling the revolutionaries of the Arab spring.
Marketers and venture investors salivate over any company promising to identify and assemble networks of like-minded consumers. Social network analysis software has become a fast-growing sector of IT services. IBM alone has spent $11bn in the past five years buying makers of such software.
A brave new networked world
By Philip Delves Broughton
http://www.ft.com/intl/cms/s/0/2fe38490-b176-11e0-9444-00144feab49a.html#axzz1SSrsjjWo
We are under siege by social networks, from Facebook to LinkedIn to school, university and even corporate alumni organisations, which are ever more aggressive and sophisticated in their networking efforts.
Facebook has more than 750m users, LinkedIn 100m, and Twitter is handling 1bn tweets a week. Technology-driven social networks have been credited with propelling the revolutionaries of the Arab spring.
Marketers and venture investors salivate over any company promising to identify and assemble networks of like-minded consumers. Social network analysis software has become a fast-growing sector of IT services. IBM alone has spent $11bn in the past five years buying makers of such software.
But there are sceptics questioning the power of these networks. Malcolm Gladwell wrote in The New Yorker last year that the impact of new forms of communication in fomenting political change was exaggerated. He distinguished between the strong ties that bind groups of revolutionaries and the weak ties that link 1m Facebook fans.
“The platforms of social media are built around weak ties,” he wrote. “Twitter is a way of following (or being followed by) people you may never have met. Facebook is a tool for efficiently managing your acquaintances, for keeping up with the people you would not otherwise be able to stay in touch with. That’s why you can have a thousand ‘friends’ on Facebook, as you never could in real life.”
While the significance of social networks for political activists may be open to question, companies are starting to find real value in mapping and analysing the strength and frequency of connections between employees and customers, and their behaviour.
Social network analysis is being used to measure job performance and forecast turnover, to rate employees for promotion, monitor their ethical standards and improve the systems for collaboration. It is now a standard diagnostic and prescriptive tool for management consultants advising companies.
By measuring who talks to whom, when and for how long, companies can uncover hidden stars. They can develop an index that reveals where real power lies – and it may not be with the people who hold the grandest titles. When the US was tracking Saddam Hussein it mapped the networks of his former chauffeurs, which led to his hideout. The Iraqi elite had no idea where he was.
Businesses also use social network analysis to help categorise customers. Nike tracks bloggers to find out whose posts most influence sales of their shoes. Telecoms companies seek out the “influencers”, those thrifty types who shop around for the best plan and take their Facebook friends and Twitter followers with them when they switch.
Rob Cross is a professor of management at the University of Virginia’s McIntire School of Commerce and a consultant on corporate social networks to companies ranging from American Express to Intel. Through web-based surveys, he maps “who creates enthusiasm within organisations and who drains it. This turns out to be wildly predictive of success or failure in innovation”.
Prof Cross also finds that social networks can create another level of work for employees. “Due to the recession, people’s workloads and spans of accountability have increased, the tools of collaboration have multiplied, and we’re seeing people can’t keep up with collaborative demands any more,” he says.
Social media tools have only made this worse. Much of his work is therefore about reducing unnecessary collaboration. Organisations can benefit from protecting key people, such as research scientists or top salespeople, from the deluge of fruitless networking and information sharing.
Social network sites, Prof Cross says, are good for creating interest groups or sharing tactical information, but in situations where you need to solve a problem it is necessary to resort to old-fashioned means of developing trust and sharing information.
At agribusiness Monsanto, executives spread across the world who were forced to implement a new global transaction system were much more productive, more quickly, when they had previous experience of collaborating.
Another useful aspect of social network sites is in reactivating dormant ties. In research published in MIT Sloan Management Review, Daniel Levin, Jorge Walter and Keith Murnighan have shown that we underestimate the power of our dormant relationships.
People we once knew well but have not seen for a while turn out to be delighted to hear from us. They also have novel insights and are happy to help. Social network sites make these dormant relationships easier to rediscover and resume, but they are only a start.
“People who haven’t seen each other in a while are delighted to e-mail, until something goes beyond expectation in a negative way,” says Prof Murnighan. “Then they might realise why the relationship was dormant.”
For substantive interactions, you still need to get on the phone or meet in person. But the lesson of his research into dormant ties, he says, is that “people move on with their lives, but they don’t forget each other. An awful lot of substance remains”.
No amount of e-mailing or Facebook poking can substitute for extended time spent with others.
Lauren Cohen and Christopher Malloy of Harvard Business School have written a series of papers on the continuing importance of old-fashioned ties in investment performance. They found mutual fund managers and sell-side analysts made much better investment returns and recommendations in companies where they had strong university alumni connections.
In the US, they found that before new regulations came into effect in 2000 to limit selective disclosure of corporate information, the return premium from the old school tie was 8.16 per cent a year. Since 2000, it has fallen to zero. But in the UK, where the rules about selective disclosure are less rigorous, the old school tie premium persists.
Prof Cohen says that the premium is explained by the many clubs and networking events laid on by universities and the informal networks and gossip shared among old friends.
People who went to the same university, she says, are “more likely to have met each other or have common acquaintances. They understand what it means if a person belonged to a certain club or participated in a specific study programme. They may know people who hired them previously. All this helps them better assess executives’ potential as leaders and business owners”.
Technology may have made everyone accessible, but it has not yet made all social networks equal.
Harnessing networks
Not all networks are equal but all have their uses. So it is important to identify the differing strength of connections within a network in order to determine how best to use it.
Weak ties Good for forming groups around shared interests, hobbies and projects as well as tactical information sharing. Twitter hashtags and Facebook groups can cluster people with similar interests, but not for long unless strong ties can be developed.
Strong ties Vital for trust-based activities and complex collaboration among remote groups. Monsanto found that executives who had worked together in the past, even though now dispersed around the world, were much more effective at complex, collaborative tasks.
Dormant ties People whom we once knew well are happy to be “reactivated”. They can be great sources of new information and perspectives.
Old school ties Alumni networks still provide access to information which can affect hiring and improve investor returns.
Subscribe to:
Posts (Atom)