In an Open-Source Society, Innovating by the Seat of Our Pants
Joici Ito The New York Times December 5, 2011
The Internet isn’t really a technology. It’s a belief system, a philosophy about the effectiveness of decentralized, bottom-up innovation. And it’s a philosophy that has begun to change how we think about creativity itself.
Almost 20 years ago, I installed on my computer a tiny piece of software called MacPPP, which connected the programs running on it to the Internet. The program immediately transformed my computer from a fancy telex machine to a device running a very early version of the graphical Web.
I was working in entertainment at the time, and I remember thinking that this connection was going to change everything. I left to join the first commercial Internet service provider in Japan, PSINet Japan, as its first chief executive. Our first serious challenge, oddly enough, was a battle over an obscure information-sharing computer protocol called X.25. Most of us laboring to build the new Internet preferred the less regulated and simpler Internet Protocol.
Until then, large intergovernmental agencies had always gathered experts to work on the technical standards that would become the DNA of the telecommunications industry, the standards to which all companies would have to build their networks and products. These researchers had produced X.25, a complex and extremely well-considered standard that seemed to anticipate every possible problem and application.
The Internet, on the other hand, was designed and deployed by small groups of researchers following the credo of one of its chief architects, David Clark: “rough consensus and running code.” Its early standards — uncomplicated, consensual — were stewarded by small organizations that resisted permission or authority. And they won: The Internet Protocol on which every connected device relies was a triumph of distributed innovation over centralized expertise.
The ethos of the Internet is that everyone should have the freedom to connect, to innovate, to program, without asking permission. No one can know the whole of the network, and by design it cannot be centrally controlled. This network was intended to be decentralized, its assets widely distributed. Today most innovation springs from small groups at its “edges.”
This technical strategy has led to the creation of a gigantic network of far-flung innovators who develop standards with one another and share the products of their work in the form of free and open-source software. The architecture of the Internet and its abundance of free software and components has driven down the cost of manufacturing, distribution and collaboration — of innovation. It used to cost millions of dollars to start a software company. Today, for little or no money, entrepreneurs are able to develop and release a “minimum viable product” and test it with real users on the Internet before they have to raise any money from investors. In their earliest iterations, Facebook, Yahoo and Google were running in dorm rooms and labs before the founders had left college or had raised outside money.
In fact, it is now usually cheaper to just try something than to sit around and try to figure out whether to try something. The product map is now often more complex and more expensive to create than trying to figure it out as you go. The compass has replaced the map, and “rough consensus and running code” has become the fundamental philosophy for the so-called lean start-up movement.
Innovators are able to prototype a new product with 3-D printers and cheap laser cutters for nearly nothing. Even complex products can be manufactured with help from supply chain companies that are making their systems available online to anybody. Today we are seeing the emergence of a community of hardware hackers and designers very reminiscent of the developers who wrote the original open standards of the Internet. An explosion of grass-roots innovation in hardware is coming — freely designed and freely shared — as it did in software.
What has been a wildly successful model for consumer Internet start-ups in Silicon Valley turns out to be an extremely good model for learning in a wide variety of fields and disciplines. The students at M.I.T.’s Media Lab experiment, create and iterate; they produce demos and prototypes, and share and collaborate with the rest of the world through the Internet and a distributed network of connections and relationships.
I don’t think education is about centralized instruction anymore; rather, it is the process establishing oneself as a node in a broad network of distributed creativity.
Neoteny, one of my favorite words, means the retention of childlike attributes in adulthood: idealism, experimentation and wonder. In this new world, not only must we behave more like children, we also must teach the next generation to retain those attributes that will allow them to be world-changing, innovative adults who will help us reinvent the future.
Joichi Ito is the director of the M.I.T. Media Lab.
Tuesday, December 6, 2011
Larry Smarr in NYTimes: An Evolution Toward a Programmable Universe
Essay Larry Smarr The New York Times December 5, 2011
Over the next 10 years, the physical world will become ever more overlaid with devices for sending and receiving information.
Already billions of processors are embedded in our smartphones, cars, appliances and buildings and the environment. These sensors can send out streams of data about their surroundings, and more and more it is anonymously transmitted to remote data centers — the “clouds” of Google, Amazon, Microsoft, Yahoo and Apple.
From these vast clouds, the companies can power apps that are “spatially aware.” For instance, Google Maps now draws on data in the cloud to sample the location and movement of cellphones in cars, producing a real-time picture of traffic congestion.
Smart electric grids are measuring our homes’ use of power; active people are tracking their heart rates; and hundreds of millions of us are uploading geo-tagged data to Flickr, Yelp, Facebook and Google Plus. As we look 10 years ahead, the fastest supercomputer (the “exascale” machine) will be composed of one billion processors, and the clouds will most likely grow to this scale as well, creating a distributed planetary computer of enormous power.
Such computational power, co-located with the gigantic storage that holds the data from all the incoming data streams, will enable faster-than-real-time simulations of many aspects of our physical world. As Mike Liebhold and his colleagues at the Institute for the Future have discussed, computing will have evolved from merely sensing local information to analyzing it to being able to control it. In this evolution, the world gradually becomes programmable.
At the California Institute for Telecommunications and Information Technology, we are using this vision to better understand the coming digital transformation of health, energy, environment and culture. We are experimenting with sensors to monitor electricity use in homes, buildings and data centers; the data can then be analyzed and used to control lighting, heating, cooling, appliances and computers to make them more energy-efficient.
It is logical that the analysis of traffic data, coupled with in-car radar and autopilot electronics, will enable software control of large numbers of robot-driven electric cars. Since buildings and transportation are major sources of greenhouse gas emissions, the sensor-aware planetary computer can be a crucial factor in reducing our carbon footprints.
The same principle applies to our bodies. I wear sensors to measure my steps, caloric burn and sleep patterns, while heart patients can wear sensors that wirelessly notify their doctors of life-threatening conditions. People will soon be able to have their genetic code and medical imaging stored in the cloud, along with charts of vital signs and detailed nutritional analysis of everything they consume.
Using this data, the planetary computer will be able to build a computational model of your body and compare your sensor stream with millions of others. Besides providing early detection of internal changes that could lead to disease, cloud-powered voice-recognition wellness coaches could provide continual personalized support on lifestyle choices, potentially staving off disease and making health care affordable for everyone.
Finally, in culture, the fine-grain streaming provided by Twitter, Facebook and Google Plus enables us to map out phenomena using “human sensors.”
For instance, a vast power failure occurred in Southern California in September; within minutes we could tell from the locations of Twitter messages saying “my power just went out” that it was widespread, long before the official announcement. Similarly, Twitter feeds from large geographic areas have been analyzed to create dynamic “social mood” or “political anger” maps, like the Google traffic maps constructed from GPS feeds.
Conceivably, the coupling of the sensor and human streams with planetary computing power will make it possible to create “social forecasts.” For good or evil, it seems inevitable that individuals, corporations, political leaders and intelligence agencies will come to use planetary computer models of social behavior to inject content into the global attention stream at just the right moment, hoping to steer the social dynamics to a desired outcome.
With the continuing exponential increase in the power of the planetary computer, one has to wonder whether we stand at the beginning of what Isaac Asimov’s “Foundation” series, more than 60 years ago, called “psychohistory.” His visionary genius Hari Seldon believed that statistical forecasting of human society’s actions would be possible with data from enough people throughout the galaxy.
In the next several decades, we will have a glimpse of whether something similar can emerge on planet Earth.
Larry Smarr is the founding director of Calit2.
Over the next 10 years, the physical world will become ever more overlaid with devices for sending and receiving information.
Already billions of processors are embedded in our smartphones, cars, appliances and buildings and the environment. These sensors can send out streams of data about their surroundings, and more and more it is anonymously transmitted to remote data centers — the “clouds” of Google, Amazon, Microsoft, Yahoo and Apple.
From these vast clouds, the companies can power apps that are “spatially aware.” For instance, Google Maps now draws on data in the cloud to sample the location and movement of cellphones in cars, producing a real-time picture of traffic congestion.
Smart electric grids are measuring our homes’ use of power; active people are tracking their heart rates; and hundreds of millions of us are uploading geo-tagged data to Flickr, Yelp, Facebook and Google Plus. As we look 10 years ahead, the fastest supercomputer (the “exascale” machine) will be composed of one billion processors, and the clouds will most likely grow to this scale as well, creating a distributed planetary computer of enormous power.
Such computational power, co-located with the gigantic storage that holds the data from all the incoming data streams, will enable faster-than-real-time simulations of many aspects of our physical world. As Mike Liebhold and his colleagues at the Institute for the Future have discussed, computing will have evolved from merely sensing local information to analyzing it to being able to control it. In this evolution, the world gradually becomes programmable.
At the California Institute for Telecommunications and Information Technology, we are using this vision to better understand the coming digital transformation of health, energy, environment and culture. We are experimenting with sensors to monitor electricity use in homes, buildings and data centers; the data can then be analyzed and used to control lighting, heating, cooling, appliances and computers to make them more energy-efficient.
It is logical that the analysis of traffic data, coupled with in-car radar and autopilot electronics, will enable software control of large numbers of robot-driven electric cars. Since buildings and transportation are major sources of greenhouse gas emissions, the sensor-aware planetary computer can be a crucial factor in reducing our carbon footprints.
The same principle applies to our bodies. I wear sensors to measure my steps, caloric burn and sleep patterns, while heart patients can wear sensors that wirelessly notify their doctors of life-threatening conditions. People will soon be able to have their genetic code and medical imaging stored in the cloud, along with charts of vital signs and detailed nutritional analysis of everything they consume.
Using this data, the planetary computer will be able to build a computational model of your body and compare your sensor stream with millions of others. Besides providing early detection of internal changes that could lead to disease, cloud-powered voice-recognition wellness coaches could provide continual personalized support on lifestyle choices, potentially staving off disease and making health care affordable for everyone.
Finally, in culture, the fine-grain streaming provided by Twitter, Facebook and Google Plus enables us to map out phenomena using “human sensors.”
For instance, a vast power failure occurred in Southern California in September; within minutes we could tell from the locations of Twitter messages saying “my power just went out” that it was widespread, long before the official announcement. Similarly, Twitter feeds from large geographic areas have been analyzed to create dynamic “social mood” or “political anger” maps, like the Google traffic maps constructed from GPS feeds.
Conceivably, the coupling of the sensor and human streams with planetary computing power will make it possible to create “social forecasts.” For good or evil, it seems inevitable that individuals, corporations, political leaders and intelligence agencies will come to use planetary computer models of social behavior to inject content into the global attention stream at just the right moment, hoping to steer the social dynamics to a desired outcome.
With the continuing exponential increase in the power of the planetary computer, one has to wonder whether we stand at the beginning of what Isaac Asimov’s “Foundation” series, more than 60 years ago, called “psychohistory.” His visionary genius Hari Seldon believed that statistical forecasting of human society’s actions would be possible with data from enough people throughout the galaxy.
In the next several decades, we will have a glimpse of whether something similar can emerge on planet Earth.
Larry Smarr is the founding director of Calit2.
Thursday, December 1, 2011
Big Data: DNA Sequencing Caught in Deluge of Data
Andrew Pollack The New York Times November 30, 2011
BGI, based in China, is the world’s largest genomics research institute, with 167 DNA sequencers producing the equivalent of 2,000 human genomes a day.
BGI churns out so much data that it often cannot transmit its results to clients or collaborators over the Internet or other communications lines because that would take weeks. Instead, it sends computer disks containing the data, via FedEx.
“It sounds like an analog solution in a digital age,” conceded Sifei He, the head of cloud computing for BGI, formerly known as the Beijing Genomics Institute. But for now, he said, there is no better way.
The field of genomics is caught in a data deluge. DNA sequencing is becoming faster and cheaper at a pace far outstripping Moore’s law, which describes the rate at which computing gets faster and cheaper.
The result is that the ability to determine DNA sequences is starting to outrun the ability of researchers to store, transmit and especially to analyze the data.
“Data handling is now the bottleneck,” said David Haussler, director of the center for biomolecular science and engineering at the University of California, Santa Cruz. “It costs more to analyze a genome than to sequence a genome.”
That could delay the day when DNA sequencing is routinely used in medicine. In only a year or two, the cost of determining a person’s complete DNA blueprint is expected to fall below $1,000. But that long-awaited threshold excludes the cost of making sense of that data, which is becoming a bigger part of the total cost as sequencing costs themselves decline.
“The real cost in the sequencing is more than just running the sequencing machine,” said Mark Gerstein, professor of biomedical informatics at Yale. “And now that is becoming more apparent.”
But the data challenges are also creating opportunities. There is demand for people trained in bioinformatics, the convergence of biology and computing. Numerous bioinformatics companies, like SoftGenetics, DNAStar, DNAnexus and NextBio, have sprung up to offer software and services to help analyze the data. EMC, a maker of data storage equipment, has found life sciences a fertile market for products that handle large amounts of information. BGI is starting a journal, GigaScience, to publish data-heavy life science papers.
“We believe the field of bioinformatics for genetic analysis will be one of the biggest areas of disruptive innovation in life science tools over the next few years,” Isaac Ro, an analyst at Goldman Sachs, wrote in a recent report.
Sequencing involves determining the order of the bases, the chemical units represented by the letters A, C, G and T, in a stretch of DNA. The cost has plummeted, particularly in the last four years, as new techniques have been introduced.
The cost of sequencing a human genome — all three billion bases of DNA in a set of human chromosomes — plunged to $10,500 last July from $8.9 million in July 2007, according to the National Human Genome Research Institute.
That is a decline by a factor of more than 800 over four years. By contrast, computing costs would have dropped by perhaps a factor of four in that time span.
The lower cost, along with increasing speed, has led to a huge increase in how much sequencing data is being produced. World capacity is now 13 quadrillion DNA bases a year, an amount that would fill a stack of DVDs two miles high, according to Michael Schatz, assistant professor of quantitative biology at the Cold Spring Harbor Laboratory on Long Island.
There will probably be 30,000 human genomes sequenced by the end of this year, up from a handful a few years ago, according to the journal Nature. And that number will rise to millions in a few years.
In a few cases, human genomes are being sequenced to help diagnose mysterious rare diseases and treat patients. But most are being sequenced as part of studies. The federally financed Cancer Genome Atlas, for instance, is sequencing the genomes of thousands of tumors and of healthy tissue from the same people, looking for genetic causes of cancer.
One near victim of the data explosion has been a federal online archive of raw sequencing data. The amount stored has more than tripled just since the beginning of the year, reaching 300 trillion DNA bases and taking up nearly 700 trillion bytes of computer memory.
Straining under the load and facing budget constraints, federal officials talked earlier this year about shutting the archive, to the dismay of researchers. It will remain open, but certain big sequencing projects will now have to pay to store their data there.
If the problem is tough for human genomes, it is far worse for the field known as metagenomics. This involves sequencing the DNA found in a particular environment, like a sample of soil or the human gut. The idea is to take a census of what microbial species are present.
E. Virginia Armbrust, who studies ocean-dwelling microscopic organisms at the University of Washington, said her lab generated 60 billion bases — as much as 20 human genomes — from just two surface water samples. It took weeks to do the sequencing, but nearly two years to then analyze the data, she said.
“There is more data that is infiltrating lots of different fields that weren’t particularly ready for that,” Professor Armbrust said. “It’s all a little overwhelming.”
The Human Microbiome Project, which is sequencing the microbial populations in the human digestive tract, has generated about a million times as much sequence data as a single human genome, said C. Titus Brown, a bioinformatics specialist at Michigan State University.
“It’s not at all clear what you do with that data,” he said. “Doing a comprehensive analysis of it is essentially impossible at the moment.”
Other scientific fields, like particle physics and astronomy, handle huge amounts of data. In those fields, however, much of the data is generated by a few huge accelerators or observatories, said Eugene Kolker, chief data officer at Seattle Children’s Hospital.
“In the life sciences, anyone can produce so much data, and it’s happening in thousands of different labs throughout the world,” he said.
Moreover, DNA is just part of the story. To truly understand biology, researchers are gathering data on the RNA, proteins and chemicals in cells. That data can be even more voluminous than data on genes. And those different types of data have to be integrated.
“We have these giant piles of data and no way to connect them” said H. Steven Wiley, a biologist at the Pacific Northwest National Laboratory. He added, “I’m sitting in front of a pile of data that we’ve been trying to analyze for the last year and a half.”
Still, many say the situation will be manageable. Jay Flatley, chief executive of Illumina, the leading supplier of sequencing machines, said he did not think information handling was a bottleneck or that it was causing people to hold off on buying new sequencers.
Researchers are increasingly turning to cloud computing so they do not have to buy so many of their own computers and disk drives.
Google might help as well.
“Google has enough capacity to do all of genomics in a day,” said Dr. Schatz of Cold Spring Harbor, who is trying to apply Google’s techniques to genomics data. Prodded by Senator Charles E. Schumer, Democrat of New York, Google is exploring cooperation with Cold Spring Harbor.
Google’s venture capital arm recently invested in DNAnexus, a bioinformatics company. DNAnexus and Google plan to host their own copy of the federal sequence archive that had once looked as if it might be closed.
The amount of data stored for a human genome will drop sharply. Sequencers produce huge amounts of raw data that then has to be analyzed and processed by software to produce the result.
With the field still young, many researchers store all the raw data, so it can be re-analyzed if better software is developed in the future.
In uncertain times, “scientists cling to their data,” said David J. Dooling, assistant director of the genome institute at Washington University in St. Louis.
But there is now so much raw data that it is becoming not feasible to re-analyze it. So researchers will increasingly store just the final results. In the case of human genomes, they might store even less — only the difference between a particular genome and some reference genome.
Professor Brown of Michigan State said: “We are going to have to come up with really clever ways to throw away data so we can see new stuff.”
BGI, based in China, is the world’s largest genomics research institute, with 167 DNA sequencers producing the equivalent of 2,000 human genomes a day.
BGI churns out so much data that it often cannot transmit its results to clients or collaborators over the Internet or other communications lines because that would take weeks. Instead, it sends computer disks containing the data, via FedEx.
“It sounds like an analog solution in a digital age,” conceded Sifei He, the head of cloud computing for BGI, formerly known as the Beijing Genomics Institute. But for now, he said, there is no better way.
The field of genomics is caught in a data deluge. DNA sequencing is becoming faster and cheaper at a pace far outstripping Moore’s law, which describes the rate at which computing gets faster and cheaper.
The result is that the ability to determine DNA sequences is starting to outrun the ability of researchers to store, transmit and especially to analyze the data.
“Data handling is now the bottleneck,” said David Haussler, director of the center for biomolecular science and engineering at the University of California, Santa Cruz. “It costs more to analyze a genome than to sequence a genome.”
That could delay the day when DNA sequencing is routinely used in medicine. In only a year or two, the cost of determining a person’s complete DNA blueprint is expected to fall below $1,000. But that long-awaited threshold excludes the cost of making sense of that data, which is becoming a bigger part of the total cost as sequencing costs themselves decline.
“The real cost in the sequencing is more than just running the sequencing machine,” said Mark Gerstein, professor of biomedical informatics at Yale. “And now that is becoming more apparent.”
But the data challenges are also creating opportunities. There is demand for people trained in bioinformatics, the convergence of biology and computing. Numerous bioinformatics companies, like SoftGenetics, DNAStar, DNAnexus and NextBio, have sprung up to offer software and services to help analyze the data. EMC, a maker of data storage equipment, has found life sciences a fertile market for products that handle large amounts of information. BGI is starting a journal, GigaScience, to publish data-heavy life science papers.
“We believe the field of bioinformatics for genetic analysis will be one of the biggest areas of disruptive innovation in life science tools over the next few years,” Isaac Ro, an analyst at Goldman Sachs, wrote in a recent report.
Sequencing involves determining the order of the bases, the chemical units represented by the letters A, C, G and T, in a stretch of DNA. The cost has plummeted, particularly in the last four years, as new techniques have been introduced.
The cost of sequencing a human genome — all three billion bases of DNA in a set of human chromosomes — plunged to $10,500 last July from $8.9 million in July 2007, according to the National Human Genome Research Institute.
That is a decline by a factor of more than 800 over four years. By contrast, computing costs would have dropped by perhaps a factor of four in that time span.
The lower cost, along with increasing speed, has led to a huge increase in how much sequencing data is being produced. World capacity is now 13 quadrillion DNA bases a year, an amount that would fill a stack of DVDs two miles high, according to Michael Schatz, assistant professor of quantitative biology at the Cold Spring Harbor Laboratory on Long Island.
There will probably be 30,000 human genomes sequenced by the end of this year, up from a handful a few years ago, according to the journal Nature. And that number will rise to millions in a few years.
In a few cases, human genomes are being sequenced to help diagnose mysterious rare diseases and treat patients. But most are being sequenced as part of studies. The federally financed Cancer Genome Atlas, for instance, is sequencing the genomes of thousands of tumors and of healthy tissue from the same people, looking for genetic causes of cancer.
One near victim of the data explosion has been a federal online archive of raw sequencing data. The amount stored has more than tripled just since the beginning of the year, reaching 300 trillion DNA bases and taking up nearly 700 trillion bytes of computer memory.
Straining under the load and facing budget constraints, federal officials talked earlier this year about shutting the archive, to the dismay of researchers. It will remain open, but certain big sequencing projects will now have to pay to store their data there.
If the problem is tough for human genomes, it is far worse for the field known as metagenomics. This involves sequencing the DNA found in a particular environment, like a sample of soil or the human gut. The idea is to take a census of what microbial species are present.
E. Virginia Armbrust, who studies ocean-dwelling microscopic organisms at the University of Washington, said her lab generated 60 billion bases — as much as 20 human genomes — from just two surface water samples. It took weeks to do the sequencing, but nearly two years to then analyze the data, she said.
“There is more data that is infiltrating lots of different fields that weren’t particularly ready for that,” Professor Armbrust said. “It’s all a little overwhelming.”
The Human Microbiome Project, which is sequencing the microbial populations in the human digestive tract, has generated about a million times as much sequence data as a single human genome, said C. Titus Brown, a bioinformatics specialist at Michigan State University.
“It’s not at all clear what you do with that data,” he said. “Doing a comprehensive analysis of it is essentially impossible at the moment.”
Other scientific fields, like particle physics and astronomy, handle huge amounts of data. In those fields, however, much of the data is generated by a few huge accelerators or observatories, said Eugene Kolker, chief data officer at Seattle Children’s Hospital.
“In the life sciences, anyone can produce so much data, and it’s happening in thousands of different labs throughout the world,” he said.
Moreover, DNA is just part of the story. To truly understand biology, researchers are gathering data on the RNA, proteins and chemicals in cells. That data can be even more voluminous than data on genes. And those different types of data have to be integrated.
“We have these giant piles of data and no way to connect them” said H. Steven Wiley, a biologist at the Pacific Northwest National Laboratory. He added, “I’m sitting in front of a pile of data that we’ve been trying to analyze for the last year and a half.”
Still, many say the situation will be manageable. Jay Flatley, chief executive of Illumina, the leading supplier of sequencing machines, said he did not think information handling was a bottleneck or that it was causing people to hold off on buying new sequencers.
Researchers are increasingly turning to cloud computing so they do not have to buy so many of their own computers and disk drives.
Google might help as well.
“Google has enough capacity to do all of genomics in a day,” said Dr. Schatz of Cold Spring Harbor, who is trying to apply Google’s techniques to genomics data. Prodded by Senator Charles E. Schumer, Democrat of New York, Google is exploring cooperation with Cold Spring Harbor.
Google’s venture capital arm recently invested in DNAnexus, a bioinformatics company. DNAnexus and Google plan to host their own copy of the federal sequence archive that had once looked as if it might be closed.
The amount of data stored for a human genome will drop sharply. Sequencers produce huge amounts of raw data that then has to be analyzed and processed by software to produce the result.
With the field still young, many researchers store all the raw data, so it can be re-analyzed if better software is developed in the future.
In uncertain times, “scientists cling to their data,” said David J. Dooling, assistant director of the genome institute at Washington University in St. Louis.
But there is now so much raw data that it is becoming not feasible to re-analyze it. So researchers will increasingly store just the final results. In the case of human genomes, they might store even less — only the difference between a particular genome and some reference genome.
Professor Brown of Michigan State said: “We are going to have to come up with really clever ways to throw away data so we can see new stuff.”
Wednesday, November 30, 2011
Paul Allen in WSJ: Why We Chose 'Open Science'
To accelerate research breakthroughs on brain diseases, the Allen Institute puts all its data online for use without fees.
Paul Allen The Wall Street Journal November 30, 2011
The Allen Institute for Brain Science in Seattle grew out of a simple question I posed in 2002 to a constellation of top people in the field: What's the most useful thing we could do to propel neuroscience forward? The consensus became our inaugural project—a comprehensive, molecular-level, three-dimensional map of the mouse brain to show precisely where every gene is active, or "expressed." It was the first step on a long road to understand how genes function in the human brain, knowledge that will point to ways to better diagnose and treat brain ailments.
A crucial aspect to this project—and others the Allen Institute has pursued over the last eight years—is an "open science" research model. Early on, we considered charging commercial users for access to our online data. From a strictly financial standpoint, it made sense to reap front-end fees and, down the line, intellectual property royalties. The revenue could cover the high costs of maintenance and development to keep the resource current and useful.
But our mission was to spark breakthroughs, and we didn't want to exclude underfunded neuroscientists who just might be the ones to make the next leap. And so we made all of our data free, with no registration required. The Institute would have no gatekeeper. Our terms-of-use agreement is about 10% as long as the one governing iTunes.
Our facility is neither the first nor the last to use a shared database to embrace "open science" and reject the competitive, single-lab R&D paradigm. Traditional research incentives—where journal publications are the coin of the realm—tend to discourage vital sharing.
In 1982, even before the dawn of the Internet, a consortium of government agencies established the open access GenBank. Maintained by a division of the National Institutes of Health (NIH), GenBank now houses the sequence data from the Human Genome Project, the inspiration for our brain mapping.
In recent years the NIH has sponsored other data-sharing portals, including the Alzheimer's Disease Neuroimaging Initiative and the Neuroscience Information Framework. Private nonprofits like the Pistoia Alliance and Sage Bionetworks are curating their own open-source repositories.
But the Allen Institute remains distinct in conducting industrial-scale big science that is fundamentally collaborative. Internally, our team of scientists and support staff works together to meet the time lines and milestones that frame each large project. The team released the initial data set from a ground-breaking human whole-brain atlas last year, and it is now midstream on a project to define the circuitry between neurons and how it affects human behavior. Most important, we generate data for the purpose of sharing it. Since opening shop in 2003, we've had 23 public releases, or about three per year. We don't wait to analyze our raw data and publish in the literature. We pour it onto the public website as soon as it passes our quality control checks. Our goal is to speed others' discoveries as much as to springboard our own future research.
The databases currently provide tens of millions of high-resolution images. The initial mouse brain atlas alone involved 600 terabytes of data, or 600 trillion bytes, more than half the total content of the Internet when we started. Since data of this volume would be of little use without effective search and navigation tools, the Institute developed a free online viewing application as well as the downloadable Brain Explorer 3D viewer, which illuminates how expressed genes are distributed throughout the brain.
Open science is a long-term and pricey proposition. It demands consistent curating, maintenance and updating of databases, and regular software and hardware upgrades. The institute offers online video tutorials on a YouTube channel and in a tutorial library. For those seeking in-person walk-throughs or forums, it hosts training workshops and user group sessions in several areas around the country each year. These services, too, are free of charge.
It is a modest cost that is paying off as the scientific community embraces the open access model. In October, the institute's suite of databases received more than 45,000 visits, from six continents and from research organizations of every stripe: universities, government laboratories, independent institutes and biotech and pharmaceutical companies. Institute brain atlases are accelerating research on the underlying biology of a broad range of diseases, from Alzheimer's and Parkinson's to autism and schizophrenia. Growing numbers of college educators, from UCLA to the Radboud University Nijmegen in the Netherlands, are building curricular modules around our online resources.
What I've concluded is that foundations and other private funders who support scientific research also can help promote wider sharing of scientific data. Before funders write a check to a university, they should ask about the researcher's policies and track record on sharing.
On the federal level, the NIH now has such strong policies on sharing data. But I'd like to see the agency do even more to put its funding where its directives are. I propose that the NIH—along with the National Science Foundation and the U.S. Department of Education—direct funding into grant awards for management and curation of existing research data of special value.
That would siphon some money for traditional research grants for new work. But I think we'd get more bang for our buck by making more data more useful to more scientists—and, by extension, to the world community that will benefit from their work.
Mr. Allen, the co-founder of Microsoft with Bill Gates, launched the nonprofit Allen Institute for Brain Science in 2003.
Paul Allen The Wall Street Journal November 30, 2011
The Allen Institute for Brain Science in Seattle grew out of a simple question I posed in 2002 to a constellation of top people in the field: What's the most useful thing we could do to propel neuroscience forward? The consensus became our inaugural project—a comprehensive, molecular-level, three-dimensional map of the mouse brain to show precisely where every gene is active, or "expressed." It was the first step on a long road to understand how genes function in the human brain, knowledge that will point to ways to better diagnose and treat brain ailments.
A crucial aspect to this project—and others the Allen Institute has pursued over the last eight years—is an "open science" research model. Early on, we considered charging commercial users for access to our online data. From a strictly financial standpoint, it made sense to reap front-end fees and, down the line, intellectual property royalties. The revenue could cover the high costs of maintenance and development to keep the resource current and useful.
But our mission was to spark breakthroughs, and we didn't want to exclude underfunded neuroscientists who just might be the ones to make the next leap. And so we made all of our data free, with no registration required. The Institute would have no gatekeeper. Our terms-of-use agreement is about 10% as long as the one governing iTunes.
Our facility is neither the first nor the last to use a shared database to embrace "open science" and reject the competitive, single-lab R&D paradigm. Traditional research incentives—where journal publications are the coin of the realm—tend to discourage vital sharing.
In 1982, even before the dawn of the Internet, a consortium of government agencies established the open access GenBank. Maintained by a division of the National Institutes of Health (NIH), GenBank now houses the sequence data from the Human Genome Project, the inspiration for our brain mapping.
In recent years the NIH has sponsored other data-sharing portals, including the Alzheimer's Disease Neuroimaging Initiative and the Neuroscience Information Framework. Private nonprofits like the Pistoia Alliance and Sage Bionetworks are curating their own open-source repositories.
But the Allen Institute remains distinct in conducting industrial-scale big science that is fundamentally collaborative. Internally, our team of scientists and support staff works together to meet the time lines and milestones that frame each large project. The team released the initial data set from a ground-breaking human whole-brain atlas last year, and it is now midstream on a project to define the circuitry between neurons and how it affects human behavior. Most important, we generate data for the purpose of sharing it. Since opening shop in 2003, we've had 23 public releases, or about three per year. We don't wait to analyze our raw data and publish in the literature. We pour it onto the public website as soon as it passes our quality control checks. Our goal is to speed others' discoveries as much as to springboard our own future research.
The databases currently provide tens of millions of high-resolution images. The initial mouse brain atlas alone involved 600 terabytes of data, or 600 trillion bytes, more than half the total content of the Internet when we started. Since data of this volume would be of little use without effective search and navigation tools, the Institute developed a free online viewing application as well as the downloadable Brain Explorer 3D viewer, which illuminates how expressed genes are distributed throughout the brain.
Open science is a long-term and pricey proposition. It demands consistent curating, maintenance and updating of databases, and regular software and hardware upgrades. The institute offers online video tutorials on a YouTube channel and in a tutorial library. For those seeking in-person walk-throughs or forums, it hosts training workshops and user group sessions in several areas around the country each year. These services, too, are free of charge.
It is a modest cost that is paying off as the scientific community embraces the open access model. In October, the institute's suite of databases received more than 45,000 visits, from six continents and from research organizations of every stripe: universities, government laboratories, independent institutes and biotech and pharmaceutical companies. Institute brain atlases are accelerating research on the underlying biology of a broad range of diseases, from Alzheimer's and Parkinson's to autism and schizophrenia. Growing numbers of college educators, from UCLA to the Radboud University Nijmegen in the Netherlands, are building curricular modules around our online resources.
What I've concluded is that foundations and other private funders who support scientific research also can help promote wider sharing of scientific data. Before funders write a check to a university, they should ask about the researcher's policies and track record on sharing.
On the federal level, the NIH now has such strong policies on sharing data. But I'd like to see the agency do even more to put its funding where its directives are. I propose that the NIH—along with the National Science Foundation and the U.S. Department of Education—direct funding into grant awards for management and curation of existing research data of special value.
That would siphon some money for traditional research grants for new work. But I think we'd get more bang for our buck by making more data more useful to more scientists—and, by extension, to the world community that will benefit from their work.
Mr. Allen, the co-founder of Microsoft with Bill Gates, launched the nonprofit Allen Institute for Brain Science in 2003.
Tuesday, November 29, 2011
NPR on Big Data
The Digital Breadcrumbs That Lead To Big Data
Yuki Noguchi NPR November 29, 2011
First of a Two Part Story
First of a Two Part Story
What do Facebook, Groupon and biotech firm Human Genome Sciences have in common? They all rely on massive amounts of data to design their products. Terabytes and even zettabytes of information about consumers or about genetic sequences can be harnessed and crunched.
The practice is called big data, and as the term suggests, it is huge in both scope and power. Analyzing big data enables anything from predicting prices to catching criminals, and has the potential to impact many industries.
One way to understand how big data works is to think about your daily life. You write an email, call your boss, pass a security camera, maybe buy a plane ticket online. Taken alone, this is disjointed, boring information. To Elizabeth Charnock, it makes up your digital character.
We have seen the industrial revolution, and we are witnessing a data revolution.
- Oren Etzioni, professor of computer science at the University of Washington
"Digital character is this idea that almost everybody these days leaves behind a giant digital breadcrumb trail," she says.
Charnock founded Cataphora, a company that can process huge amounts of this sort of data about employees to determine patterns. She says those patterns can predict everything from a person's mood to their skill as a manager to a person's inclination to commit fraud.
Take rogue trader Jerome Kerviel, who cost his French bank billions of dollars in losses.
"His cell phone bill was literally an order of magnitude larger than any of his coworkers — why?" Charnock asks. "Well, because he wanted to put less things in writing. He almost never took vacation, even though French people love to take vacation."
Charnock says Kerviel also circumvented usual trading and communication protocols.
If you've got your eye on that brand new camera with all the features you never imagined you'd need, how do you know if it's time to buy? Decide.com is a prediction tool that tells you when gadget prices are likely to rise, stay the same or fall – and if you should wait a few weeks to buy the rumored newer model.
Decide collects prices of over 100,000 electronic products every day from hundreds of online retailers. It also searches technology blogs for rumors of upcoming new releases, adding up to over 25 GB of data per day.
Decide's four computer science Ph.D's create algorithms to mine the data and predict whether prices will go up or down, similar to what the finance industry has done for years to forecast stock prices. But now that data storage has become so cheap, other businesses can get in the game.
The practice is called big data, and as the term suggests, it is huge in both scope and power. Analyzing big data enables anything from predicting prices to catching criminals, and has the potential to impact many industries.
One way to understand how big data works is to think about your daily life. You write an email, call your boss, pass a security camera, maybe buy a plane ticket online. Taken alone, this is disjointed, boring information. To Elizabeth Charnock, it makes up your digital character.
We have seen the industrial revolution, and we are witnessing a data revolution.
- Oren Etzioni, professor of computer science at the University of Washington
"Digital character is this idea that almost everybody these days leaves behind a giant digital breadcrumb trail," she says.
Charnock founded Cataphora, a company that can process huge amounts of this sort of data about employees to determine patterns. She says those patterns can predict everything from a person's mood to their skill as a manager to a person's inclination to commit fraud.
Take rogue trader Jerome Kerviel, who cost his French bank billions of dollars in losses.
"His cell phone bill was literally an order of magnitude larger than any of his coworkers — why?" Charnock asks. "Well, because he wanted to put less things in writing. He almost never took vacation, even though French people love to take vacation."
Charnock says Kerviel also circumvented usual trading and communication protocols.
If you've got your eye on that brand new camera with all the features you never imagined you'd need, how do you know if it's time to buy? Decide.com is a prediction tool that tells you when gadget prices are likely to rise, stay the same or fall – and if you should wait a few weeks to buy the rumored newer model.
Decide collects prices of over 100,000 electronic products every day from hundreds of online retailers. It also searches technology blogs for rumors of upcoming new releases, adding up to over 25 GB of data per day.
The data is sent to Amazon's cloud storage and processed with the help of Hadoop software, which is used for many big data projects to organize the information and minimize mistakes.
Decide's four computer science Ph.D's create algorithms to mine the data and predict whether prices will go up or down, similar to what the finance industry has done for years to forecast stock prices. But now that data storage has become so cheap, other businesses can get in the game.
In addition to crunching the numbers, Decide analyzes thousands of blog posts and press releases to see if any rumors have surfaced. Sources that were reliable in the past are listened to more closely.
All in all, Decide has nearly 100 terabytes of data to analyze and end up with a simple prediction: Should you buy now or wait for prices to drop
—Sara Carothers, Stephanie d'Otreppe/NPR
But big data is not just about connecting dots to detect crime. The ability to process so much information and process it so quickly makes all kinds of things possible that weren't before. So LinkedIn finds jobs or people you might like to know about, and biotech companies can analyze gene sequences in billions of combinations to design drugs.
Data analytics itself is not new. Two decades ago, Wall Street hired teams of physicists to analyze investments. But in the last couple of years, computing, storage and bandwidth capacity have become so cheap that it has altered the scale of what's possible.
Now, with very little money, a gifted student or a small startup can design big-data applications.
"Everywhere you look, there's an opportunity to collect more data and then apply a statistical or mathematical approach to understanding what's happening," says Chris Kemp, chief executive officer of Nebula, a firm that provides storage and computing capacity for other companies to be able to process their big data applications.
Kemp says ultimately big data will give consumers better tools so they can do a better job of predicting things like prices, such as whether an airfare is likely to go up or down. Farmers can do a better job of insuring their crops if they can forecast the weather with greater accuracy.
Oren Etzioni, a professor of computer science at the University of Washington, says this trend is fueling intense demand for mathematics and computing talent.
"We have seen the industrial revolution, and we are witnessing a data revolution," Etzioni says.
He's started three big-data companies. One of them, Decide.com, employs four Ph.D.s to design better programs to forecast prices on consumer electronics.
Etzioni says a good data scientist can write algorithms that filter data, understand what it's telling you, and then graphically represent it. The end result is like getting a bird's-eye view of a vast territory of information.
Big data can, and occasionally does, go wrong. Comic examples of that include mismatched recommendations, like "My TiVo thinks I'm gay." "But think about a company divulging your Web surfing history with your name attached and you begin to get a sense of how big data opens the door to new possibilities of security or privacy breaches.
James Slavet, a venture capitalist at Greylock Partners, says his firm invests in companies that use big data creatively and responsibly. He says data does not stand in for human judgment.
"They do use it to make the judgment more sound, more objective and to hopefully lead to better decision making," he says.
—Sara Carothers, Stephanie d'Otreppe/NPR
"Any one of those things, you kind of say, 'So what?' But what we look for is a number of them that on the surface perhaps don't seem to be related but all seem to be happening at the same time," she says.
Charnock says had the French bank analyzed that data, it might have flagged the rogue trader earlier.But big data is not just about connecting dots to detect crime. The ability to process so much information and process it so quickly makes all kinds of things possible that weren't before. So LinkedIn finds jobs or people you might like to know about, and biotech companies can analyze gene sequences in billions of combinations to design drugs.
Data analytics itself is not new. Two decades ago, Wall Street hired teams of physicists to analyze investments. But in the last couple of years, computing, storage and bandwidth capacity have become so cheap that it has altered the scale of what's possible.
Now, with very little money, a gifted student or a small startup can design big-data applications.
"Everywhere you look, there's an opportunity to collect more data and then apply a statistical or mathematical approach to understanding what's happening," says Chris Kemp, chief executive officer of Nebula, a firm that provides storage and computing capacity for other companies to be able to process their big data applications.
Kemp says ultimately big data will give consumers better tools so they can do a better job of predicting things like prices, such as whether an airfare is likely to go up or down. Farmers can do a better job of insuring their crops if they can forecast the weather with greater accuracy.
Oren Etzioni, a professor of computer science at the University of Washington, says this trend is fueling intense demand for mathematics and computing talent.
"We have seen the industrial revolution, and we are witnessing a data revolution," Etzioni says.
He's started three big-data companies. One of them, Decide.com, employs four Ph.D.s to design better programs to forecast prices on consumer electronics.
Etzioni says a good data scientist can write algorithms that filter data, understand what it's telling you, and then graphically represent it. The end result is like getting a bird's-eye view of a vast territory of information.
Big data can, and occasionally does, go wrong. Comic examples of that include mismatched recommendations, like "My TiVo thinks I'm gay." "But think about a company divulging your Web surfing history with your name attached and you begin to get a sense of how big data opens the door to new possibilities of security or privacy breaches.
James Slavet, a venture capitalist at Greylock Partners, says his firm invests in companies that use big data creatively and responsibly. He says data does not stand in for human judgment.
"They do use it to make the judgment more sound, more objective and to hopefully lead to better decision making," he says.
Slavet calls big data a tectonic shift, one that will continue to affect many things we do for decades to come.
Listen to the Story (and see the Graphs) at http://www.npr.org/2011/11/29/142521910/the-digital-breadcrumbs-that-lead-to-big-data
Tuesday, November 22, 2011
Big Data: The news forecast (Wired)
Tom Cheshire Wired December 2011
See gallery of illustrations at http://www.wired.co.uk/magazine/archive/2011/12/features/the-news-forecast/viewgallery#!image-number=1
In September 2010 the Yemen ministry of industry announced a national strategy to combat food shortages. The UN Food Price Index had reached consecutive record highs in the previous few months. Yemen had also suffered flooding, which had killed around 100 people and disrupted farming. The strategy, which included a review of existing subsidies and the development of food-for-work programmes, proved ineffective. By December 2010, concerned that protests over rising food prices were starting to grow in Tunisia, Yemen's president, Ali Abdullah Saleh, halved income tax and ordered the government to control the prices of basic commodities.
By late January, though, thousands of protesters had taken to the streets to demand Saleh's resignation, brandishing flatbread with baked-in slogans and wearing the food as helmets. The clashes continued into February and grew more violent as the UN's Food Price Index reached an all-time high. On March 18, 45 protesters were killed when an unidentified gunman opened fire. Six days later, the government fought a battle with al-Qaeda gunmen in the province of Abyan and Marib, killing 15. The same day, 10,000 protesters gathered in the capital city Sana'a. Saleh said that he would accept the opposition's transition plan that day, but clung on for another month. He finally quit Yemen for Saudi Arabia after a bomb planted in the presidential compound exploded, killing seven people; Saleh suffered 40 per cent burns, shrapnel wounds and internal bleeding. In total, Human Rights Watch estimates that 233 protesters were killed on the streets. Three months later, Saleh unexpectedly flew back to Yemen; 100 more protesters and tribesmen were killed in the first five days of his return and the situation remains unresolved.
A year earlier, on January 12, 2010, a tech startup posted an article on its blog: "Yemen heading for disaster in 2010?" The author, "Ninja Shoes", wrote: "Based on the information we've gathered, Yemen will likely experience food shortages and torrential floods in 2010. This combination of natural disasters, propensity for famine and malnutrition, and challenges with Islamic radicals and terrorists, make it a hot spot for conflict in the future."
The 20 employees of Recorded Future aren't foreign-policy experts. They aren't traders either, but if you'd started using Recorded Future's predictions to buy US stocks on January 1, 2009, you would have made an annual return of 56.69 per cent. (The S&P 500 had an annualised return of 17.22 per cent over the same period.) Between May 13 and August 5 this year, as markets behaved with vertiginous abandon, their strategy returned 10.4 per cent; in contrast, the S&P 500 lost 9.9 per cent of its value. They're data experts: computer scientists, statisticians and experts in linguistics. And in the data, they think, lies the future.
All Recorded Future's predictions, whatever the field, are based on publicly available information -- news articles, government sites, financial reports, tweets -- fed into the company's own algorithms. The result, it claims, is a "new tool that allows you to visualise the future" -- one that is changing how government intelligence agencies gather information and how giant hedge funds place bets. On its website, Recorded Future states: "We don't grant interviews and we don't issue press releases." But behind closed doors, the company is developing the technology that has been described be one tech blog as an "information weapon".
The company, cofounded by Christopher Ahlberg, an entrepreneur who sold his first business for $195 million and served in the Swedish special forces, has $8.5 million in funding. Its first two investors were Google and the CIA. Recorded Future counts US government agencies, banks and hedge funds among the clients paying million-dollar contracts. But its true ambition is to organise all the data on the internet for similar predictive analysis -- to make the future calculable.
Recorded Future's main office is in Gothenburg, Sweden. On a drab morning in May, trams clang past a metal door that doesn't bear the company's name. Two flights of stairs lead to a wooden door, with a discreet sticker label-gunned above the letterbox in caps: "RECORDED FUTURE". The rooms date from the 17th century; they're airy and bright with high ceilings and intricate plaster mouldings. Eight employees work here on the technical aspects of the system. The company also has offices in Boston, New York and Arlington, Virginia -- ten minutes' drive from the Pentagon, 15 from Langley.
"Yemen took four or five months longer than we predicted," says Ahlberg, 43, sitting on a sofa in a small meeting room. Before Wired visited, he warned over the telephone: "You won't get a government agency out of my mouth. Dude, if I do that, they're coming to take my kids." In person, he's tall, with hair cropped short, and is quick to laugh. The telephone caveat still stands, but Ahlberg is willing to talk for the first time about what exactly it is his company does and why Google, intelligence agencies and hedge funds are all so interested.
Ahlberg was born in September 1968, in Kungälv, a town 40 minutes' drive north of Gothenburg. His father was a captain on merchant ships, his mother taught French and English in Sweden. In his first year at secondary school, he created a drawing program on his Sinclair Spectrum called Art CAD ("like an early version of Photoshop") and sold individual copies by advertising it in the local paper. After school, he wanted to study computer science but first had to complete military service in 1987. He chose the Lapplands Jägarregemente special forces, and began training for a hypothetical Russian invasion: "They would come in from Finland and go to Norway; we were supposed to cut them off in the middle. We were supposed to do what the Iraqis are doing now, guerrilla warfare. But we were master cross-country skiers."
Ahlberg then went onto take his degree at Chalmers University of Technology in Gothenburg. As a post doc, he travelled to the University of Maryland to work as a visiting researcher at the Human-Computer Interaction Lab for two summers. During the first visit, at 23, he co-authored a published paper with the director of the lab; the second summer, he co-wrote two, about the new field of data visualisation. Ahlberg returned to Sweden, finishing his PhD in four years instead of six, but it was his work at Maryland that formed the basis for his first company, Spotfire. Launched in 1996, the business created visualisation tools for business intelligence; in 2007, it was bought by Tibco for $195 million(£125m). "I didn't have to work anymore," says Ahlberg. "But I can't stop."
Spotfire had helped businesses visualise internal databases. After the sale, "we started hanging around in coffee shops in Boston and New York", says Ahlberg. "We thought: what's the most interesting data source out there? And it's nebulous, but the web is the most interesting dataset there is on the planet. Instead of just corporate databases, let's think about the web as my data source." Ahlberg began talking with Staffan Truvé, who had supervised his PhD and started Spotfire with him. "It was in the back of my head that, as humans, we had generally started to become better at predicting things," says 48-year-old Truvé. "Your car tells you that you need to change your oil in 200km, or there is a sign saying your bus is coming in five minutes. These tiny predictive signals are popping up everywhere." Ahlberg was excited: "So then the premise becomes that the web has predictive power. How can we harvest that?"
A few isolated, eye-catching examples have shown the prognostic possibilities in such data. In 2008, Google showed search queries could accurately predict the spread of flu in the US up to two weeks before the federal Centers for Disease Control. In his book, Super Crunchers, Ian Ayers claimed that creditcard companies can predict with 98 per cent accuracy whether you'll divorce, based on your purchases -- Google's Marissa Mayer even quoted the statistic at SXSW 2011 (however, in a recent statement, Visa denied that it monitored such data or made any such conclusions, saying the claim was "inaccurate and wrong"). One recent study has shown Twitter to be 88.67 per cent accurate in predicting the Dow Jones three days in advance. This July, financier Paul Hawtin founded a London hedge fund that is based entirely on social media. And in September, a researcher from the University of Illinois fed the Nautilus supercomputer with 100 million news items, much like Recorded Future does, and "anticipated" the Arab Spring and the killing of Osama Bin Laden, albeit retrospectively -- a prediction of the past. But for Ahlberg to develop a tool that could create predictions for any input, from finance to terrorism, would be much harder. Recorded Future would not only have to index the internet, but also understand and interpret it.
The first generation of search engines, such as Lycos and Alta Vista, used traditional text search to deliver web pages, deploying their own algorithms, but essentially looking at individual documents in isolation. Google changed this in 1998. Its PageRank algorithm analysed the links between web pages, promoting those that had more links pointing to them from other sites. Recorded Future is part of the third generation: instead of explicit link analysis, it examines implicit links -- what it calls "invisible links" between documents that refer to the same entities or events. It does this by separating the documents and their content from what they talk about, identifying canonical entities and events that exist outside of the article.
"What matters is that it's freaking complicated," says Ahlberg. In practice, Recorded Future harvests 25,000 data sources as RSS feeds, which could include Companies House and US Securities and Exchange Commission filings, a New York Times article, Twitter and Facebook posts, obscure blogs (there's one on Norwegian salmon fishing) or transcripts from earnings calls or political speeches -- "just a flood of stuff", says Ahlberg. It does the same for Chinese and Arabic sources. "Then we look for entities -- people, places, technologies; and events -- a murder, a bomb explosion, a person moving from A to B, product launches."
This linguistic analysis is "really tough", according to Truvé. Because Recorded Future takes sources from all over the internet, rather than a particular data set, "the data is not so nice". "We could have built a perfect data set around Pfizer, say, or Barack Obama," says Ahlberg. "It's harder then to think of the big picture. So we tried to make this ambitious." Recorded Future currently uses two separate algorithms, one proprietary, one licensed, to analyse language; the staffers in Gothenburg are tweaking them continually to see which works better. But the result is that Recorded Future knows who Nicolas Sarkozy is, say: that he's the president of France, he's the husband of Carla Bruni, he's 1.65m tall in his socks, he travelled to Deauville for the G8 summit in May. If you Google "president of France", you'll get two Wikipedia pages on "president of France" then " Nicolas Sarkozy". Useful, but Google doesn't know how the two, Sarkozy and the presidency, are actually related; it's just searching for pages linking to the terms.
Recorded Future ranks all these canonical entities and events, based on the number of references to them, the credibility of the document or document source and several other factors, such as the co-occurrence of different events and entities in the same or in related documents, to create a "momentum" score. Positive or negative sentiment is added to this score. For example, searching big pharma in general will tell you that over the next five years, nine of the world's 15 best-selling medicines will lose patent protection -- the event earns a high momentum score because it is backed by 13 news items from 12 sources -- or that, specifically, Inhibitex, a biopharma business, will need cash in November 2011 if it plans to fund the Phase 2b development of a new drug internally, based on five items from five sources.
Recorded Future isn't the only company attempting to bring hardcore linguistic analysis to a larger audience. Wolfram Alpha is a search engine that can understand a query such as "nuclear explosions in China" and deliver relevant information such as maps and kilotonnes per explosion, although it's culled from "tame" data curated by the company itself. And IBM didn't develop Watson just to school humans on Jeopardy; it's actually a huge research project dedicated to processing questions asked in natural language, based on four terabytes of structured and unstructured data sets, including the full text of Wikipedia. "There are any number of offerings coming on to the market now," says Colin Shearer, senior vice president at SPSS, a predictive-analytics company owned by IBM. One of those is Quid, a two-year-old, 45-strong business founded in 2008. "Human activity has never left an information trail like it does today," says Bob Goodson, its founder. "If only we could harness the intelligence that's locked in the information, we could build systems to understand the world better, and therefore make better decisions." Quid includes Microsoft among its customers. With $15 million in investment, it aims to be the next Bloomberg in business intelligence.
Where Recorded Future goes beyond mere analysis of open data, though, is by adding the "time and space" dimension of the documents -- "references to when and where an event has taken place, or when and where it will take place," says Truvé, "since many documents actually refer to events expected to take place in the future." Using RSS streams allows Recorded Future to have a publishing time as an anchor point for this temporal analysis, which means it can deal with difficult expressions such as "next week", "in three months' time" or "in two quarters". This may sound simple, but it's crucial: the time and space analysis is the first way Recorded Future can make predictions about the future -- by aggregating weighted opinions about the likely timing of future events using algorithmic crowdsourcing. On top of that, it uses statistical models to predict future happenings based on historical records of similar chains of events. "The secret sauce is not dependent upon one ingredient," says Truvé. "It's a combination."
On April 1, 2009, a few months after Ahlberg and Truvé began testing this combination, Ahlberg met Rich Miner, the co-creator of the Android mobile operating system (with Andy Rubin) and a partner at Google Ventures, at the Starbucks on Harvard Square, Massachusetts. Miner was impressed: "We believed there was predictive power in the information contained in the web," he says. "If you can organise that information temporally, then you can look at past and present, and infer things from the future. That's pretty unique so far from Recorded Future." The CIA thought so, too.
In the 40s the allies routinely bombed rail bridges to disrupt supply lines into Nazi-occupied France. After a raid, though, the Royal Air Force couldn't fly reconnaissance missions over the targets as they were considered too risky, so it didn't know if a bridge had been destroyed. The Special Operations Executive (SOE), however, came up with a novel strategy for finding out. By monitoring the daily prices of oranges on sale at various fruit stalls Paris, SOE agents dropped behind enemy lines were able to tell which supply chains had been affected. (Germans embedded in London were doing the same thing; unfortunately for the Nazis, they were under the control of SOE and were fed false information.) This is the differ- ence between information and intelligence: information is the price of oranges, intelli- gence is knowing which supply chain has been affected. This openly available, "free" infor- mation, when it's turned into intelligence, becomes extremely valuable.
"Open-source intelligence has always been crucial, but for most of the cold war it was neglected by western intelligence agencies," says Calder Walton, a research associate at Cambridge University and author of the book Empire of Secrets, to be published in 2013. "That was the archetypal intelligence war: intelligence necessarily involved information that couldn't be gained from any other source -- human agents or telephone tapping." That doesn't mean covert intelligence was more effective, though: Daniel Moynihan, a former US senator, compared CIA reports gathered from secret sources with Soviet documents recovered after the fall of the Berlin Wall and found they significantly overestimated Soviet capabilities. But he discovered that western think tanks using publicly available material, such as the RAND Corporation, were much more accurate. US diplomat George Kennan estimated in 1997 that "95 per cent of what we need to know about foreign countries could very well be obtained by the careful and competent study of perfectly legitimate sources of information open and available to us".
"All of this has changed since the collapse of the Soviet Union," says Walton. "Open-source intelligence has boomed in recent years -- especially since 9/11." At a conference in 2008, Michael Hayden, then director of the CIA, said: "Open-source intelligence contributes to national security in unique and valuable ways virtually every day." Stephen Mercado, an ana- lyst in the CIA directorate of science and technology, estimates that 80 per cent of all valuable intelligence now comes from open sources. In January 2011, Sir Gus O'Donnell, head of the UKcivil service, told the Chilcot inquiry into the invasion of Iraq: "I have strongly and always been of the view that we probably underestimated open source [intelligence]." Open source is the big growth area in intelligence and every western agency is looking for the tools to give it an edge.
Ahlberg refuses to discuss his company's work with the CIA, or even whether there is work with the CIA. In-Q-Tel (IQT) is the CIA's investment arm (mission statement: "Identifies, adapts and delivers innovative technological solutions to support the missions of the Central Intelligence Agency"). It invests only in startup companies that will "provide strong, near-term advantages (within 36 months) to the IC [intelligence community]." IQT doesn't invest without the US secret intelligence services in mind. It backed Recorded Future with slightly less than $2.5 million.
Stephen Davidson, an investor at IQT who sits on Recorded Future's board, refused to comment; a spokesperson for IQT said that "while we are pleased to have Recorded Future as part of the IQT portfolio, we will respectfully decline to provide additional information about our investment". Does Ahlberg know what intelligence purposes Recorded Future is put to? "We would not know about those things," he says, folding his arms. "At this stage, I don't even want to know what people are doing with some of these things." He points out that IQT is "an independent company; at least to my knowledge theycan't force any [government agency] to use it." Truvé, though, says Recorded Future is working with 17 or 18 intelligence agencies. Another board member, Roger Ehrenberg, used to run a $6 billion hedge fund for Deutsche Bank before setting up his own firm, IA Ventures. According to Ehrenberg, In-Q-Tel is "actively involved" with Recorded Future. "Fundamentally, they look to invest in companies where they know they have a customer within the government," he says. "It's not just the CIA." Chris Holden, who works in Recorded Future's Arlington office, admitted to wired (with some understatement) that "we have a little bit of work with the federal government". Holden says that Recorded Future is being used to identify technologies the US government may invest in, such as nanotechnology in body armour. "It's not all super secret stuff necessarily." So, does having IQT as an investor mean thatRecorded Future is beholden to the US government, even if it is a private company? "We are an independent company," repeats Ahlberg. "Neither the US government, nor Google, nor hedge funds nor banks have ever tried to make us do anything. And frankly, you're sitting here with a bunch of Swedes. There's no way in hell you could get them to do anything bad."
Still, it's possible to identify examples of how one might use Recorded Future for open-source intelligence. Take the al-Qaeda leadership after Bin Laden's death: who would fill the vacuum? Recorded Future ran a search. Ayman al-Zawahiri, a founding member of Egypt's Islamic Jihad militant group, and long considered by the US government to be Bin Laden's right-hand man, showed some significant spikes in recorded and discussed activity in the last 12 months,especially when he called for military backing of Libyan rebels, suggesting al-Qaeda could fill a power vacuum in that country. But al-Zawahiri's sentiment score was extremely negative, to the degree that conspiracy theories were emerging that he was responsible for disclosing Bin Laden's location to the US. Saif al-Adel, a senior al-Qaeda commander, was attracting attention back in October 2010, written about as "the new face of al-Qaeda in 2011". Recorded Future concluded that it was "clear that Said al-Adel has been routed in Pakistan for some time now and appears to be embedded in the political structure of al-Qaeda"; his momentum score was high. They also found that Libyan Abu Yahya al-Libi (described by a former CIA analyst as an "insurgent-theologian"), offered access to one of the most volatile regions on the globe right now, based on his current likely location, which al-Qaeda might consider a useful foothold. Finally, they looked at Anwar al-Awlaki, a Yemeni-American imam who posted pro-al-Qaeda/anti-western YouTube videos, and ran a blog and Facebook page. Clearly of interest given the attempted US drone strike to kill him days after Bin Laden's death, he started building momentum in late March and April with reports that he was urging on the Arab Spring protests. Al- Awlaki was killed in Yemen dur- ing a US drone attack on September 30 this year.
Recorded Future concluded that multiple players will rise to prominence regionally, that al-Qaeda could split around al-Zawahiri, and that al-Qaeda sees advantage to be taken in the Arab pro-democracy protests. A couple of months later, al-Zawahiri was confirmed, although experts were sceptical about whether he could unite the membership in Saudi Arabia and the Gulf States behind him.
Ahlberg is more willing to talk about how Recorded Future is being used in finance. "If you take our momentum score, and look across S&P 500 companies, can you predict the liquidity or the stock volume of those companies over time?" asks Ahlberg. "It turns out you can." Stock that is being talked about and is in investors' attention is, of course, more likely to be traded: "It's much easier to prove volume than direction, whether a stock is going up or down." So Recorded Future takes momentum and combines it with sentiment -- whether a company is mentioned in a positive light -- and derives a score. Taking these news bursts across the S&P 500, it can sort them into ten different groups, from high to low. "Then you say, every day, 'I am going to own what is in the top and short what is in the bottom.' You are making lots of small picks on a daily basis - this strategy turns over the portfolio 63 per cent every day."
Running predictive tests on data from January 2009 to January 2011, Recorded Future showed that its top decile has a beta (a measure of risk in portfolio) of 1.08 -- fairly low -- and a statistically significant annualised continuous alpha (a risk-adjusted measure of active return on investment) of +16 per cent. The bottom two deciles had a high beta (1.37 and 1.34, respectively) but with statistically significant negative alphas, at -42 per cent and -26 per cent annually. "Constructing hedged portfolios out of the securities in these deciles provides some compelling trading strategies," says Evan Sparks, an analyst at Recorded Future.
Beyond high-frequency trading strategies, the company says it can predict stock shifts on the basis of one-day events, separated into scheduled events and speculative events. "The theory that if something is written saying, 'on Friday so and so will release earnings', that should be priced into the market immediately," says Ahlberg. "In reality it is not." Recorded Future took 19,000 such events and asked what happens to the stock price. On average, as stocks come into those scheduled events, the prices rise; coming out of them they fall five base points either way. "It's like finding a roulette wheel that is skewed."
Another way is to examine the next two weeks of a particular business's future and look for certain events. One is insiders selling stock. You may think this would be a good time to sell; in fact, insiders often sell just after stocks have already peaked. So Recorded Future looks for data that can be combined with this knowledge. If an insider sells stock after a management lay-off, stock falls on average 1.5 per cent. Expand this event to a whole market and "You have 2,000 events within 2011," says Ahlberg. "By turning it into a big data screen, I have created my own skewed roulette wheel I can consistently bet on." Chris Malloy is an associate professor at Harvard Business School who specialises in behavioural finance. He's played with Recorded Future's data: "I haven't seen anything with that ability. It's pretty neat -- no one's doing that. The predictability is certainly good."
What Recorded Future can't forecast are "black swan events", which are by definition unpredictable and undirected. "You can look at what happens afterwards, though," says Ahlberg. He takes the example of a natural disaster. "Start looking at how other countries behave. After a natural disaster, the US will travel there every time, the UK does it 50 per cent of the time, Iran will do it every single time, China never really does." China did, though, after the 2010 Chilean earthquake. Two months later, it announced a new trade agreement. China didn't travel to Haiti: no trade agreement followed. But it did after Pakistan was hit by flood, soon announcing a $10 billion deal. "We're looking for those historical patterns and using them to predict what might happen," says Ahlberg.
Recorded Future's hedge-fund clients are only slightly less secretive than theCIA. Ehrenberg says a handful of Wall Street hedge funds and banks are using the technology: "Recorded Future is a high-value signal, relative to conventional quantitative-analysis trading signals. People are making money." Josh Holden, CEO of Fina Technologies, which creates algorithms for high-frequency quant trading by hedge funds, says that Recorded Future's client base "is closely guarded. But there are more than a few firms using it.
Sandfire AG is a Swiss consultancy in the public and private security sectors, and a client of Recorded Future. "It helps us keep track of travel routes of high-level decision-makers," says Felix Juhl, a senior partner. "A state visit by a high-ranking politician may be followed by specific corporate activities. Keeping track of travel routes can serve as an early warning."
Ahlberg says Recorded Future now earns revenues in the millions of dollars from a client base of less than 100, but which includes governments, hedge funds, big banks, watchdogs and consultancies. This select clientele place a high value on the distilled insight the company provides. Ahlberg sees a big opportunity: "Even within what we have started around finance and intelligence, there is no reason why we couldn't build another $100 million-revenue company within a small set of years." But Recorded Future plans on being more than just a profitable business tool. Ahlberg is expanding its indexes: he eventually wants every piece of data on the planet streaming live through his company's algorithms. The ultimate goal? "We want to organise the world -- and the internet -- for analysis." What Ahlberg doesn't say, perhaps deliberately, is that ever more data will likely lead to ever more accurate predictions. "It's dangerous to start talking about predicting the future," he says. "We're trying to play that down."
Tom Cheshire is assistant editor at wired. He wrote about the Ariane 5 rocket in 10.11
See gallery of illustrations at http://www.wired.co.uk/magazine/archive/2011/12/features/the-news-forecast/viewgallery#!image-number=1
In September 2010 the Yemen ministry of industry announced a national strategy to combat food shortages. The UN Food Price Index had reached consecutive record highs in the previous few months. Yemen had also suffered flooding, which had killed around 100 people and disrupted farming. The strategy, which included a review of existing subsidies and the development of food-for-work programmes, proved ineffective. By December 2010, concerned that protests over rising food prices were starting to grow in Tunisia, Yemen's president, Ali Abdullah Saleh, halved income tax and ordered the government to control the prices of basic commodities.
By late January, though, thousands of protesters had taken to the streets to demand Saleh's resignation, brandishing flatbread with baked-in slogans and wearing the food as helmets. The clashes continued into February and grew more violent as the UN's Food Price Index reached an all-time high. On March 18, 45 protesters were killed when an unidentified gunman opened fire. Six days later, the government fought a battle with al-Qaeda gunmen in the province of Abyan and Marib, killing 15. The same day, 10,000 protesters gathered in the capital city Sana'a. Saleh said that he would accept the opposition's transition plan that day, but clung on for another month. He finally quit Yemen for Saudi Arabia after a bomb planted in the presidential compound exploded, killing seven people; Saleh suffered 40 per cent burns, shrapnel wounds and internal bleeding. In total, Human Rights Watch estimates that 233 protesters were killed on the streets. Three months later, Saleh unexpectedly flew back to Yemen; 100 more protesters and tribesmen were killed in the first five days of his return and the situation remains unresolved.
A year earlier, on January 12, 2010, a tech startup posted an article on its blog: "Yemen heading for disaster in 2010?" The author, "Ninja Shoes", wrote: "Based on the information we've gathered, Yemen will likely experience food shortages and torrential floods in 2010. This combination of natural disasters, propensity for famine and malnutrition, and challenges with Islamic radicals and terrorists, make it a hot spot for conflict in the future."
The 20 employees of Recorded Future aren't foreign-policy experts. They aren't traders either, but if you'd started using Recorded Future's predictions to buy US stocks on January 1, 2009, you would have made an annual return of 56.69 per cent. (The S&P 500 had an annualised return of 17.22 per cent over the same period.) Between May 13 and August 5 this year, as markets behaved with vertiginous abandon, their strategy returned 10.4 per cent; in contrast, the S&P 500 lost 9.9 per cent of its value. They're data experts: computer scientists, statisticians and experts in linguistics. And in the data, they think, lies the future.
All Recorded Future's predictions, whatever the field, are based on publicly available information -- news articles, government sites, financial reports, tweets -- fed into the company's own algorithms. The result, it claims, is a "new tool that allows you to visualise the future" -- one that is changing how government intelligence agencies gather information and how giant hedge funds place bets. On its website, Recorded Future states: "We don't grant interviews and we don't issue press releases." But behind closed doors, the company is developing the technology that has been described be one tech blog as an "information weapon".
The company, cofounded by Christopher Ahlberg, an entrepreneur who sold his first business for $195 million and served in the Swedish special forces, has $8.5 million in funding. Its first two investors were Google and the CIA. Recorded Future counts US government agencies, banks and hedge funds among the clients paying million-dollar contracts. But its true ambition is to organise all the data on the internet for similar predictive analysis -- to make the future calculable.
Recorded Future's main office is in Gothenburg, Sweden. On a drab morning in May, trams clang past a metal door that doesn't bear the company's name. Two flights of stairs lead to a wooden door, with a discreet sticker label-gunned above the letterbox in caps: "RECORDED FUTURE". The rooms date from the 17th century; they're airy and bright with high ceilings and intricate plaster mouldings. Eight employees work here on the technical aspects of the system. The company also has offices in Boston, New York and Arlington, Virginia -- ten minutes' drive from the Pentagon, 15 from Langley.
"Yemen took four or five months longer than we predicted," says Ahlberg, 43, sitting on a sofa in a small meeting room. Before Wired visited, he warned over the telephone: "You won't get a government agency out of my mouth. Dude, if I do that, they're coming to take my kids." In person, he's tall, with hair cropped short, and is quick to laugh. The telephone caveat still stands, but Ahlberg is willing to talk for the first time about what exactly it is his company does and why Google, intelligence agencies and hedge funds are all so interested.
Ahlberg was born in September 1968, in Kungälv, a town 40 minutes' drive north of Gothenburg. His father was a captain on merchant ships, his mother taught French and English in Sweden. In his first year at secondary school, he created a drawing program on his Sinclair Spectrum called Art CAD ("like an early version of Photoshop") and sold individual copies by advertising it in the local paper. After school, he wanted to study computer science but first had to complete military service in 1987. He chose the Lapplands Jägarregemente special forces, and began training for a hypothetical Russian invasion: "They would come in from Finland and go to Norway; we were supposed to cut them off in the middle. We were supposed to do what the Iraqis are doing now, guerrilla warfare. But we were master cross-country skiers."
Ahlberg then went onto take his degree at Chalmers University of Technology in Gothenburg. As a post doc, he travelled to the University of Maryland to work as a visiting researcher at the Human-Computer Interaction Lab for two summers. During the first visit, at 23, he co-authored a published paper with the director of the lab; the second summer, he co-wrote two, about the new field of data visualisation. Ahlberg returned to Sweden, finishing his PhD in four years instead of six, but it was his work at Maryland that formed the basis for his first company, Spotfire. Launched in 1996, the business created visualisation tools for business intelligence; in 2007, it was bought by Tibco for $195 million(£125m). "I didn't have to work anymore," says Ahlberg. "But I can't stop."
Spotfire had helped businesses visualise internal databases. After the sale, "we started hanging around in coffee shops in Boston and New York", says Ahlberg. "We thought: what's the most interesting data source out there? And it's nebulous, but the web is the most interesting dataset there is on the planet. Instead of just corporate databases, let's think about the web as my data source." Ahlberg began talking with Staffan Truvé, who had supervised his PhD and started Spotfire with him. "It was in the back of my head that, as humans, we had generally started to become better at predicting things," says 48-year-old Truvé. "Your car tells you that you need to change your oil in 200km, or there is a sign saying your bus is coming in five minutes. These tiny predictive signals are popping up everywhere." Ahlberg was excited: "So then the premise becomes that the web has predictive power. How can we harvest that?"
A few isolated, eye-catching examples have shown the prognostic possibilities in such data. In 2008, Google showed search queries could accurately predict the spread of flu in the US up to two weeks before the federal Centers for Disease Control. In his book, Super Crunchers, Ian Ayers claimed that creditcard companies can predict with 98 per cent accuracy whether you'll divorce, based on your purchases -- Google's Marissa Mayer even quoted the statistic at SXSW 2011 (however, in a recent statement, Visa denied that it monitored such data or made any such conclusions, saying the claim was "inaccurate and wrong"). One recent study has shown Twitter to be 88.67 per cent accurate in predicting the Dow Jones three days in advance. This July, financier Paul Hawtin founded a London hedge fund that is based entirely on social media. And in September, a researcher from the University of Illinois fed the Nautilus supercomputer with 100 million news items, much like Recorded Future does, and "anticipated" the Arab Spring and the killing of Osama Bin Laden, albeit retrospectively -- a prediction of the past. But for Ahlberg to develop a tool that could create predictions for any input, from finance to terrorism, would be much harder. Recorded Future would not only have to index the internet, but also understand and interpret it.
The first generation of search engines, such as Lycos and Alta Vista, used traditional text search to deliver web pages, deploying their own algorithms, but essentially looking at individual documents in isolation. Google changed this in 1998. Its PageRank algorithm analysed the links between web pages, promoting those that had more links pointing to them from other sites. Recorded Future is part of the third generation: instead of explicit link analysis, it examines implicit links -- what it calls "invisible links" between documents that refer to the same entities or events. It does this by separating the documents and their content from what they talk about, identifying canonical entities and events that exist outside of the article.
"What matters is that it's freaking complicated," says Ahlberg. In practice, Recorded Future harvests 25,000 data sources as RSS feeds, which could include Companies House and US Securities and Exchange Commission filings, a New York Times article, Twitter and Facebook posts, obscure blogs (there's one on Norwegian salmon fishing) or transcripts from earnings calls or political speeches -- "just a flood of stuff", says Ahlberg. It does the same for Chinese and Arabic sources. "Then we look for entities -- people, places, technologies; and events -- a murder, a bomb explosion, a person moving from A to B, product launches."
This linguistic analysis is "really tough", according to Truvé. Because Recorded Future takes sources from all over the internet, rather than a particular data set, "the data is not so nice". "We could have built a perfect data set around Pfizer, say, or Barack Obama," says Ahlberg. "It's harder then to think of the big picture. So we tried to make this ambitious." Recorded Future currently uses two separate algorithms, one proprietary, one licensed, to analyse language; the staffers in Gothenburg are tweaking them continually to see which works better. But the result is that Recorded Future knows who Nicolas Sarkozy is, say: that he's the president of France, he's the husband of Carla Bruni, he's 1.65m tall in his socks, he travelled to Deauville for the G8 summit in May. If you Google "president of France", you'll get two Wikipedia pages on "president of France" then " Nicolas Sarkozy". Useful, but Google doesn't know how the two, Sarkozy and the presidency, are actually related; it's just searching for pages linking to the terms.
Recorded Future ranks all these canonical entities and events, based on the number of references to them, the credibility of the document or document source and several other factors, such as the co-occurrence of different events and entities in the same or in related documents, to create a "momentum" score. Positive or negative sentiment is added to this score. For example, searching big pharma in general will tell you that over the next five years, nine of the world's 15 best-selling medicines will lose patent protection -- the event earns a high momentum score because it is backed by 13 news items from 12 sources -- or that, specifically, Inhibitex, a biopharma business, will need cash in November 2011 if it plans to fund the Phase 2b development of a new drug internally, based on five items from five sources.
Recorded Future isn't the only company attempting to bring hardcore linguistic analysis to a larger audience. Wolfram Alpha is a search engine that can understand a query such as "nuclear explosions in China" and deliver relevant information such as maps and kilotonnes per explosion, although it's culled from "tame" data curated by the company itself. And IBM didn't develop Watson just to school humans on Jeopardy; it's actually a huge research project dedicated to processing questions asked in natural language, based on four terabytes of structured and unstructured data sets, including the full text of Wikipedia. "There are any number of offerings coming on to the market now," says Colin Shearer, senior vice president at SPSS, a predictive-analytics company owned by IBM. One of those is Quid, a two-year-old, 45-strong business founded in 2008. "Human activity has never left an information trail like it does today," says Bob Goodson, its founder. "If only we could harness the intelligence that's locked in the information, we could build systems to understand the world better, and therefore make better decisions." Quid includes Microsoft among its customers. With $15 million in investment, it aims to be the next Bloomberg in business intelligence.
Where Recorded Future goes beyond mere analysis of open data, though, is by adding the "time and space" dimension of the documents -- "references to when and where an event has taken place, or when and where it will take place," says Truvé, "since many documents actually refer to events expected to take place in the future." Using RSS streams allows Recorded Future to have a publishing time as an anchor point for this temporal analysis, which means it can deal with difficult expressions such as "next week", "in three months' time" or "in two quarters". This may sound simple, but it's crucial: the time and space analysis is the first way Recorded Future can make predictions about the future -- by aggregating weighted opinions about the likely timing of future events using algorithmic crowdsourcing. On top of that, it uses statistical models to predict future happenings based on historical records of similar chains of events. "The secret sauce is not dependent upon one ingredient," says Truvé. "It's a combination."
On April 1, 2009, a few months after Ahlberg and Truvé began testing this combination, Ahlberg met Rich Miner, the co-creator of the Android mobile operating system (with Andy Rubin) and a partner at Google Ventures, at the Starbucks on Harvard Square, Massachusetts. Miner was impressed: "We believed there was predictive power in the information contained in the web," he says. "If you can organise that information temporally, then you can look at past and present, and infer things from the future. That's pretty unique so far from Recorded Future." The CIA thought so, too.
In the 40s the allies routinely bombed rail bridges to disrupt supply lines into Nazi-occupied France. After a raid, though, the Royal Air Force couldn't fly reconnaissance missions over the targets as they were considered too risky, so it didn't know if a bridge had been destroyed. The Special Operations Executive (SOE), however, came up with a novel strategy for finding out. By monitoring the daily prices of oranges on sale at various fruit stalls Paris, SOE agents dropped behind enemy lines were able to tell which supply chains had been affected. (Germans embedded in London were doing the same thing; unfortunately for the Nazis, they were under the control of SOE and were fed false information.) This is the differ- ence between information and intelligence: information is the price of oranges, intelli- gence is knowing which supply chain has been affected. This openly available, "free" infor- mation, when it's turned into intelligence, becomes extremely valuable.
"Open-source intelligence has always been crucial, but for most of the cold war it was neglected by western intelligence agencies," says Calder Walton, a research associate at Cambridge University and author of the book Empire of Secrets, to be published in 2013. "That was the archetypal intelligence war: intelligence necessarily involved information that couldn't be gained from any other source -- human agents or telephone tapping." That doesn't mean covert intelligence was more effective, though: Daniel Moynihan, a former US senator, compared CIA reports gathered from secret sources with Soviet documents recovered after the fall of the Berlin Wall and found they significantly overestimated Soviet capabilities. But he discovered that western think tanks using publicly available material, such as the RAND Corporation, were much more accurate. US diplomat George Kennan estimated in 1997 that "95 per cent of what we need to know about foreign countries could very well be obtained by the careful and competent study of perfectly legitimate sources of information open and available to us".
"All of this has changed since the collapse of the Soviet Union," says Walton. "Open-source intelligence has boomed in recent years -- especially since 9/11." At a conference in 2008, Michael Hayden, then director of the CIA, said: "Open-source intelligence contributes to national security in unique and valuable ways virtually every day." Stephen Mercado, an ana- lyst in the CIA directorate of science and technology, estimates that 80 per cent of all valuable intelligence now comes from open sources. In January 2011, Sir Gus O'Donnell, head of the UKcivil service, told the Chilcot inquiry into the invasion of Iraq: "I have strongly and always been of the view that we probably underestimated open source [intelligence]." Open source is the big growth area in intelligence and every western agency is looking for the tools to give it an edge.
Ahlberg refuses to discuss his company's work with the CIA, or even whether there is work with the CIA. In-Q-Tel (IQT) is the CIA's investment arm (mission statement: "Identifies, adapts and delivers innovative technological solutions to support the missions of the Central Intelligence Agency"). It invests only in startup companies that will "provide strong, near-term advantages (within 36 months) to the IC [intelligence community]." IQT doesn't invest without the US secret intelligence services in mind. It backed Recorded Future with slightly less than $2.5 million.
Stephen Davidson, an investor at IQT who sits on Recorded Future's board, refused to comment; a spokesperson for IQT said that "while we are pleased to have Recorded Future as part of the IQT portfolio, we will respectfully decline to provide additional information about our investment". Does Ahlberg know what intelligence purposes Recorded Future is put to? "We would not know about those things," he says, folding his arms. "At this stage, I don't even want to know what people are doing with some of these things." He points out that IQT is "an independent company; at least to my knowledge theycan't force any [government agency] to use it." Truvé, though, says Recorded Future is working with 17 or 18 intelligence agencies. Another board member, Roger Ehrenberg, used to run a $6 billion hedge fund for Deutsche Bank before setting up his own firm, IA Ventures. According to Ehrenberg, In-Q-Tel is "actively involved" with Recorded Future. "Fundamentally, they look to invest in companies where they know they have a customer within the government," he says. "It's not just the CIA." Chris Holden, who works in Recorded Future's Arlington office, admitted to wired (with some understatement) that "we have a little bit of work with the federal government". Holden says that Recorded Future is being used to identify technologies the US government may invest in, such as nanotechnology in body armour. "It's not all super secret stuff necessarily." So, does having IQT as an investor mean thatRecorded Future is beholden to the US government, even if it is a private company? "We are an independent company," repeats Ahlberg. "Neither the US government, nor Google, nor hedge funds nor banks have ever tried to make us do anything. And frankly, you're sitting here with a bunch of Swedes. There's no way in hell you could get them to do anything bad."
Still, it's possible to identify examples of how one might use Recorded Future for open-source intelligence. Take the al-Qaeda leadership after Bin Laden's death: who would fill the vacuum? Recorded Future ran a search. Ayman al-Zawahiri, a founding member of Egypt's Islamic Jihad militant group, and long considered by the US government to be Bin Laden's right-hand man, showed some significant spikes in recorded and discussed activity in the last 12 months,especially when he called for military backing of Libyan rebels, suggesting al-Qaeda could fill a power vacuum in that country. But al-Zawahiri's sentiment score was extremely negative, to the degree that conspiracy theories were emerging that he was responsible for disclosing Bin Laden's location to the US. Saif al-Adel, a senior al-Qaeda commander, was attracting attention back in October 2010, written about as "the new face of al-Qaeda in 2011". Recorded Future concluded that it was "clear that Said al-Adel has been routed in Pakistan for some time now and appears to be embedded in the political structure of al-Qaeda"; his momentum score was high. They also found that Libyan Abu Yahya al-Libi (described by a former CIA analyst as an "insurgent-theologian"), offered access to one of the most volatile regions on the globe right now, based on his current likely location, which al-Qaeda might consider a useful foothold. Finally, they looked at Anwar al-Awlaki, a Yemeni-American imam who posted pro-al-Qaeda/anti-western YouTube videos, and ran a blog and Facebook page. Clearly of interest given the attempted US drone strike to kill him days after Bin Laden's death, he started building momentum in late March and April with reports that he was urging on the Arab Spring protests. Al- Awlaki was killed in Yemen dur- ing a US drone attack on September 30 this year.
Recorded Future concluded that multiple players will rise to prominence regionally, that al-Qaeda could split around al-Zawahiri, and that al-Qaeda sees advantage to be taken in the Arab pro-democracy protests. A couple of months later, al-Zawahiri was confirmed, although experts were sceptical about whether he could unite the membership in Saudi Arabia and the Gulf States behind him.
Ahlberg is more willing to talk about how Recorded Future is being used in finance. "If you take our momentum score, and look across S&P 500 companies, can you predict the liquidity or the stock volume of those companies over time?" asks Ahlberg. "It turns out you can." Stock that is being talked about and is in investors' attention is, of course, more likely to be traded: "It's much easier to prove volume than direction, whether a stock is going up or down." So Recorded Future takes momentum and combines it with sentiment -- whether a company is mentioned in a positive light -- and derives a score. Taking these news bursts across the S&P 500, it can sort them into ten different groups, from high to low. "Then you say, every day, 'I am going to own what is in the top and short what is in the bottom.' You are making lots of small picks on a daily basis - this strategy turns over the portfolio 63 per cent every day."
Running predictive tests on data from January 2009 to January 2011, Recorded Future showed that its top decile has a beta (a measure of risk in portfolio) of 1.08 -- fairly low -- and a statistically significant annualised continuous alpha (a risk-adjusted measure of active return on investment) of +16 per cent. The bottom two deciles had a high beta (1.37 and 1.34, respectively) but with statistically significant negative alphas, at -42 per cent and -26 per cent annually. "Constructing hedged portfolios out of the securities in these deciles provides some compelling trading strategies," says Evan Sparks, an analyst at Recorded Future.
Beyond high-frequency trading strategies, the company says it can predict stock shifts on the basis of one-day events, separated into scheduled events and speculative events. "The theory that if something is written saying, 'on Friday so and so will release earnings', that should be priced into the market immediately," says Ahlberg. "In reality it is not." Recorded Future took 19,000 such events and asked what happens to the stock price. On average, as stocks come into those scheduled events, the prices rise; coming out of them they fall five base points either way. "It's like finding a roulette wheel that is skewed."
Another way is to examine the next two weeks of a particular business's future and look for certain events. One is insiders selling stock. You may think this would be a good time to sell; in fact, insiders often sell just after stocks have already peaked. So Recorded Future looks for data that can be combined with this knowledge. If an insider sells stock after a management lay-off, stock falls on average 1.5 per cent. Expand this event to a whole market and "You have 2,000 events within 2011," says Ahlberg. "By turning it into a big data screen, I have created my own skewed roulette wheel I can consistently bet on." Chris Malloy is an associate professor at Harvard Business School who specialises in behavioural finance. He's played with Recorded Future's data: "I haven't seen anything with that ability. It's pretty neat -- no one's doing that. The predictability is certainly good."
What Recorded Future can't forecast are "black swan events", which are by definition unpredictable and undirected. "You can look at what happens afterwards, though," says Ahlberg. He takes the example of a natural disaster. "Start looking at how other countries behave. After a natural disaster, the US will travel there every time, the UK does it 50 per cent of the time, Iran will do it every single time, China never really does." China did, though, after the 2010 Chilean earthquake. Two months later, it announced a new trade agreement. China didn't travel to Haiti: no trade agreement followed. But it did after Pakistan was hit by flood, soon announcing a $10 billion deal. "We're looking for those historical patterns and using them to predict what might happen," says Ahlberg.
Recorded Future's hedge-fund clients are only slightly less secretive than theCIA. Ehrenberg says a handful of Wall Street hedge funds and banks are using the technology: "Recorded Future is a high-value signal, relative to conventional quantitative-analysis trading signals. People are making money." Josh Holden, CEO of Fina Technologies, which creates algorithms for high-frequency quant trading by hedge funds, says that Recorded Future's client base "is closely guarded. But there are more than a few firms using it.
Sandfire AG is a Swiss consultancy in the public and private security sectors, and a client of Recorded Future. "It helps us keep track of travel routes of high-level decision-makers," says Felix Juhl, a senior partner. "A state visit by a high-ranking politician may be followed by specific corporate activities. Keeping track of travel routes can serve as an early warning."
Ahlberg says Recorded Future now earns revenues in the millions of dollars from a client base of less than 100, but which includes governments, hedge funds, big banks, watchdogs and consultancies. This select clientele place a high value on the distilled insight the company provides. Ahlberg sees a big opportunity: "Even within what we have started around finance and intelligence, there is no reason why we couldn't build another $100 million-revenue company within a small set of years." But Recorded Future plans on being more than just a profitable business tool. Ahlberg is expanding its indexes: he eventually wants every piece of data on the planet streaming live through his company's algorithms. The ultimate goal? "We want to organise the world -- and the internet -- for analysis." What Ahlberg doesn't say, perhaps deliberately, is that ever more data will likely lead to ever more accurate predictions. "It's dangerous to start talking about predicting the future," he says. "We're trying to play that down."
Tom Cheshire is assistant editor at wired. He wrote about the Ariane 5 rocket in 10.11
Monday, November 21, 2011
Brynjolfson in Atlantic: The Big Data Boom Is the Innovation Story of Our Time
See also: Where great ideas really come from. A special report
Erik Brynjolfsson and Andrew McAfee The Atlantic November, 21 2011, 9:50 AM ET
The data revolution has turned customers into unwitting business consultants, as our purchases and searches are tracked to improve everything from websites to delivery routes
In the 1670s, in Delft, Netherlands, a scientist named Anton van Leeuwenhoek did something many scientists had done for 100 years before him. He built a microscope.
This microscope was different, but it was not extraordinary. Like so many inventions, he borrowed and tweaked his predecessors' ingenuity. But when he looked through this microscope, he found things that did seem extraordinary. He called them "animalcules," microbes in water droplets and human blood that ultimately provided the foundation for the germ theory of disease and eventually inspired a host of medicines and treatments.
The Leeuwenhoek discovery is crucial to our understanding of innovation, not only because it changed the face of biochemistry, but also because it represents a fundamental theme of discovery.
Breakthroughs in innovation often rely on breakthroughs in measurement.
THE DATA BOOM
Today businesses can measure their activities and customer relationships with unprecedented precision. As a result, they are awash with data. This is particularly evident in the digital economy, where clickstream data give precisely targeted and real-time insights into consumer behavior.
In turn, customers are acting as unwitting business consultants for these companies. Our purchases, searches, and online activities are being tracked to improve everything from websites to delivery routes and drug manufacturing.
Anyone with access to a Web browser can get summaries of billions of keyword searches, and this information is highly predictive of present and future economic activity, such as housing purchases and prices. Mobile phones, automobiles, factory automation systems and other devices are routinely instrumented to generate streams of data on their activities, making possible an emerging field of "reality mining" to analyze this information.
Manufacturers and retailers use radio-frequency identification (RFID) tags to deliver terabits of data on inventories and supplier interactions and then feed this information into analytical models to optimize and reinvent their business processes. Much of this information is generated for free, by computers, and sits unused, at least initially. A few years after installing a large enterprise resource planning system, it is common for companies to purchase a "business intelligence" module to try to make use of the flood of data that they now have on their operations. As Ron Kohavi at Microsoft memorably put it, objective, fine-grained data are replacing HiPPOs (Highest Paid Person's Opinions) as the basis for decision-making at more and more companies. For example:
-- Enologix has used this approach to help Gallo vineyards accurately predict the wine ratings that Robert Parker would give to various new wines
-- UPS has mined data on truck delivery times to develop a new routing method
-- Match.com as even developed new algorithms for matching men and women for dates
For each innovation, analysts drew on new measurement technologies to supplant human experts who relied more on intuition. However, for all its strengths, measurements have a shortcoming. They cannot determine causality. (A simple example: Shoe sizes and readings scores are correlated for school children, but one does not cause the other; instead, they both reflect a third variable, which is age.) Fortunately, science has a second powerful tool designed precisely to address questions of causality.
That tool is called experimentation.
AN EXPERIMENT EVERY SECOND
Science has been dominated by the experimental approach for nearly 400 years. Running controlled experiments is the gold standard for sorting out cause and effect. But experimentation has been difficult for businesses throughout history because of cost, speed and convenience. It is only recently that businesses have learned to run real-time experiments on their customers. The key enabler was the Web.
Consider two "born-digital" companies, Amazon and Google. A central part of Amazon's research strategy is a program of "A-B" experiments where it develops two versions of its website and offers them to matched samples of customers. Using this method, Amazon might test a new recommendation engine for books, a new service feature, a different check-out process, or simply a different layout or design. Amazon sometimes gets sufficient data within just a few hours to see a statistically significant difference.
This ability to rapidly test ideas fundamentally changes the company's mindset and approach to innovation. Rather than agonize for months over a choice, or model hypothetical scenarios, the company simply asks the customers and get an answer in real time.
According to Google economist Hal Varian, his company is running on the order of 100-200 experiments on any given day, as they test new products and services, new algorithms and alternative designs. An iterative review process aggregates findings and frequently leads to further rounds of more targeted experimentation.
At the same time, Google's competitors, partners, customers and third party consultants are doing their own experiments, creating a complex, interacting ecosystem that demands continuous innovation. While Google currently dominates the market for web search, it is unlikely that it would have any market share at all if it still relied on the original, unmodified PageRank algorithm that Larry Page and Sergey Brin developed in 1998.
HOW DATA TURNED AROUND A CASINO Greg Linden, who led one set of experiments at Amazon, describes the emerging experimentation philosophy succinctly: "To find high impact experiments, you need to try a lot of things. Genius is born from a thousand failures. In each failed test, you learn something that helps you find something that will work. Constant, continuous, ubiquitous experimentation is the most important thing."
These words echo the approach of innovators since Thomas Edison, but IT has made it possible to apply it to a much broader class of business challenges and significantly compress the "hypothesis-to-experiment" cycle time.
While web-based companies have been particularly aggressive in using business experiments to drive innovation, other industries are getting in the game. Caesar's Entertainment (formerly Harrah's), the hotel and casino company, transformed itself from a 2nd-tier casino to an industry leader in large part because of the culture of experimentation introduced by CEO Gary Loveman.
When Loveman, an economics PhD from MIT and former Harvard Business School professor, arrived at the company, he found that it was already gathering a great deal of data about its customer interactions with existing information systems and programs such as its Total Rewards loyalty card. However, it wasn't using these data to develop improved processes, products and services. After becoming CEO, he developed strategies to continually tests new promotions, price points, services, workflow, employee incentive plans and casino layouts using controlled experiments.
Widespread business experimentation has required a fundamental change in the corporate culture. As Loveman puts it "There are two things that will get you fired here: stealing from the company, or running an experiment without a properly designed control group."
***
While passive data gathering can be useful, measurement is far more valuable when coupled with conscious, active experimentation and sharing of insights. Likewise, the value of undertaking the experiments themselves is proportionately greater if the organization can capitalize on those experiments in more locations and at greater scale. In combination, these practices constitute a new kind of "R&D" that draws on the strengths of digitization to speed innovation.
Subscribe to:
Posts (Atom)