To accelerate research breakthroughs on brain diseases, the Allen Institute puts all its data online for use without fees.
Paul Allen The Wall Street Journal November 30, 2011
The Allen Institute for Brain Science in Seattle grew out of a simple question I posed in 2002 to a constellation of top people in the field: What's the most useful thing we could do to propel neuroscience forward? The consensus became our inaugural project—a comprehensive, molecular-level, three-dimensional map of the mouse brain to show precisely where every gene is active, or "expressed." It was the first step on a long road to understand how genes function in the human brain, knowledge that will point to ways to better diagnose and treat brain ailments.
A crucial aspect to this project—and others the Allen Institute has pursued over the last eight years—is an "open science" research model. Early on, we considered charging commercial users for access to our online data. From a strictly financial standpoint, it made sense to reap front-end fees and, down the line, intellectual property royalties. The revenue could cover the high costs of maintenance and development to keep the resource current and useful.
But our mission was to spark breakthroughs, and we didn't want to exclude underfunded neuroscientists who just might be the ones to make the next leap. And so we made all of our data free, with no registration required. The Institute would have no gatekeeper. Our terms-of-use agreement is about 10% as long as the one governing iTunes.
Our facility is neither the first nor the last to use a shared database to embrace "open science" and reject the competitive, single-lab R&D paradigm. Traditional research incentives—where journal publications are the coin of the realm—tend to discourage vital sharing.
In 1982, even before the dawn of the Internet, a consortium of government agencies established the open access GenBank. Maintained by a division of the National Institutes of Health (NIH), GenBank now houses the sequence data from the Human Genome Project, the inspiration for our brain mapping.
In recent years the NIH has sponsored other data-sharing portals, including the Alzheimer's Disease Neuroimaging Initiative and the Neuroscience Information Framework. Private nonprofits like the Pistoia Alliance and Sage Bionetworks are curating their own open-source repositories.
But the Allen Institute remains distinct in conducting industrial-scale big science that is fundamentally collaborative. Internally, our team of scientists and support staff works together to meet the time lines and milestones that frame each large project. The team released the initial data set from a ground-breaking human whole-brain atlas last year, and it is now midstream on a project to define the circuitry between neurons and how it affects human behavior. Most important, we generate data for the purpose of sharing it. Since opening shop in 2003, we've had 23 public releases, or about three per year. We don't wait to analyze our raw data and publish in the literature. We pour it onto the public website as soon as it passes our quality control checks. Our goal is to speed others' discoveries as much as to springboard our own future research.
The databases currently provide tens of millions of high-resolution images. The initial mouse brain atlas alone involved 600 terabytes of data, or 600 trillion bytes, more than half the total content of the Internet when we started. Since data of this volume would be of little use without effective search and navigation tools, the Institute developed a free online viewing application as well as the downloadable Brain Explorer 3D viewer, which illuminates how expressed genes are distributed throughout the brain.
Open science is a long-term and pricey proposition. It demands consistent curating, maintenance and updating of databases, and regular software and hardware upgrades. The institute offers online video tutorials on a YouTube channel and in a tutorial library. For those seeking in-person walk-throughs or forums, it hosts training workshops and user group sessions in several areas around the country each year. These services, too, are free of charge.
It is a modest cost that is paying off as the scientific community embraces the open access model. In October, the institute's suite of databases received more than 45,000 visits, from six continents and from research organizations of every stripe: universities, government laboratories, independent institutes and biotech and pharmaceutical companies. Institute brain atlases are accelerating research on the underlying biology of a broad range of diseases, from Alzheimer's and Parkinson's to autism and schizophrenia. Growing numbers of college educators, from UCLA to the Radboud University Nijmegen in the Netherlands, are building curricular modules around our online resources.
What I've concluded is that foundations and other private funders who support scientific research also can help promote wider sharing of scientific data. Before funders write a check to a university, they should ask about the researcher's policies and track record on sharing.
On the federal level, the NIH now has such strong policies on sharing data. But I'd like to see the agency do even more to put its funding where its directives are. I propose that the NIH—along with the National Science Foundation and the U.S. Department of Education—direct funding into grant awards for management and curation of existing research data of special value.
That would siphon some money for traditional research grants for new work. But I think we'd get more bang for our buck by making more data more useful to more scientists—and, by extension, to the world community that will benefit from their work.
Mr. Allen, the co-founder of Microsoft with Bill Gates, launched the nonprofit Allen Institute for Brain Science in 2003.
Wednesday, November 30, 2011
Tuesday, November 29, 2011
NPR on Big Data
The Digital Breadcrumbs That Lead To Big Data
Yuki Noguchi NPR November 29, 2011
First of a Two Part Story
First of a Two Part Story
What do Facebook, Groupon and biotech firm Human Genome Sciences have in common? They all rely on massive amounts of data to design their products. Terabytes and even zettabytes of information about consumers or about genetic sequences can be harnessed and crunched.
The practice is called big data, and as the term suggests, it is huge in both scope and power. Analyzing big data enables anything from predicting prices to catching criminals, and has the potential to impact many industries.
One way to understand how big data works is to think about your daily life. You write an email, call your boss, pass a security camera, maybe buy a plane ticket online. Taken alone, this is disjointed, boring information. To Elizabeth Charnock, it makes up your digital character.
We have seen the industrial revolution, and we are witnessing a data revolution.
- Oren Etzioni, professor of computer science at the University of Washington
"Digital character is this idea that almost everybody these days leaves behind a giant digital breadcrumb trail," she says.
Charnock founded Cataphora, a company that can process huge amounts of this sort of data about employees to determine patterns. She says those patterns can predict everything from a person's mood to their skill as a manager to a person's inclination to commit fraud.
Take rogue trader Jerome Kerviel, who cost his French bank billions of dollars in losses.
"His cell phone bill was literally an order of magnitude larger than any of his coworkers — why?" Charnock asks. "Well, because he wanted to put less things in writing. He almost never took vacation, even though French people love to take vacation."
Charnock says Kerviel also circumvented usual trading and communication protocols.
If you've got your eye on that brand new camera with all the features you never imagined you'd need, how do you know if it's time to buy? Decide.com is a prediction tool that tells you when gadget prices are likely to rise, stay the same or fall – and if you should wait a few weeks to buy the rumored newer model.
Decide collects prices of over 100,000 electronic products every day from hundreds of online retailers. It also searches technology blogs for rumors of upcoming new releases, adding up to over 25 GB of data per day.
Decide's four computer science Ph.D's create algorithms to mine the data and predict whether prices will go up or down, similar to what the finance industry has done for years to forecast stock prices. But now that data storage has become so cheap, other businesses can get in the game.
The practice is called big data, and as the term suggests, it is huge in both scope and power. Analyzing big data enables anything from predicting prices to catching criminals, and has the potential to impact many industries.
One way to understand how big data works is to think about your daily life. You write an email, call your boss, pass a security camera, maybe buy a plane ticket online. Taken alone, this is disjointed, boring information. To Elizabeth Charnock, it makes up your digital character.
We have seen the industrial revolution, and we are witnessing a data revolution.
- Oren Etzioni, professor of computer science at the University of Washington
"Digital character is this idea that almost everybody these days leaves behind a giant digital breadcrumb trail," she says.
Charnock founded Cataphora, a company that can process huge amounts of this sort of data about employees to determine patterns. She says those patterns can predict everything from a person's mood to their skill as a manager to a person's inclination to commit fraud.
Take rogue trader Jerome Kerviel, who cost his French bank billions of dollars in losses.
"His cell phone bill was literally an order of magnitude larger than any of his coworkers — why?" Charnock asks. "Well, because he wanted to put less things in writing. He almost never took vacation, even though French people love to take vacation."
Charnock says Kerviel also circumvented usual trading and communication protocols.
If you've got your eye on that brand new camera with all the features you never imagined you'd need, how do you know if it's time to buy? Decide.com is a prediction tool that tells you when gadget prices are likely to rise, stay the same or fall – and if you should wait a few weeks to buy the rumored newer model.
Decide collects prices of over 100,000 electronic products every day from hundreds of online retailers. It also searches technology blogs for rumors of upcoming new releases, adding up to over 25 GB of data per day.
The data is sent to Amazon's cloud storage and processed with the help of Hadoop software, which is used for many big data projects to organize the information and minimize mistakes.
Decide's four computer science Ph.D's create algorithms to mine the data and predict whether prices will go up or down, similar to what the finance industry has done for years to forecast stock prices. But now that data storage has become so cheap, other businesses can get in the game.
In addition to crunching the numbers, Decide analyzes thousands of blog posts and press releases to see if any rumors have surfaced. Sources that were reliable in the past are listened to more closely.
All in all, Decide has nearly 100 terabytes of data to analyze and end up with a simple prediction: Should you buy now or wait for prices to drop
—Sara Carothers, Stephanie d'Otreppe/NPR
But big data is not just about connecting dots to detect crime. The ability to process so much information and process it so quickly makes all kinds of things possible that weren't before. So LinkedIn finds jobs or people you might like to know about, and biotech companies can analyze gene sequences in billions of combinations to design drugs.
Data analytics itself is not new. Two decades ago, Wall Street hired teams of physicists to analyze investments. But in the last couple of years, computing, storage and bandwidth capacity have become so cheap that it has altered the scale of what's possible.
Now, with very little money, a gifted student or a small startup can design big-data applications.
"Everywhere you look, there's an opportunity to collect more data and then apply a statistical or mathematical approach to understanding what's happening," says Chris Kemp, chief executive officer of Nebula, a firm that provides storage and computing capacity for other companies to be able to process their big data applications.
Kemp says ultimately big data will give consumers better tools so they can do a better job of predicting things like prices, such as whether an airfare is likely to go up or down. Farmers can do a better job of insuring their crops if they can forecast the weather with greater accuracy.
Oren Etzioni, a professor of computer science at the University of Washington, says this trend is fueling intense demand for mathematics and computing talent.
"We have seen the industrial revolution, and we are witnessing a data revolution," Etzioni says.
He's started three big-data companies. One of them, Decide.com, employs four Ph.D.s to design better programs to forecast prices on consumer electronics.
Etzioni says a good data scientist can write algorithms that filter data, understand what it's telling you, and then graphically represent it. The end result is like getting a bird's-eye view of a vast territory of information.
Big data can, and occasionally does, go wrong. Comic examples of that include mismatched recommendations, like "My TiVo thinks I'm gay." "But think about a company divulging your Web surfing history with your name attached and you begin to get a sense of how big data opens the door to new possibilities of security or privacy breaches.
James Slavet, a venture capitalist at Greylock Partners, says his firm invests in companies that use big data creatively and responsibly. He says data does not stand in for human judgment.
"They do use it to make the judgment more sound, more objective and to hopefully lead to better decision making," he says.
—Sara Carothers, Stephanie d'Otreppe/NPR
"Any one of those things, you kind of say, 'So what?' But what we look for is a number of them that on the surface perhaps don't seem to be related but all seem to be happening at the same time," she says.
Charnock says had the French bank analyzed that data, it might have flagged the rogue trader earlier.But big data is not just about connecting dots to detect crime. The ability to process so much information and process it so quickly makes all kinds of things possible that weren't before. So LinkedIn finds jobs or people you might like to know about, and biotech companies can analyze gene sequences in billions of combinations to design drugs.
Data analytics itself is not new. Two decades ago, Wall Street hired teams of physicists to analyze investments. But in the last couple of years, computing, storage and bandwidth capacity have become so cheap that it has altered the scale of what's possible.
Now, with very little money, a gifted student or a small startup can design big-data applications.
"Everywhere you look, there's an opportunity to collect more data and then apply a statistical or mathematical approach to understanding what's happening," says Chris Kemp, chief executive officer of Nebula, a firm that provides storage and computing capacity for other companies to be able to process their big data applications.
Kemp says ultimately big data will give consumers better tools so they can do a better job of predicting things like prices, such as whether an airfare is likely to go up or down. Farmers can do a better job of insuring their crops if they can forecast the weather with greater accuracy.
Oren Etzioni, a professor of computer science at the University of Washington, says this trend is fueling intense demand for mathematics and computing talent.
"We have seen the industrial revolution, and we are witnessing a data revolution," Etzioni says.
He's started three big-data companies. One of them, Decide.com, employs four Ph.D.s to design better programs to forecast prices on consumer electronics.
Etzioni says a good data scientist can write algorithms that filter data, understand what it's telling you, and then graphically represent it. The end result is like getting a bird's-eye view of a vast territory of information.
Big data can, and occasionally does, go wrong. Comic examples of that include mismatched recommendations, like "My TiVo thinks I'm gay." "But think about a company divulging your Web surfing history with your name attached and you begin to get a sense of how big data opens the door to new possibilities of security or privacy breaches.
James Slavet, a venture capitalist at Greylock Partners, says his firm invests in companies that use big data creatively and responsibly. He says data does not stand in for human judgment.
"They do use it to make the judgment more sound, more objective and to hopefully lead to better decision making," he says.
Slavet calls big data a tectonic shift, one that will continue to affect many things we do for decades to come.
Listen to the Story (and see the Graphs) at http://www.npr.org/2011/11/29/142521910/the-digital-breadcrumbs-that-lead-to-big-data
Tuesday, November 22, 2011
Big Data: The news forecast (Wired)
Tom Cheshire Wired December 2011
See gallery of illustrations at http://www.wired.co.uk/magazine/archive/2011/12/features/the-news-forecast/viewgallery#!image-number=1
In September 2010 the Yemen ministry of industry announced a national strategy to combat food shortages. The UN Food Price Index had reached consecutive record highs in the previous few months. Yemen had also suffered flooding, which had killed around 100 people and disrupted farming. The strategy, which included a review of existing subsidies and the development of food-for-work programmes, proved ineffective. By December 2010, concerned that protests over rising food prices were starting to grow in Tunisia, Yemen's president, Ali Abdullah Saleh, halved income tax and ordered the government to control the prices of basic commodities.
By late January, though, thousands of protesters had taken to the streets to demand Saleh's resignation, brandishing flatbread with baked-in slogans and wearing the food as helmets. The clashes continued into February and grew more violent as the UN's Food Price Index reached an all-time high. On March 18, 45 protesters were killed when an unidentified gunman opened fire. Six days later, the government fought a battle with al-Qaeda gunmen in the province of Abyan and Marib, killing 15. The same day, 10,000 protesters gathered in the capital city Sana'a. Saleh said that he would accept the opposition's transition plan that day, but clung on for another month. He finally quit Yemen for Saudi Arabia after a bomb planted in the presidential compound exploded, killing seven people; Saleh suffered 40 per cent burns, shrapnel wounds and internal bleeding. In total, Human Rights Watch estimates that 233 protesters were killed on the streets. Three months later, Saleh unexpectedly flew back to Yemen; 100 more protesters and tribesmen were killed in the first five days of his return and the situation remains unresolved.
A year earlier, on January 12, 2010, a tech startup posted an article on its blog: "Yemen heading for disaster in 2010?" The author, "Ninja Shoes", wrote: "Based on the information we've gathered, Yemen will likely experience food shortages and torrential floods in 2010. This combination of natural disasters, propensity for famine and malnutrition, and challenges with Islamic radicals and terrorists, make it a hot spot for conflict in the future."
The 20 employees of Recorded Future aren't foreign-policy experts. They aren't traders either, but if you'd started using Recorded Future's predictions to buy US stocks on January 1, 2009, you would have made an annual return of 56.69 per cent. (The S&P 500 had an annualised return of 17.22 per cent over the same period.) Between May 13 and August 5 this year, as markets behaved with vertiginous abandon, their strategy returned 10.4 per cent; in contrast, the S&P 500 lost 9.9 per cent of its value. They're data experts: computer scientists, statisticians and experts in linguistics. And in the data, they think, lies the future.
All Recorded Future's predictions, whatever the field, are based on publicly available information -- news articles, government sites, financial reports, tweets -- fed into the company's own algorithms. The result, it claims, is a "new tool that allows you to visualise the future" -- one that is changing how government intelligence agencies gather information and how giant hedge funds place bets. On its website, Recorded Future states: "We don't grant interviews and we don't issue press releases." But behind closed doors, the company is developing the technology that has been described be one tech blog as an "information weapon".
The company, cofounded by Christopher Ahlberg, an entrepreneur who sold his first business for $195 million and served in the Swedish special forces, has $8.5 million in funding. Its first two investors were Google and the CIA. Recorded Future counts US government agencies, banks and hedge funds among the clients paying million-dollar contracts. But its true ambition is to organise all the data on the internet for similar predictive analysis -- to make the future calculable.
Recorded Future's main office is in Gothenburg, Sweden. On a drab morning in May, trams clang past a metal door that doesn't bear the company's name. Two flights of stairs lead to a wooden door, with a discreet sticker label-gunned above the letterbox in caps: "RECORDED FUTURE". The rooms date from the 17th century; they're airy and bright with high ceilings and intricate plaster mouldings. Eight employees work here on the technical aspects of the system. The company also has offices in Boston, New York and Arlington, Virginia -- ten minutes' drive from the Pentagon, 15 from Langley.
"Yemen took four or five months longer than we predicted," says Ahlberg, 43, sitting on a sofa in a small meeting room. Before Wired visited, he warned over the telephone: "You won't get a government agency out of my mouth. Dude, if I do that, they're coming to take my kids." In person, he's tall, with hair cropped short, and is quick to laugh. The telephone caveat still stands, but Ahlberg is willing to talk for the first time about what exactly it is his company does and why Google, intelligence agencies and hedge funds are all so interested.
Ahlberg was born in September 1968, in Kungälv, a town 40 minutes' drive north of Gothenburg. His father was a captain on merchant ships, his mother taught French and English in Sweden. In his first year at secondary school, he created a drawing program on his Sinclair Spectrum called Art CAD ("like an early version of Photoshop") and sold individual copies by advertising it in the local paper. After school, he wanted to study computer science but first had to complete military service in 1987. He chose the Lapplands Jägarregemente special forces, and began training for a hypothetical Russian invasion: "They would come in from Finland and go to Norway; we were supposed to cut them off in the middle. We were supposed to do what the Iraqis are doing now, guerrilla warfare. But we were master cross-country skiers."
Ahlberg then went onto take his degree at Chalmers University of Technology in Gothenburg. As a post doc, he travelled to the University of Maryland to work as a visiting researcher at the Human-Computer Interaction Lab for two summers. During the first visit, at 23, he co-authored a published paper with the director of the lab; the second summer, he co-wrote two, about the new field of data visualisation. Ahlberg returned to Sweden, finishing his PhD in four years instead of six, but it was his work at Maryland that formed the basis for his first company, Spotfire. Launched in 1996, the business created visualisation tools for business intelligence; in 2007, it was bought by Tibco for $195 million(£125m). "I didn't have to work anymore," says Ahlberg. "But I can't stop."
Spotfire had helped businesses visualise internal databases. After the sale, "we started hanging around in coffee shops in Boston and New York", says Ahlberg. "We thought: what's the most interesting data source out there? And it's nebulous, but the web is the most interesting dataset there is on the planet. Instead of just corporate databases, let's think about the web as my data source." Ahlberg began talking with Staffan Truvé, who had supervised his PhD and started Spotfire with him. "It was in the back of my head that, as humans, we had generally started to become better at predicting things," says 48-year-old Truvé. "Your car tells you that you need to change your oil in 200km, or there is a sign saying your bus is coming in five minutes. These tiny predictive signals are popping up everywhere." Ahlberg was excited: "So then the premise becomes that the web has predictive power. How can we harvest that?"
A few isolated, eye-catching examples have shown the prognostic possibilities in such data. In 2008, Google showed search queries could accurately predict the spread of flu in the US up to two weeks before the federal Centers for Disease Control. In his book, Super Crunchers, Ian Ayers claimed that creditcard companies can predict with 98 per cent accuracy whether you'll divorce, based on your purchases -- Google's Marissa Mayer even quoted the statistic at SXSW 2011 (however, in a recent statement, Visa denied that it monitored such data or made any such conclusions, saying the claim was "inaccurate and wrong"). One recent study has shown Twitter to be 88.67 per cent accurate in predicting the Dow Jones three days in advance. This July, financier Paul Hawtin founded a London hedge fund that is based entirely on social media. And in September, a researcher from the University of Illinois fed the Nautilus supercomputer with 100 million news items, much like Recorded Future does, and "anticipated" the Arab Spring and the killing of Osama Bin Laden, albeit retrospectively -- a prediction of the past. But for Ahlberg to develop a tool that could create predictions for any input, from finance to terrorism, would be much harder. Recorded Future would not only have to index the internet, but also understand and interpret it.
The first generation of search engines, such as Lycos and Alta Vista, used traditional text search to deliver web pages, deploying their own algorithms, but essentially looking at individual documents in isolation. Google changed this in 1998. Its PageRank algorithm analysed the links between web pages, promoting those that had more links pointing to them from other sites. Recorded Future is part of the third generation: instead of explicit link analysis, it examines implicit links -- what it calls "invisible links" between documents that refer to the same entities or events. It does this by separating the documents and their content from what they talk about, identifying canonical entities and events that exist outside of the article.
"What matters is that it's freaking complicated," says Ahlberg. In practice, Recorded Future harvests 25,000 data sources as RSS feeds, which could include Companies House and US Securities and Exchange Commission filings, a New York Times article, Twitter and Facebook posts, obscure blogs (there's one on Norwegian salmon fishing) or transcripts from earnings calls or political speeches -- "just a flood of stuff", says Ahlberg. It does the same for Chinese and Arabic sources. "Then we look for entities -- people, places, technologies; and events -- a murder, a bomb explosion, a person moving from A to B, product launches."
This linguistic analysis is "really tough", according to Truvé. Because Recorded Future takes sources from all over the internet, rather than a particular data set, "the data is not so nice". "We could have built a perfect data set around Pfizer, say, or Barack Obama," says Ahlberg. "It's harder then to think of the big picture. So we tried to make this ambitious." Recorded Future currently uses two separate algorithms, one proprietary, one licensed, to analyse language; the staffers in Gothenburg are tweaking them continually to see which works better. But the result is that Recorded Future knows who Nicolas Sarkozy is, say: that he's the president of France, he's the husband of Carla Bruni, he's 1.65m tall in his socks, he travelled to Deauville for the G8 summit in May. If you Google "president of France", you'll get two Wikipedia pages on "president of France" then " Nicolas Sarkozy". Useful, but Google doesn't know how the two, Sarkozy and the presidency, are actually related; it's just searching for pages linking to the terms.
Recorded Future ranks all these canonical entities and events, based on the number of references to them, the credibility of the document or document source and several other factors, such as the co-occurrence of different events and entities in the same or in related documents, to create a "momentum" score. Positive or negative sentiment is added to this score. For example, searching big pharma in general will tell you that over the next five years, nine of the world's 15 best-selling medicines will lose patent protection -- the event earns a high momentum score because it is backed by 13 news items from 12 sources -- or that, specifically, Inhibitex, a biopharma business, will need cash in November 2011 if it plans to fund the Phase 2b development of a new drug internally, based on five items from five sources.
Recorded Future isn't the only company attempting to bring hardcore linguistic analysis to a larger audience. Wolfram Alpha is a search engine that can understand a query such as "nuclear explosions in China" and deliver relevant information such as maps and kilotonnes per explosion, although it's culled from "tame" data curated by the company itself. And IBM didn't develop Watson just to school humans on Jeopardy; it's actually a huge research project dedicated to processing questions asked in natural language, based on four terabytes of structured and unstructured data sets, including the full text of Wikipedia. "There are any number of offerings coming on to the market now," says Colin Shearer, senior vice president at SPSS, a predictive-analytics company owned by IBM. One of those is Quid, a two-year-old, 45-strong business founded in 2008. "Human activity has never left an information trail like it does today," says Bob Goodson, its founder. "If only we could harness the intelligence that's locked in the information, we could build systems to understand the world better, and therefore make better decisions." Quid includes Microsoft among its customers. With $15 million in investment, it aims to be the next Bloomberg in business intelligence.
Where Recorded Future goes beyond mere analysis of open data, though, is by adding the "time and space" dimension of the documents -- "references to when and where an event has taken place, or when and where it will take place," says Truvé, "since many documents actually refer to events expected to take place in the future." Using RSS streams allows Recorded Future to have a publishing time as an anchor point for this temporal analysis, which means it can deal with difficult expressions such as "next week", "in three months' time" or "in two quarters". This may sound simple, but it's crucial: the time and space analysis is the first way Recorded Future can make predictions about the future -- by aggregating weighted opinions about the likely timing of future events using algorithmic crowdsourcing. On top of that, it uses statistical models to predict future happenings based on historical records of similar chains of events. "The secret sauce is not dependent upon one ingredient," says Truvé. "It's a combination."
On April 1, 2009, a few months after Ahlberg and Truvé began testing this combination, Ahlberg met Rich Miner, the co-creator of the Android mobile operating system (with Andy Rubin) and a partner at Google Ventures, at the Starbucks on Harvard Square, Massachusetts. Miner was impressed: "We believed there was predictive power in the information contained in the web," he says. "If you can organise that information temporally, then you can look at past and present, and infer things from the future. That's pretty unique so far from Recorded Future." The CIA thought so, too.
In the 40s the allies routinely bombed rail bridges to disrupt supply lines into Nazi-occupied France. After a raid, though, the Royal Air Force couldn't fly reconnaissance missions over the targets as they were considered too risky, so it didn't know if a bridge had been destroyed. The Special Operations Executive (SOE), however, came up with a novel strategy for finding out. By monitoring the daily prices of oranges on sale at various fruit stalls Paris, SOE agents dropped behind enemy lines were able to tell which supply chains had been affected. (Germans embedded in London were doing the same thing; unfortunately for the Nazis, they were under the control of SOE and were fed false information.) This is the differ- ence between information and intelligence: information is the price of oranges, intelli- gence is knowing which supply chain has been affected. This openly available, "free" infor- mation, when it's turned into intelligence, becomes extremely valuable.
"Open-source intelligence has always been crucial, but for most of the cold war it was neglected by western intelligence agencies," says Calder Walton, a research associate at Cambridge University and author of the book Empire of Secrets, to be published in 2013. "That was the archetypal intelligence war: intelligence necessarily involved information that couldn't be gained from any other source -- human agents or telephone tapping." That doesn't mean covert intelligence was more effective, though: Daniel Moynihan, a former US senator, compared CIA reports gathered from secret sources with Soviet documents recovered after the fall of the Berlin Wall and found they significantly overestimated Soviet capabilities. But he discovered that western think tanks using publicly available material, such as the RAND Corporation, were much more accurate. US diplomat George Kennan estimated in 1997 that "95 per cent of what we need to know about foreign countries could very well be obtained by the careful and competent study of perfectly legitimate sources of information open and available to us".
"All of this has changed since the collapse of the Soviet Union," says Walton. "Open-source intelligence has boomed in recent years -- especially since 9/11." At a conference in 2008, Michael Hayden, then director of the CIA, said: "Open-source intelligence contributes to national security in unique and valuable ways virtually every day." Stephen Mercado, an ana- lyst in the CIA directorate of science and technology, estimates that 80 per cent of all valuable intelligence now comes from open sources. In January 2011, Sir Gus O'Donnell, head of the UKcivil service, told the Chilcot inquiry into the invasion of Iraq: "I have strongly and always been of the view that we probably underestimated open source [intelligence]." Open source is the big growth area in intelligence and every western agency is looking for the tools to give it an edge.
Ahlberg refuses to discuss his company's work with the CIA, or even whether there is work with the CIA. In-Q-Tel (IQT) is the CIA's investment arm (mission statement: "Identifies, adapts and delivers innovative technological solutions to support the missions of the Central Intelligence Agency"). It invests only in startup companies that will "provide strong, near-term advantages (within 36 months) to the IC [intelligence community]." IQT doesn't invest without the US secret intelligence services in mind. It backed Recorded Future with slightly less than $2.5 million.
Stephen Davidson, an investor at IQT who sits on Recorded Future's board, refused to comment; a spokesperson for IQT said that "while we are pleased to have Recorded Future as part of the IQT portfolio, we will respectfully decline to provide additional information about our investment". Does Ahlberg know what intelligence purposes Recorded Future is put to? "We would not know about those things," he says, folding his arms. "At this stage, I don't even want to know what people are doing with some of these things." He points out that IQT is "an independent company; at least to my knowledge theycan't force any [government agency] to use it." Truvé, though, says Recorded Future is working with 17 or 18 intelligence agencies. Another board member, Roger Ehrenberg, used to run a $6 billion hedge fund for Deutsche Bank before setting up his own firm, IA Ventures. According to Ehrenberg, In-Q-Tel is "actively involved" with Recorded Future. "Fundamentally, they look to invest in companies where they know they have a customer within the government," he says. "It's not just the CIA." Chris Holden, who works in Recorded Future's Arlington office, admitted to wired (with some understatement) that "we have a little bit of work with the federal government". Holden says that Recorded Future is being used to identify technologies the US government may invest in, such as nanotechnology in body armour. "It's not all super secret stuff necessarily." So, does having IQT as an investor mean thatRecorded Future is beholden to the US government, even if it is a private company? "We are an independent company," repeats Ahlberg. "Neither the US government, nor Google, nor hedge funds nor banks have ever tried to make us do anything. And frankly, you're sitting here with a bunch of Swedes. There's no way in hell you could get them to do anything bad."
Still, it's possible to identify examples of how one might use Recorded Future for open-source intelligence. Take the al-Qaeda leadership after Bin Laden's death: who would fill the vacuum? Recorded Future ran a search. Ayman al-Zawahiri, a founding member of Egypt's Islamic Jihad militant group, and long considered by the US government to be Bin Laden's right-hand man, showed some significant spikes in recorded and discussed activity in the last 12 months,especially when he called for military backing of Libyan rebels, suggesting al-Qaeda could fill a power vacuum in that country. But al-Zawahiri's sentiment score was extremely negative, to the degree that conspiracy theories were emerging that he was responsible for disclosing Bin Laden's location to the US. Saif al-Adel, a senior al-Qaeda commander, was attracting attention back in October 2010, written about as "the new face of al-Qaeda in 2011". Recorded Future concluded that it was "clear that Said al-Adel has been routed in Pakistan for some time now and appears to be embedded in the political structure of al-Qaeda"; his momentum score was high. They also found that Libyan Abu Yahya al-Libi (described by a former CIA analyst as an "insurgent-theologian"), offered access to one of the most volatile regions on the globe right now, based on his current likely location, which al-Qaeda might consider a useful foothold. Finally, they looked at Anwar al-Awlaki, a Yemeni-American imam who posted pro-al-Qaeda/anti-western YouTube videos, and ran a blog and Facebook page. Clearly of interest given the attempted US drone strike to kill him days after Bin Laden's death, he started building momentum in late March and April with reports that he was urging on the Arab Spring protests. Al- Awlaki was killed in Yemen dur- ing a US drone attack on September 30 this year.
Recorded Future concluded that multiple players will rise to prominence regionally, that al-Qaeda could split around al-Zawahiri, and that al-Qaeda sees advantage to be taken in the Arab pro-democracy protests. A couple of months later, al-Zawahiri was confirmed, although experts were sceptical about whether he could unite the membership in Saudi Arabia and the Gulf States behind him.
Ahlberg is more willing to talk about how Recorded Future is being used in finance. "If you take our momentum score, and look across S&P 500 companies, can you predict the liquidity or the stock volume of those companies over time?" asks Ahlberg. "It turns out you can." Stock that is being talked about and is in investors' attention is, of course, more likely to be traded: "It's much easier to prove volume than direction, whether a stock is going up or down." So Recorded Future takes momentum and combines it with sentiment -- whether a company is mentioned in a positive light -- and derives a score. Taking these news bursts across the S&P 500, it can sort them into ten different groups, from high to low. "Then you say, every day, 'I am going to own what is in the top and short what is in the bottom.' You are making lots of small picks on a daily basis - this strategy turns over the portfolio 63 per cent every day."
Running predictive tests on data from January 2009 to January 2011, Recorded Future showed that its top decile has a beta (a measure of risk in portfolio) of 1.08 -- fairly low -- and a statistically significant annualised continuous alpha (a risk-adjusted measure of active return on investment) of +16 per cent. The bottom two deciles had a high beta (1.37 and 1.34, respectively) but with statistically significant negative alphas, at -42 per cent and -26 per cent annually. "Constructing hedged portfolios out of the securities in these deciles provides some compelling trading strategies," says Evan Sparks, an analyst at Recorded Future.
Beyond high-frequency trading strategies, the company says it can predict stock shifts on the basis of one-day events, separated into scheduled events and speculative events. "The theory that if something is written saying, 'on Friday so and so will release earnings', that should be priced into the market immediately," says Ahlberg. "In reality it is not." Recorded Future took 19,000 such events and asked what happens to the stock price. On average, as stocks come into those scheduled events, the prices rise; coming out of them they fall five base points either way. "It's like finding a roulette wheel that is skewed."
Another way is to examine the next two weeks of a particular business's future and look for certain events. One is insiders selling stock. You may think this would be a good time to sell; in fact, insiders often sell just after stocks have already peaked. So Recorded Future looks for data that can be combined with this knowledge. If an insider sells stock after a management lay-off, stock falls on average 1.5 per cent. Expand this event to a whole market and "You have 2,000 events within 2011," says Ahlberg. "By turning it into a big data screen, I have created my own skewed roulette wheel I can consistently bet on." Chris Malloy is an associate professor at Harvard Business School who specialises in behavioural finance. He's played with Recorded Future's data: "I haven't seen anything with that ability. It's pretty neat -- no one's doing that. The predictability is certainly good."
What Recorded Future can't forecast are "black swan events", which are by definition unpredictable and undirected. "You can look at what happens afterwards, though," says Ahlberg. He takes the example of a natural disaster. "Start looking at how other countries behave. After a natural disaster, the US will travel there every time, the UK does it 50 per cent of the time, Iran will do it every single time, China never really does." China did, though, after the 2010 Chilean earthquake. Two months later, it announced a new trade agreement. China didn't travel to Haiti: no trade agreement followed. But it did after Pakistan was hit by flood, soon announcing a $10 billion deal. "We're looking for those historical patterns and using them to predict what might happen," says Ahlberg.
Recorded Future's hedge-fund clients are only slightly less secretive than theCIA. Ehrenberg says a handful of Wall Street hedge funds and banks are using the technology: "Recorded Future is a high-value signal, relative to conventional quantitative-analysis trading signals. People are making money." Josh Holden, CEO of Fina Technologies, which creates algorithms for high-frequency quant trading by hedge funds, says that Recorded Future's client base "is closely guarded. But there are more than a few firms using it.
Sandfire AG is a Swiss consultancy in the public and private security sectors, and a client of Recorded Future. "It helps us keep track of travel routes of high-level decision-makers," says Felix Juhl, a senior partner. "A state visit by a high-ranking politician may be followed by specific corporate activities. Keeping track of travel routes can serve as an early warning."
Ahlberg says Recorded Future now earns revenues in the millions of dollars from a client base of less than 100, but which includes governments, hedge funds, big banks, watchdogs and consultancies. This select clientele place a high value on the distilled insight the company provides. Ahlberg sees a big opportunity: "Even within what we have started around finance and intelligence, there is no reason why we couldn't build another $100 million-revenue company within a small set of years." But Recorded Future plans on being more than just a profitable business tool. Ahlberg is expanding its indexes: he eventually wants every piece of data on the planet streaming live through his company's algorithms. The ultimate goal? "We want to organise the world -- and the internet -- for analysis." What Ahlberg doesn't say, perhaps deliberately, is that ever more data will likely lead to ever more accurate predictions. "It's dangerous to start talking about predicting the future," he says. "We're trying to play that down."
Tom Cheshire is assistant editor at wired. He wrote about the Ariane 5 rocket in 10.11
See gallery of illustrations at http://www.wired.co.uk/magazine/archive/2011/12/features/the-news-forecast/viewgallery#!image-number=1
In September 2010 the Yemen ministry of industry announced a national strategy to combat food shortages. The UN Food Price Index had reached consecutive record highs in the previous few months. Yemen had also suffered flooding, which had killed around 100 people and disrupted farming. The strategy, which included a review of existing subsidies and the development of food-for-work programmes, proved ineffective. By December 2010, concerned that protests over rising food prices were starting to grow in Tunisia, Yemen's president, Ali Abdullah Saleh, halved income tax and ordered the government to control the prices of basic commodities.
By late January, though, thousands of protesters had taken to the streets to demand Saleh's resignation, brandishing flatbread with baked-in slogans and wearing the food as helmets. The clashes continued into February and grew more violent as the UN's Food Price Index reached an all-time high. On March 18, 45 protesters were killed when an unidentified gunman opened fire. Six days later, the government fought a battle with al-Qaeda gunmen in the province of Abyan and Marib, killing 15. The same day, 10,000 protesters gathered in the capital city Sana'a. Saleh said that he would accept the opposition's transition plan that day, but clung on for another month. He finally quit Yemen for Saudi Arabia after a bomb planted in the presidential compound exploded, killing seven people; Saleh suffered 40 per cent burns, shrapnel wounds and internal bleeding. In total, Human Rights Watch estimates that 233 protesters were killed on the streets. Three months later, Saleh unexpectedly flew back to Yemen; 100 more protesters and tribesmen were killed in the first five days of his return and the situation remains unresolved.
A year earlier, on January 12, 2010, a tech startup posted an article on its blog: "Yemen heading for disaster in 2010?" The author, "Ninja Shoes", wrote: "Based on the information we've gathered, Yemen will likely experience food shortages and torrential floods in 2010. This combination of natural disasters, propensity for famine and malnutrition, and challenges with Islamic radicals and terrorists, make it a hot spot for conflict in the future."
The 20 employees of Recorded Future aren't foreign-policy experts. They aren't traders either, but if you'd started using Recorded Future's predictions to buy US stocks on January 1, 2009, you would have made an annual return of 56.69 per cent. (The S&P 500 had an annualised return of 17.22 per cent over the same period.) Between May 13 and August 5 this year, as markets behaved with vertiginous abandon, their strategy returned 10.4 per cent; in contrast, the S&P 500 lost 9.9 per cent of its value. They're data experts: computer scientists, statisticians and experts in linguistics. And in the data, they think, lies the future.
All Recorded Future's predictions, whatever the field, are based on publicly available information -- news articles, government sites, financial reports, tweets -- fed into the company's own algorithms. The result, it claims, is a "new tool that allows you to visualise the future" -- one that is changing how government intelligence agencies gather information and how giant hedge funds place bets. On its website, Recorded Future states: "We don't grant interviews and we don't issue press releases." But behind closed doors, the company is developing the technology that has been described be one tech blog as an "information weapon".
The company, cofounded by Christopher Ahlberg, an entrepreneur who sold his first business for $195 million and served in the Swedish special forces, has $8.5 million in funding. Its first two investors were Google and the CIA. Recorded Future counts US government agencies, banks and hedge funds among the clients paying million-dollar contracts. But its true ambition is to organise all the data on the internet for similar predictive analysis -- to make the future calculable.
Recorded Future's main office is in Gothenburg, Sweden. On a drab morning in May, trams clang past a metal door that doesn't bear the company's name. Two flights of stairs lead to a wooden door, with a discreet sticker label-gunned above the letterbox in caps: "RECORDED FUTURE". The rooms date from the 17th century; they're airy and bright with high ceilings and intricate plaster mouldings. Eight employees work here on the technical aspects of the system. The company also has offices in Boston, New York and Arlington, Virginia -- ten minutes' drive from the Pentagon, 15 from Langley.
"Yemen took four or five months longer than we predicted," says Ahlberg, 43, sitting on a sofa in a small meeting room. Before Wired visited, he warned over the telephone: "You won't get a government agency out of my mouth. Dude, if I do that, they're coming to take my kids." In person, he's tall, with hair cropped short, and is quick to laugh. The telephone caveat still stands, but Ahlberg is willing to talk for the first time about what exactly it is his company does and why Google, intelligence agencies and hedge funds are all so interested.
Ahlberg was born in September 1968, in Kungälv, a town 40 minutes' drive north of Gothenburg. His father was a captain on merchant ships, his mother taught French and English in Sweden. In his first year at secondary school, he created a drawing program on his Sinclair Spectrum called Art CAD ("like an early version of Photoshop") and sold individual copies by advertising it in the local paper. After school, he wanted to study computer science but first had to complete military service in 1987. He chose the Lapplands Jägarregemente special forces, and began training for a hypothetical Russian invasion: "They would come in from Finland and go to Norway; we were supposed to cut them off in the middle. We were supposed to do what the Iraqis are doing now, guerrilla warfare. But we were master cross-country skiers."
Ahlberg then went onto take his degree at Chalmers University of Technology in Gothenburg. As a post doc, he travelled to the University of Maryland to work as a visiting researcher at the Human-Computer Interaction Lab for two summers. During the first visit, at 23, he co-authored a published paper with the director of the lab; the second summer, he co-wrote two, about the new field of data visualisation. Ahlberg returned to Sweden, finishing his PhD in four years instead of six, but it was his work at Maryland that formed the basis for his first company, Spotfire. Launched in 1996, the business created visualisation tools for business intelligence; in 2007, it was bought by Tibco for $195 million(£125m). "I didn't have to work anymore," says Ahlberg. "But I can't stop."
Spotfire had helped businesses visualise internal databases. After the sale, "we started hanging around in coffee shops in Boston and New York", says Ahlberg. "We thought: what's the most interesting data source out there? And it's nebulous, but the web is the most interesting dataset there is on the planet. Instead of just corporate databases, let's think about the web as my data source." Ahlberg began talking with Staffan Truvé, who had supervised his PhD and started Spotfire with him. "It was in the back of my head that, as humans, we had generally started to become better at predicting things," says 48-year-old Truvé. "Your car tells you that you need to change your oil in 200km, or there is a sign saying your bus is coming in five minutes. These tiny predictive signals are popping up everywhere." Ahlberg was excited: "So then the premise becomes that the web has predictive power. How can we harvest that?"
A few isolated, eye-catching examples have shown the prognostic possibilities in such data. In 2008, Google showed search queries could accurately predict the spread of flu in the US up to two weeks before the federal Centers for Disease Control. In his book, Super Crunchers, Ian Ayers claimed that creditcard companies can predict with 98 per cent accuracy whether you'll divorce, based on your purchases -- Google's Marissa Mayer even quoted the statistic at SXSW 2011 (however, in a recent statement, Visa denied that it monitored such data or made any such conclusions, saying the claim was "inaccurate and wrong"). One recent study has shown Twitter to be 88.67 per cent accurate in predicting the Dow Jones three days in advance. This July, financier Paul Hawtin founded a London hedge fund that is based entirely on social media. And in September, a researcher from the University of Illinois fed the Nautilus supercomputer with 100 million news items, much like Recorded Future does, and "anticipated" the Arab Spring and the killing of Osama Bin Laden, albeit retrospectively -- a prediction of the past. But for Ahlberg to develop a tool that could create predictions for any input, from finance to terrorism, would be much harder. Recorded Future would not only have to index the internet, but also understand and interpret it.
The first generation of search engines, such as Lycos and Alta Vista, used traditional text search to deliver web pages, deploying their own algorithms, but essentially looking at individual documents in isolation. Google changed this in 1998. Its PageRank algorithm analysed the links between web pages, promoting those that had more links pointing to them from other sites. Recorded Future is part of the third generation: instead of explicit link analysis, it examines implicit links -- what it calls "invisible links" between documents that refer to the same entities or events. It does this by separating the documents and their content from what they talk about, identifying canonical entities and events that exist outside of the article.
"What matters is that it's freaking complicated," says Ahlberg. In practice, Recorded Future harvests 25,000 data sources as RSS feeds, which could include Companies House and US Securities and Exchange Commission filings, a New York Times article, Twitter and Facebook posts, obscure blogs (there's one on Norwegian salmon fishing) or transcripts from earnings calls or political speeches -- "just a flood of stuff", says Ahlberg. It does the same for Chinese and Arabic sources. "Then we look for entities -- people, places, technologies; and events -- a murder, a bomb explosion, a person moving from A to B, product launches."
This linguistic analysis is "really tough", according to Truvé. Because Recorded Future takes sources from all over the internet, rather than a particular data set, "the data is not so nice". "We could have built a perfect data set around Pfizer, say, or Barack Obama," says Ahlberg. "It's harder then to think of the big picture. So we tried to make this ambitious." Recorded Future currently uses two separate algorithms, one proprietary, one licensed, to analyse language; the staffers in Gothenburg are tweaking them continually to see which works better. But the result is that Recorded Future knows who Nicolas Sarkozy is, say: that he's the president of France, he's the husband of Carla Bruni, he's 1.65m tall in his socks, he travelled to Deauville for the G8 summit in May. If you Google "president of France", you'll get two Wikipedia pages on "president of France" then " Nicolas Sarkozy". Useful, but Google doesn't know how the two, Sarkozy and the presidency, are actually related; it's just searching for pages linking to the terms.
Recorded Future ranks all these canonical entities and events, based on the number of references to them, the credibility of the document or document source and several other factors, such as the co-occurrence of different events and entities in the same or in related documents, to create a "momentum" score. Positive or negative sentiment is added to this score. For example, searching big pharma in general will tell you that over the next five years, nine of the world's 15 best-selling medicines will lose patent protection -- the event earns a high momentum score because it is backed by 13 news items from 12 sources -- or that, specifically, Inhibitex, a biopharma business, will need cash in November 2011 if it plans to fund the Phase 2b development of a new drug internally, based on five items from five sources.
Recorded Future isn't the only company attempting to bring hardcore linguistic analysis to a larger audience. Wolfram Alpha is a search engine that can understand a query such as "nuclear explosions in China" and deliver relevant information such as maps and kilotonnes per explosion, although it's culled from "tame" data curated by the company itself. And IBM didn't develop Watson just to school humans on Jeopardy; it's actually a huge research project dedicated to processing questions asked in natural language, based on four terabytes of structured and unstructured data sets, including the full text of Wikipedia. "There are any number of offerings coming on to the market now," says Colin Shearer, senior vice president at SPSS, a predictive-analytics company owned by IBM. One of those is Quid, a two-year-old, 45-strong business founded in 2008. "Human activity has never left an information trail like it does today," says Bob Goodson, its founder. "If only we could harness the intelligence that's locked in the information, we could build systems to understand the world better, and therefore make better decisions." Quid includes Microsoft among its customers. With $15 million in investment, it aims to be the next Bloomberg in business intelligence.
Where Recorded Future goes beyond mere analysis of open data, though, is by adding the "time and space" dimension of the documents -- "references to when and where an event has taken place, or when and where it will take place," says Truvé, "since many documents actually refer to events expected to take place in the future." Using RSS streams allows Recorded Future to have a publishing time as an anchor point for this temporal analysis, which means it can deal with difficult expressions such as "next week", "in three months' time" or "in two quarters". This may sound simple, but it's crucial: the time and space analysis is the first way Recorded Future can make predictions about the future -- by aggregating weighted opinions about the likely timing of future events using algorithmic crowdsourcing. On top of that, it uses statistical models to predict future happenings based on historical records of similar chains of events. "The secret sauce is not dependent upon one ingredient," says Truvé. "It's a combination."
On April 1, 2009, a few months after Ahlberg and Truvé began testing this combination, Ahlberg met Rich Miner, the co-creator of the Android mobile operating system (with Andy Rubin) and a partner at Google Ventures, at the Starbucks on Harvard Square, Massachusetts. Miner was impressed: "We believed there was predictive power in the information contained in the web," he says. "If you can organise that information temporally, then you can look at past and present, and infer things from the future. That's pretty unique so far from Recorded Future." The CIA thought so, too.
In the 40s the allies routinely bombed rail bridges to disrupt supply lines into Nazi-occupied France. After a raid, though, the Royal Air Force couldn't fly reconnaissance missions over the targets as they were considered too risky, so it didn't know if a bridge had been destroyed. The Special Operations Executive (SOE), however, came up with a novel strategy for finding out. By monitoring the daily prices of oranges on sale at various fruit stalls Paris, SOE agents dropped behind enemy lines were able to tell which supply chains had been affected. (Germans embedded in London were doing the same thing; unfortunately for the Nazis, they were under the control of SOE and were fed false information.) This is the differ- ence between information and intelligence: information is the price of oranges, intelli- gence is knowing which supply chain has been affected. This openly available, "free" infor- mation, when it's turned into intelligence, becomes extremely valuable.
"Open-source intelligence has always been crucial, but for most of the cold war it was neglected by western intelligence agencies," says Calder Walton, a research associate at Cambridge University and author of the book Empire of Secrets, to be published in 2013. "That was the archetypal intelligence war: intelligence necessarily involved information that couldn't be gained from any other source -- human agents or telephone tapping." That doesn't mean covert intelligence was more effective, though: Daniel Moynihan, a former US senator, compared CIA reports gathered from secret sources with Soviet documents recovered after the fall of the Berlin Wall and found they significantly overestimated Soviet capabilities. But he discovered that western think tanks using publicly available material, such as the RAND Corporation, were much more accurate. US diplomat George Kennan estimated in 1997 that "95 per cent of what we need to know about foreign countries could very well be obtained by the careful and competent study of perfectly legitimate sources of information open and available to us".
"All of this has changed since the collapse of the Soviet Union," says Walton. "Open-source intelligence has boomed in recent years -- especially since 9/11." At a conference in 2008, Michael Hayden, then director of the CIA, said: "Open-source intelligence contributes to national security in unique and valuable ways virtually every day." Stephen Mercado, an ana- lyst in the CIA directorate of science and technology, estimates that 80 per cent of all valuable intelligence now comes from open sources. In January 2011, Sir Gus O'Donnell, head of the UKcivil service, told the Chilcot inquiry into the invasion of Iraq: "I have strongly and always been of the view that we probably underestimated open source [intelligence]." Open source is the big growth area in intelligence and every western agency is looking for the tools to give it an edge.
Ahlberg refuses to discuss his company's work with the CIA, or even whether there is work with the CIA. In-Q-Tel (IQT) is the CIA's investment arm (mission statement: "Identifies, adapts and delivers innovative technological solutions to support the missions of the Central Intelligence Agency"). It invests only in startup companies that will "provide strong, near-term advantages (within 36 months) to the IC [intelligence community]." IQT doesn't invest without the US secret intelligence services in mind. It backed Recorded Future with slightly less than $2.5 million.
Stephen Davidson, an investor at IQT who sits on Recorded Future's board, refused to comment; a spokesperson for IQT said that "while we are pleased to have Recorded Future as part of the IQT portfolio, we will respectfully decline to provide additional information about our investment". Does Ahlberg know what intelligence purposes Recorded Future is put to? "We would not know about those things," he says, folding his arms. "At this stage, I don't even want to know what people are doing with some of these things." He points out that IQT is "an independent company; at least to my knowledge theycan't force any [government agency] to use it." Truvé, though, says Recorded Future is working with 17 or 18 intelligence agencies. Another board member, Roger Ehrenberg, used to run a $6 billion hedge fund for Deutsche Bank before setting up his own firm, IA Ventures. According to Ehrenberg, In-Q-Tel is "actively involved" with Recorded Future. "Fundamentally, they look to invest in companies where they know they have a customer within the government," he says. "It's not just the CIA." Chris Holden, who works in Recorded Future's Arlington office, admitted to wired (with some understatement) that "we have a little bit of work with the federal government". Holden says that Recorded Future is being used to identify technologies the US government may invest in, such as nanotechnology in body armour. "It's not all super secret stuff necessarily." So, does having IQT as an investor mean thatRecorded Future is beholden to the US government, even if it is a private company? "We are an independent company," repeats Ahlberg. "Neither the US government, nor Google, nor hedge funds nor banks have ever tried to make us do anything. And frankly, you're sitting here with a bunch of Swedes. There's no way in hell you could get them to do anything bad."
Still, it's possible to identify examples of how one might use Recorded Future for open-source intelligence. Take the al-Qaeda leadership after Bin Laden's death: who would fill the vacuum? Recorded Future ran a search. Ayman al-Zawahiri, a founding member of Egypt's Islamic Jihad militant group, and long considered by the US government to be Bin Laden's right-hand man, showed some significant spikes in recorded and discussed activity in the last 12 months,especially when he called for military backing of Libyan rebels, suggesting al-Qaeda could fill a power vacuum in that country. But al-Zawahiri's sentiment score was extremely negative, to the degree that conspiracy theories were emerging that he was responsible for disclosing Bin Laden's location to the US. Saif al-Adel, a senior al-Qaeda commander, was attracting attention back in October 2010, written about as "the new face of al-Qaeda in 2011". Recorded Future concluded that it was "clear that Said al-Adel has been routed in Pakistan for some time now and appears to be embedded in the political structure of al-Qaeda"; his momentum score was high. They also found that Libyan Abu Yahya al-Libi (described by a former CIA analyst as an "insurgent-theologian"), offered access to one of the most volatile regions on the globe right now, based on his current likely location, which al-Qaeda might consider a useful foothold. Finally, they looked at Anwar al-Awlaki, a Yemeni-American imam who posted pro-al-Qaeda/anti-western YouTube videos, and ran a blog and Facebook page. Clearly of interest given the attempted US drone strike to kill him days after Bin Laden's death, he started building momentum in late March and April with reports that he was urging on the Arab Spring protests. Al- Awlaki was killed in Yemen dur- ing a US drone attack on September 30 this year.
Recorded Future concluded that multiple players will rise to prominence regionally, that al-Qaeda could split around al-Zawahiri, and that al-Qaeda sees advantage to be taken in the Arab pro-democracy protests. A couple of months later, al-Zawahiri was confirmed, although experts were sceptical about whether he could unite the membership in Saudi Arabia and the Gulf States behind him.
Ahlberg is more willing to talk about how Recorded Future is being used in finance. "If you take our momentum score, and look across S&P 500 companies, can you predict the liquidity or the stock volume of those companies over time?" asks Ahlberg. "It turns out you can." Stock that is being talked about and is in investors' attention is, of course, more likely to be traded: "It's much easier to prove volume than direction, whether a stock is going up or down." So Recorded Future takes momentum and combines it with sentiment -- whether a company is mentioned in a positive light -- and derives a score. Taking these news bursts across the S&P 500, it can sort them into ten different groups, from high to low. "Then you say, every day, 'I am going to own what is in the top and short what is in the bottom.' You are making lots of small picks on a daily basis - this strategy turns over the portfolio 63 per cent every day."
Running predictive tests on data from January 2009 to January 2011, Recorded Future showed that its top decile has a beta (a measure of risk in portfolio) of 1.08 -- fairly low -- and a statistically significant annualised continuous alpha (a risk-adjusted measure of active return on investment) of +16 per cent. The bottom two deciles had a high beta (1.37 and 1.34, respectively) but with statistically significant negative alphas, at -42 per cent and -26 per cent annually. "Constructing hedged portfolios out of the securities in these deciles provides some compelling trading strategies," says Evan Sparks, an analyst at Recorded Future.
Beyond high-frequency trading strategies, the company says it can predict stock shifts on the basis of one-day events, separated into scheduled events and speculative events. "The theory that if something is written saying, 'on Friday so and so will release earnings', that should be priced into the market immediately," says Ahlberg. "In reality it is not." Recorded Future took 19,000 such events and asked what happens to the stock price. On average, as stocks come into those scheduled events, the prices rise; coming out of them they fall five base points either way. "It's like finding a roulette wheel that is skewed."
Another way is to examine the next two weeks of a particular business's future and look for certain events. One is insiders selling stock. You may think this would be a good time to sell; in fact, insiders often sell just after stocks have already peaked. So Recorded Future looks for data that can be combined with this knowledge. If an insider sells stock after a management lay-off, stock falls on average 1.5 per cent. Expand this event to a whole market and "You have 2,000 events within 2011," says Ahlberg. "By turning it into a big data screen, I have created my own skewed roulette wheel I can consistently bet on." Chris Malloy is an associate professor at Harvard Business School who specialises in behavioural finance. He's played with Recorded Future's data: "I haven't seen anything with that ability. It's pretty neat -- no one's doing that. The predictability is certainly good."
What Recorded Future can't forecast are "black swan events", which are by definition unpredictable and undirected. "You can look at what happens afterwards, though," says Ahlberg. He takes the example of a natural disaster. "Start looking at how other countries behave. After a natural disaster, the US will travel there every time, the UK does it 50 per cent of the time, Iran will do it every single time, China never really does." China did, though, after the 2010 Chilean earthquake. Two months later, it announced a new trade agreement. China didn't travel to Haiti: no trade agreement followed. But it did after Pakistan was hit by flood, soon announcing a $10 billion deal. "We're looking for those historical patterns and using them to predict what might happen," says Ahlberg.
Recorded Future's hedge-fund clients are only slightly less secretive than theCIA. Ehrenberg says a handful of Wall Street hedge funds and banks are using the technology: "Recorded Future is a high-value signal, relative to conventional quantitative-analysis trading signals. People are making money." Josh Holden, CEO of Fina Technologies, which creates algorithms for high-frequency quant trading by hedge funds, says that Recorded Future's client base "is closely guarded. But there are more than a few firms using it.
Sandfire AG is a Swiss consultancy in the public and private security sectors, and a client of Recorded Future. "It helps us keep track of travel routes of high-level decision-makers," says Felix Juhl, a senior partner. "A state visit by a high-ranking politician may be followed by specific corporate activities. Keeping track of travel routes can serve as an early warning."
Ahlberg says Recorded Future now earns revenues in the millions of dollars from a client base of less than 100, but which includes governments, hedge funds, big banks, watchdogs and consultancies. This select clientele place a high value on the distilled insight the company provides. Ahlberg sees a big opportunity: "Even within what we have started around finance and intelligence, there is no reason why we couldn't build another $100 million-revenue company within a small set of years." But Recorded Future plans on being more than just a profitable business tool. Ahlberg is expanding its indexes: he eventually wants every piece of data on the planet streaming live through his company's algorithms. The ultimate goal? "We want to organise the world -- and the internet -- for analysis." What Ahlberg doesn't say, perhaps deliberately, is that ever more data will likely lead to ever more accurate predictions. "It's dangerous to start talking about predicting the future," he says. "We're trying to play that down."
Tom Cheshire is assistant editor at wired. He wrote about the Ariane 5 rocket in 10.11
Monday, November 21, 2011
Brynjolfson in Atlantic: The Big Data Boom Is the Innovation Story of Our Time
See also: Where great ideas really come from. A special report
Erik Brynjolfsson and Andrew McAfee The Atlantic November, 21 2011, 9:50 AM ET
The data revolution has turned customers into unwitting business consultants, as our purchases and searches are tracked to improve everything from websites to delivery routes
In the 1670s, in Delft, Netherlands, a scientist named Anton van Leeuwenhoek did something many scientists had done for 100 years before him. He built a microscope.
This microscope was different, but it was not extraordinary. Like so many inventions, he borrowed and tweaked his predecessors' ingenuity. But when he looked through this microscope, he found things that did seem extraordinary. He called them "animalcules," microbes in water droplets and human blood that ultimately provided the foundation for the germ theory of disease and eventually inspired a host of medicines and treatments.
The Leeuwenhoek discovery is crucial to our understanding of innovation, not only because it changed the face of biochemistry, but also because it represents a fundamental theme of discovery.
Breakthroughs in innovation often rely on breakthroughs in measurement.
THE DATA BOOM
Today businesses can measure their activities and customer relationships with unprecedented precision. As a result, they are awash with data. This is particularly evident in the digital economy, where clickstream data give precisely targeted and real-time insights into consumer behavior.
In turn, customers are acting as unwitting business consultants for these companies. Our purchases, searches, and online activities are being tracked to improve everything from websites to delivery routes and drug manufacturing.
Anyone with access to a Web browser can get summaries of billions of keyword searches, and this information is highly predictive of present and future economic activity, such as housing purchases and prices. Mobile phones, automobiles, factory automation systems and other devices are routinely instrumented to generate streams of data on their activities, making possible an emerging field of "reality mining" to analyze this information.
Manufacturers and retailers use radio-frequency identification (RFID) tags to deliver terabits of data on inventories and supplier interactions and then feed this information into analytical models to optimize and reinvent their business processes. Much of this information is generated for free, by computers, and sits unused, at least initially. A few years after installing a large enterprise resource planning system, it is common for companies to purchase a "business intelligence" module to try to make use of the flood of data that they now have on their operations. As Ron Kohavi at Microsoft memorably put it, objective, fine-grained data are replacing HiPPOs (Highest Paid Person's Opinions) as the basis for decision-making at more and more companies. For example:
-- Enologix has used this approach to help Gallo vineyards accurately predict the wine ratings that Robert Parker would give to various new wines
-- UPS has mined data on truck delivery times to develop a new routing method
-- Match.com as even developed new algorithms for matching men and women for dates
For each innovation, analysts drew on new measurement technologies to supplant human experts who relied more on intuition. However, for all its strengths, measurements have a shortcoming. They cannot determine causality. (A simple example: Shoe sizes and readings scores are correlated for school children, but one does not cause the other; instead, they both reflect a third variable, which is age.) Fortunately, science has a second powerful tool designed precisely to address questions of causality.
That tool is called experimentation.
AN EXPERIMENT EVERY SECOND
Science has been dominated by the experimental approach for nearly 400 years. Running controlled experiments is the gold standard for sorting out cause and effect. But experimentation has been difficult for businesses throughout history because of cost, speed and convenience. It is only recently that businesses have learned to run real-time experiments on their customers. The key enabler was the Web.
Consider two "born-digital" companies, Amazon and Google. A central part of Amazon's research strategy is a program of "A-B" experiments where it develops two versions of its website and offers them to matched samples of customers. Using this method, Amazon might test a new recommendation engine for books, a new service feature, a different check-out process, or simply a different layout or design. Amazon sometimes gets sufficient data within just a few hours to see a statistically significant difference.
This ability to rapidly test ideas fundamentally changes the company's mindset and approach to innovation. Rather than agonize for months over a choice, or model hypothetical scenarios, the company simply asks the customers and get an answer in real time.
According to Google economist Hal Varian, his company is running on the order of 100-200 experiments on any given day, as they test new products and services, new algorithms and alternative designs. An iterative review process aggregates findings and frequently leads to further rounds of more targeted experimentation.
At the same time, Google's competitors, partners, customers and third party consultants are doing their own experiments, creating a complex, interacting ecosystem that demands continuous innovation. While Google currently dominates the market for web search, it is unlikely that it would have any market share at all if it still relied on the original, unmodified PageRank algorithm that Larry Page and Sergey Brin developed in 1998.
HOW DATA TURNED AROUND A CASINO Greg Linden, who led one set of experiments at Amazon, describes the emerging experimentation philosophy succinctly: "To find high impact experiments, you need to try a lot of things. Genius is born from a thousand failures. In each failed test, you learn something that helps you find something that will work. Constant, continuous, ubiquitous experimentation is the most important thing."
These words echo the approach of innovators since Thomas Edison, but IT has made it possible to apply it to a much broader class of business challenges and significantly compress the "hypothesis-to-experiment" cycle time.
While web-based companies have been particularly aggressive in using business experiments to drive innovation, other industries are getting in the game. Caesar's Entertainment (formerly Harrah's), the hotel and casino company, transformed itself from a 2nd-tier casino to an industry leader in large part because of the culture of experimentation introduced by CEO Gary Loveman.
When Loveman, an economics PhD from MIT and former Harvard Business School professor, arrived at the company, he found that it was already gathering a great deal of data about its customer interactions with existing information systems and programs such as its Total Rewards loyalty card. However, it wasn't using these data to develop improved processes, products and services. After becoming CEO, he developed strategies to continually tests new promotions, price points, services, workflow, employee incentive plans and casino layouts using controlled experiments.
Widespread business experimentation has required a fundamental change in the corporate culture. As Loveman puts it "There are two things that will get you fired here: stealing from the company, or running an experiment without a properly designed control group."
***
While passive data gathering can be useful, measurement is far more valuable when coupled with conscious, active experimentation and sharing of insights. Likewise, the value of undertaking the experiments themselves is proportionately greater if the organization can capitalize on those experiments in more locations and at greater scale. In combination, these practices constitute a new kind of "R&D" that draws on the strengths of digitization to speed innovation.
How crowdsourcing is changing science
Gareth Cook The Boston Globe November 11, 211
At the end of the 19th century, a team of British archeologists happened upon what is now one of the world's most treasured trash dumps.
The site, situated west of the main course of the Nile, about five days journey south of Memphis, lay near the city of Oxyrhynchus. Garbage mounds are always a sweet target for those interested in the past, but what made the Oxyrhynchus dump special was its exceptional dryness. The water table lay deep; it never rained. And this meant that the 2,000-year-old papyrus in the mounds, and the text inscribed on it, were remarkably well preserved.
Eventually some half a million pieces of papyrus were drawn from the desert and shipped back to Oxford University, where generations of scholars have been painstakingly transcribing and translating them. The manuscripts are rich, fascinating, and varied. The texts include lost comedies by the great Athenian playwright Menander, and the controversial Gospel of Thomas, along with glimpses of daily life — personal notes, receipts for the purchase of donkeys and dates — and the occasional scrap of sex magic.
The pace, however, has been glacial. After a hundred-plus years, scholars have been able to work through only about 15 percent of the collection. The finish line appeared to lie centuries in the future.
But a few months ago, the papyrologists tried something bold. They put up a website, called Ancient Lives, with a game that allowed members of the public to help transcribe the ancient Greek at home by identifying images from the papyrus. Help began pouring in. In the short time the site has been running, people have contributed 4 million transcriptions. They have helped identify Thucydides, Aristophanes, Plutarch's "On the Cleverness of Animals," and more.
Ancient Lives is part of a new approach to the conduct of modern scholarship, called crowd science or citizen science. The idea is to unlock thorny research projects by tapping the time and enthusiasm of the general public. In just the last few years, crowd science projects have generated notable contributions to fields as disparate as ecology, AIDS research, and astronomy. The approach has already accelerated research in a handful of specialized fields. And it may also accomplish something else: breaking down some of the old divisions between the highly educated mandarins of the academy and the curious amateurs out in the world.
"It may seem intimidating when we say you are going to help transcribe ancient Greek papyri, but it's all about pattern recognition, and the brain excels at pattern recognition," says James Brusuelas, an Oxford classicist who is part of the Ancient Lives team. "The reaction has been fantastic."
One reason for the sudden turn to crowd science is that it offers an imaginative answer to a central problem of 21st-century science: too much information. Oxford's scholars had an overwhelming load of work given them, in the form of a desert trove. More often, though, scientists are themselves creating floods of data that they simply don't have the hours to interpret. Every night, robotic telescopes relentlessly track the sky, pouring terabytes of images into hard drive farms. From biological labs come rivers of genetic code. And in many other fields — from high energy physics to environmental science — researchers are puzzling over how to handle the sudden embarrassment of riches.
For now, the new citizen science has touched only the tiniest fraction of the research conducted around the world. But its early successes, which have shocked even the architects of the approach, suggest that over time pro-am collaborations hold the potential to alter the landscape of science in important ways, harnessing countless able brains to do work that was once the province of a few overwhelmed experts. And as it does, it also offers an uncomfortable insight: There are ways that the structure of modern science may actually be limiting what we can learn.
The idea of recruiting amateur scientists has roots that go back at least a century. In 1900, in the early days of the American conservation movement, ornithologist Frank Chapman organized a Christmas bird census. Teams of avid birders collected observations from Toronto to Baldwin, La.: the American black duck, the red-breasted nuthatch, the common grackle, and 86 other species. It was an unprecedented one-day data dump. The Christmas Bird Count has become an Audubon tradition, with about 60,000 people going out every year, and the data it has generated through the years have proved invaluable to researchers.
Today there are firefly counts, herring counts, and ladybug counts. One can help track spiders or bats or coral reefs. A new iPhone app called Noah (for Networked Organisms and Habitats) allows users to snap pictures of species they come across and share the information with researchers and others. A similar British effort, called iSpot, led to the discovery of two species that had not been recorded before in England, according to a report by the BBC. Some projects use networks of observers to monitor the timing of natural events, such as the arrival of hummingbirds, or the budding of flowers, which provide information on the planet's changing climate. None of these projects would be possible without countless amateurs willing to serve as devoted foot soldiers across the planet.
The advent of the Internet has also opened up a new possibility: that the interested public could offer scholars more than help gathering data. In the best-known early example, they offered up their computers: 1999 saw the launch of SETI@home, an example of "distributed computing" in which volunteers downloaded software so their idling computers could help crunch radio-telescope data for signs of alien life.
More recently, though, has come a truly fascinating turn: the move from people volunteering their computers' down time, to people volunteering their brains' down time — from distributed computing to distributed thinking. Oxford University astronomer Chris Lintott says that his own involvement dates back to a 2007 conversation he had over a pint at the Royal Oak, a traditional watering hole for Oxford astronomers where tables are crammed into small rooms with old fireplaces and ancient wood beams. A student, Kevin Schawinski, had recently finished the exhausting task of categorizing 50,000 galaxy images for a project. As they spoke, though, it became clear that that wasn't nearly enough: What the project really required to succeed was to categorize a million galaxies.
"One look at Kevin's face," says Lintott, "suggested we should find an alternative method."
This led them to create Galaxy Zoo in 2007. The site provided a simple tutorial that trained people to classify galaxies by their appearance, and then served up images that astronomers had not yet categorized. Galaxy Zoo was so popular that soon after it launched, the servers literally caught fire from all the activity. A schoolteacher sitting in an apartment in the southeast of the Netherlands discovered a strange green cloud that had never been observed before. The astronomical data from the project have been used in a growing list of scientific publications.
The approach was so successful that Lintott and the other organizers decided to expand it to other areas, including solar explosions and climate change, under the name Zooniverse. (The Ancient Lives project uses the Zooniverse website.) Meanwhile, many other scientists and organizations are jumping in: One popular website, scienceforcitizens.net, lists more than 400 projects, and the site's founder says she expects to hit 1,000 within a year.
What marks this as an important milestone in the history of science is the new way it harnesses the power of the mind. There are many tasks that are beyond the grasp of even today's computers, particularly those which involve interpreting complex images. Like identifying cancer cells. Or categorizing galaxies. Or picking out letters of ancient Greek, written in a faded ink with a fast, messy hand, without breaks between words. The Internet, it turns out, is a brilliant way to feed those problems into an array of the planet's true supercomputers — human brains.A recent discovery highlights the sophis
tication of the work volunteers can do. Biologists are keenly interested in the three-dimensional shapes assumed by protein molecules inside the human body. Proteins are intimately involved in many aspects of life, but they fold into shapes that can be very difficult to predict, even given their precise chemical makeup. Protein-folding is a roadblock that holds up research into many diseases.
So a team of scientists at the University of Washington created a game called FoldIt, which gives players an image of a protein molecule and video game-like tools for folding the molecule. As the energy required to maintain the molecule in a particular shape drops — meaning it's closer to nature's solution — a player's score increases. FoldIt is a potentially addictive game that requires excellent spatial reasoning. Some players excelled at it — indeed, some became whizzes, and the researchers put their skills to work on unsolved problems. In September, the scientists announced that a team of its players had deciphered the folding of a protein important in AIDS research.
In a paper describing the result for Nature Structural and Molecular Biology, the scientists argued to their colleagues that a line had been crossed: "Although much attention has recently been given to the potential of crowdsourcing and game playing, this is the first instance we are aware of in which online gamers solved a longstanding scientific problem."
FoldIt is the most impressive demonstration yet that the public can make genuine contributions to scientific projects. But its success also stands as a potent critique of the way that the scientific enterprise is currently organized.
Science is, for the most part, a closed society organized into little fiefdoms of highly trained specialists, which means only a few minds engage with any given problem. Before FoldIt, for example, a problem in protein folding was the exclusive province of a relatively small number of experts — even though, it is now clear, there are real contributions to be made by 13-year-old video gamers.
The system is shaped in part by the force of tradition, but the larger challenge is that most scientific data is proprietary. A scientist works long and hard to generate original data, and then expects to reap the reward in the form of publishing the first research paper to describe some new phenomenon. She is not going to want share this data with others, particularly strangers, any more than say, an investigative reporter would want to share his notes before a story has been written. Harnessing 1,000 people requires sending your data out into the world — something that science is loath to do. The scientist's interest in keeping things private and getting credit, in other words, is directly opposed to society's interest in tackling some problems with a hive of the best minds.
There are exceptions, such as large astronomical and biological data sets that are available for anyone to work with. But the last 10 years have seen a boom in technology that allows large numbers of people to do amazing, cooperative things with information, and the scientific establishment has taken only baby steps toward figuring out ways to share it productively, according to Michael Nielsen, a former theoretical physicist and author of "Reinventing Discovery: The New Era of Networked Science."
To encourage this shift, the federal government, which funds the lion's share of the country's research, has been pressuring scientists to work more cooperatively, and share more of what they find faster. And there is a nascent effort within academia to identify ways that scientists might be recognized for their contributions to the community as a whole, beyond the publication of their individual discoveries.
"It is essential that scientists be rewarded when they share," says Nielsen.
It's a difficult problem, and Nielsen says he expects the real rewards of networked science to be tallied over decades, not years. Even if science becomes more open, there are also practical limitations: It takes a certain brilliance, and a lot of work, to recognize problems that can be shared with a crowd, and set up the systems needed for strangers to work together productively. It is not always clear when this tactic will move a project forward, or slow it down.
With time, though, one might expect a new type of scientist to emerge: one who is especially adept at recognizing problems, and designing projects, that tap the brilliance of a dispersed and motley team, whoever they may be.
Science is driven forward by discovery, and we appear to stand at the beginning of a democratization of discovery. An ordinary person can be the one who realizes that a long arm of a protein probably tucks itself just so; a woman who never went to college can provide the crucial transcription that reveals a spidery script to be a love poem from 2,000 years in the past. Nobody can say where the movement will go, but among the new pioneers of crowd science, there is a palpable sense that they have just happened upon a powerful, poorly understood new resource.
"We have used," says Lintott, "just a tiny fraction of the human attention span that goes into an episode of Jerry Springer."
Gareth Cook is a Globe columnist, a Pulitzer Prize-winning journalist, and a former editor of Ideas. He can be reached at cook@globe.com. Follow him on Twitter @garethideas.
At the end of the 19th century, a team of British archeologists happened upon what is now one of the world's most treasured trash dumps.
The site, situated west of the main course of the Nile, about five days journey south of Memphis, lay near the city of Oxyrhynchus. Garbage mounds are always a sweet target for those interested in the past, but what made the Oxyrhynchus dump special was its exceptional dryness. The water table lay deep; it never rained. And this meant that the 2,000-year-old papyrus in the mounds, and the text inscribed on it, were remarkably well preserved.
Eventually some half a million pieces of papyrus were drawn from the desert and shipped back to Oxford University, where generations of scholars have been painstakingly transcribing and translating them. The manuscripts are rich, fascinating, and varied. The texts include lost comedies by the great Athenian playwright Menander, and the controversial Gospel of Thomas, along with glimpses of daily life — personal notes, receipts for the purchase of donkeys and dates — and the occasional scrap of sex magic.
The pace, however, has been glacial. After a hundred-plus years, scholars have been able to work through only about 15 percent of the collection. The finish line appeared to lie centuries in the future.
But a few months ago, the papyrologists tried something bold. They put up a website, called Ancient Lives, with a game that allowed members of the public to help transcribe the ancient Greek at home by identifying images from the papyrus. Help began pouring in. In the short time the site has been running, people have contributed 4 million transcriptions. They have helped identify Thucydides, Aristophanes, Plutarch's "On the Cleverness of Animals," and more.
Ancient Lives is part of a new approach to the conduct of modern scholarship, called crowd science or citizen science. The idea is to unlock thorny research projects by tapping the time and enthusiasm of the general public. In just the last few years, crowd science projects have generated notable contributions to fields as disparate as ecology, AIDS research, and astronomy. The approach has already accelerated research in a handful of specialized fields. And it may also accomplish something else: breaking down some of the old divisions between the highly educated mandarins of the academy and the curious amateurs out in the world.
"It may seem intimidating when we say you are going to help transcribe ancient Greek papyri, but it's all about pattern recognition, and the brain excels at pattern recognition," says James Brusuelas, an Oxford classicist who is part of the Ancient Lives team. "The reaction has been fantastic."
One reason for the sudden turn to crowd science is that it offers an imaginative answer to a central problem of 21st-century science: too much information. Oxford's scholars had an overwhelming load of work given them, in the form of a desert trove. More often, though, scientists are themselves creating floods of data that they simply don't have the hours to interpret. Every night, robotic telescopes relentlessly track the sky, pouring terabytes of images into hard drive farms. From biological labs come rivers of genetic code. And in many other fields — from high energy physics to environmental science — researchers are puzzling over how to handle the sudden embarrassment of riches.
For now, the new citizen science has touched only the tiniest fraction of the research conducted around the world. But its early successes, which have shocked even the architects of the approach, suggest that over time pro-am collaborations hold the potential to alter the landscape of science in important ways, harnessing countless able brains to do work that was once the province of a few overwhelmed experts. And as it does, it also offers an uncomfortable insight: There are ways that the structure of modern science may actually be limiting what we can learn.
The idea of recruiting amateur scientists has roots that go back at least a century. In 1900, in the early days of the American conservation movement, ornithologist Frank Chapman organized a Christmas bird census. Teams of avid birders collected observations from Toronto to Baldwin, La.: the American black duck, the red-breasted nuthatch, the common grackle, and 86 other species. It was an unprecedented one-day data dump. The Christmas Bird Count has become an Audubon tradition, with about 60,000 people going out every year, and the data it has generated through the years have proved invaluable to researchers.
Today there are firefly counts, herring counts, and ladybug counts. One can help track spiders or bats or coral reefs. A new iPhone app called Noah (for Networked Organisms and Habitats) allows users to snap pictures of species they come across and share the information with researchers and others. A similar British effort, called iSpot, led to the discovery of two species that had not been recorded before in England, according to a report by the BBC. Some projects use networks of observers to monitor the timing of natural events, such as the arrival of hummingbirds, or the budding of flowers, which provide information on the planet's changing climate. None of these projects would be possible without countless amateurs willing to serve as devoted foot soldiers across the planet.
The advent of the Internet has also opened up a new possibility: that the interested public could offer scholars more than help gathering data. In the best-known early example, they offered up their computers: 1999 saw the launch of SETI@home, an example of "distributed computing" in which volunteers downloaded software so their idling computers could help crunch radio-telescope data for signs of alien life.
More recently, though, has come a truly fascinating turn: the move from people volunteering their computers' down time, to people volunteering their brains' down time — from distributed computing to distributed thinking. Oxford University astronomer Chris Lintott says that his own involvement dates back to a 2007 conversation he had over a pint at the Royal Oak, a traditional watering hole for Oxford astronomers where tables are crammed into small rooms with old fireplaces and ancient wood beams. A student, Kevin Schawinski, had recently finished the exhausting task of categorizing 50,000 galaxy images for a project. As they spoke, though, it became clear that that wasn't nearly enough: What the project really required to succeed was to categorize a million galaxies.
"One look at Kevin's face," says Lintott, "suggested we should find an alternative method."
This led them to create Galaxy Zoo in 2007. The site provided a simple tutorial that trained people to classify galaxies by their appearance, and then served up images that astronomers had not yet categorized. Galaxy Zoo was so popular that soon after it launched, the servers literally caught fire from all the activity. A schoolteacher sitting in an apartment in the southeast of the Netherlands discovered a strange green cloud that had never been observed before. The astronomical data from the project have been used in a growing list of scientific publications.
The approach was so successful that Lintott and the other organizers decided to expand it to other areas, including solar explosions and climate change, under the name Zooniverse. (The Ancient Lives project uses the Zooniverse website.) Meanwhile, many other scientists and organizations are jumping in: One popular website, scienceforcitizens.net, lists more than 400 projects, and the site's founder says she expects to hit 1,000 within a year.
What marks this as an important milestone in the history of science is the new way it harnesses the power of the mind. There are many tasks that are beyond the grasp of even today's computers, particularly those which involve interpreting complex images. Like identifying cancer cells. Or categorizing galaxies. Or picking out letters of ancient Greek, written in a faded ink with a fast, messy hand, without breaks between words. The Internet, it turns out, is a brilliant way to feed those problems into an array of the planet's true supercomputers — human brains.A recent discovery highlights the sophis
tication of the work volunteers can do. Biologists are keenly interested in the three-dimensional shapes assumed by protein molecules inside the human body. Proteins are intimately involved in many aspects of life, but they fold into shapes that can be very difficult to predict, even given their precise chemical makeup. Protein-folding is a roadblock that holds up research into many diseases.
So a team of scientists at the University of Washington created a game called FoldIt, which gives players an image of a protein molecule and video game-like tools for folding the molecule. As the energy required to maintain the molecule in a particular shape drops — meaning it's closer to nature's solution — a player's score increases. FoldIt is a potentially addictive game that requires excellent spatial reasoning. Some players excelled at it — indeed, some became whizzes, and the researchers put their skills to work on unsolved problems. In September, the scientists announced that a team of its players had deciphered the folding of a protein important in AIDS research.
In a paper describing the result for Nature Structural and Molecular Biology, the scientists argued to their colleagues that a line had been crossed: "Although much attention has recently been given to the potential of crowdsourcing and game playing, this is the first instance we are aware of in which online gamers solved a longstanding scientific problem."
FoldIt is the most impressive demonstration yet that the public can make genuine contributions to scientific projects. But its success also stands as a potent critique of the way that the scientific enterprise is currently organized.
Science is, for the most part, a closed society organized into little fiefdoms of highly trained specialists, which means only a few minds engage with any given problem. Before FoldIt, for example, a problem in protein folding was the exclusive province of a relatively small number of experts — even though, it is now clear, there are real contributions to be made by 13-year-old video gamers.
The system is shaped in part by the force of tradition, but the larger challenge is that most scientific data is proprietary. A scientist works long and hard to generate original data, and then expects to reap the reward in the form of publishing the first research paper to describe some new phenomenon. She is not going to want share this data with others, particularly strangers, any more than say, an investigative reporter would want to share his notes before a story has been written. Harnessing 1,000 people requires sending your data out into the world — something that science is loath to do. The scientist's interest in keeping things private and getting credit, in other words, is directly opposed to society's interest in tackling some problems with a hive of the best minds.
There are exceptions, such as large astronomical and biological data sets that are available for anyone to work with. But the last 10 years have seen a boom in technology that allows large numbers of people to do amazing, cooperative things with information, and the scientific establishment has taken only baby steps toward figuring out ways to share it productively, according to Michael Nielsen, a former theoretical physicist and author of "Reinventing Discovery: The New Era of Networked Science."
To encourage this shift, the federal government, which funds the lion's share of the country's research, has been pressuring scientists to work more cooperatively, and share more of what they find faster. And there is a nascent effort within academia to identify ways that scientists might be recognized for their contributions to the community as a whole, beyond the publication of their individual discoveries.
"It is essential that scientists be rewarded when they share," says Nielsen.
It's a difficult problem, and Nielsen says he expects the real rewards of networked science to be tallied over decades, not years. Even if science becomes more open, there are also practical limitations: It takes a certain brilliance, and a lot of work, to recognize problems that can be shared with a crowd, and set up the systems needed for strangers to work together productively. It is not always clear when this tactic will move a project forward, or slow it down.
With time, though, one might expect a new type of scientist to emerge: one who is especially adept at recognizing problems, and designing projects, that tap the brilliance of a dispersed and motley team, whoever they may be.
Science is driven forward by discovery, and we appear to stand at the beginning of a democratization of discovery. An ordinary person can be the one who realizes that a long arm of a protein probably tucks itself just so; a woman who never went to college can provide the crucial transcription that reveals a spidery script to be a love poem from 2,000 years in the past. Nobody can say where the movement will go, but among the new pioneers of crowd science, there is a palpable sense that they have just happened upon a powerful, poorly understood new resource.
"We have used," says Lintott, "just a tiny fraction of the human attention span that goes into an episode of Jerry Springer."
Gareth Cook is a Globe columnist, a Pulitzer Prize-winning journalist, and a former editor of Ideas. He can be reached at cook@globe.com. Follow him on Twitter @garethideas.
Friday, November 18, 2011
Big Data Video/Stats
The Economist online November 18, 2011, 16:09
Drowning in numbers
See http://www.economist.com/blogs/dailychart/2011/11/big-data-0
Digital data will flood the planet—and help us understand it better
More from The World in 2012
Drowning in numbers
See http://www.economist.com/blogs/dailychart/2011/11/big-data-0
Digital data will flood the planet—and help us understand it better
More from The World in 2012
Thursday, November 17, 2011
Can Big Data Fix Healthcare?
Colin Hill Forbes November 17, 2011
What will healthcare look like in the year 2020? One thing is certain: we can’t afford its current trajectory. Left unchecked, our $2.6 trillion in annual spending will grow to $4.6 trillion by 2020, one-fifth of GDP. With almost 80 million Baby Boomers approaching retirement, economists forecast these trends will likely bankrupt Medicare and Medicaid in the near future. And while healthcare reform ignites a number of important changes, alone it does not resolve our issues. It’s critical we fix our system now.
Growth in Literature
Over the past 50 years medicine has grown dramatically. We now have more preventive, diagnostic, and treatment alternatives than ever before, with more being developed all the time.
This proliferation has been accompanied by an explosion in literature. Fully 35 percent of the 20 million articles indexed in MEDLINE were published in the last 10 years, with the annual pace approaching one million articles.
Despite this vast body of literature, the healthcare system remains starved for evidence of what works. While we have many treatment alternatives, in many cases we do not have much of an idea where best to apply them. Treatments that work well for some work poorly – or worse – for others.
We must learn to distinguish what interventions work, and for whom.
Limited Knowledge
Unfortunately, we have critical limitations in our ability to evaluate treatment effectiveness. First, and despite its volume, medical literature is of variable quality and has limited generalizability. Only a small fraction of studies compare the effectiveness of different treatments, and few evaluate effectiveness in real-world settings.
Evidence of what does work is increasingly overturned. A recent study found that 13 percent of articles concerning a clinical practice published in the New England Journal of Medicine in 2009 were reversals of previous findings.
And clinical practice guidelines, whose goal is to synthesize research into evidence for use by clinicians, don’t fare much better. Their quality is highly variable and, in some instances, quite poor – even the best have a limited shelf life. A 2001 review of guidelines estimated that half had become outdated in less than six years. It is unlikely that the situation has improved since then.
Clearly, and despite all the effort and expense, healthcare remains one of our nation’s most well-endowed, yet-poorly-informed, industries, with an approach to creating evidence clearly inadequate to the needs of practitioners and patients.
We must start producing better evidence faster and on a large scale. Before we can reduce costs and deliver meaningful improvements in outcomes, we must have meaningful evidence. Without it, we can never know what works, and for whom.
New Sources of Information
New types and sources of health care data have become available – or soon will – and in overwhelming quantity. The federal government is investing $20 billion in Electronic Health Records; industry is developing new electronic transaction standards; and innovators like PatientsLikeMe, 23andMe, Fitbit and Zeo are helping people generate and share their own data. The era of Big Data in healthcare has arrived.
Can Big Data Fix Healthcare? A recent McKinsey report called Big Data, “the next frontier for innovation, competition and productivity.” The Aspen Institute reported on the “promise and perils” of it. The Economist issued a special report about it. O’Reilly Media hosted two conferences on it this year alone. In all of these, Big Data’s opportunity to transform health care was featured prominently.
Are these expectations justified? Can Big Data fix healthcare? What analytic technologies will be required to actually deliver on Big Data’s promise and discover what works? What other promises need to be met for this to become a reality, and what will reality look like when evidence becomes available at the push of a button?
My name is Colin Hill. I am CEO and co-founder of GNS Healthcare, a healthcare analytics company focused on using observational data directly, and at-scale, to create evidence of what works for whom in healthcare. We’ll take on these questions and more in subsequent entries. Welcome to the conversation!
What will healthcare look like in the year 2020? One thing is certain: we can’t afford its current trajectory. Left unchecked, our $2.6 trillion in annual spending will grow to $4.6 trillion by 2020, one-fifth of GDP. With almost 80 million Baby Boomers approaching retirement, economists forecast these trends will likely bankrupt Medicare and Medicaid in the near future. And while healthcare reform ignites a number of important changes, alone it does not resolve our issues. It’s critical we fix our system now.
Growth in Literature
Over the past 50 years medicine has grown dramatically. We now have more preventive, diagnostic, and treatment alternatives than ever before, with more being developed all the time.
This proliferation has been accompanied by an explosion in literature. Fully 35 percent of the 20 million articles indexed in MEDLINE were published in the last 10 years, with the annual pace approaching one million articles.
Despite this vast body of literature, the healthcare system remains starved for evidence of what works. While we have many treatment alternatives, in many cases we do not have much of an idea where best to apply them. Treatments that work well for some work poorly – or worse – for others.
We must learn to distinguish what interventions work, and for whom.
Limited Knowledge
Unfortunately, we have critical limitations in our ability to evaluate treatment effectiveness. First, and despite its volume, medical literature is of variable quality and has limited generalizability. Only a small fraction of studies compare the effectiveness of different treatments, and few evaluate effectiveness in real-world settings.
Evidence of what does work is increasingly overturned. A recent study found that 13 percent of articles concerning a clinical practice published in the New England Journal of Medicine in 2009 were reversals of previous findings.
And clinical practice guidelines, whose goal is to synthesize research into evidence for use by clinicians, don’t fare much better. Their quality is highly variable and, in some instances, quite poor – even the best have a limited shelf life. A 2001 review of guidelines estimated that half had become outdated in less than six years. It is unlikely that the situation has improved since then.
Clearly, and despite all the effort and expense, healthcare remains one of our nation’s most well-endowed, yet-poorly-informed, industries, with an approach to creating evidence clearly inadequate to the needs of practitioners and patients.
We must start producing better evidence faster and on a large scale. Before we can reduce costs and deliver meaningful improvements in outcomes, we must have meaningful evidence. Without it, we can never know what works, and for whom.
New Sources of Information
New types and sources of health care data have become available – or soon will – and in overwhelming quantity. The federal government is investing $20 billion in Electronic Health Records; industry is developing new electronic transaction standards; and innovators like PatientsLikeMe, 23andMe, Fitbit and Zeo are helping people generate and share their own data. The era of Big Data in healthcare has arrived.
Can Big Data Fix Healthcare? A recent McKinsey report called Big Data, “the next frontier for innovation, competition and productivity.” The Aspen Institute reported on the “promise and perils” of it. The Economist issued a special report about it. O’Reilly Media hosted two conferences on it this year alone. In all of these, Big Data’s opportunity to transform health care was featured prominently.
Are these expectations justified? Can Big Data fix healthcare? What analytic technologies will be required to actually deliver on Big Data’s promise and discover what works? What other promises need to be met for this to become a reality, and what will reality look like when evidence becomes available at the push of a button?
My name is Colin Hill. I am CEO and co-founder of GNS Healthcare, a healthcare analytics company focused on using observational data directly, and at-scale, to create evidence of what works for whom in healthcare. We’ll take on these questions and more in subsequent entries. Welcome to the conversation!
Subscribe to:
Posts (Atom)