Showing posts with label Global Innovation. Show all posts
Showing posts with label Global Innovation. Show all posts
Thursday, October 18, 2012
IBM's Watson Is Learning Its Way To Saving Lives
A few years ago, IBM’s new computer was a game-playing curiosity. Now Watson is poised to change the way human beings make decisions about medicine, finance, and work.
Jon Gertner, Fast Company, October 15, 2012.
The woman was gravely ill. Her name was Ms. Yamato. Thirty-seven years old, born in Osaka, Japan, she had never smoked, and yet there it was anyway: a spot on her lung.
A doctor had already performed a bronchoscopy and had made the diagnosis of cancer. Then he referred the patient to Mark Kris, an oncologist at Memorial Sloan-Kettering Cancer Center in New York. Seated alongside me in his office on the Upper East Side of Manhattan, Kris is showing me Ms. Yamato's electronic medical record on an iPad. "I'm preparing for the first visit," he explains, swiping the screen to show what that entails. He's interested in running at least two tests on the patient. The first is an MRI, to find out if the cancer has spread to her brain. The second involves a deeper diagnostic regimen. Lung cancer tumors are not all the same; there are thousands of variations. So a test that examines the mutations within a tumor will be crucial, he says. It so happens that cancer patients born in East Asia who have never smoked often have a particular mutation that responds well to a medication by the name of Erlotinib. That may be the case here. One can hope.
Over the past year, IBM executives have come to believe that Watson represents the first machine of the third computer age.
The woman is not real. She happens to be a character within an app that IBM has created for Watson, its new computer. Watson's special talent, its reason for being, is a singular ability to grasp the intricacies of human language and answer exceedingly difficult questions. You may have heard about Watson already. Back in 2007, a group of computer engineers at IBM's research labs in upstate New York began building the machine--named for IBM's founder, Thomas J. Watson--with the goal of creating a question-and-answer technology that would be more authoritative and powerful than anything on the planet. The initial objective of the Watson group was simple: to win in the game show Jeopardy!, something Watson famously achieved in February 2011. Yet the group had a far more important goal: to turn Watson into a business, hopefully one of some scale. So starting in late 2009, a business development team at IBM began holding meetings outside the company in an effort to understand the ultimate worth of this new technology. No doubt it could be a business one day. But what kind of business?
"The first thing that hit us about Watson," recalls John Kelly, IBM's chief of research, "was that this thing could be applied almost anywhere." Early on, IBM executives decided to focus on a field in which Watson could have a notable social impact while also proving its ability to master a complex body of knowledge. The team chose medicine. They believed Watson could help doctors make diagnoses and, even more important, select treatments. Specifically, they thought Watson could be the perfect tool to chart the complex decision trees that cancer specialists like Kris negotiate every day as they weigh treatment options that might involve radiation, surgery, and any of countless chemotherapy drugs. Watson can ingest more data in a day than any human could in a lifetime. It can read all of the world's medical journals in less time than it takes a physician to drink a cup of coffee. All at once, it can peruse patient histories; keep an eye on the latest drug trials; stay apprised of the potency of new therapies; and hew closely to state-of-the-art guidelines that help doctors choose the best treatments. Watson never goes on vacation. And it never forgets a fact. On the contrary, it keeps learning.
This fall, after six months of teaching their treatment guidelines to Watson, the doctors at Sloan-Kettering will begin testing the IBM machine on real patients. The Ms. Yamato app shows how it will work. After Kris inputs the results of her medical tests, Watson begins deliberating. "It's going through its algorithms," Kris says as we stare at the iPad. "It's seeing where the data sends it today." On the screen, a colorful globe spins. In a few seconds, Watson offers three possible courses of chemotherapy, charted as bars with varying levels of confidence--one choice above 90% and two above 80%. "Watson doesn't give you the answer," Kris says. "It gives you a range of answers." Then it's up to Kris to make the call. He regards the options on the screen and wonders how they might change if Ms. Yamato happened to develop a common symptom: hemoptysis, or coughing up blood.
"Let's try that," he says. He inputs the information and shows me the result approvingly. Watson has dropped one drug from the top chemo regimen. That's just what Kris would have done.
To make sense of all this--that is, to gauge both the value of Watson to a hospital like Sloan-Kettering and its potential to change forever the worlds of medicine and business--you could follow two different paths. You might consider Watson's evolutionary promise. Watson can almost certainly generate huge administrative benefits. Already, one large health insurer--Indiana-based Wellpoint--has begun using a Watson computer in its Virginia data center to speed along the authorization for medical procedures. Usually, authorizations are evaluated by a team of trained nurses and can sometimes take weeks to come through. Watsonizing the process would speed it up--a boon for a doctor like Kris, who now must wait while assistants exchange faxes with insurers before he can get clearance for any expensive tests.
Kris shows me what happens when Watson's treatment plan calls for an MRI. A button pops up on his screen to ask for preauthorization. "I just click that," he says, and it's done instantly.
I ask him what if Watson's request is denied.
Kris seems amused by the question. Watson has already consulted the latest medical literature, and it's been trained by the best cancer doctors in the world. "Who is the authority that is going to trump that?" he asks. Insurers balk at paying for unnecessary procedures; Watson's expert opinion essentially guarantees the necessity.
But the more intriguing path is the second one--a consideration of Watson's potential to do something revolutionary. This is the trail that captivates Kris. Eventually, he thinks, Watson could provide any doctor anywhere with the world's best second opinion. A physician in a community hospital in the Midwest, or at a remote medical center in China, could have instant access to everything that the medical field's best oncologists--people like Kris and his colleagues at Sloan-Kettering--have taught Watson. What is more, Watson will be able to excavate facts beyond the ken of Sloan-Kettering's current lineup of specialists. As Kris says, "We could ask Watson: What is the best treatment for this rare condition based on all of Sloan-Kettering's records?" It could then go through several years of cancer cases looking for the most successful outcomes. In time, it could even look at hospital records from around the world. As Manoj Saxena, the IBM executive now in charge of commercializing Watson, tells me: "It's like being able to take a knowledge worker--cancer specialist, nurse, bond trader, portfolio manager, whatever--and equip that person with the best knowledge, and have it available at their fingertips." As Watson evolves, Saxena believes, these knowledge banks will significantly alter how, and how well, humans make decisions.
Within a few years, for instance, Watson may be reaching well beyond oncology to assist patients suffering from any chronic disease and help general practitioners make diagnoses in their offices. Ultimately, Saxena believes, Watson could play an essential role in the diagnosis and treatment of mental health; in the financial services industry, where Citibank is testing it now; and in education. It could become the world's smartest dietitian.
How Watson Works
IBM's Watson computer begins trials in the health care industry this fall. The initial goal is to help oncologists make better decisions for cancer treatment; eventually, the computer will also aid in the diagnosis and treatment of other chronic diseases.
1. For well over a year, the Watson computers have been "trained" in science and medicine. Technicians feed Watson medical textbooks and journals, patient histories, and treatment guidelines.
2. At Memorial Sloan-Kettering Cancer Center in New York, doctors have begun using a Watson appon a tablet to access the computer through the cloud. The doctor logs in to Watson and begins to input data and ask questions.
3. When the oncologist queries Watson about a course of treatment for a lung or breast cancer patient, the computer--with its ability to understand natural language--notes keywords in the query, such as the particular type of cancer and the genomic variant of the tumor.
4. Watson then springs into action, using its massively parallel processors to review millions of pages of text in seconds. It explores the patient's medical history, medications, and other existing conditions. It then combines this information with recent data from the patient's medical tests and may comb through studies of patient groups at Sloan-Kettering who have had similar types of cancer. It also reviews doctors' and nurses' notes, recent medical research, journal articles, and treatment guidelines.
5. Watson then generates hypotheses for treatment. On the tablet app, these appear as separate options with varying levels of confidence. For instance, Watson might score one treatment option--a combination of chemotherapy drugs--with a 95% confidence level, suggesting it would be the most sensible path. It might also highlight options with lower scores as alternative treatment courses. The doctor then weighs the options and makes the call.
Saxena now commands a team of about 200 people who are working to adapt Watson's skills for various IBM clients. He and I are discussing his progress over lunch one day near IBM's upstate New York headquarters when he leans back and tells me that after creating two successful tech startups, both of which he sold (the second to IBM), his current job is far and away the most meaningful endeavor of his life. Those startups, he confides, were exciting, important. "But this," he says of the Watson rollout, "this is stuff that is going to change the course of history."
Over the past year, IBM executives have come to believe that Watson represents the first machine of the third computer age, a category now referred to within the company as cognitive computing. As Kelly describes it, the first generation of computers were tabulating machines that added up figures. "The second generation," he says, "were the programmable systems--the mainframe, the first IBM 360, PCs, all the computers we have today." Now, Kelly believes, we've arrived at the cognitive moment--a moment of true artificial intelligence. These computers, such as Watson, can recognize important content within language, both written and spoken. They do not ask us to communicate with them in their coded language; they speak ours. And perhaps most important, they can learn, so they improve without constant human instruction.
Siri, on the iPhone, might be considered an elementary example. Watson is industrial strength. "Computers do numerical calculations, they move data around, and they've been doing that forever," David Ferrucci, the IBM researcher who commanded the team that built the first Watson computer, tells me one day at IBM's research labs. "When I think about Watson, it's interpreting the information in human terms. It's saying: What does this mean to me? And that's a big deal." Also significant is how Watson renders an answer. Unlike its responses in Jeopardy!, in the real world it will perform as it did for Kris at Sloan-Kettering--by giving not a single solution but a range of probable solutions, each backed up by Watson's evidence and ranked by its level of confidence. In the lingo of computer science, that makes the machine probabilistic rather than deterministic. One might say this trait gives Watson a humanizing glow of humility and diminishes concerns that it marks a stride toward a computer-led dystopia. Watson, in IBM's marketing schema, is here to help with our questions, rather than solve them. In the case of medicine, it--for Watson is not really a he--is here to support doctors, not replace them.
The Watson of today is not precisely the same machine that won in Jeopardy! IBM has fine-tuned its software and algorithms for medical applications (or, in the case of Citibank, financial services applications). Watson has shrunk, too, from a row of about a dozen server racks that would have filled a small bedroom to an assemblage about the size of a double-door refrigerator. But for all the concentrated power, it doesn't look like anything special. Its sleek black servers are standard IBM Power 750s. You could wander around Watson and regard its blinking lights, as I did on a quiet midsummer afternoon at IBM's research labs, and not think something unusual is happening inside it. But there is. The way Watson solves problems--or, rather, the way it looks for answers, simultaneously sending out thousands of inquiries in all directions and then scoring the evidence it collects--is different from how other computers work. One person at IBM likens Watson's process to (1) gathering hundreds or thousands of possible solutions from a vast data bank, (2) pouring them into a giant funnel, (3) stirring with a dash of algorithms, and (4) letting only the best drip out of the bottom.
At the moment, a half-dozen Watsons are scattered around the country. Some are on the premises of IBM clients, as with the insurer Wellpoint, while others are cloud based, which is how hospitals such as Sloan-Kettering will access Watson. "Effectively, there's no limit to how many Watsons there can be," Bernie Meyerson, IBM's VP of innovation, tells me. Watson is a creation of software, not hardware. "That's the beauty of it," he says.
Watson is different from big servers and mainframes in other ways, too. The best computers of today have the extraordinary processing power needed to create, say, complex supply chains for building a new automobile or planning a satellite launch. These machines are good at manipulating the vast amounts of clearly defined data--numbers and facts--known as structured information. But most of the world's information is more ambiguous and less precise and lies beyond their reckoning. "We now have this proliferation of what we call Big Data," Saxena, Watson's business manager, tells me, referring to the flood of information created by our computers, our electronic sensors, and ourselves. "Ninety percent of the world's information was created in the last two years," he says. "But 80% of that 90% is unstructured or semistructured information, like doctor's notes or product reviews on Amazon." This near infinitude also includes tweets, blogs, emails--all the noise and scribble of modern life. So any company that aspired to manage the data of all the world's businesses would today be able to analyze only a small part of it. Watson, though, is a genius at reading unstructured information. And it's precisely this facility that explains why IBM sees such a rich business opportunity here.
It likewise explains why medicine is a logical first choice. While some health information is indeed structured--think of blood-pressure readings or cholesterol counts--the vast majority is unstructured. This cache includes textbooks, medical journals, patient records, and nurse and doctor evaluations. In fact, medicine embodies so much unstructured information that its proliferation has, by the account of many medical professionals, far outstripped the ability of doctors to keep up. Neither better training nor continuing education could ever wholly remedy this problem. When I meet with Herbert Chase, a professor of clinical medicine at Columbia University who consulted with IBM during the early stages of the Watson project, he says it is "not humanly possible" for a busy doctor to keep abreast of the current literature.
One result of information overload is a high rate of misdiagnosis and consequently incorrect treatment. By some estimates, Saxena tells me, 20% of initial diagnoses of cancer are eventually altered. "Imagine the implications of cancer care if there is a one in five chance that for the next six months whatever therapy they're giving you is wrong," he says.
Deciding on a course of treatment is even tougher than making a diagnosis. "It's still possible for a doctor to know the ways that people get sick," says Chase, who is also a kidney specialist. "But what is unmanageable, and what has been for decades, is knowing what the best option is today." Some applications now available to doctors are meant to alleviate this problem; one popular web-based tool is named Isabel. But Watson, in Chase's view, reaches a different level of sophistication. "I'll give you an example of a test we thought up for Watson," he tells me one day in his Manhattan office. "A patient was pregnant, had Lyme disease, and was also allergic to penicillin. And Watson came up with a drug. The first thing I thought was, Watson made a mistake. That drug can't be given to someone allergic to penicillin." But Chase was wrong, not Watson. "My knowledge was about five years old," he says. "And in the past couple of years, all the muckety-mucks had reviewed all the studies and had concluded yes, you can give that drug to someone who's allergic to penicillin."
To Chase, this proves a point: If you're a patient, you don't want to believe your doctor doesn't know everything. But he or she doesn't, and can't. At its best, the dispensation of treatment is inefficient today. "At its worst," Chase says, "it's subpar, incorrect, wrong therapy," and doesn't reach the standard of care to which his profession aspires. "As you can imagine," he adds, "this is not something we like talking about."
Last year, IBM turned 100 years old, which sets it apart from West Coast counterparts like Amazon, Apple, Google, HP, and Microsoft--all younger and ostensibly the tech world's leading innovators. To delve into IBM's recent research, though, is to wonder if our perception of technological leadership sometimes suffers from the distortions of branding and familiarity. We use iPhones and search engines and laser printers every day. But IBM's technologies are lodged deeper within the infrastructure of daily life; you're tapping into them whenever you send an email, for instance, or log on to a website. IBM has been granted more patents than any other company in the world for 19 years in a row. Yet since getting out of the laptop business in 2004, it has not produced a single product that it sells directly to the consumer.
If you're a patient, you don't want to believe your doctor doesn't know everything. But he or she doesn't, and can't.
To understand how Watson figures into the company's culture of ideas, or to see how it represents the kind of large-scale innovation that arguably lies beyond the capabilities of any startup, it helps to understand what the company actually does these days. IBM has operations in 172 countries and an organizational chart that resembles a vast Soviet bureaucracy. It employs about 433,000 men and women. Though IBM still sells hardware--big mainframe computers, silicon chips, and supercomputers--mainly it makes money selling software and consulting services to businesses and governments. The company's strategy has been validated of late by its performance: IBM's stock price has been on an upward trek for the past five years, and its winning streak has attracted the likes of Warren Buffett, who last year decided the company merited an investment of $10.7 billion. Meanwhile, as one of the few global titans to invest staggering sums on R&D ($6 billion to $7 billion a year), IBM maintains one of the world's last great industrial laboratories. At its main research center in Yorktown Heights, New York, a jet-age dream of glass curtain walls and rusticated stone designed by the Finnish-American architect Eero Saarinen, IBM employs the bulk of what is likely the world's largest mathematics department, with 300 members. If you're looking for a new PC design, you're out of luck here. But if you're shopping around for a new or better algorithm, IBM can build you one.
Not everyone is impressed by the direction of IBM's management. A relentless focus on earnings and cost cutting has led to a significant offshoring of domestic jobs, and a vocal corps of disillusioned or laid-off IBMers regularly take to the web to lament that the company's best days are behind it. IBM has also had its share of technological stumbles, apparently bungling several high-profile government contracts in recent years (in Texas and Indiana, for example) that left the company embroiled in disagreements with unhappy clients. And though these flare-ups may be uncommon, the company otherwise rarely quickens the pulse, with a long-standing reputation for being slow, steady, reliable, and maybe a little dull. IBM doesn't have big growth spikes or ballyhooed product launches; rather, it has plodding, long-term client contracts built around its ability to help optimize, say, a company's global IT services or a public utility's electrical grid. The corporation moves along like a supertanker. "IBM's annual revenue base is huge--$100 billion," says Toni Sacconaghi, a technology analyst for Sanford C. Bernstein. "So to move the needle is tough. It's hard to find big new products."
The managers and engineers keep looking anyway. One way IBM tries to infuse the troops with a sense of mission is through its periodic attempts to create for itself a Grand Challenge, such as the construction of Deep Blue, a chess-playing computer, or, more recently, Watson. The Grand Challenges are focused and expensive efforts--IBM will not verify Watson's cost, but estimates put the sum between $100 million and $1 billion--to push the company beyond the competition.
Watson's origins can arguably be traced back some years to a more modest annual initiative IBM calls the Global Technology Outlook, or GTO. Anyone at IBM can contribute to the outlook, and most of the results are eventually made public. The GTO tries to identify future business opportunities by putting a spotlight on various technology trends. A while ago, the IBM outlook pointed to analytics as a potentially huge field. Not long after, then-CEO (and current chairman) Sam Palmisano green-lighted IBM's acquisition of about $16 billion in smaller companies that had computer technologies to do this kind of work--essentially, to comb through vast stores of data, both structured and unstructured, and help extract nuggets from the global corporate babel.
Like Big Data or cloud computing, analytics is one of those contemporary catchphrases that everyone talks about but no one pauses to define. Bernie Meyerson, IBM's VP of innovation, argues that the great promise of analytics is not just to spot trends or glean information for boosting sales but to use computers and software to change the future. "Analytics is the capability to see what no human can," he says. Recently, at a public event, Meyerson was asked if IBM missed out by not building a tablet to compete with the iPad. He responded that as part of its Smarter Cities Initiative, IBM had just spent several years gathering all of the data on car transportation in Singapore; it then fed the data into a model it had built to predict the time and location of traffic jams. "We know from history what happens in Singapore if you slow the lights down in one direction by three seconds, and how to tweak the model so the jam never happens," he told his questioner. "And so there will be a traffic jam that never occurs because we can predict what happens 20 minutes from now, because we can take enough Big Data and crunch it, and do analytics on it. So we're predicting the future, and changing it. And you're asking me if I'm worried about a tablet?"
Watson, too, fits into Meyerson's conception of analytics, though it aims to change not the future of a traffic jam but of illness and investing. And by all indications, that tantalizing promise is not lost on the business community. "I have my shoulder against the door," Saxena tells me. He means he is turning clients away--something I heard from several other sources, too--until IBM executives feel confident Watson has proved its credibility at places like Wellpoint and Sloan-Kettering. Saxena seems certain that Watson will be a multibillion-dollar business, though he will only go so far as to say that by 2015, IBM will have annual revenues of about $16 billion from its analytics portfolio, of which Watson will be a part. When I put the question of Watson's potential to John Kelly, IBM's chief of research, he says: "It's like asking, at the very beginning, How big will the PC industry be?"
Kelly notes that the business model for Watson is still to be determined. He isn't sure whether selling Watson as a computer or marketing it as a service will make the most sense. But he feels he has time to decide. None of IBM's competitors, more than a year after the Jeopardy! victory, has announced a Q&A technology like Watson. "I think we have a huge lead," Kelly tells me. "When people realize this is not a one-off game machine but a new era of computing, then you'll see other companies tripling down to catch up."
I asked a number of people, both within IBM and outside of it, whether other organizations could have built this machine first. The consensus was probably not. The reasons did not precisely connect to IBM's technological capabilities--Google and Microsoft have plenty of computer prodigies in their ranks too. Rather, it was the combination of assets at IBM that made the difference. The company had its vast corporate lab, huge sums it was ready to invest, a profound expertise in hardware as well as software, and a collaborative culture that brought in lots of help from academia. And crucially, it had its business clients. In this respect, being a company that doesn't cater to consumers has advantages. Watson is only as bright as its teachers. Without the staff at Sloan-Kettering, where doctors like Mark Kris teach it oncology, Watson would not be nearly so smart. In fact, it might be kinda dumb. Or it might get all sorts of things wrong, like Siri does, except you'll be looking not for a pizza parlor but for a tumor.
From the start, the team that originally built Watson under David Ferrucci has worked out of a big room on the second floor of IBM's Hawthorne Labs in Westchester County, New York. Hawthorne is a large glass cube of a building situated about 30 miles north of New York City. Inside the Watson work space are five fake wood-grained tables, each home to a group of computer engineers who sit around and alternately immerse themselves in their screens or break to discuss coding with a neighbor. The mood here is sober. The staffers bring water bottles, not junk food. These aren't the unlined faces you'll see at a startup. Indeed, Ferrucci, who sits off to the side, is a suburban dad who looks like he'd be just as comfortable standing in front of a grill with a basting brush as he is overseeing his team. The walls here are covered with huge whiteboards crammed with the hieroglyphics of computer science. Overhead lights cast the room in gloomy fluorescence. The place has the neglected feel of a finished basement in a 1970s-era subdivision.
In early fall, the Watson team, now about 45 strong, began moving its work to a gleaming new space in IBM's main Yorktown Heights research laboratory--a promotion that reflects their importance as they support Saxena's much larger business development group while simultaneously working on the next iteration of Watson, known as Watson 2.0. One of the team's goals is to make Watson adaptable enough so that it doesn't require several dozen people spending a year to get it ready for every new application, such as medicine or financial services. But a more immediate project is to help Watson through the U.S. Medical Licensing Examination, the complex test all med-school graduates must take before practicing. If it passes, says Ferrucci, "that doesn't mean I can have a computer be a doctor." But IBM would gain what he calls "a crisp metric" that proves Watson has a real proficiency in medicine. The credential would no doubt help Watson's standing with health insurers, doctors, and patients, too. Passing the licensing exam is a difficult task--far harder than winning at Jeopardy!--but in early September, Ferrucci seemed pleased by the results. The computer is doing "interestingly well," he said. He sounded confident that Dr. Watson will ace the test by year's end.
Harder to intuit is how soon afterward Watson will infiltrate society. When I ask Jaime Carbonell, a computer science professor at Carnegie Mellon, he says he has no doubt the impact of Watson will be significant. "But I don't think there will be one moment of, 'Now we have it and yesterday we didn't,'" Carbonell remarks. "It will take time to permeate. Like cell phones, which were big, clumsy things you could barely carry at first." Was there a year, or month, or day, he asks, when cell phones began to change the world? "I can't think of when that was," he says. "But now we can't do without them."
Such is the course of technology: Electronic tools initially available only to the elite grow ever faster, smaller, cheaper. Kelly tells me he believes that eventually Watson will shrink to the size of a handheld device. Randy Katz, a computer science professor at UC Berkeley, sees a more approachable Watson, too. "Can the person in the street ask Watson a question now? No, he can't," says Katz. "But in five or 10 years, will there be systems like that--like Siri, but much better? I think the answer is yes."
In many of my conversations at IBM, the talk often drifts to applications of Watson. All sorts of intriguing scenarios are presented to me--for instance, that Watson will soon analyze not just words but images, such as MRIs and EKGs. Or it will diagnose a spider bite on a child's arm in a crop field in Africa, transmitted via smartphone by his worried father to a U.S. hospital. One afternoon, Saxena suggests this one: When you think you're coming down with the flu, Watson will be able to discern, before you even arrive at the doctor's office, that it might be a ragweed allergy, based on your medical record (you've had the same symptoms twice before at this time of year); your symptoms (gleaned from the insurance claim and diagnostic information in journals); and recent news (it just read an article in the Austin-American Statesman on a ragweed outbreak near your hometown).
It all sounds amazing. It's also speculative. Watson has not yet saved a life or a dollar of medical costs, or added anything, really, to IBM's bottom line. It has not yet faced its resistors--doctors who may find the technology objectionable and slow its adoption. It has not yet, as Saxena believes it will, changed the course of history. It has only won a television game show.
Still, Saxena predicts the computer will begin to scale up dramatically late next year. "By then," he says, "we will have built the technology, demonstrated it, built the tooling and methods around it. We will have the recipe book, and then we'll just push it out." But he will only have reached the end of Watson's beginning.
A version of this article appears in the November 2012 issue of Fast Company.
Monday, August 13, 2012
FT: Big data is watching you
Gillian Tett, The Financial Times, August 10, 2012
Information from mobile devices is not just changing the way the western world lives
During the past few years, there has been one constant in my life. Wherever I have travelled, in America or anywhere else, I have carried a mobile phone. Usually, the device is so ubiquitous, I do not think about this habit at all (except when I lose my phone and panic).
But last week, I took part in a seminar organised by America’s Brookings Institution and Blum Center to discuss development and global economics. And now I am looking at that mobile phone with fresh eyes. For what became clear in discussions with aid workers, healthcare officials and US diplomats is that those oft-ignored mobile devices are not just changing the way the western world lives – but changing the lives of poor societies, too. This, in turn, has some intriguing potential to reshape parts of how the global development business is done.
These days, there are about 2.5 billion people in emerging markets countries who own a mobile phone. In places such as the Philippines, Mexico and South Africa, mobile phone coverage is nearly 100 per cent of the population, while in Uganda it is 85 per cent. That has not only left people better connected than before – which has big political and commercial implications – it has also made their movements, habits and ideas far more transparent. And that is significant, given that it has often been extremely hard to monitor poor societies in the past, particularly when they are scattered over large regions.
Consider what happened two-and-a-half years ago when the Haitian earthquake struck. The population scattered when the tremors hit, leaving aid agencies scrambling to work out where to send help. Traditionally, they could only have done this by flying over the affected areas, or travelling on the ground. But some researchers at Columbia University and the Karolinska Institute took a different tack: they started tracking the Sim cards inside mobile phones owned by Haitians, to work out where their owners were located or moving. That helped them to “accurately analyse the destination of more than 600,000 people who were displaced from Port au Prince”, as a UN report says. Then, when a cholera epidemic hit Haiti later, the same researchers tracked the Sim cards again, to put medicine in the correct locations – and prevent the disease from spreading.
Aid groups are not just tracking those physical phones; they are also starting to watch levels of mobile phone usage and patterns of bill payment, too. If this suddenly changes, it can indicate rising levels of economic distress, far more accurately than, say, GDP data. Inside the UN, the secretary general is now launching a project called Global Pulse to screen some of the 2.5 quintillion bytes of so-called “big data” being generated each day around the world, including on social media sites such as Twitter and Facebook. These sites are strikingly popular in parts of the emerging markets world; Indonesia, for example, has one of the most Twitter-addicted populations on the planet. Thus if the UN (or anyone else) spots a sudden increase in certain keywords, this can also provide an early warning of distress. References to food or ethnic strife, for example, may indicate the onset of famine or civil unrest. Similarly, medical researchers have learnt in the past couple of years that social media references to infection area are powerful early warning signal of epidemics – and more timely than official alerts from government doctors.
Such developments are – unsurprisingly – controversial. For just as the spread of social media has sparked a blizzard of concern about privacy in the west, some observers worry about the dark side of this technological revolution in the emerging markets as well. Not everyone who may want to track these data is benign. Facebook might allow activists to express opposition to governments (as in the Arab spring), but social media data could also help repressive governments monitor their populations. Companies can use the data, too; there are initiatives under way to use it to develop credit scores for the poor.
But such concerns are not deterring the UN. On the contrary, Robert Kirkpatrick, a former IT expert who now runs the UN’s Global Pulse unit, argues that we should treat those 2.5 quintillion bytes of big data as an international common good. He dreams of using these data to create the social media equivalent of “metereological stations”, which can test the winds of public debate, spot economic trends and predict looming problems in a beneficial way. Even if this idea sounds far-fetched, economists can already use this information to track how economies are developing in poor regions of the world with much more precision and timeliness than ever before.
That mobile phone in my pocket, in other words, does not just connect me to my friends. It is now part of a shared human experience – and database – that spans the globe, and which is growing in depth and power each day. And that has implications most of us have barely begun to understand. It is both a sobering and exciting thought, whether you are now sitting on a holiday beach, in a humdrum office – or anywhere else in the world.
Information from mobile devices is not just changing the way the western world lives
During the past few years, there has been one constant in my life. Wherever I have travelled, in America or anywhere else, I have carried a mobile phone. Usually, the device is so ubiquitous, I do not think about this habit at all (except when I lose my phone and panic).
But last week, I took part in a seminar organised by America’s Brookings Institution and Blum Center to discuss development and global economics. And now I am looking at that mobile phone with fresh eyes. For what became clear in discussions with aid workers, healthcare officials and US diplomats is that those oft-ignored mobile devices are not just changing the way the western world lives – but changing the lives of poor societies, too. This, in turn, has some intriguing potential to reshape parts of how the global development business is done.
These days, there are about 2.5 billion people in emerging markets countries who own a mobile phone. In places such as the Philippines, Mexico and South Africa, mobile phone coverage is nearly 100 per cent of the population, while in Uganda it is 85 per cent. That has not only left people better connected than before – which has big political and commercial implications – it has also made their movements, habits and ideas far more transparent. And that is significant, given that it has often been extremely hard to monitor poor societies in the past, particularly when they are scattered over large regions.
Consider what happened two-and-a-half years ago when the Haitian earthquake struck. The population scattered when the tremors hit, leaving aid agencies scrambling to work out where to send help. Traditionally, they could only have done this by flying over the affected areas, or travelling on the ground. But some researchers at Columbia University and the Karolinska Institute took a different tack: they started tracking the Sim cards inside mobile phones owned by Haitians, to work out where their owners were located or moving. That helped them to “accurately analyse the destination of more than 600,000 people who were displaced from Port au Prince”, as a UN report says. Then, when a cholera epidemic hit Haiti later, the same researchers tracked the Sim cards again, to put medicine in the correct locations – and prevent the disease from spreading.
Aid groups are not just tracking those physical phones; they are also starting to watch levels of mobile phone usage and patterns of bill payment, too. If this suddenly changes, it can indicate rising levels of economic distress, far more accurately than, say, GDP data. Inside the UN, the secretary general is now launching a project called Global Pulse to screen some of the 2.5 quintillion bytes of so-called “big data” being generated each day around the world, including on social media sites such as Twitter and Facebook. These sites are strikingly popular in parts of the emerging markets world; Indonesia, for example, has one of the most Twitter-addicted populations on the planet. Thus if the UN (or anyone else) spots a sudden increase in certain keywords, this can also provide an early warning of distress. References to food or ethnic strife, for example, may indicate the onset of famine or civil unrest. Similarly, medical researchers have learnt in the past couple of years that social media references to infection area are powerful early warning signal of epidemics – and more timely than official alerts from government doctors.
Such developments are – unsurprisingly – controversial. For just as the spread of social media has sparked a blizzard of concern about privacy in the west, some observers worry about the dark side of this technological revolution in the emerging markets as well. Not everyone who may want to track these data is benign. Facebook might allow activists to express opposition to governments (as in the Arab spring), but social media data could also help repressive governments monitor their populations. Companies can use the data, too; there are initiatives under way to use it to develop credit scores for the poor.
But such concerns are not deterring the UN. On the contrary, Robert Kirkpatrick, a former IT expert who now runs the UN’s Global Pulse unit, argues that we should treat those 2.5 quintillion bytes of big data as an international common good. He dreams of using these data to create the social media equivalent of “metereological stations”, which can test the winds of public debate, spot economic trends and predict looming problems in a beneficial way. Even if this idea sounds far-fetched, economists can already use this information to track how economies are developing in poor regions of the world with much more precision and timeliness than ever before.
That mobile phone in my pocket, in other words, does not just connect me to my friends. It is now part of a shared human experience – and database – that spans the globe, and which is growing in depth and power each day. And that has implications most of us have barely begun to understand. It is both a sobering and exciting thought, whether you are now sitting on a holiday beach, in a humdrum office – or anywhere else in the world.
Tuesday, July 10, 2012
The Power of Information
What’s the big deal about the information economy anyway? Surely, information has played a role in commerce since ancient times. What’s changed?
One reason for the confusion is that we’re not used to making the distinction between information and knowledge. Go to Istanbul’s ancient Grand Bazaar and you’ll instantly grasp that the traders know a lot, but not much that they can easily share even if they want to. That, it turns out, makes all the difference.
The emergence of information as a storable, fungible entity is transforming our economy and our society in ways that we scarcely realize. It’s making us richer, smarter and even healthier. What’s more, its impact is accelerating, so we’ll see a lot more change in the coming decades than we have in the past. In fact, we’re just getting started.
A Theory of Information
If you had to put a date on it, the digital age truly began in 1948. It was in that year that two seismic events happened, both at Bell Labs. The first and the more famous was the invention of the transistor, which forms the basis for our present (albeit soon to be defunct) digital technology paradigm.
The lesser known, but in some ways more important event, was Claude Shannon’s groundbreaking paper A Mathematical Theory of Communication, which launched information theory out of thin air, seemingly with no precursor.
Although not widely publicized or understood at the time, it has become central to all the digital technology we use today.
The basic idea was that information isn’t a function of content, but rather a lack of randomness, which can be broken down to a single unit – a choice between two alternatives. Much like a coin toss which lacks information while in the air, but takes on a level of certainty when it lands, information arises when ambiguity disappears.
He called this unit, a “binary digit” or a bit and much like the pound, quart, meter or liter, has become such a basic unit of measurement that it’s hard to imagine our modern world without it.
Storage and Transfer
In the late 1940’s, Shannon’s colleague at Bell Labs, Richard Hamming became frustrated that computer errors continually ruined his work. He built upon Shannon’s paper and created Hamming code, small bits of extra information that would allow computers to detect and correct errors.
Of course, that increased the amount of information that needed to be processed. No problem, information theory also shows us how to compress information by eliminating redundancies. Common technologies that we have come to use everyday, like JPEG and MP3 are, in fact compression techniques that have their roots in Shannon’s 1948 paper.
It is storage and transfer that make the information age so different from what we knew in the past. In contrast to the knowledge that a bazaar trader possesses, computers can duplicate and transfer information an infinite number of times, with as little error as we choose.
Accelerating Returns
The ability to store and transfer information efficiently has an interesting side effect – we can improve at an exponential pace. Returns to our efforts not only increase, they accelerate. The most famous example is Moore’s Law, which says that processing speeds double about every 18 months.
Look at the chart above and notice the logarithmic scale. From 2000 to 2010, the number of transistors on a single chip increased from 10 million to a billion. That’s a hundred times more than the increase the previous decade and nearly a million times more than the decade before that. In the next ten years, we can expect it to increase 100 -fold again.
What’s even more amazing and also of paramount importance, is that the principle isn’t exclusive to processing, but applies to every facet of information technology, from storage to bandwidth to power consumption, everywhere you look, efficiency continues to improve exponentially.
The Information Invasion
Here’s where it gets really interesting. As information technology becomes more widely deployed, the information content of other products and services increases and they begin to follow the same exponential trends. Take a look at genome sequencing:
As sequencing genomes became less of a pure biological science and more of an information science, the pace of advancement changed drastically enabling a whole new field of bioinformatics that will revolutionize medicine.
Similar trends are being played out across almost every industry you can think of. In manufacturing, new technologies like 3D printing and (eventually) programmable matter are creating what The Economist calls a third industrial revolution. Even when you go and buy a box of cereal at Wal-Mart, a good portion of the retail price is made up of informationally dense logistics.
Ray Kurzweil spoke volumes when he said that in the future “all technologies will essentially become information technologies, including energy.”
From Belief to Observation
As the informational content of products and services increases and returns accelerate, lots of good things happen. Science fiction becomes engineered fact. The unthinkable will become commonplace. Incomes will rise while poverty falls. Seemingly intractable problems will be solved and dire needs will be met.
Yet there is also quite a bit that is unsettling. As technological cycles shorten, business models will have shorter life spans. The internalized experiences that we have come to regard as intuition will fail us more often. Our ability to plan will diminish and the need to experiment (and fail) will increase.
And that’s what’s disconcerting. This new information economy doesn’t run on beliefs or even, to a certain extent on ambition, but algorithms which, powered by ever more abundant processing power, test and accept or discard a dizzying multitude of possibilities, the results of which can be retrieved and recombined with other experiments.
The power of information means that we are no longer required to believe, only to imagine, test and observe.
Wednesday, May 30, 2012
Mary Meeker's Annual Internet Trends Report
The Wall Street Journal, May 30, 2012
Mary Meeker, the former Internet analyst known during the dot-com boom as the “Queen of the Net,” has released her annual (and always massive) slide deck on the latest Internet trends.
The 112 slides are heavily focused on rapid mobile adoption and runs through a number of examples of business models that are being re-invented thanks to fast-evolving devices, better connectivity and new interfaces.
Meeker, who joined venture-capital firm Kleiner Perkins from Morgan Stanley last year, cuts to the hear of Kleiner Perkins’ push into what it calls the “third wave of innovation,” combining social networking, mobile and e-commerce.
Friday, May 18, 2012
Googling cancer: search algorithms can scan disease for patient risk
Winter C, Kristiansen G, Kersting S, Roy J, Aust D, et al. (2012) Google Goes Cancer: Improving Outcome Prediction for Cancer Patients by Network-Based Ranking of Marker Genes. PLoS Comput Biol 8(5)
The algorithm Google uses to rank search results can now scan cancers to see which molecules best reveal the risks patients face, researchers have found, Txchnologist reports.
NetRank
The algorithm Google uses to rank which results pop up first in search queries, PageRank, orders results based on how other web pages are connected to them via hyperlinks.
Researchers modified PageRank to develop NetRank, which scans how genes and proteins in a cell are similarly connected through a network of interactions with their neighbors — “‘friends’ in the social network analogy,” said researcher Christof Winter, a medical doctor and computational biologist at Lund University in Sweden.
The investigators focused on pancreatic cancer, the most common form of which, pancreatic ductal adenocarcinoma, accounts for approximately 130,000 deaths each year in Europe and the United States. Very few tests exist to find out a prognosis for the disease — how it might progress, whether a patient might live or die.
The researchers used NetRank on about 20,000 proteins to see which ones were the best indicators for survival. They identified seven proteins that could help assess how aggressive a patient’s tumor is and guide clinicians to decide if the prognosis was worth trying chemotherapy or not.
As to how accurate prognoses based on these seven markers were, roughly speaking, “our markers are right in two-thirds of cases, and wrong in one-third,” Winter said. These markers were 6 to 9 percent more accurate at prognoses compared with those relying on conventional clinical parameters.
In addition to improving prognoses of cancer, this research could also help identify new targets to help destroy tumors.
Currently Winter and his colleagues are analyzing DNA and RNA data from breast cancer. The hope is “to develop a DNA-based prognostic blood test for breast cancer patients,” he said.
Sunday, January 22, 2012
Google: Figuring out how much the Web is worth
Dorothy Chou, Senior Policy Analyst at Google, January 20, 2012
Today we're launching a website called Value of the Web to collect research that sheds new light on how the Internet affects our world. It's available in English, French, German, Russian and Spanish and currently features studies that focus on 17 different regions, the value of cloud computing in Europe and the value of search around the world. While we can't use industrial metrics to fully capture the Web's contributions to our information society, as my teammate Jonathan pointed out, these reports are the best existing efforts to quantify the Internet's contributions to the economy and society thus far.
The value that is calculated in the reports ranges from the GDP contribution of the firms who provide the essential hardware and software to power the Internet, to jobs that are created due to the low cost of IT for small businesses enabled by cloud computing.
With two billion people online today and another five billion set to join them in the next 20 years, studies predict that the Internet's contributions will be large, increasing and distributed across sectors and people in the global economy.
For example, McKinsey found that Internet search, in its broadest form, accounts for $780 billion in value across the globe each year. And only 4% of that total goes to search companies—the rest goes to consumers and corporations who harness search in order to improve the way they find and use information every day.
In some cases, the reports show the enormous potential of getting more businesses online if governments take steps to encourage commercial use of the Internet or increasing access to broadband.
In other cases, the findings project exponential growth for economies that are already engaging in e-commerce online. The Boston Consulting Group predicts that by 2015, at least 10% of the British economy will be Internet-based. Universal broadband access and creating new business models that capture consumer surplus could increase the value added by the Internet by roughly £43 billion, which is just less than half what the British government spends on education today. If similar measures are adopted by the Japanese government, small businesses alone will contribute an additional ¥5 trillion to the Japanese economy in the next five years—not to mention the ripple effects.
We hope the site will become a central repository for insight derived from new measurements and data that move toward a more complete understanding of the Web's impact. In order to fully harness the power of this medium, we need to start using these numbers to illuminate policy decisions and light a pathway for innovation.
We'll continue developing the site by adding more improvements over time, including more languages and content. Check back frequently for updates or choose to subscribe for alerts via email.
posted by
Today we're launching a website called Value of the Web to collect research that sheds new light on how the Internet affects our world. It's available in English, French, German, Russian and Spanish and currently features studies that focus on 17 different regions, the value of cloud computing in Europe and the value of search around the world. While we can't use industrial metrics to fully capture the Web's contributions to our information society, as my teammate Jonathan pointed out, these reports are the best existing efforts to quantify the Internet's contributions to the economy and society thus far.
The value that is calculated in the reports ranges from the GDP contribution of the firms who provide the essential hardware and software to power the Internet, to jobs that are created due to the low cost of IT for small businesses enabled by cloud computing.
With two billion people online today and another five billion set to join them in the next 20 years, studies predict that the Internet's contributions will be large, increasing and distributed across sectors and people in the global economy.
For example, McKinsey found that Internet search, in its broadest form, accounts for $780 billion in value across the globe each year. And only 4% of that total goes to search companies—the rest goes to consumers and corporations who harness search in order to improve the way they find and use information every day.
In some cases, the reports show the enormous potential of getting more businesses online if governments take steps to encourage commercial use of the Internet or increasing access to broadband.
In other cases, the findings project exponential growth for economies that are already engaging in e-commerce online. The Boston Consulting Group predicts that by 2015, at least 10% of the British economy will be Internet-based. Universal broadband access and creating new business models that capture consumer surplus could increase the value added by the Internet by roughly £43 billion, which is just less than half what the British government spends on education today. If similar measures are adopted by the Japanese government, small businesses alone will contribute an additional ¥5 trillion to the Japanese economy in the next five years—not to mention the ripple effects.
We hope the site will become a central repository for insight derived from new measurements and data that move toward a more complete understanding of the Web's impact. In order to fully harness the power of this medium, we need to start using these numbers to illuminate policy decisions and light a pathway for innovation.
We'll continue developing the site by adding more improvements over time, including more languages and content. Check back frequently for updates or choose to subscribe for alerts via email.
posted by
Monday, December 12, 2011
Big Data and Europe: Turning government data into gold
European Commission Press Release December 12, 2011
Internet Governance
From the Release: “The Commission has launched an Open Data Strategy for Europe, which is expected to deliver a EUR 40 billion boost to the EU's economy each year. Europe's public administrations are sitting on a goldmine of unrealised economic potential: the large volumes of information collected by numerous public authorities and services.
Member States such as the United Kingdom and France are already demonstrating this value. The strategy to lift performance EU-wide is three-fold: firstly the Commission will lead by example, opening its vaults of information to the public for free through a new data portal. Secondly, a level playing field for open data across the EU will be established. Finally, these new measures are backed by the EUR 100 million which will be granted in 2011-2013 to fund research into improved data-handling technologies.”
See also:
http://ec.europa.eu/information_society/policy/psi/index_en.htm
http://www.zdnet.co.uk/news/regulation/2011/12/12/reuse-of-public-data-to-get-easier-under-new-eu-rules-40094628/
Internet Governance
From the Release: “The Commission has launched an Open Data Strategy for Europe, which is expected to deliver a EUR 40 billion boost to the EU's economy each year. Europe's public administrations are sitting on a goldmine of unrealised economic potential: the large volumes of information collected by numerous public authorities and services.
Member States such as the United Kingdom and France are already demonstrating this value. The strategy to lift performance EU-wide is three-fold: firstly the Commission will lead by example, opening its vaults of information to the public for free through a new data portal. Secondly, a level playing field for open data across the EU will be established. Finally, these new measures are backed by the EUR 100 million which will be granted in 2011-2013 to fund research into improved data-handling technologies.”
See also:
http://ec.europa.eu/information_society/policy/psi/index_en.htm
http://www.zdnet.co.uk/news/regulation/2011/12/12/reuse-of-public-data-to-get-easier-under-new-eu-rules-40094628/
Monday, October 17, 2011
Jeff Jonas et al on Big Data
Big Data: Advancing the Art of Analytics
Federal NewsRadio Monday - 10/17/2011, 12:39pm ET
October 20th, 2010 at 11 AM
The application of knowledge discovery within the cloud is immensely powerful, but not inbuilt. We are collectively moving past the question of "what is cloud computing", and swiftly moving towards "how does the cloud enable advanced analysis against massive volumes of data?" With industry and government leveraging multiple clouds, how do we successfully share and search large collections of data across systems, departments, and geographies? Organizations will continue to discuss and better understand the analytic power and economies of cloud computing, in the sense of data storage, sharing, and management; but we are quickly discovering that creating knowledge from data is more than just a discussion of technology.
It's a discussion of what can be accomplished when massive data and cloud computing efficiencies combine to make advanced analysis and innovation possible.
Listen at http://www.federalnewsradio.com/?nid=240&sid=2080808
Panelists:
Michael Byrne- Geographic Information Officer, Federal Communications CommissionJeff Jonas- Chief Scientist, IBM Entity Analytics GroupDavid Mihelcic- Chief Technology Officer, Defense Information Systems AgencyChris Nissen- National Security Analysis Group, MITREMike Olson- Chief Executive Officer, Cloudera
Moderator: Chris Kelly - Senior Vice President, Booz Allen Hamilton
Federal NewsRadio Monday - 10/17/2011, 12:39pm ET
October 20th, 2010 at 11 AM
The application of knowledge discovery within the cloud is immensely powerful, but not inbuilt. We are collectively moving past the question of "what is cloud computing", and swiftly moving towards "how does the cloud enable advanced analysis against massive volumes of data?" With industry and government leveraging multiple clouds, how do we successfully share and search large collections of data across systems, departments, and geographies? Organizations will continue to discuss and better understand the analytic power and economies of cloud computing, in the sense of data storage, sharing, and management; but we are quickly discovering that creating knowledge from data is more than just a discussion of technology.
It's a discussion of what can be accomplished when massive data and cloud computing efficiencies combine to make advanced analysis and innovation possible.
Listen at http://www.federalnewsradio.com/?nid=240&sid=2080808
Panelists:
Michael Byrne- Geographic Information Officer, Federal Communications CommissionJeff Jonas- Chief Scientist, IBM Entity Analytics GroupDavid Mihelcic- Chief Technology Officer, Defense Information Systems AgencyChris Nissen- National Security Analysis Group, MITREMike Olson- Chief Executive Officer, Cloudera
Moderator: Chris Kelly - Senior Vice President, Booz Allen Hamilton
Tuesday, October 11, 2011
Big Data in the Dirt (and the Cloud)
Quentin Hardy The New York Times October 11, 2011, 12:01 am
Big data, the term for scanning loads of information for possibly profitable patterns, is a growing sector of corporate technology. Mostly people think in terms of online behavior, like mouse clicks, LinkedIn affiliations and Amazon shopping choices. But other big databases in the real world, lying around for years, are there to exploit.
A company called the Climate Corporation was formed in 2006 by two former Google employees who wanted to make use of the vast amount of free data published by the National Weather Service on heat and precipitation patterns around the country. At first they called the company WeatherBill, and used the data to sell insurance to businesses that depended heavily on the weather, from ski resorts and miniature golf courses to house painters and farmers.
It did pretty well, raising more than $50 million from the likes of Google Ventures, Khosla Ventures, and Allen & Company. The problem was, it was hard to sell insurance policies to so many little businesses, even using an online shopping model. People like having their insurance explained. The answer was to get even more data, and focus on the agriculture market through the same sales force that sells federal crop insurance.
“We took 60 years of crop yield data, and 14 terabytes of information on soil types, every two square miles for the United States, from the Department of Agriculture,” says David Friedberg, chief executive of the Climate Corporation, a name WeatherBill started using Tuesday. “We match that with the weather information for one million points the government scans with Doppler radar — this huge national infrastructure for storm warnings — and make predictions for the effect on corn, soybeans and winter wheat.”
The product, insurance against things like drought, too much rain at the planting or the harvest, or an early freeze, is sold through 10,000 agents nationwide. The Climate Corporation, which also added Byron Dorgan, the former senator from North Dakota, to its board on Tuesday, will very likely get into insurance for specialty crops like tomatoes and grapes, which do not have federal insurance.
Like the weather information, the data on soils was free for the taking. The hard and expensive part is turning the data into a product. Mr. Friedberg was an early member of the corporate development team at Google. The co-founder, Siraj Khaliq, worked in distributed computing, which involves apportioning big data computing problems across multiple machines. He works as the Climate Corporation’s chief technical officer. Out of the staff of 60 in the company’s San Francisco office (another 30 work in the field) about 12 have doctorates, in areas like environmental science and applied mathematics.
“They like that this is a real-world problem, not just clicks on a Web site,” Mr. Friedberg says.
He figures that the Climate Corporation is one of the world’s largest users of MapReduce, an increasingly popular software technique for making sense of very large data systems. The number crunching is performed on Amazon.com’s Amazon Web Services computers.
The Climate Corporation is working with data intended to judge how different crops will react to certain soils, water and heat. It might be valuable to commodities traders as well, but Mr. Friedberg figures the better business is to expand in farming. Besides the other crops, he is looking at offering the service in Canada and Brazil, or anywhere else that he can get decent long-term data. It’s unlikely he’ll get the quality he got from the federal government, for a price anywhere near “free.”
Big data, the term for scanning loads of information for possibly profitable patterns, is a growing sector of corporate technology. Mostly people think in terms of online behavior, like mouse clicks, LinkedIn affiliations and Amazon shopping choices. But other big databases in the real world, lying around for years, are there to exploit.
A company called the Climate Corporation was formed in 2006 by two former Google employees who wanted to make use of the vast amount of free data published by the National Weather Service on heat and precipitation patterns around the country. At first they called the company WeatherBill, and used the data to sell insurance to businesses that depended heavily on the weather, from ski resorts and miniature golf courses to house painters and farmers.
It did pretty well, raising more than $50 million from the likes of Google Ventures, Khosla Ventures, and Allen & Company. The problem was, it was hard to sell insurance policies to so many little businesses, even using an online shopping model. People like having their insurance explained. The answer was to get even more data, and focus on the agriculture market through the same sales force that sells federal crop insurance.
“We took 60 years of crop yield data, and 14 terabytes of information on soil types, every two square miles for the United States, from the Department of Agriculture,” says David Friedberg, chief executive of the Climate Corporation, a name WeatherBill started using Tuesday. “We match that with the weather information for one million points the government scans with Doppler radar — this huge national infrastructure for storm warnings — and make predictions for the effect on corn, soybeans and winter wheat.”
The product, insurance against things like drought, too much rain at the planting or the harvest, or an early freeze, is sold through 10,000 agents nationwide. The Climate Corporation, which also added Byron Dorgan, the former senator from North Dakota, to its board on Tuesday, will very likely get into insurance for specialty crops like tomatoes and grapes, which do not have federal insurance.
Like the weather information, the data on soils was free for the taking. The hard and expensive part is turning the data into a product. Mr. Friedberg was an early member of the corporate development team at Google. The co-founder, Siraj Khaliq, worked in distributed computing, which involves apportioning big data computing problems across multiple machines. He works as the Climate Corporation’s chief technical officer. Out of the staff of 60 in the company’s San Francisco office (another 30 work in the field) about 12 have doctorates, in areas like environmental science and applied mathematics.
“They like that this is a real-world problem, not just clicks on a Web site,” Mr. Friedberg says.
He figures that the Climate Corporation is one of the world’s largest users of MapReduce, an increasingly popular software technique for making sense of very large data systems. The number crunching is performed on Amazon.com’s Amazon Web Services computers.
The Climate Corporation is working with data intended to judge how different crops will react to certain soils, water and heat. It might be valuable to commodities traders as well, but Mr. Friedberg figures the better business is to expand in farming. Besides the other crops, he is looking at offering the service in Canada and Brazil, or anywhere else that he can get decent long-term data. It’s unlikely he’ll get the quality he got from the federal government, for a price anywhere near “free.”
Subscribe to:
Posts (Atom)

