Monday, May 30, 2011
CTO Amazon: Data Without Limits
This year NEXT is all about Data Love. ...Data is the resource for the digital value creation and fuel for the economy. Today, data is what electricity has been for the industrial age…
Business developers, marketing experts and agency managers are faced with the challenge to create new applications out of the ever-growing data stream with added value for the consumer. In our data-driven economy, the consumer is in the focus point of consideration. Because his behaviour determines who wins, what lasts and what will be sold. Data is the crucial driver to develop relevant products and services for the consumer.
(http://nextconf.eu/next11/programme/)
Data Without Limits by Werner Vogels, CTO Amazon
Talk at http://video.nextconf.eu/video/1880845/data-without-limits
ABOUT THE SPEAKER:
Werner Vogels, Vice President & CTO at Amazon.com, is responsible for driving the company's technology vision.
Prior to joining Amazon, he worked as a researcher at Cornell University where he was a principal investigator in several research projects that target the scalability and robustness of mission-critical enterprise computing systems. He has held positions of VP of Technology and CTO in companies that handled the transition of academic technology into industry.
Werner holds a Ph.D. from the Vrije Universiteit in Amsterdam and has authored many articles for journals and conferences, most of them on distributed systems technologies for enterprise computing. He was named the 2008 CTO of the Year by Information Week for his contributions to making Cloud Computing a reality. For his unique style in engaging customers, media and the general public, he received the 2009 Media Momentum Personality of Award.
See also
Stay Healthy: Big Data and Our Bodies by David Rowan, Wired UK
http://video.nextconf.eu/video/1879024/stay-healthy-big-data-and-our
Thursday, May 26, 2011
What Big Data Needs: A Code of Ethical Practices
Thursday, May 26, 2011
Four key principles that companies should follow if they hope to analyze customers' data without alienating them.
By Jeffrey F. Rayport
http://www.technologyreview.com/printer_friendly_article.aspx?id=37548
In this era of Big Data, there is little that cannot be tracked in our online lives—or even in our offline lives. Consider one new Silicon Valley venture, called Color: it aims to make use of GPS devices in mobile phones, combined with built-in gyroscopes and accelerometers, to parse streams of photos that users take and thus pinpoint their locations. By watching as these users share photos and analyzing aspects of the pictures, as well as ambient sounds picked up by the microphone in each handset, Color aims to show not only where they are, but also whom they are with. While this kind of service might prove attractive to customers interested in tapping into mobile social networks, it also could creep out even ardent technophiles.
Color illustrates a stark reality: companies are steadily gaining new ways to capture information about us. They now have the technology to make sense of massive amounts of unstructured data, using natural language processing, machine learning, and software architectures such as Hadoop, which handles high volumes of simultaneous search queries. Messy data of this kind, long relegated to data warehouses, is now the target of data mining. So is the information generated by social networks—user profiles and posts. Its quantity is staggering: a recent report from the market intelligence firm IDC estimates that in 2009 stored information totaled 0.8 zetabytes, the equivalent of 800 billion gigabytes. IDC predicts that by 2020, 35 zetabytes of information will be stored globally. Much of that will be customer information. As the store of data grows, the analytics available to draw inferences from it will only become more sophisticated.
It's no wonder that there are calls for corporations to create positions such as chief privacy officer, chief safety officer, and chief data officer, or that American and European legislators have been considering several kinds of privacy measures. In one bipartisan effort, Senators John McCain and John Kerry have proposed the Consumer Privacy Bill of Rights Act of 2011, which aims, in part, to restrict what online companies can do with customer data. Senator Jay Rockefeller has proposed his own piece of legislation, the Do-Not-Track Online Act of 2011. The European Union's Article 29 Working Group is addressing similar concerns.
In the private sector, the Digital Advertising Alliance has sought to get ahead of such rule-making by introducing its own privacy framework to assure the security and safety of customer information. Its Self-Regulatory Program for Online Behavior Advertising comes on the heels of several incidents: Epsilon's admission that hackers gained access to customer information from clients such as CitiGroup, Target, and Walgreen's; Sony's revelation that its PlayStation platform failed to safeguard the account information of up to 100 million customers; and Apple's confirmation that it uses an unencrypted file stored in iTunes accounts to track movements of individual iPhone users in the physical world.
For all the privacy concerns, the online economy creates enormous value by using customer information. In 2009, according an ad industry study cited by the Wall Street Journal, the average price of an untargeted ad online was $1.98 per thousand views. The average price of a targeted ad was $4.12 per thousand. We used to measure the success of websites as if they were portals—by how much traffic they could muster. Now we measure them as social networks—by how much they know about their users. This is why Wal-Mart recently acquired Kosmix, a Silicon Valley startup that filters and finds meaning in vast streams of Twitter messages. Other retailers, along with digital players such as Facebook and Yahoo, are using the technology of another startup, Cloudera, to sort through enormous quantities of behavioral information compiled over years (sometimes decades) in search of insights based on patterns that only machines can fathom. Intelligence generated in these ways can lead to better games from companies like Zynga and better advertising from your favorite brands. David Moore, the CEO of 24/7 Real Media, argues that when an ad is targeted properly, "it ceases to be an ad; it becomes important information."
The opportunity for profit helps explain the rise of dozens of data exchanges, data marts, predictive analytic engines, and other intermediaries. It's also why players such as Google, Facebook, and Zynga, among many others, are finding ways to aggregate ever more information about users. Facebook provides but one example of how extensive this kind of tracking can be. Its seemingly innocuous "Like" button has become ubiquitous online. Click on one of these buttons, and you can instantly share something that pleases you with your friends. But simply visit a page with a "Like" button on it while you're logged in to Facebook, and Facebook can track what you do there. The first aspect sounds great for consenting adults; the latter is more than a little unsettling. Facebook is hardly alone. A company called Lotame helps target online advertising by placing tags (sometimes known as beacons) on browsers to monitor what users are typing on any Web page they might view.
The potential dark side of Big Data suggests the need for a code of ethical principles. Here are some proposals for how to structure them.
Clarity on Practices: When data is being collected, let users know about it—in real time. Such disclosure would address the issue of hidden files and unauthorized tracking. Giving users access to what a company knows about them could go a long way toward building trust. Google has done this already. If you want to know what Google knows about you, go to www.google.com/ads/preferences, and you can see both the data it has collected and the inferences it, and third parties, have drawn from what you've done.
Simplicity of Settings: One way to avoid an Orwellian nightmare is to give users a chance to figure out for themselves what level of privacy they really want. In theory, Facebook does this. In practice, as Nick Bilton reported recently in the New York Times, Facebook's privacy policy has more words (5,830) than the United States Constitution (4,543, not counting the amendments). But that's just the tip of the iceberg. Try changing your privacy settings, and you will encounter over 50 privacy toggles giving rise to over 170 privacy options.
Privacy by Design: Some argue that neither clarity nor simplicity is sufficient. Ann Cavoukian, privacy commissioner for the province of Ontario, coined the phrase "privacy by design" to propose that organizations incorporate privacy protections into everything they do. This does not mean Web and mobile businesses collect no customer information. It simply means they make customer privacy a guiding principle, right from the start. Microsoft, which in 2006 issued a report called "Privacy Guidelines for Developing Software Products and Services," has embraced this principle, using a renewed emphasis on privacy as a way to differentiate itself; the latest version of Internet Explorer, IE9, lets users activate features that can block third-party ads and content.
Exchange of Value: Walk into a local Starbucks, and you're likely to feel flattered if a barista remembers your name and favorite beverage. Something similar applies on the Web: the more a service provider knows about you, the greater the chance that you'll like the service. Radical transparency could make it easier for digital businesses to show customers what they will get in exchange for sharing their personal information. That's what Netflix did in running a public competition offering third-party developers a $1 million award for creating the most effective movie recommendation engine. It was an open acknowledgement that Netflix was using users' movie-viewing histories to provide increasingly targeted, and thus more useful, recommendations.
These principles are by no means exhaustive, but they begin to outline how companies might realize the value of Big Data and mitigate its risks. Adopting such principles would also get ahead of policymakers' well-intentioned but often misguided efforts to rule the digital economy. That said, perhaps the most important rule is one that goes without saying, something akin to the Golden Rule: "Do unto the data of others as you would have them do unto yours." That kind of thinking might go a long way toward creating the kind of digital world we want-and deserve.
Jeffrey F. Rayport specializes in analyzing the strategic implications of digital technologies for business and organizational design. He is a managing partner of MarketspaceNext, a strategic advisory firm; an operating partner at Castanea Partners; and a former faculty member at Harvard Business School. Carine Carmy contributed research to this article.
Monday, May 23, 2011
Surowiecki/New Yorker on economic data
A Billion Prices Now
By James Surowiecki May 30, 2011 The Financial Page
Between official government statistics, industry surveys, and Wall Street forecasts, it often seems like we're drowning in data, often of uncertain value. But consider the alternative. In the early years of the Great Depression, it was clear that things were awful, but the government had few good figures to go on; there was no official G.D.P. number, and no solid information about unemployment. As a result, policymakers persistently underestimated the severity of the crisis. In June of 1930, relying on some anecdotal evidence of an upturn, Herbert Hoover announced, "The Depression is over."
And in his State of the Union address that December he said that two and a half million Americans were unemployed. But, as Hoover acknowledged, that number was eight months old. At the time of the speech, five million people were out of work, and a hundred thousand more were losing their jobs every week. Washington was making policy in the dark.
The government learned from experience, though. In 1934, a team of economists came up with the first measurement of national income, which developed into what we now know as G.D.P. The late thirties saw a more rigorous and systematic collection of unemployment data. And in the years after the Second World War—which accelerated the trend toward quantifying things—the amount of economic information available to policymakers grew exponentially. Today, our picture of the economy is more detailed and sophisticated than ever, and that makes it easier for businesses and the government to react quickly to changes in the economy.
And yet our picture of what's going on is far from perfect. The government continues to track inflation, for instance, by gathering price data much as it did in the nineteen-fifties: it surveys consumers by phone to see where they buy, surveys businesses to see how much they charge, checks out shopping malls to price goods. This leaves out consumers who have only cell phones, and it probably overstates inflation by not fully accounting for things like the impact of big-box stores. The larger problem, though, is the time it takes: the Consumer Price Index's figures don't come out until a month after the fact. In turbulent times, that's too slow.
A new venture called the Billion Prices Project may help change that. The B.P.P., which was designed by the M.I.T. economists Alberto Cavallo and Roberto Rigobon, gathers price data not via survey but, rather, by continuously scouring the Web for prices of online goods around the world. (In the U.S., it collects more than half a million prices daily—five times the number that the government looks at.) Using this information, Cavallo and Rigobon have succeeded in building what amounts to the first real-time inflation index. The B.P.P. tells us what's happening now, not what was happening a month ago. For instance, after Lehman Brothers went under, in September, 2008, the project's data showed that businesses started cutting prices almost immediately, which suggested that demand had collapsed. The government's numbers, by contrast, didn't show this deflationary pressure until that November. This year, there's been a mild uptick in annual inflation, and again the B.P.P. detected the new trend before the Consumer Price Index did. That kind of early heads-up could help governments make more timely decisions.
The B.P.P. can also help keep governments honest. In much of the world, as a 2010 study of developing countries found, governments regularly manipulate economic data—downplaying inflation, overstating job growth, and the like. The B.P.P. makes that more difficult by providing an independent check on the official numbers. And, while there's no evidence that this kind of thing happens in the U.S., if you're a conspiracy theorist you can now look to the B.P.P. rather than to the regular inflation number to see what's happening.
The B.P.P. doesn't offer a complete picture. In particular, it doesn't cover most services. But, even if it's unlikely to replace the C.P.I. anytime soon, it will almost certainly make the C.P.I. better. Indeed, it wouldn't be surprising if the Bureau of Labor Statistics, which has already said that some of its methods are out of date, begins to move toward a real-time model. The B.P.P., in its rough-and-ready way, is part of a data revolution. Cheap computing power and the Internet have made it possible for companies to do what previously only the government could. The Case-Shiller Index of home resale prices, for instance, has become the benchmark for U.S. house prices. The most concrete result of all this is that policymakers will be better informed than ever before. It also means that they'll have fewer excuses when they mess up.
That's the catch, of course. Giving policymakers more information doesn't mean that they'll believe it or act on it. In the years leading up to the financial crisis of 2008, after all, there was a lot that people didn't know, but the fundamental problems were obvious: housing prices were rising too fast and banks were flinging loans at unqualified borrowers with reckless abandon. Yet the Federal Reserve and banking regulators failed to take any action to try to pierce the bubble before it brought the economy crashing down. These days, all the available numbers, including the C.P.I. and the Billion Prices Project, suggest that inflation is under control. Still, politicians and some Fed members are fretting that a huge price spike may be imminent, and are pushing to make monetary policy tighter, even in the face of unacceptably high unemployment. An enormous amount has been done lately to make sure that policymakers have the numbers they need. The question is whether they'll use them. ♦
'Jeopardy!'-winning computer delving into medicine
A doctor who is helping to prepare IBM's Watson computer system for work as a medical tool says such blog entries may be included in Watson's database.
Watson is best known for handily defeating the world's best "Jeopardy!" players on TV earlier this year. IBM says Watson, with its ability to understand plain language, can digest questions about a person's symptoms and medical history and quickly suggest diagnoses and treatments.
The company is still perhaps two years from marketing a medical Watson, and it says no prices have been established. But it envisions several uses, including a doctor simply speaking into a handheld device to get answers at a patient's bedside.
Watson won't be the first such product on the medical market, however, and one rival company says it isn't impressed.
At a recent demonstration for The Associated Press, Watson was gradually given information about a fictional patient with an eye problem. As more clues were unveiled — blurred vision, family history of arthritis, Connecticut residence — Watson's suggested diagnoses evolved from uveitis to Behcet's disease to Lyme disease. It gave the final diagnosis a 73 percent confidence rating.
"You do get eye problems in Lyme disease but it's not common," Dr. Herbert Chase said. "You can't fool Watson."
For "Jeopardy!" Watson was fed encyclopedias, dictionaries, books, news, and movie scripts. For health care, it's on a diet of medical textbooks and journals. It could also link to the electronic health records that the federal government wants hospitals to maintain. Medical students are peppering it with sample questions to help train it.
Chase, a Columbia University medical school professor, says anecdotal information — such as personal blogs from medical websites — may also be included.
"What people say about their treatment ... it's not to be ignored just because it's anecdotal," Chase said. "We certainly listen when our patients talk to us, and that's anecdotal." Chase and other experts say cramming Watson with the latest medical information will help with a major problem in modern health care: information overload.
"For at least 30 years it's been clear that it's not possible for us to know everything," he said. "Every day, doctors have questions they can't find the answers to. Even if you sit down at a search engine, it's so labor intensive and it takes so long to find answers."
Carl Kesselman, director of the Health Informatics Center at the University of Southern California, says the "deluge of information" is a significant problem.
"Advances in medicine are increasing rapidly: genomics, specialized drugs, off-label uses, increasingly finer-grained classifications of disease," said Kesselman, who is not involved with the Watson project. "The ability to ask 'Jeopardy!'-style questions and get that kind of information retrieval, to sort through all the stuff out there and point you to the latest literature, would be of potentially huge value."
Michael Yuan, chief scientist at Ringful Health, a medical consulting company in Austin, Texas, that has worked with IBM, cited a 1999 study of 103 doctors that found they fielded more than 1,100 questions a day, of which 64 percent were never answered.
"That's a huge potential for people to make mistakes," he said. "Watson is the type of solution that can really reduce that."
In "Jeopardy!" Watson was asked for one correct answer, whether it was answering questions about Sir Christopher Wren, the Lion of Nimrud or the Church Lady from "Saturday Night Live."
But in its medical guise, when presented a set of symptoms, Watson offers several possible diagnoses, ranked in order of its confidence.
"In medicine, we don't want one answer, we want a list of options," Chase said.
Kesselman said having options might help doctors accept a computer's findings.
"Will a physician ever blindly accept a diagnosis coming out of a computer? I don't think that will happen anytime soon," he said.
Chase said seeing more than one choice might also help doctors move away from what he called "anchoring," or getting too attached to a diagnosis.
"If a person has a 95 percent chance of having disease X, there's still a one-in-20 chance that they have something else," he said. "We often forget what's in that 5 percent. But Watson won't."
The treatment application works much like the diagnosis application. In the demonstration, Watson first suggested the antibiotic doxycycline for treating Lyme disease, then switched to cefuroxime when told the patient was pregnant and allergic to penicillin.
Chase said Watson will know the latest treatment guidelines — which are complex and often updated — "and can see if they're not being met."
"You have to match the right treatment with each unique patient," Chase said. "You can't treat everybody with high blood pressure the same way — a 75-year-old man with prostate cancer who felt dizzy last week and a 32-year-old woman."
Yuan said Watson's influence will depend on "how widely it is adopted."
"You have to wonder if a hospital is going to plunk down a couple of million dollars," he said.
IBM's Dan Pelino, general manager for global health care, said clients won't have to buy a complete Watson system. He said possible future uses include:
An existing private medical database known as Isabel is already used by some multi-hospital health systems. Co-founder Jason Maude of Isabel Healthcare said that from what he's heard about IBM's plans for Watson, "It's kind of what we've had for about 10 years."
An online demonstration of Isabel showed similarities to the Watson model — symptoms are entered, and the computer searches through a database for a possible diagnosis. Maude, who named Isabel for a daughter who escaped a serious misdiagnosis as a child, says Isabel's database has been "tuned and honed" over time.
He said prices for using Isabel range from a few thousand dollars a year for a family practice to as much as $400,000 for a health system.
Pelino said Watson is much faster and Chase said Watson is better at understanding non-medical terms.
"Watson knows that 'difficulty swallowing' is 'dysphagia,'" he said.
Isabel has been used at the Orlando Health hospital network in Florida since last fall, and "has had its successes," said Dr. Jay Falk, chief academic medical officer. He said less experienced doctors use it under the guidance of senior clinicians "who can make some judgments about the likelihood of what's given on the list of diagnoses."
"There's no question that there's a need for a tool that will help in this regard," Falk said. "Whether Isabel itself is the answer is unclear." Overall, he said, "We're enjoying learning with it."
IBM said Watson can answer some medical questions in the same few moments it took on "Jeopardy!" Yuan noted studies have shown that "If it takes more than two minutes, it won't get used."
As on "Jeopardy!" — where Watson identified Toronto as a U.S. city and Picasso as an art period — the computer occasionally bungles a medical question.
"I think once we were asking what type of drug we should use and the answer was a person's name," Chase said. "In fairness, I think it was a person associated with the drug."
And of course there are things Watson cannot do. It won't know a patient's appetite for risk, for example, or feelings about end-of-life treatment.
"That's why you have to emphasize that the decisions aren't coming from the computer, they're coming from the patient," Chase said.
Chase's suggestion that medical blogs be included may have something to do with his own medical history.
Several years ago, fighting a cholesterol problem, he took Lipitor and was soon plagued with insomnia. He suspected a connection but found nothing in textbooks or journals.
"I go to the blogosphere, and it was like, 'You moron, don't take Lipitor before you go to bed because you'll never sleep again!'
"Now it's five years later, and if you Google Lipitor and insomnia, it's all over the place," Chase said.
Copyright © 2011 The Associated Press. All rights reserved.
Friday, May 13, 2011
New Ways to Exploit Raw Data May Bring Surge of Innovation, a Study Says
Math majors, rejoice. Businesses are going to need tens of thousands of you in the coming years as companies grapple with a growing mountain of data.
Data is a vital raw material of the information economy, much as coal and iron ore were in the Industrial Revolution. But the business world is just beginning to learn how to process it all.
The current data surge is coming from sophisticated computer tracking of shipments, sales, suppliers and customers, as well as e-mail, Web traffic and social network comments. The quantity of business data doubles every 1.2 years, by one estimate.
Mining and analyzing these big new data sets can open the door to a new wave of innovation, accelerating productivity and economic growth. Some economists, academics and business executives see an opportunity to move beyond the payoff of the first stage of the Internet, which combined computing and low-cost communications to automate all kinds of commercial transactions.
The next stage, they say, will exploit Internet-scale data sets to discover new businesses and predict consumer behavior and market shifts.
Others are skeptical of the “big data” thesis. They see limited potential beyond a few marquee examples, like Google in Internet search and online advertising.
The McKinsey Global Institute, the research arm of the consulting firm, is coming down on the side of the optimists in a lengthy study to be published on Friday. The report, based on nine months of work is “Big Data: The Next Frontier for Innovation, Competition and Productivity.” It makes estimates of the potential benefits from deploying data-harvesting technologies and skills.
The McKinsey research unit, for example, says the value to the health care system in the United States could be $300 billion a year, and that American retailers could increase their operating profit margins by 60 percent.
But the study also identifies challenges. One hurdle is a talent and skills gap. The United States alone, McKinsey projects, will need 140,000 to 190,000 more people with “deep analytical” skills, typically experts in statistical methods and data-analysis technologies.
McKinsey says the nation will also need 1.5 million more data-literate managers, whether retrained or hired. The report points to the need for a sweeping change in business to adapt a new way of managing and making decisions that relies more on data analysis. Managers, according to the McKinsey researchers, must grasp the principles of data analytics and be able to ask the right questions.
“Every manager will really have to understand something about statistics and experimental design going forward,” said Michael Chui, a senior fellow at the McKinsey Global Institute.
The study estimates that the use of personal location data could save consumers worldwide more than $600 billion annually by 2020. Computers determine users’ whereabouts by tracking their mobile devices, like cellphones. The study cites smartphone location services including Foursquare and Loopt, for locating friends, and ones for finding nearby stores and restaurants.
But the biggest single consumer benefit, the study says, is going to come from time and fuel savings from location-based services — tapping into real-time traffic and weather data — that help drivers avoid congestion and suggest alternative routes. The location tracking, McKinsey says, will work either from drivers’ mobile phones or GPS systems in cars.
Personal location data raises privacy concerns. Both Google and Apple, for example, have faced protests recently for collecting location data without most users’ knowledge. The McKinsey report says such services should require that users have a choice and opt-in to use them, but the report does not deal with privacy issues in detail.
The sizable projected payoff for consumers, some experts say, is not surprising. “Much of the benefit of innovation always flows to consumers,” said Martin Baily, an economist at the Brookings Institution, who was an adviser on the study. “So the large consumer surplus makes sense.”
In health care, the biggest slice of the $300 billion gain is expected to come from more effectively using data to inform treatment decisions. The tools include clinical decision support to assist doctors, and comparative effectiveness research to make more informed decisions on drug therapy.
For example, the Department of Veterans Affairs and Kaiser Permanente save millions of dollars a year in treating many patients with high cholesterol with generic statins instead of branded statins, like Lipitor. But such tailored treatments require electronic health records for tracking results, and most of the nation’s hospitals and physicians still use paper records.
Skeptics say the economic payoff from harnessing big data sets is mostly wishful thinking so far. The nation’s technology-assisted increase in productivity began in 1995 and continued through 2004, having trailed off since, despite investments in data analytics.
“The big dividend mostly hasn’t arrived yet,” said Tyler Cowen, an economist at George Mason University.
The McKinsey authors say that the big-data trend is just getting under way. It will take years, they say, before the gains show up in the economic statistics, just as it did for computers to prove they were engines of productivity.
“But it’s clear that data is an important factor of production now,” said James Manyika, a director of the McKinsey Global Institute.
Wednesday, May 11, 2011
4 Trends Shaping the Emerging "Superfluid" Economy
Friday, May 6, 2011
Facebook Could Be Planning a Visual Dashboard of Your Life
By Christopher Mims
Ever wondered just how much coffee you drank last year, or which movies you saw, and when? New Web and mobile apps make it possible to track, and visualize, this personal information graphically, and the trend could be set to expand dramatically.
This is because Facebook recently acquired one of the leading personal-data-tracking mobile apps and hired its creators. The social-networking giant could be gearing up to offer users ways to chart the minutiae of their lives with personalized infographics.
Nick Felton and Ryan Case, two New York-based designers, have pioneered turning the mundane contours of an everyday life into a kind of visual narrative. Each year, Felton publishes an "annual report" on his own life: an infographic that charts out his habits and lifestyle in great detail.
Felton and Case have also created a mobile app, called Daytum, that lets users gather personal data and represent it using infographics. Daytum already has 80,000 users, whose pages provide a detailed snapshot of everything from coffee drinking habits to baseball stadium visits. The app gives users the ability to easily record their own information, whatever it might be, and display it in an attractive manner, whether or not they are a designer.
Daytum is part of a larger trend in tracking personal information. But traditional personal tracking applications tend to revolve around medical data, sleep schedules, and the like. In Felton's creative visualizations, even something as mundane as how many concerts he attended in the past year becomes a kind of art. "I think there's storytelling potential in data," he says.
Felton says he can't talk about what he'll be doing at Facebook, but says, "Clearly, companies like Facebook recognize the value of the kind of work we were doing."
At Facebook, users already engage in countless acts of data entry, so it's possible that the data Felton will be visualizing will already be available. Automated data gathering through smart phones—especially location data—provides even more data to mine.
Eventually—with users' permission—this kind of personal information could be mined by marketers and advertisers. Ted Morgan, CEO of the geolocation software company Skyhook, compares the trend to the way advertisers currently track some TV viewers' watching habits. In the future, he says, tracking data will be "like a Nielsen rating box for your life. It will track where you go and what you do. [Advertisers are] going to pay people to do this."
One company is already exploring this possibility. Locately offers users promotions and discounts if they agree to opt in to its mobile data-gathering network. This lets the company gather data on where people go, what they do, and what they buy; the company sells that data to businesses who want to use it for market research and advertising.
Gathering detailed personal data can produce surprising insights, says Felton. "The way people describe themselves is not really in line with their true behavior," he notes. For example, users who track what TV shows they actually watch may find that they spend more time on shows they don't identify as their favorites.
Perhaps this could lead to a whole new kind of friend discovery, one based not on our expressed interests, but on our actual interests. Picture a beefed-up version of Facebook's "people you might also know" feature informed not just by who you're connected to, but what behavior you have in common.
"It could be shared affinities that are not recognized by either [party]," says Felton. The downside, of course, is that people who are a lot like us often drive us crazy. "You might hate them," says Felton. "Isn't that part of what annoys us about our families?"
Copyright Technology Review 2011.
FT: A binary goldmine
Take an app launched recently by Color, one of the most ambitious and best financed of the crop of start-ups that has sprung up in Silicon Valley to cash in on the smartphone boom. Pictures taken by users are mixed into streams with those taken by others who are nearby, or with whom users are often in contact, building ad hoc social networks.
The software taps deeply into handsets, drawing on components such as Global Positioning System chips, gyroscopes and accelerometers to pinpoint where they are, how fast they are moving and which way up they are being held. The lighting conditions in pictures taken with the gadgets, along with the digital “fingerprints” of surrounding noises coming through their microphones, provides other useful crumbs of information.
Thus informed, Color can work out precisely who the user is walking down the street with, says Bill Nguyen, the serial entrepreneur behind the company.
Such innovations are the tip of a data iceberg. Smartphones, social networks and other accoutrements of modern digital life are generating vast new data sets that are revving up the digital economy.
Accompanying all this is a trend that has given the technology lexicon a new term: big data. Rather than sampling only small parts of the digital data deluge, modern companies have a new option: they can study all of it.
But the falling cost of technology and the generation of much more digital information have opened the field to a much wider group of businesses and made it possible to make more informed judgments about customer behaviour.
While “big data” has become the buzzword, a better description would be “messy data”, says Roger Ehrenberg of IA Ventures, an early-stage investor. Harvesting, cleaning up and organising raw data in a way that it can be processed is a large part of the battle, he says.
This has been complicated further by the big growth in unstructured data – information, such as text, that is not organised in a way that a computer can easily process. With the volume of user-generated text and video growing rapidly, this has become one of the main focuses of technological development.
Chief among the new tools are natural language processing, which enables a computer to extract meaning from text, and machine learning, the feedback loops through which computers can test their conclusions on large amounts of data in order progressively to refine their results.
Subjecting large data sets to analysis has also been made easier by two of the forces that have reshaped information technology more widely: the spread of low-cost, standardised computer hardware and the emergence of open-source software.
This has created a cheap computing platform for new technologies such as Hadoop – a piece of software architecture that is designed to handle massive amounts of data. The idea was based on breakthroughs at Google, which needed to find ways to conduct large volumes of intensive web searches simultaneously. It has since been taken up by companies including Facebook and Yahoo.
The rise of cloud computing – which centralises storage and processing power in larger data centres – has also brought big data within the reach of more companies. By tapping into the cloud computing services offered by Amazon, say, a company such as Color can get instant access to all the analytical power it needs without needing to take on the fixed costs of buying its own servers, says D.J. Patil, chief product officer at the IT start-up.
It is also stoking simmering privacy concerns. When Steve Jobs, Apple chief executive, was forced to apologise last week over the handling of data about the location of iPhone and iPad owners, it touched a raw public nerve and resulted in immediate Congressional hearings in Washington.
Color says it plans to use the information it collects to create new services for its customers. By combining it with data from social networks, says DJ Patil, chief product officer, it can tell its users: “Here are people who are near you, and here is how you might know them.” He says the company has no plans, at least for now, to use the information for other purposes, such as sending targeted advertising to customers.
However, with the tide of digital information rising fast – and more sophisticated ways being found to make business use of it – many companies are already being drawn into the new world of sophisticated data collection and analysis.
Some are using it to tailor their own products more precisely to the preferences of their users; others to target advertising of their products more accurately. Some are also selling the data they gather from their customers to the brokers and aggregators who act as middlemen in data markets that have sprung up to recycle such information.
A new consensus is needed to govern the use of this increasingly valuable commodity, says Michele Luzi of management consultancy Bain & Company, which conducted a study for the World Economic Forum on the issue. “Ultimately, you have to have a system of rights,” he says – something that balances the valid, but often conflicting, interests of individuals, governments and businesses.
While lawmakers and regulators on both sides of the Atlantic are becoming more exercised, such an agreement – not to mention the infrastructure and regulations to support it – remains some way off.
Meanwhile, as the analysis of digital information develops, the traditional management virtues of gut instinct and seat-of-the-pants decision-making are being replaced by reliance on intensive number-crunching and the objective testing of multiple potential courses of action.
For business leaders, “the big skill in future will be to ask the right question”, says Tim O’Reilly, a technology commentator and publisher.
Besides smartphones, new sources of data include social networks, blogs and other sources of user-generated content; sensors collecting everything from traffic patterns to a user’s heart rhythm; and click streams generated by people spending an increasing amount of their lives online.
Much of the information is in unstructured form. It has never been collated in a traditional relational database, where it could be queried at will. Without techniques to harvest, verify and analyse it – often in real time – valuable commercial signals are lost in the noise.
It sometimes takes the analysis of massive data sets to detect useful patterns, says Michael Olson. His California start-up, Cloudera, is commercialising the type of technology used by companies such as Facebook and Yahoo to crunch through vast bodies of information. Retailers, for instance, might learn far more from the 10 years’ worth of customer data they can now analyse in one go than from the more limited runs to which they were once restricted, he says.
. . .
Companies born in the digital age are often highly attuned to the possibilities presented by these untapped reservoirs of digital information. Like Color, they place data collection and analysis at their core, and build their business processes on their skills in these areas.
Better-established companies are jumping on the bandwagon. Giant US retailer Walmart is the latest to join the fray, last month buying Kosmix, a Silicon Valley company that filters the deluge of messages on Twitter. Walmart’s understanding of its customers has hitherto been limited to data about purchasing histories and browsing habits, says Ms Ranzetta, who is a member of the Kosmix board. In future, it will be able to tap into information on their personal preferences and interests as well.
Such mining of Twitter and other social sites is being used in a wide range of industries. Roger Ehrenberg, a former hedge fund manager who invests in technology start-ups, says demand is high among financial traders for help with assembling masses of data, or for refining them so that “what’s coming through is a signal, rather than a raw feed”.
Filtering tweets in real time for practical information is one of the most challenging of these tasks, both Mr Ehrenberg and Ms Ranzetta say. Often, it is only when data from such sources are combined with other information that their value emerges.
For instance, combining details of senior management moves revealed in companies’ regulatory filings with changes to profiles on LinkedIn, the business networking site for professionals, may yield valuable insights into what is happening inside companies, says Mr Ehrenberg.
Crunching through vast data sets can reveal patterns in fields far removed from the financial markets that would not otherwise be visible. The result is the rise of techniques such as behavioural clustering (grouping people on the basis of common behavioural characteristics, rather than more traditional demographics) and look-alike marketing (marketing to a particular user based on previous successes in marketing to others with similar profiles).
Comparing people in this way may have many uses. Analysing the detailed financial behaviour of very large groups of customers over a protracted period, for instance, could give banks a clue as to which are most likely to default next, says Mr Olson at Cloudera.
Such uses of predictive analytics – a marriage of statistical modelling and data mining – raise troubling questions. Is it fair, for instance, to judge a person merely on a prediction of their future behaviour?
And what are the long-term consequences of using such analyses to categorise people ever more narrowly, shaping the types of information and advertising they are fed online? Will this lead to a form of digital determinism, in which it becomes hard to escape a life that has been preordained by some giant bank, retailer or government department?
While the use of these techniques is still in its infancy, the digital crumbs of personal information left scattered across the web are already being swept up and used with surprising results.
“If I examine any new data set, the chances are I can find something in that data that has predictive value,” says Frank Rotman, a former head of analytics at Capital One, a US financial company that was a pioneer in the field. He says existing laws about how credit decisions are made, along with current social norms, place limits on how this information is used.
The rules are laxer, however, when it comes to how credit is marketed in the first place. And, ultimately, the opacity of this largely unregulated field makes it hard to tell exactly which signals from the digital morass are being used to inform business or government decisions that have a direct bearing on many lives.
“Where it gets murky and scary is the stuff that’s being sucked out of the social system, where you have no idea how it is being used,” says Chris Larsen, chief executive of Prosper, a web service through which individuals lend to each other directly.
Tuesday, April 26, 2011
When there’s no such thing as too much information
INFORMATION overload is a headache for individuals and a huge challenge for businesses. Companies are swimming, if not drowning, in wave after wave of data — from increasingly sophisticated computer tracking of shipments, sales, suppliers and customers, as well as e-mail, Web traffic and social-network comments. These Internet-era technologies, by one estimate, are doubling the quantity of business data every 1.2 years.
Yet the data explosion is also an enormous opportunity. In a modern economy, information should be the prime asset — the raw material of new products and services, smarter decisions, competitive advantage for companies, and greater growth and productivity.
Is there any real evidence of a "data payoff" across the corporate world? It has taken a while, but new research led by Erik Brynjolfsson, an economist at the Sloan School of Management at the Massachusetts Institute of Technology, suggests that the beginnings are now visible.
Mr. Brynjolfsson and his colleagues, Lorin Hitt, a professor at the Wharton School of the University of Pennsylvania, and Heekyung Kim, a graduate student at M.I.T., studied 179 large companies. Those that adopted "data-driven decision making" achieved productivity that was 5 to 6 percent higher than could be explained by other factors, including how much the companies invested in technology, the researchers said.
In the study, based on a survey and follow-up interviews, data-driven decision making was defined not only by collecting data, but also by how it is used — or not — in making crucial decisions, like whether to create a new product or service.
The central distinction, according to Mr. Brynjolfsson, is between decisions based mainly on "data and analysis" and on the traditional management arts of "experience and intuition."A 5 percent increase in output and productivity, he says, is significant enough to separate winners from losers in most industries.The companies that are guided by data analysis, Mr. Brynjolfsson says, are "harbingers of a trend in how managers make decisions." "And it has huge implications for competitiveness and growth," he adds. The research is not yet published, but it was presented at an academic conference this month.
The conclusion that companies that rely heavily on data analysis are likely to outperform others is not new. Notably, Thomas H. Davenport, a professor of information technology and management at Babson College, has made that point, and his most recent book, with Jeanne G. Harris and Robert Morison, is "Analytics at Work: Smarter Decisions, Better Results" (Harvard Business Press, 2010).
And companies like Google, whose search and advertising business is based on exploiting and organizing online information, are testimony to the power of intelligent data sifting.But the new research appears to be broader and to apply economic measurement to the impact of data-led decision making in a way not done before. "To the best of our knowledge," Mr. Brynjolfsson says, "this is the first quantitative evidence of the anecdotes we're been hearing about."
Mr. Brynjolfsson emphasizes that the spread of such decision making is just getting started, even though the data surge began at least a decade ago. That pattern is familiar in history. The productivity payoff from a new technology comes only when people adopt new management skills and new ways of working.
The electric motor, for example, was introduced in the early 1880s. But that technology did not generate discernible productivity gains until the 1920s. It took that long for the use of motors to spread, and for businesses to reorganize work around the mass-production assembly line, the efficiency breakthrough of its day.
The story was much the same with computers. By 1987, the personal computer revolution was more than a decade old, when Robert M. Solow, an economist and Nobel laureate, dryly observed, "You can see the computer age everywhere but in the productivity statistics."
It was not until 1995 that productivity in the American economy really started to pick up. The Internet married computing to low-cost communications, opening the door to automating all kinds of commercial transactions. The gains continued through 2004, well after the dot-com bubble burst and investment in technology plummeted.
The technology absorption lag accounts for the delayed productivity benefits, observes Robert J. Gordon, an economist at Northwestern University."It's never pure technology that makes the difference," Mr. Gordon says. "It's reorganizing things — how work is done. And technology does allow new forms of organization."
Since 2004, productivity has slowed again. Historically, Mr. Gordon notes, productivity wanes when innovation based on fundamental new technologies runs out. The steam engine and railroads fueled the first industrial revolution, he says; the second was powered by electricity and the internal combustion engine. The Internet, according to Mr. Gordon, qualifies as the third industrial revolution — but one that will prove far more short-lived than the previous two. "I think we're seeing hints that we're running through inventions of the Internet revolution," he says.
STILL, the software industry is making a big bet that the data-driven decision making described in Mr. Brynjolfsson's research is the wave of the future. The drive to help companies find meaningful patterns in the data that engulfs them has created a fast-growing industry in what is known as "business intelligence" or "analytics" software and services. Major technology companies — I.B.M., Oracle, SAP and Microsoft — have collectively spent more than $25 billion buying up specialist companies in the field.
I.B.M. alone says it has spent $14 billion on 25 companies that focus on data analytics. That business now employs 8,000 consultants and 200 mathematicians. I.B.M. said last week that it expected its analytics business to grow to $16 billion by 2015.
"The biggest change facing corporations is the explosion of data," says David Grossman, a technology analyst at Stifel Nicolaus. "The best business is in helping customers analyze and manage all that data."
Tuesday, April 19, 2011
How new Internet standards will finally deliver a mobile revolution
As the Web experience evolves, smartphones may soon live up to their name, and every business’s mobile strategy will grow in importance.
The rate at which developers are writing apps and consumers buying them is dizzying, and ingrained behavior can be hard to change. Web-centricity may raise security fears among users because programs are no longer installed on specific devices and because data are stored remotely. And there could be fragmentation issues with both the standard and the browsers—after all, existing ones, such as Google’s Chrome, Microsoft’s Internet Explorer, and Mozilla’s Firefox, don’t all treat the current standard, HTML4, the same way.2
Software developers. Application developers currently pay a fee of up to 30 percent to device makers, telecommunications operators, or operating-system developers whenever an application is sold to a consumer. In a Web-centric world, developers can avoid these intermediaries: not only can the same application be sold across all devices but anyone can set up a Web store and sell directly to users. Google, for instance, is already charging application developers a distribution fee of about 5 percent through its Chrome Web store.3 In addition, the emergence of an open platform will probably motivate bigger enterprise software companies to introduce—and quickly—mobile-based programs for managing customer relationships, marketing, and supply chains.