Showing posts with label Point of view. Show all posts
Showing posts with label Point of view. Show all posts
Thursday, January 3, 2013
Are We All Being Fooled by Big Data?
Michael Moritz, Linked In, January 03, 2013
When, on a summer Sunday morning in 1987, three hundred thousand people crammed onto the central span of San Francisco’s Golden Gate Bridge, they came perilously close to participating in the largest accident in American history. The bridge's engineers had made copious calculations and had designed it to sway nearly 28 feet and shoulder the burden of hundreds of vehicles. But nobody had ever predicted that a gigantic crowd of pedestrians, attracted by the fiftieth anniversary of its opening, would be stuck between its towering pylons unable to move in any direction. As a result, the bridge flattened out and came within whiskers of straining every last fiber of its vermilion superstructure.
The consequences of faulty data, wonky forecasts, ill-conceived opinions, loose predictions, incorrect assumptions and, in the case of the Golden Gate Bridge, an improbable event form the backbone of Nate Silver’s absorbing new book, The Signal and the Noise: Why Most Predictions Fail but Some Don’t. This book, written by the voice behind the popular election forecasting blog, FiveThirtyEight, now licensed by the New York Times, is a reminder that while data doesn’t lie, it does allow people to deceive themselves and others. In some cases it's a question of the bigger the data, the grander the deception.
These days our entire lives revolve around predictions. Government departments project the cost of health exchanges, the rate of economic growth, next year’s crop yields, the future birth rate and the arms buildup of unfriendly countries. Websites and retailers anticipate what we want to find and buy; oil companies gauge the best sites for drilling; pharmaceutical companies assess the probable efficacy of molecules on a disease; while, in the background, the bobble-heads on television incessantly spew out largely irrelevant and inaccurate forecasts. In the meantime, we busy ourselves with personal projections. How long will our commute take? When will the turkey be golden? How much will the price of a stock rise? What will the future value be of a law degree?
Some of these forecasts are surprisingly accurate while others are shockingly dismal. Silver, who has become the Woody Allen of statisticians, explains the reasons. Like many others, the 34-year-old Silver became fascinated with numbers because of a boyhood devotion to baseball. Unlike his peers, Silver – after a brief and frustrating spell as a consultant – instinctively returned to the challenges of numbers. He took up internet poker (only to eventually discover that the odds were not in his favor) and, also started to unravel the riddles presented by data.
There are events that – at least on the surface – defy forecast: things that are so outlandish or improbable that, for most people at one time or another, they seem inconceivable. Think of Pearl Harbor, 9/11, Fukushima, a black President of the United States or Apple as the world’s most valuable company. Yet all, to varying extents, were possible to predict if people had been able to separate the important from the trivial (a.k.a. the signal from the noise) and make the giant leap of faith which converts the improbable into the possible. While Silver provides a supple assessment of the reasons we struggle to comprehend these sorts of possibilities, the majority of his book is devoted to an often-hilarious account of how we deal with more mundane challenges.
About ten years ago, Silver developed a system for predicting the performance of batters and hitters for Baseball Prospectus. The exercise helped him develop his approach to predictions. It is no coincidence that Silver fastened on both baseball and politics. In each pursuit there is an enormous trove of accurate, historical information. The baseball fiend can immerse himself in minutiae such as hits, on-base percentages and pitches thrown, while the political junkie can stare at votes recorded, demographic shifts and polling results. Silver gradually discovered that in baseball the data, while essential, could be made richer with the judicious application of human judgment. This must have come as a reassuring endorsement for baseball scouts whose usefulness had been much maligned in the years following the publication of Moneyball, Michael Lewis’ much-read book about the way data had helped Billy Beane transform the Oakland A’s. After all, it is difficult for a machine to measure the determination, pluck, grit (and wandering eye or fondness for drink) of a baseball player.
The same goes for politics, the field in which Silver made his reputation with his accurate predictions about the 2008 races (which he subsequently burnished in 2012). Here he bases many of his predictions on the averages of poll results conducted by others. This, he has discovered, provides more accurate forecasts for election nights than reliance on a single pollster, no matter how sterling the reputation. When Silver does stray from the received wisdom, he does so with caution and says, “The further I move away from consensus, the stronger my evidence has got to be … that I have things right.” This is an observation worth dwelling upon because it helps explain why most people have such trouble making the correct decision about an unconventional selection or the path less trodden. Making a decision frowned upon by a committee or a popular opinion is a lonely place to be.
Accurate information married with human judgment is the best ally for the prognosticator. This explains why some forecasts, such as those for hurricanes, are so good and others, such as economic predictions, are so poor. Thanks to a knowledge of past catastrophes, satellite photography, weather balloons and airplanes that fly into the eye of the storms, the National Hurricane Center can predict the path and severity of hurricanes with remarkable certainty several days in advance of when they collide with land. This information, enhanced by the analysis of scientists, has improved the National Hurricane Center’s forecasting accuracy by 350% in the past 25 years. The fact that 1,833 people died when Hurricane Katrina swamped New Orleans is not because of faulty forecasting, but mainly because the city’s Mayoral office hesitated about ordering a compulsory emergency evacuation until it was too late. According to Silver, weather forecasting for the subsequent two or three days, at least as promulgated by the National Weather Service (before it falls into the buffoonish hands of the local TV weathermen for whom ratings are more important than accuracy), is also something that can be counted on.
Economic forecasting is another matter. Part of the reason that predictions about hurricanes and the weather have improved is that scientists, mathematicians and programmers can build computer models from accurate molecular data of cloud formations. The same is not true for the economy where attempts to capture every calorie of economic endeavor are much harder. Even the U.S. government – irrespective of whether a Democrat or Republican is at the helm – has proved woefully inept at forecasting overall GDP growth let alone more refined measures. It’s not uncommon for economic forecasters to fail to predict recessions even after they are already underway. It’s a wonder that any bank or company bothers to keep an economist on the payroll. They all might be better off employing the descendants of Carnac the Magnificent, the soothsayer from the East once played by Johnny Carson.
While economists have plenty of excuses, the same does not go for the rating agencies that, prior to the housing collapse, so conspicuously labeled the thousands of mortgages they bundled together as relatively riskless. Even if you are prepared to accept that officials at S&P, Moody’s and Fitch were merely guilty of a failure of judgment – as opposed to criminal collusion – they made the colossal mistake of not recognizing the consequences of uncertainty (a risk that is hard to measure): the close correlation between all these mortgages. They did not understand that they had designed a monstrous, nationwide pileup of concrete, glass and wood. It’s no coincidence that these same rating agencies are today all involved in designing a future economic calamity: the implosion of municipal, state and corporate pension obligations. In this case they are even more culpable because they are willfully ignoring copious amounts of stock market data which, if heeded, would instantly catapult these pension systems into default.
If an economist might deserve some pity, it’s a teaspoonful compared to what should be given to those charged with making an accurate prediction of the timing and strength of an earthquake. These hapless devils don’t have accurate pictures of geological formations dozens of miles below the earth’s crust or reams of data supplied by probes latched to different striations. Nonetheless, in our data-drenched age, the geologist is still somehow expected to provide certainty about a cataclysmic event that may last a matter of seconds. That’s especially true in Italy, where, in the wake of the 2009 quake that killed over 300 people in the central Italian town of L’Aquila, six scientists and a government official were found guilty of manslaughter and sentenced to six years in jail for not protecting their neighborhood. The sentencing magistrates, like the Japanese in the ninth century, must just believe that earthquakes can be accurately predicted from the behavior of catfish.
Thursday, December 6, 2012
Ethan Roeder (Big Data Czar for Obama) in NYT: I Am Not Big Brother
Ethan Roeder, The New York Times, December 6, 2012
I’VE grown accustomed to reading inaccurate accounts of my day job. I’m in political data.
If I’m not spying on private citizens through the security cam in the parking garage, I’m probably sifting through their garbage for discarded pages from their diaries or deploying billions of spambots to crack into their e-mail. Reading what others muse about my profession is the opposite of my middle-school experience: people with only superficial information about me make a bunch of assumptions to fill in what’s missing and decide that I’m an all-knowing super-genius.
Sadly for me, this is a bunch of malarkey. You may chafe at how much the online world knows about you, but campaigns don’t know anything more about your online behavior than any retailer, news outlet or savvy blogger.
There are two categories of online data: information users provide explicitly, and stuff they communicate implicitly through their behavior. The explicit data includes e-mails and comments that users share directly. The implicit data comes from “click tracking,” which tells a campaign what buttons are getting pressed and how often. Combined, these two categories of data allow a campaign to put together an online experience that will resonate with as many people as possible, but also to customize the experience so that you are more likely to encounter content that’s relevant to you.
At times it might seem like sorcery to the recipient of a targeted e-mail, but it’s just a product of two simple factors: remembering who you are and remembering what you like.
In the offline world, which is my personal area of expertise, campaigns don’t know much more today than they did 44 years ago. In 1968, George Romney, then a candidate for the Republican presidential nomination, made headlines for using a newfangled “secret weapon” — a voter file. As The Times explained that winter, it is “an electronic data bank” that “contains the only really accurate, up-to-date roster of enrolled New Hampshire Republicans any candidate here has ever compiled, plus pertinent information about all of them.” Imagine how freaked out those New Hampshire Republicans must have been.
Virtually all of the offline data that people like me traffic in is boring, basic and publicly available. Want to know the year of birth for everyone who is registered to vote in Ohio? Just Google “Ohio voter file download.” There you go. I was born in 1976. Now we’re even.
How do we predict whether people are going to vote or not? We look at the voter file. It tells us how often a person votes, although not for whom. Not all strategists agree about how to interpret this information, but the source of the data is no secret.
What’s really new in politics today is not the data itself but how campaigns make sense of it. Cheaper and more plentiful computing power allows campaigns to process far more information than ever before to look for patterns, trends and correlations.
The science of modeling is a modern-day application of a practice that has been around for nearly 200 years: polling. Pollsters ask voters whom they support for president and how strongly. Campaigns then take demographic information about these voters into account in order to make assumptions about the entire population of a given state. The mechanics are exactly the same for public polls and internal campaign analyses. The difference is that the campaigns use statistical techniques to apply these assumptions to individual records in the voter file rather than stopping short and simply assuming that entire sections of the electorate will behave identically.
Contemporary data practice also frees campaigns from having to make assumptions about voters in the first place. In 2011 and 2012, the Obama campaign, with the help of more than two million volunteers, had more than 24 million conversations with voters. Online tools gave Obama supporters resources to help them play a crucial role in their neighborhoods, and a series of “share your story” pages on the campaign Web site provided a venue for voters to communicate directly with the campaign in long form.
All of this feedback doesn’t neatly boil down to a “yes” or “no” in a database — and why should it? Numerous avenues of listening, combined with the digital capacity to hold on to qualitative feedback, make campaigns aware of the differences among voters’ motivations, attitudes, protestations — not just their demographics and voting history. In a nation of over 200 million eligible voters, technology is allowing campaigns to finally see through the fog of the crowd and engage voters one by one.
In other words, there is no giant blue computer sitting on the 101st floor of a sleek skyscraper, surrounded by bubbling tubes of illuminated liquid, spitting out the manifest destiny of America’s voters. Campaigns are moving away from the meaningless labels of pollsters and newsweeklies — “Nascar dads” and “waitress moms” — and moving toward treating each voter as a separate person.
In 2012 you didn’t just have to be an African-American from Akron or a suburban married female age 45 to 54. More and more, the information age allows people to be complicated, contradictory and unique. New technologies and an abundance of data may rattle the senses, but they are also bringing a fresh appreciation of the value of the individual to American politics.
Ethan Roeder was the data director of Obama for America.
Wednesday, September 5, 2012
REINVENTING SOCIETY IN THE WAKE OF BIG DATA
A Conversation with Alex (Sandy) Pentland Edge Video (24-Minutes)
ALEX 'SANDY' PENTLAND is a pioneer in big data, computational social science, mobile and health systems, and technology for developing countries. He is one of the most-cited computer scientists in the world and was named by Forbes as one of the world's seven most powerful data scientists. He currently directs MIT's Human Dynamics Laboratory and the MIT Media Lab Entrepreneurship Program, and advises the World Economic Forum, Nissan Motor Corporation, and a variety of start-up firms.
Recently I seem to have become MIT's Big Data guy, with people like Tim O'Reilly and "Forbes" calling me one of the seven most powerful data scientists in the world. I'm not sure what all of that means, but I have a distinctive view about Big Data, so maybe it is something that people want to hear.
I believe that the power of Big Data is that it is information about people's behavior instead of information about their beliefs. It's about the behavior of customers, employees, and prospects for your new business. It's not about the things you post on Facebook, and it's not about your searches on Google, which is what most people think about, and it's not data from internal company processes and RFIDs. This sort of Big Data comes from things like location data off of your cell phone or credit card, it's the little data breadcrumbs that you leave behind you as you move around in the world.
"What those breadcrumbs tell is the story of your life. It tells what you've chosen to do. That's very different than what you put on Facebook. What you put on Facebook is what you would like to tell people, edited according to the standards of the day. Who you actually are is determined by where you spend time, and which things you buy. Big data is increasingly about real behavior, and by analyzing this sort of data, scientists can tell an enormous amount about you. They can tell whether you are the sort of person who will pay back loans. They can tell you if you're likely to get diabetes."
Sandy Pentland's EDGE Profile page: http://edge.org/memberbio/alex_(sandy)_pentland
Permalink: http://www.edge.org/conversation/reinventing-society-in-the-wake-of-big-data
[ED. NOTE: Part of the ongoing series "COMPUTATIONAL SOCIAL SCIENCE @ Edge"
http://edge.org/event/special/computational-social-science]
Wednesday, August 1, 2012
VC Bill Davidow: Virtual Reality Is Addictive and Unhealthy
William H. Davidow, IEEE Spectrum, August 2012
In my days as an engineer, I ran the microprocessor division at Intel Corp. I then became a venture capitalist, investing in companies that built semiconductors, computers, networking systems, and Internet-related services. I focused on products that helped businesses run more effectively and gave little thought to how they might affect our minds, social interactions, and governance.
That lapse now comes home to me as I see people walking down the street, eyes fixed on the screens of their mobile phones, ears plugged into their iPods, oblivious to their surroundings…to reality itself. They are not managing their tools; their tools are managing them. Tools now make the rules, and we struggle to keep up.
I’ve spent my career developing and financing the companies that supply these profoundly powerful tools. For the most part, I thought of them as harmless, and I believed my job was simply to make the tools better so that others would use them to improve the world. Only in recent years have I become aware of and concerned about their serious side effects. And so I have decided to study them and do my best to explain those effects to the world. Here’s what I’ve learned.
First, it wasn’t always this way. Our relationship with tools dates back millions of years, and anthropologists still debate whether it was the intelligence of human-apes that enabled them to create tools or the creation of tools that enabled them to become intelligent. In any case, everyone agrees that after those first tools had been created, our ancestors’ intelligence coevolved with the tools. In the process our forebears’ jaws became weaker, their digestive systems slighter, and their brains heavier. Chimpanzees, genetically close to us though they are, have bodies two to five times as strong as ours on a relative basis and brains about a quarter as big. In humans, energy that would have gone into other organs instead is used to run energy-hungry brains. And those brains, augmented by tools, more than make up for any diminishment in guts and muscle. Indeed, it’s been a great evolutionary tradeoff: There are 7 billion people but only a few hundred thousand chimpanzees.
In the distant past our tools improved slowly enough to allow our minds, our bodies, our family structures, and our political organizations to keep up. The earliest stone tools are about 2.6 million years old. As those and other tools became more refined and sophisticated, our bodies and minds changed to take advantage of their power. This adaptation was spread over more than a hundred thousand generations.
Our social structures evolved along with the tools. Some 10 000 years ago, tribes of roaming hunter-gatherers began to stay in one place to raise crops. Agriculture made cities possible, and with cities came the arts and commerce. As transportation improved and cities grew, it became important to control distant places that supplied food and raw materials. About 6000 years ago, ancient city-states such as Uruk emerged in Mesopotamia and governed the surrounding countryside. Millennia later, Athens, the largest of the Greek city-states, controlled about 2500 square kilometers and most of the Aegean Sea. So far, so good.
But with the invention of movable type in medieval times, followed by other improved ways of connecting—better ships and roads, trains, planes, automobiles, and the Internet—technology raced ahead of us. In our own time the most striking example of the acceleration of technology has been Moore’s Law, under which the density of transistors on chips has doubled every 18 months. The performance of single processors long rose in tandem with the transistor count, and even after that relationship stalled in the mid-2000s, the switch to devices using many processors kept the performance curve pointing upward.
This exponential rise in capability has greatly augmented the pooling of knowledge from different sources to achieve the creative synthesis described by the 19thcentury mathematician and philosopher Henri Poincaré: “To create consists precisely in not making useless combinations, and in making those which are useful and which are only in a small minority…. Among chosen combinations, the most fertile will often be those formed of elements drawn from domains which are far apart.”
Drawn in part from the Internet, the newly created knowledge gets deposited back on the Internet, increasing its scope and accelerating the development of technology. Burgeoning knowledge in turn drives rapid change—it advances technology, transforms business institutions, and changes how markets work and how people interact. Governments, social
institutions, and our brains struggle to keep up.
Even the nature of change itself has changed. Living creatures started out evolving in one dimension defined by the physical world and another defined by the biological world. Then came humans, who added a third dimension—the artificial one engendered by tools and technology. Now, with the widespread use of the Internet, a fourth dimension has been added—that of virtual reality, or cyberspace. It is indeed appropriate to consider this last dimension as real and distinct from the tools and technologies of the past, because however fast those things may have changed, the rate of change in the virtual space is much faster. It took a lot of time to build physical infrastructure—railroads, highways, bridges, skyscrapers, and so forth. But in virtual space, entire new infrastructures can arise overnight, as Google and Facebook have proved.
I now believe that our minds, bodies, businesses, governments, and social institutions are no longer capable of coping with the rapid rate of change. And it is obvious that this change is indeed more rapid than any comparable change that came before.
Think of the many years it took Barnes & Noble to build its retail chain of U.S. bookstores. The company set up its first bookstore in 1917, and by 2010 it was operating 717 stores. It took time for the company to find the proper locations, lease them, and stock them with inventory, and still more time to build them into viable businesses. The company was limited in what it could do because only certain physical locations were suitable for retail stores.
Compare that long history to the rise of Amazon.com, which started in 1994 and was operating in virtual space throughout the United States by the next year, putting a bookstore in every home that had an Internet connection.
Barnes & Noble responded with an online strategy of its own, one that now gives pride of place to sales of e-books, and the company continues to fight on; meanwhile, competitor Borders was liquidated. Both constitute sterling examples of the “creative destruction” of capitalism, as the economist Joseph Schumpeter put it. But the fact that entire business models can come and go that fast is extraordinary. It also indicates the challenges that rapid change presents to other institutions.
An example is our financial institutions, transformed in the past 20 years by radical innovations, such as the introduction of high-speed trading—in which computers trade securities with other computers—and the immensely complex new investment vehicles known as derivatives. These innovations, enabled by improvements in computing power and telecommunications, have made markets more volatile and played leading roles in the recent worldwide economic meltdown. And our regulatory framework has failed to keep up.
In 2005 about 80 percent of the shares of companies listed on the New York Stock Exchange were traded on its floor or on its proprietary digital system; today only 20 percent are traded there. Most of the other shares are traded on alternative systems that have cropped up, systems in which many of the old rules and regulations no longer apply. High-frequency traders have exploited these lightly regulated trading systems. In the United States, computer-driven algorithms now execute 60 percent of all stock trades. One result was the Flash Crash on 6 May 2010, when the Dow Jones Industrial Average fell 1000 points in a matter of minutes. Although it recovered 600 points of the loss minutes later, the episode shows us just how little we understand about high finance and how vulnerable we are to its vagaries.
European stock exchanges have become vulnerable as well. The Better Alternative Trading System, for example, has taken market share away from the London Stock Exchange. Xavier Rolet, CEO of the London Stock Exchange, has tried to counter this threat by diversifying the exchange’s business. Whether this strategy can slow the decline of the London exchange is unclear.
The Internet has also made it easier to participate in the over-the-counter market in derivatives—which are essentially side bets on the value of assets. In 2000, the notional value of these OTC derivatives was US $60 trillion; by 2007 it had risen to $600 trillion. Losses on these derivatives played a major role in the 2008 financial crisis and are still causing problems today. In late 2009, a little-known company called Markit Group created the iTraxx SovX index, which made it easier to use derivatives to place bets on the possibility of a Greek default. As the cost of insuring Greek debt based on the iTraxx SovX rose, investors shunned Greek bonds, making it harder for the country to borrow. Of course, the Greek economy was in dire straits anyway, but those problems were aggravated by the new strategies of speculation that technology has made possible.
Regulatory reform has been a case of too little, too late. Basel III, an agreement concluded in 2011 by the banking supervisors of 10 major industrialized countries, is too weak. And in the United States, the 2319page Dodd-Frank legislation will probably prove to be too complex to achieve its goal of averting another meltdown. Ultimately, the only way to deal with financial innovation in a virtual world is through international regulatory systems based on commonly accepted principles. But as a former head of the Bank for International Settlements told me, good luck with that. The tools are making the rules.
Subscribe to:
Posts (Atom)