Sunday, November 18, 2012
You Can't Say That on the Internet
Evgeny Morozov, The New York Times, November 16, 2012
A BASTION of openness and counterculture, Silicon Valley imagines itself as the un-Chick-fil-A. But its hyper-tolerant facade often masks deeply conservative, outdated norms that digital culture discreetly imposes on billions of technology users worldwide.
What is the vehicle for this new prudishness? Dour, one-dimensional algorithms, the mathematical constructs that automatically determine the limits of what is culturally acceptable.
Consider just a few recent kerfuffles. In early September, The New Yorker found its Facebook page blocked for violating the site’s nudity and sex standards. Its offense: a cartoon of Adam and Eve in the Garden of Eden. Eve’s bared nipples failed Facebook’s decency test.
That’s right — a venerable publication that still spells “re-elect” as “reëlect” is less puritan than a Californian start-up that wants to “make the world more open.”
And fighting obscenity can be good for business. Impermium, a Silicon Valley company that helps Web sites deal with unwanted reader comments, has begun marketing technology that identifies “all kinds of harmful content — such as violence, racism, flagrant profanity, and hate speech — and allows site owners to act on it in real-time, before it reaches readers.” Impermium will police the readers — but who will police Impermium?
Apple, too, has strayed from its iconoclastic roots. When Naomi Wolf’s latest book, “Vagina: A New Biography,” went on sale in its iBooks store, Apple turned “Vagina” into “V****a.” After numerous complaints, Apple restored the title, but who knows how many other books are still affected?
True, these books are still on sale. Unlike the good old United States Post Office, which once confiscated “Lady Chatterley’s Lover” and other books it deemed too lewd, Silicon Valley does not engage in direct censorship. What it does, though, is present ideas and terms that have gained public acceptance as something to be ashamed of. Silicon Valley doesn’t just reflect social norms — it actively shapes them in ways that are, for the most part, imperceptible.
The proliferation of the Autocomplete function on popular Web sites is a case in point. Nominally, all it does is complete your search query — on YouTube, on Google, on Amazon — before you’ve finished typing, using an algorithm to predict what you’re most likely typing. A nifty feature — but it, too, reinforces primness.
How so? Consider George Carlin’s classic comedy routine “Seven Words You Can Never Say on Television.” See how many of those words would autocomplete on your favorite Web site. In my case, YouTube would autocomplete none. Amazon almost none (it also hates “penis” and “vagina”). Of Carlin’s seven words, Google would autocomplete only “piss.”
Until recently, even the word “bisexual” wouldn’t autocomplete at Google; it’s only this past August that Google, after many complaints, began to autocomplete some, but not all, queries for that term. In 2010, the hacker magazine 2600 published a long blacklist of similar words. While I didn’t verify all 400 of them on Google, a few that I did try — like “swastika” and “Lolita” — failed to autocomplete. Is Nabokov not trending in Mountain View? Alas, these algorithms are not particularly bright: unable to distinguish between Nabokov’s novel and child pornography, they assume you want the latter.
Why won’t tech companies let us freely use terms that already enjoy wide circulation and legitimacy? Do they fashion themselves as our new guardians? Are they too greedy to correct their algorithms’ mistakes?
Thanks to Silicon Valley, our public life is undergoing a transformation. Accompanying this digital metamorphosis is the emergence of new, algorithmic gatekeepers, who, unlike the gatekeepers of the previous era — journalists, publishers, editors — don’t flaunt their cultural authority. They may even be unaware of it themselves, eager to deploy algorithms for fun and profit.
Many of these gatekeepers remain invisible — until something goes wrong. Thus, in early September, the online livestream from the Hugo Awards, the Oscars of the science fiction world, was interrupted with a cryptic copyright warning, right before the popular author Neil Gaiman was to deliver an acceptance speech.
Apparently, Ustream — the site streaming the ceremony — was using the services of another company to determine whether its streamed videos violated any copyrights. The partner company draws on a very large video archive to see, in real time, if what’s being streamed matches anything in its collection. Somehow, the celebratory video that preceded Mr. Gaiman’s speech tripped a copyright match, and the feed was cut off, even though the organizers had all the requisite permissions (and, under the doctrine of fair use, probably didn’t need them anyway).
The limitations of algorithmic gatekeeping are on full display here. How do you teach the idea of “fair use” to an algorithm? Context matters, and there’s no rule book here; that’s why we have courts. From the perspective of sticky, amorphous human culture, semi-automation — pairing up humans with algorithms — beats full automation. Sometimes, gaps are productive. But will profit-driven Silicon Valley ever acknowledge this insight?
Our reputations are increasingly at the mercy of algorithms, too. No one knows this better than Bettina Wulff, the former German first lady who has sued Google for autocompleting searches for her name with words like “escort” and “prostitute.” Ms. Wulff insists that Google’s algorithms spread false rumors about her; Google says that the suggested terms are just an “algorithmically generated result of objective factors, including the popularity of the entered search terms.”
Google’s defense would sound tenable if its own algorithms weren’t so easy to trick. In 2010, the marketing expert Brent Payne paid an army of assistants to search for “Brent Payne manipulated this.” Soon anyone typing “Brent P” into Google would see that phrase in their autocomplete suggestions. After Mr. Payne publicized his experiment, Google removed that particular suggestion, but how many similar cases have gone undetected? What is “objective” about such algorithmic “truths”?
Quaint prudishness, excessive enforcement of copyright, unneeded damage to our reputations: algorithmic gatekeeping is exacting a high toll on our public life. Instead of treating algorithms as a natural, objective reflection of reality, we must take them apart and closely examine each line of code.
Can we do it without hurting Silicon Valley’s business model? The world of finance, facing a similar problem, offers a clue. After several disasters caused by algorithmic trading earlier this year, authorities in Hong Kong and Australia drafted proposals to establish regular independent audits of the design, development and modifications of computer systems used in such trades. Why couldn’t auditors do the same to Google?
Silicon Valley wouldn’t have to disclose its proprietary algorithms, only share them with the auditors. A drastic measure? Perhaps. But it’s one that is proportional to the growing clout technology companies have in reshaping not only our economy but also our culture.
Obviously, Silicon Valley won’t develop or embrace similar norms overnight. However, instead of accepting this new reality as a fait accompli, we must ensure that, in pursuing greater profits, our new algorithmic gatekeepers are forced to accept the idea that their culture-defining function comes with great responsibility.
The author of the forthcoming book “To Save Everything, Click Here: The Folly of Technological Solutionism.”
Thaler: Applause for the Numbers Machine
Richard H. Thaler, The New York Times, November 18, 2012
THE biggest winners on Election Day weren’t politicians; they were numbers folks.
Computer scientists, behavioral scientists, statisticians and everyone who works with data should be proud. They told us who was going to win, but they also helped to make many of those victories happen.
Three groups of geeks deserve the love they rarely receive: people who run political polls, those who analyze the polls and those who figure out how to help campaigns connect with voters.
Many people doubted the accuracy of political polling this year. Part of the skepticism was based on the wide range of predictions, with some showing President Obama in the lead, and others Mitt Romney. But there were additional, structural reasons to worry whether pollsters would be able to find representative samples of voters.
One problem is that people are harder to reach on the telephone these days. About a third of voters no longer have a land line, and many of those who have them don’t pick up calls from strangers. So modern polling companies have to work harder to find voters willing to answer questions, then have to guess which of these respondents will actually show up and vote.
So it may come as a surprise that, collectively, polling companies did quite well during this election season. Although there was a small tendency for the pollsters to overestimate Mr. Romney’s share of the vote, a simple average of the polls in swing states produced a very accurate prediction of the Electoral College outcome. Notably, the most accurate polls tended to be done via the Internet, many by companies new to this field. That’s geek victory No. 1.
This relatively accurate polling data provided the raw material for the second group of election pioneers: poll analysts like Nate Silver, who writes the FiveThirtyEight blog for The New York Times, as well as Simon Jackman at Stanford, Sam Wang at Princeton and Drew Linzer at Emory University.
What do poll analysts do? They are like the meteorologists who forecast hurricanes. Data for meteorologists comes from satellites and other tracking stations; data for the poll analysts comes from polling companies. The analysts’ job is to take the often conflicting data from the polls and explain what it all means.
Worry about the reliability of the polling data led to widespread skepticism, or even outright hostility, toward poll analysts. The phrase “garbage in, garbage out” was one of the more polite criticisms bouncing around the Internet in the days before the election.
Because the polls were not, in fact, garbage, the first job of a poll analyst was quite easy: to average the results of the various polls, weighing more reliable and recent polls more heavily and correcting for known biases. (Some polls consistently project higher voter shares for one party or the other.)
A harder but more valuable task is to help readers translate the polling data into forecasts of the probability of victory. In Florida, where the final polls showed essentially a tie, according to Mr. Silver’s weighting method, it’s easy to see why he said the chance of either candidate winning the state was 50 percent. Ultimately, President Obama would very narrowly carry the state.
But what about North Carolina, where Mr. Silver projected that Mitt Romney would get 50.6 percent of the vote and President Obama, 48.9 percent? Looking at that very small difference, what probability would you have assigned to a Romney victory in that state?
Most people would guess something very close to 50-50. But not a good numbers guy. By looking back at previous elections with polling data this close, Mr. Silver estimated that Mr. Romney’s chances of winning North Carolina were 74 percent, a number that may seem surprisingly high. (Mr. Romney won the state.)
The slightly larger but still seemingly tiny lead that the president held in Ohio, another swing state, led poll analysts to predict that the chance of an Obama victory in Ohio was around 90 percent. And because Mr. Romney would have to win several such states with small Obama leads in order to prevail in the Electoral College, the analysts ended up with similarly high degrees of confidence in an overall Obama victory. They ended up predicting the Electoral College outcome almost exactly right, especially if you consider the final outcome in Florida to be a virtual tie, as they had projected.
Pundits making forecasts, some of whom had mocked the poll analysts, didn’t fare as well, and many failed miserably. George F. Will predicted that Mr. Romney would win 321 electoral votes, which turned out to be very close to President Obama’s actual total of 332. Jim Cramer from CNBC was nearly as wrong in the opposite direction, projecting that the president would win 440 electoral votes.
There is a lesson here. When it comes to assessing the chances of some complicated combination of events, gut feelings are pretty much useless. Pundits are no better at forecasting election outcomes than they would be at predicting the final path of a hurricane. Smart pundits should consider either abandoning this activity, or consulting with the geeks before rendering their guesses.
The third set of folks who deserve recognition in this election cycle were a group of young people working in a windowless room at Obama headquarters, affectionately known as the cave. They were part of the effort by the numbers-oriented campaign manager, Jim Messina, to maximize turnout.
THERE are two basic parts of an election campaign. The first comes under the category of messaging — deciding what a candidate should say and what ads to run. Most of the commentary we read about elections focuses on this component.
The second part is turnout, and in some ways is even more important. Here is a simple bit of math that you don’t have to be a geek to understand: It doesn’t matter which candidate a person prefers unless that person shows up and votes.
Pundits will debate for eternity which campaign did a better job of communicating its message, but there is no doubt which campaign won the turnout contest. Young, black and Hispanic voters all turned out in higher numbers than expected, and they often supported President Obama.
Much was made of the big Obama advantage in field offices in swing states. But those field offices would have been little good to the campaign without modern tools to find potential voters, have them register and encourage them to vote. In the weeks leading up to the election, the Obama canvassers had accurate lists of potential voters and field-tested scripts for their contacts with voters. This explains in part why Democrats were such heavy users of early voting.
By contrast, Project Orca, a get-out-the-vote computer program for the Romney campaign that wasn’t designed to be used until Election Day, reportedly had some bugs.
There should be something reassuring about this Obama campaign efficiency to all Americans, even those who supported Mr. Romney based on his success in business. When it came to the business of running a campaign, it was the former professor and community organizer who had the more technologically savvy organization and made more effective use of its resources, including geek power.
Richard H. Thaler is a professor of economics and behavioral science at the Booth School of Business at the University of Chicago. He was an informal adviser to the Obama campaign.
Saturday, November 17, 2012
Beware the Smart Campaign
Zeynep Tufekci, The New York Times, November 16, 2012
“I AM not a number. I am a free man!” was the famous cry of prisoner Number Six, who could never escape his Kafkaesque village on the 1960s television show “The Prisoner.” This is a prescient cry for an era when numbers follow us everywhere. Jim Messina, the victorious Obama campaign manager, probably agrees that you are not a number. That’s because you are four numbers.
The Obama campaign assigned all potential swing-state voters one number, on a scale of 1 to 100, that represented the likelihood that they would support Mr. Obama, and another number for the prospect that they would show up at the polls. A third metric evaluated the odds that an Obama supporter who was an inconsistent voter could be nudged to the polls, and a fourth score estimated how persuadable someone was by a conversation on a particular issue (which was, of course, also determined by crunching more numbers).
Mr. Messina is understandably proud of his team, which included an unprecedented number of data analysts and social scientists. As a social scientist and a former computer programmer, I enjoy the recognition my kind are getting. But I am nervous about what these powerful tools may mean for the health of our democracy, especially since we know so little about it all.
For all the bragging on the winning side — and an explicit coveting of these methods on the losing side — there are many unanswered questions. What data, exactly, do campaigns have on voters? How exactly do they use it? What rights, if any, do voters have over this data, which may detail their online browsing habits, consumer purchases and social media footprints?
How did Mr. Obama win? The message and the candidate matter, of course; it’s easier to persuade voters if your policies are more popular and your candidate more appealing. But a modern winning campaign requires more. As Mr. Messina explained, his campaign made an “unparalleled” $100 million investment in technology, demanded “data on everything,” “measured everything” and ran 66,000 computer simulations every day. In contrast, Mitt Romney’s campaign’s data operations were lagging, buggy and nowhere as sophisticated. A senior Romney aide described the shock he experienced in seeing the Obama campaign turn out “voters they never even knew existed.” And that kind of ability matters: while Mr. Obama did win decisively, the size of his lead in four states that determined the outcome, Florida, Ohio, Virginia and Colorado, was about 400,000 votes — or about 1.2 percent of the eligible voters.
The confluence of marketing and politics goes back a long way. A blizzard of direct mail engineered by political consultants is credited with defeating President Harry S. Truman’s national health care proposal after World War II. The new methods, however, are not just better direct mail. Noxious TV ads and slick mailers are like machetes compared with the scalpels of social-science-based big-data. The crude methods may still work to soften the ground and drown out other voices, but in the end they are still very big sticks. Sometimes they kill the patient — just ask swing-state voters about the TV ads they were bombarded with.
The scalpels, on the other hand, can be precise and effective in a quiet, un-public way. They take persuasion into a private, invisible realm. Misleading TV ads can be countered and fact-checked. A misleading message sent in just the kind of e-mail you will open or ad you will click on remains hidden from challenge by the other campaign or the media. Or someone who visits evangelical Web sites might be carefully shielded from messages about gay rights, and someone who has hostile views toward environmentalism may receive messages stroking that sentiment even if the broader campaign woos the green vote elsewhere.
What I really worry about, though, is that these new methods are more effective in manipulating people. Social scientists increasingly understand that much of our decision making is irrational and emotional. For example, the Obama campaign used pictures of the president’s family at every opportunity. This was no accident. The campaign field-tested this as early as 2007 through a rigorous randomized experiment, the kind used in clinical trials for medical drugs, and settled on the winning combination of image, message and button placement. I agree that his family is wonderful and his daughters are cute. But an increasing role of “likability” factors, which we now understand better how to manipulate, is not good for democracy.
These methods will also end up empowering better-financed campaigns. The databases are expensive, the algorithms are proprietary, the results of experiments by campaigns are secret, and the analytics require special expertise. The Democrats have an early advantage partly because academics and data analysts tend to be Democrats. Money will solve that problem. This will shift power in both parties even more toward the richer campaigns and may well be the final nail in the coffin of public financing for presidential campaigns.
What is to be done? Campaigns should make public every outreach message so we at least know what they are saying. These messages can be placed in a public database like campaign contributions so the other side can be aware of, and have the right to respond to, false claims. Political access to proprietary databases should be regulated to provide an even playing field.
I’m not claiming that the Obama campaign used these methods to mislead. However, the fact that the winning campaign’s “chief data scientist” was previously employed to “maximize the efficiency of supermarket sales promotions” does not thrill me. You should be worried even if your candidate is — for the moment — better at these methods. Democracy should not just be about how to persuade people to vote for one candidate over another by any means necessary.
Zeynep Tufekci is a fellow at the Center for Information Technology Policy at Princeton University.
Obama's Approach to Big Data: Do As I Say, Not As I Do
Politicians' Policy Decisions May Stymie Tools That Got Them Elected
Kate Kaye, Ad Age, November 16, 2012
One of the keys to success for President Barack Obama's reelection bid was its masterful use of data. But lost in the hype is this: The administration supports a browser-based do not track system that, if pervasive, would throw a wrench into the data-collection tactics that empowered the campaign.
Even today BarackObama.com features data-tracking cookies from several online ad and analytics firms.
The Mitt Romney and Obama campaigns spent hundreds of thousands of dollars in 2012 on data and related services to enhance their own voter contact information, inform their online and offline messaging and target ads. At the same time, Congress is inspecting the practices of firms that buy, sell and filter consumer data for corporate marketers.
"The Obama administration and the GOP should confront head-on the privacy issues raised by [their] far-reaching use of digital profiling and targeting data," argued privacy advocate Jeffrey Chester, founder of the Center for Digital Democracy. "It would be unfortunate for the administration's work to advance Do Not Track and other key safeguards if they failed to tackle the use of powerful data targeting technologies by political campaigns."
Industry and privacy wonks actually agree
It's a rare occurrence, but both Mr. Chester and the ad industry are in agreement on one thing: They both appreciate the attention the Obama data machine is getting. Privacy groups want to raise awareness of data collection and usage in the hopes of generating public support for curbing what they see as an increasingly infiltrative violation of personal privacy by marketers and the mushrooming data industry.
"Protecting the privacy of consumers and citizens should require policymakers from both sides to confront the civil liberties implications of what has been unleashed," added Mr. Chester, noting that the 2012 campaigns should divulge what data they collected, how they targeted ads and what will happen to the information now that the election is over.
Industry players, especially their Capitol Hill lobbyists, aim to convince legislators that the very data practices some of them criticize are helping them and their colleagues win races.
"Big data isn't going to help Todd Aken," said Mike Zaneis, general counsel of the Interactive Advertising Bureau, referring to the disgraced Congressman from Missouri who lost his Senate campaign after claiming women can ward off pregnancy resulting from "legitimate rape." Continued Mr. Zaneis, "But the Obama campaign used a lot of online data and a tremendous amount of offline data to go precinct-by-precinct to get-out-the-vote."
Third-party tags
More than a week after the election, BarackObama.com houses an array of third-party tags that track users for ad targeting and campaign and site analytics. Yesterday, around fifteen ad company tags were surfaced by Evidon's Ghostery software, including tags from BlueKai, which calls itself a "big data activation solution," and Appnexus, which among other things allows advertisers to use a variety of user behavioral data to target ads to those users on Facebook.
Both the Obama and Romney campaigns used social-media-widget and data provider ShareThis to target fundraising ads and identify issues and trends swing state voters were interested in, according to ShareThis CEO Kurt Abrahamson. The company tracks when people visit web pages and share them on Twitter, Facebook, LinkedIn or other popular social sites and allows advertisers to target ads using that anonymized information.
Clashing goals of campaigning and governing
Data tracking tools and techniques that have helped legislators on both sides of the aisle build supporter lists, generate donations and get out the vote could be stymied by a do-not-track browser standard or restrictive privacy legislation.
In February, the Federal Trade Commission and the ad industry announced they'd work together with browser companies to develop a DNT standard. At the same time, the U.S. Commerce Department introduced a consumer privacy bill of rights that guided companies to provide individual control over data collection, better data security measures, and transparency of data use, and also called for "a reasonable amount of data collection by companies." Secretary of Commerce John Bryson said at the time the department would work with Congress to implement the privacy bill of righs -- which some deem to be supportive of industry's self-regulatory approach -- through legislation.
The Digital Advertising Alliance, a large coalition of ad industry trade groups, has conducted an "ongoing dialogue with the FTC as recently as yesterday to figure out how to implement the [DNT] standard," said Stu Ingis, counsel to the DAA, on Wednesday. The DAA oversees the industry's Ad Choices program, which allows people to opt-out from online ad targeting through display ads that include the group's small triangular symbol. It's not entirely clear whether the FTC is confident that the DAA's self-regulatory program is enough to protect consumer privacy.
As reported by Politico earlier this month, FTC Chairman Jon Leibowitz said, "If by the end of the year or early next year, we haven't seen a real Do Not Track option for consumers, I suspect the commission will go back and think about whether we want to endorse legislation." Mr. Leibowitz is expected by beltway insiders to step down at the end of the year, and some believe his goal to finalize a DNT standard before he leaves is pressurizing the situation.
A free pass for political data?
Enter the Bipartisan Congressional Privacy Caucus. The group recently received responses to inquiries into several data firms that manage and analyze, and in some cases buy and sell, online and offline consumer data. Nine firms -- Acxiom, Epsilon, Equifax, Experian, Harte-Hanks, Intelius, Fair Isaac, Merkle, and Meredith Corp. -- submitted lengthy and often vague answers to a series of questions about their data businesses and practices.
"Many questions about how these data brokers operate have been left unanswered, particularly how they analyze personal information to categorize and rate consumers," said lawmakers in a joint statement regarding the companies' responses.
Absent from the list of data firms questioned were similar companies that deal mainly in voter file and political information that is often enhanced with consumer demographic, shopping and other data. For instance, NGP Van, the Democratic data powerhouse favored by the Obama team was not part of the inquiry. The Obama campaign and DNC spent hundreds of thousands of dollars with NGP Van this election cycle alone. The firm matches its voter data with data from TargetSmart, which offers "the richest set of consumer and interest data, allowing the most sophisticated targeting," according to the NGP Van site.
Other political data firms left out of the inquiry include Catalist, another Democratic data firm; Campaign Grid, which offers Republican data and online ad targeting; and Aristotle, a well-established non-partisan political data company. People involved with the congressional inquiry deny that political data firms were left off the list for any strategic reason.
In a press release about the data broker responses, the Privacy Caucus stated it "will push for whatever steps are necessary to make sure Americans know how this industry operates and are granted control over their own information."
Rep. Ed Markey, a Democrat from Massachusetts and Caucus co-chair, has sponsored a Do Not Track Kids Act and a mobile privacy bill.
Observers don't expect a privacy bill to be passed anytime soon; if that does happen, it may not apply to political campaigns or groups anyway. For instance, political messages are exempt from CAN-SPAM laws, and political organizations are not restricted by the Do Not Call Registry.
"Often when data laws are being proposed and put forward, the politicians exempt themselves," said Don Hinman, senior VP for data strategy at Epsilon, which gets some of its data from political advertisers but mainly is a purveyor of consumer information.
Mr. Ingis considers it exemption for political messages to be a first amendment issue. "It would be very hard for such a limitation on political messages to be restricted. . . . and I think that would have been true in the context of Do Not Call if they would have gone there," he said.
Tuesday, November 13, 2012
How 'Social Intelligence' Can Guide Decisions
By offering decision makers rich real-time data, social media is giving some companies fresh strategic insight.
Martin Harrysson, Estelle Metayer, and Hugo Sarrazin, McKinsey Quarterly, November 2012
In many companies, marketers have been first movers in social media, tapping into it for insights on how consumers think and behave. As social technologies mature and organizations become convinced of their power, we believe they will take on a broader role: informing competitive strategy. In particular, social media should help companies overcome some limits of old-school intelligence gathering, which typically involves collecting information from a range of public and propriety sources, distilling insights using time-tested analytic methods, and creating reports for internal company “clients” often “siloed” by function or business unit.
Today, many people who have expert knowledge and shape perceptions about markets are freely exchanging data and viewpoints through social platforms. By identifying and engaging these players, employing potent Web-focused analytics to draw strategic meaning from social-media data, and channeling this information to people within the organization who need and want it, companies can develop a “social intelligence” that is forward looking, global in scope, and capable of playing out in real time.
This isn’t to suggest that “social” will entirely displace current methods of intelligence gathering. But it should emerge as a strong complement. As it does, social-intelligence literacy will become a critical asset for C-level executives and board members seeking the best possible basis for their decisions.
In this article, we explore four distinct ways social technologies can augment the intelligence-gathering approaches of companies. As Exhibit 1 makes clear, social media has little effect on some aspects of the intelligence cycle—in particular, the need to identify priorities for exploration and decision making over the next 6 to 12 months, as well as the use of assembled information to make unbiased decisions. But social technologies can play a surprisingly central role in how information is sourced, collected, analyzed, and distributed.
Wednesday, November 7, 2012
Andrew McAfee : Let the Crowd Fix Your Product's Bugs
Andrew McAfee, Harvard Business Review Blog, November 6, 2012
I'm starting to come to the conclusion that of all the myths businesses and their leaders tell themselves, one of the most harmful is that they know where the expertise is. The more I learn about the results from crowdsourcing and open innovation efforts, the more I believe that the smart strategy is to expose your problems and challenges to as many people as possible and let them show you what they can do. Here's my most recent example of the power of this approach.
The online startup Kaggle assembles a diverse group of people from around the world to work on tough problems submitted by organizations. The company runs data science competitions, where the goal is to arrive at a better prediction than the submitting organization's starting 'baseline' prediction. Results from these contests are striking in a couple ways. For one thing, improvements over the baseline are usually substantial. In one case, Allstate submitted a dataset of vehicle characteristics and asked the Kaggle community to predict which of them would have later personal liability claims filed against them. The contest lasted approximately three months, and drew in more than 100 contestants. The winning prediction was more than 270% better than the insurance company's baseline.
Another interesting fact is that the majority of Kaggle contests are won by people who are marginal to the domain of the challenge — who, for example, made the best prediction about hospital readmission rates despite having no experience in health care — and so would not have been consulted as part of any traditional search for solutions. In many cases, these demonstrably capable and successful data scientists acquired their expertise in new and decidedly digital ways.
Between February and September of 2012 Kaggle hosted two competitions sponsored by the Hewlett Foundation about computer grading of student essays. Improvements in this area are important because essays are better at capturing student learning than multiple choice questions, but much more expensive to grade when human raters are used. So automatic grading of written answers would both improve the quality of testing and lower its cost. Kaggle and Hewlett worked with many education experts to set up the competitions, and as they were preparing to launch some of these people were worried.
The first contest was to consist of two rounds. Eleven established educational testing companies would compete against each other in the first, with members of Kaggle's community of data scientists invited to join in, individually or in teams, in the second. The experts were worried that the Kaggle crowd would simply not be competitive. After all, each of the testing companies had been working on automatic grading for some time, and had devoted substantial resources to the problem. Their hundreds of man years of accumulated experience and expertise seemed like an insurmountable advantage over a bunch of novices.
They needn't have worried. Many of the 'novices' drawn to the challenge outperformed all of the testing companies in the essay competition, and came closer to the consensus score of the human graders than did any of the humans themselves. The surprises continued when Kaggle investigated who the top performers were. In both competitions, none of the top three finishers had any previous significant experience with either essay grading or natural language processing. And in the second competition, none of the top three finishers had any formal training in artificial intelligence beyond a free online course offered by Stanford AI faculty and open to anyone in the world who wanted to take it. And people all over the world did, and learned a lot from it. The top three individual finishers were from, respectively, America, Slovenia, and Singapore.
Businesses certainly know where a lot of the relevant expertise is in any situation, but results like those from Kaggle show me that they certainly don't know where all of it is. As the open source software advocate Eric Raymond famously observed, with enough eyeballs all bugs are shallow. So why not expose your tough problems to as many eyeballs as possible?
Tuesday, November 6, 2012
How Big Data Could Determine the Winner of Today's Election
Tarun Wadhwa, Forbes, November 6, 2012
If your favorite soda is Diet Dr. Pepper, the chances are that you’ll be supporting Mitt Romney. Pepsi drinker? You’re most likely voting for Barack Obama. If you drink Mountain Dew, you probably don’t care either way.
These types of conclusions may seem simplistic and superficial, but both campaigns are betting that they will be the key to deciding who the next President of the United States is.
It’s more than what you drink, what you shop for, who your friends are, what websites you visit: all reveal clues to your political leanings. Campaigns have entered the era of “Big Data”—they target voters based on scraps of information they gather from unlikely places.
Thanks to the rise of mobile technology and social media, the number of records collected by data brokers on voter behavior has tripled—from 300 pieces in 2004 to more than 900 pieces today.
Campaigns care about your personal life
Voters used to be the ones obsessing over details of a candidate’s personal life. Now the tables have turned. Campaigns research the personal lives of the voter.
Micro-targeting, a technique that delivers ads based on the personal traits of a voter, was once considered impossible. But in 2004, it was recognized for helping George W. Bush defeat John Kerry. Now it is used by almost every campaign.
Because of the intricacies of our electoral system, a relatively small group of people ends up deciding the outcome of elections. In the 2000 Presidential campaign, hundreds of millions of dollars was spent on reaching just 7 percent of voters—fewer than 8 million people. Even a small advantage in mobilizing potential voters in a swing state can determine the difference between a win and a loss.
In this election cycle, more than $3 billion dollars has been spent on broadcast-television advertising, which has remained the dominant form of political communication for the last fifty years. But times are changing. Television purchases are no longer as effective as they used to be. A study showed that 88% of voters with DVRs skip ads and that 45% use something other than live TV as their primary mode for viewing videos. These proportions are even higher in younger demographics.
The next frontier: digital behavioral advertising
Just as television advertising revolutionized the field in the 1960s, this election will likely mark digital-behavioral advertising as the next frontier in voter outreach.
As a nation, we are already divided along partisan lines. We access different media, each with its own messaging and focus. Now we will receive different messages depending on who we are. Zac Moffatt, digital director for Mitt Romney’s campaign, said to The New York Times that “two people in the same house could get different messages,” and that “not only would the message change, the type of content would change.”
In an article for Stanford Law Review, Daniel Kreiss, a journalism professor at University of North Carolina, Chapel Hill, explains how this can have negative long-term consequences for democratic participation. With so much sensitive personal information in so many hands, there are risks of data breaches and unauthorized disclosure.
Citizens may hesitate to engage in political discussion on line for fear of being tagged and put into a marketing database. And the high cost of political data and consulting activities may make it difficult for less affluent candidates to compete effectively. Perhaps most worrying, campaigns may “redline” an electorate (by ignoring voters who won’t be sympathetic to their views because a model deems them unworthy of investment).
Political targeting – what’s next?
Sophisticated modeling and targeting will become commonplace at every step of the political process. NGOs, interest groups, and candidates for local office will be the next to adopt these methods.
United in Purpose, an evangelical Christian non-profit, is currently using such technology to assign points to voters based on whether they like NASCAR or fishing, and whether they are on anti-abortion or traditional marriage lists. If these voters have a score of over 600 points, they are considered “serious about their faith”. They will be contacted if they have not registered to vote.
Many voters would be surprised to learn that their interactions with both campaigns are being recorded and analyzed using technology similar to what Target uses to determine whether teenage girls are pregnant. When voters do learn what their candidates are doing, as many as 86 percent want this to stop. They regard it as an invasion of privacy. Yet these types of activities are legally considered political speech, so there are hardly any restrictions in place.
What is most worrisome is that there is no easy way to opt out of these databases, or to limit what information is collected about you, or how it is used. Sadly, we can’t “de-friend” or “unfollow” the politicians.
Saturday, November 3, 2012
Google Now: Behind the Predictive Future of Search
How Google learned to un-fragment itself and create the next big thing
Dieter Bohn, The Verge, October 29, 2012
For decades, visions of the future have played with the magical possibilities of computers: they'll know where you are, what you want, and can access all the world's information with a simple voice prompt. That vision hasn't come to pass, yet, but features like Apple's Siri and Google Now offer a keyhole peek into a near future reality where your phone is more "Personal Assistant" than "Bar bet settler." The difference is that the former actually understands what you need while the latter is a blunt search instrument.\
Google Now is one more baby step in that direction. Introduced this past June with Android 4.1 "Jelly Bean," it's designed to ambiently give you information you might need before you ask for it. To pull off that ambitious goal, Google takes advantage of multiple parts of the company: comprehensive search results, robust speech recognition, and most of all Google's surprisingly deep understanding of who you are and what you want to know.
With Android 4.2, launching alongside the Nexus 4 and Nexus 10 on November 13th, Google has updated the feature with new information cards in new categories. And yet, the amount of engineering effort that makes Google Now possible is out of proportion to what it does — it's a massive, cross-company effort for what seems like a relatively small product. That difference is a clue. Google Now isn't important for what it does, well, "now," but the building blocks are there for a radically different kind of platform in the future.
We sat down with the teams responsible for some of the technology that went into Google Now to find out what makes it tick today and discover some hints about what it could be in the future.
A deeper understanding
You may not be familiar with Google Now, primarily because it's only available on the sliver of Android devices running Jelly Bean (and up) — a situation that sadly won't change with the latest version. It's essentially an app that combines two important functions: voice search and "cards" that bubble up relevant information on a contextual basis.
Actually, Google Now technically only refers to the ambient information part of the equation, a branding kerfuffle that distinguishes it from Apple's Siri product yet still causes confusion. Those cards might contain local restaurants, the traffic on your commute home, or when your flight is about to take off. They appear automatically as Google tries to guess the information you'll need at any given moment.
While it seems like a relatively simple service, it's only really possible because of the massive amount of computational power Google can leverage alongside the massive amount of data Google knows about you thanks to your searches. It's "precisely what Google is best at," Android's director of product management, Hugo Barra, tells us. "It really feels like we’ve been working on Google Now for the past ten years. Because Google Now touches every back-end of Google, every different web service that’s been developed over the last ten years or so is part of this service."
The breadth of that backend and the simple cards it enables is what makes Google Now so intriguing as a product. One of Barra’s favorite examples is a voice search for something that pulls from all those multiple sources and turns it into a comprehensible and useful result. Searching for “Directions to the museum with the William Paley exhibition” causes Google to 1) find that exhibition, 2) understand you care about the museum where it is being shown, 3) know your location, and finally 4) present you with a simple map card to the museum itself along with a button to immediately get directions.
Taking all of that complex data and turning it into a relatively simple and useful interface is a gargantuan undertaking, but Google has started with a somewhat small set of categories for the types of cards it shows. With Jelly Bean, you'd see calendar alerts, weather, flight times, sports scores, transit directions, local restaurants, and a few more categories of information.
Even within that limited set of data, Google has to make choices about which cards to show you and when. It uses a few different signals — location, time, and all of your recent searches figuring prominently among them — to decide what to show you in any given moment. "It’s essentially a ranking problem, and it’s a very complicated one," according to Barra, but Google has perhaps more experience at solving ranking problems than any other company after years of delivering search results.
In my experience, Google is able to get you the "right" information you want a relatively small percentage of the time, but that low hit rate doesn't actually hurt the experience all that much. That's mainly thanks to the fairly small number of categories cards fall into, but also to the fact that when Google Now gets it right, it really feels magical. The sort of thing you might manually search for — like your commute time home — is simply waiting for you.
With the latest update, Google is expanding Now into new categories, increasing the different kinds of information it's able to provide. The new additions aren't radically ambitious, but that's in fitting with the overall feel of Google Now. What it shows you is more about serendipitous information than structured data.
The first category involved Gmail integration. With your permission, Google will keep an eye on your inbox and recognize flight confirmations, hotel reservations, restaurant bookings, event tickets, and package tracking emails. It will take that knowledge and give you a relevant card when appropriate — say, giving you your hotel information when you land in the right city or letting you know when it's time to leave for a concert.
The new features are part of Google’s growing efforts to provide relevant results based on the knowledge it’s accumulated about you. As search gets better, so do people’s expectations for what it provides. “Of course Google’s going to access more than just the public information on the web,” Scott Huffman, Engineering Director for Search Quality at Google tells us, “Google’s going to know when my flight is, whether my package has gotten here yet and where my wife is and how long it’s going to take her to get home this afternoon. [...] Of course, Google knows that stuff.” If you’re willing to opt in to letting Google know so much about you — and increasingly, opting in is the default — then Google wants to return the favor by using that information to your benefit. It requires you to trust Google quite a bit, but the company hopes that your trust will be rewarded.
These new cards are actually similar to a feature that Google added to its web search results this past August, both in content and in style. That's probably not an accident — if you assume Google has already won the battle for search, the next battle is giving you information before you even search for it. When it comes to deciding which data to give you, Barra tells us that Google has "a pipeline [...], possibly in the hundreds of cards” from its many engineering teams. Rather than flood users with all of those new cards, Google is taking a slow and steady approach to adding those new features — if only because right now it can only add those cards with a software update.
Some of the other new categories of cards are relatively minor additions: stocks, news, local concerts, movies, and local attractions. It also has a basic exercise tracking card that utilizes the phone’s accelerometer and location data: every month it will let you know how far you've walked or biked and also tell you how it compared to the month previous. Another new card lets you know that you're near a "photo opportunity," as Product Management Director Baris Gultekin told us. It uses data from Google's Panaramio service, noting when you're close to a place that has a "high density of pictures taken at a spot." You can see photos that were taken at the landmark and, Google hopes, take one yourself.
Neural networks
Just as Google Now's ambient information is backed by a massive and unseen engineering effort, Google's voice search is a simple feature that belies the effort that goes behind it. Huffman points out that getting voice search right actually involves more than just turning spoken words into textual queries, "speech recognition, natural language understanding, and understanding entities and knowledge in the world [all] really have to come together."
Voice search is the sort of feature that we take for granted on smartphones — Apple’s Siri and even Windows Phone both use the feature to offer up search results that go beyond basic web searches. What used to be a "hey neat" kind of feature is increasingly becoming an expected feature, and Google is well aware of that, "As you make search better, people’s expectations go up." To meet those expectations, Google is attacking all three of the areas Huffman delineated in equal measure.
Speech recognition is a very difficult problem to solve, as anybody who has dealt with voice search knows all too well. Recently, Google has changed its approach to making it work in a fundamental way, replacing a system that was the result of years of effort with a new framework for understanding the spoken word. Google has shifted to using a neural network that's much more effective at understanding speech.
A neural network is a computer system that behaves a bit like the actual neurons in your brain do.
Essentially, the computer is designed with layers of software-based "neurons" that do the same thing actual neurons do: take input in and "fire" off to other neurons based on the data they receive. Over the summer, the results of research led by Google Fellow Jeff Dean's on neural networks made some waves: Google had taught a computer to recognize cats in videos. The interesting part is that the neural network essentially created the concept of "cat" on its own without direct human intervention.
Here's how it works: The first layer of neurons looks for very simple things, like angled lines or colors. If it sees something that matches, it fires off a signal. There's then a second layer of neurons, which simply pays attention to sets of neurons firing from the first layer. As you add in more and more layers with the same behavior, you essentially add in layers of conceptual abstraction until, at the very top layer, there's a neuron that has trained itself to recognize cats 15.8 percent of the time.
Google's research scientists took this method and essentially applied it directly to speech recognition, fellow researcher Vincent Vanhoucke told us. "We picked up the kind of work that Jeff’s team was doing and just changed the input of the system." Google used the neural network at a very basic level of speech recognition: understanding and interpreting the basic sounds of speech: phonemes.
The approach "led to about between 20 to 25 percent reduction in the error rate in our system," according to Vanhoucke. The neural network turned out to be exceptionally good at solving what used to be very thorny problems in speech recognition. Accounting for "different environments, [...] different accents, different tones of voice, different pitches, different background noise, different microphones, [...], people talking in the background, different audio conditions" became much easier because the network was able to automatically learn how to account for each situation.
Knowledge Graph
Just understanding the words you've spoken isn't enough, obviously. Just as a neural network trades in increasing layers of abstraction, Google itself needs to move beyond basic web queries. In a very real way, Google is trying to get its computers to actually understand what it is you're asking them. Part of that comes from a relatively new initiative called the "Knowledge Graph," the company's effort to compile a database of "entities" in the world.
Today, Google's servers are aware of 500 million such entities, and "knowing" those things means that the company is able to act on them in interesting ways. For example, if you search for a “Tom Cruise,” Google knows you’re referring to a person instead of a vacation and can then tell you specific facts about him instead of simply crawling the web for related words. In truth, Google only knows those details because it is so adept at crawling the web — but the additional layer of abstraction created by putting that information into the structured Knowledge Graph means that Google can do more with search results. It "allows voice search, in some sense, [to] give me something to talk about," says Huffman. In Tom Cruise example, returning an “entity” instead of just a search result means you can contextually ask for more information, like “what movies has he been in?” or “how tall is he, really?”
Having something to talk about and talking to somebody are two different things, and with regard to the latter Google is again taking a Google-esque approach. As opposed to Apple's Siri, which you could say has a distinct personality, Huffman says that Google has "shied away from the idea of kind of a human persona for search or for the entity that you’re interacting with and instead tried to go for, in some sense, ‘hey, you’re interacting with all of Google.’"
The Google that you're interacting with in Google Now is very different than the Google you used even a year ago. The company's products have often felt fragmented, serving small niches and launched without feeling fully thought-through — and then in too many cases simply killed off. That may have been a function of the fact that Google is so large and does so much — but Google Now is a sign that all the different parts of Google are finally working together in a cohesive way.
"Google Now actually started as a twenty percent project," Barra told us. Google famously encourages its employees to work on "side projects" for some portion of their time, and what's interesting about Google Now is that although it started two years ago as one of these side projects, it's become a catalyst for integrating so many different parts of Google. Barra tells us that “we literally have dozens of teams working with us right now,” and the achievement with Google Now is that it feels like those teams are integrated, not fragmented.
In a single app, the company has combined its latest technologies: voice search that understands speech like a human brain, knowledge of real-world entities, a (somewhat creepy) understanding of who and where you are, and most of all its expertise at ranking information. Google has taken all of that and turned it into an interesting and sometimes useful feature, but if you look closely you can see that it's more than just a feature, it's a beta test for the future.
Crowdsourcing Medical Treatments
Samahope Crowdsources Simple, Life-Saving Surgeries For The Poor
Ellen McGirt, Fast Company, November 2, 2012
Veteran social entrepreneur Leila Janah of Samasource recently co-launched a new project to crowdfund medical treatment for the very poor. Think of it as Kiva for surgery.
“I just started bawling,” Leila Janah is telling me about a trip she took to Sierra Leone earlier this year. "I’m usually pretty steely as a matter of course. But I’ve been crying a lot more lately.”
Janah, the founder of Samasource, a nonprofit organization that brings paid digital work to very poor women and youth, is no stranger to harsh realities. She studies them for a living.
But a sweet teenaged girl named Tiangay Kaiwo moved Janah to tears. Kaiwo was waiting for surgery at the government hospital in Bo to repair the extensive damage to her body after her teacher brutally raped her. Traditional tribal remedies involving herbs and a bath in boiling water exacerbated her condition. Kaiwo had been living with a painful rectovaginal fistula for over a year, but best as Janah could tell, the rapes began when she was 12 years old. Janah met dozens of girls and women with similar stories, all needing life-changing surgeries that nobody could afford. “Nobody even to hold their hands, to tell them that this wasn’t their fault,” Janah said.
When the slice of the market you are trying to corner is as troubled as the one that Janah is, then tears are clearly a rational response. But after tears, at least if you’re built like Janah, comes action.
Enter her latest project, Samahope, an experiment in crowdfunding medical treatment--like burn care--for the very poor. Think of it as Kiva for surgery. “There are millions of people who need corrective surgeries that we take for granted in the West,” she says. But the very poor often need types of care, like fistula surgery, that are now wholly unfamiliar and largely unnecessary in the developed world. (You can contribute to the development of the site via their Indiegogo campaign here.)
The site launched last month and has already funded a handful of surgeries; there are over 70 profiles on the site. If you believe in the premise of the League of Extraordinary Women, then the business case is clear: If you get one girl back on her feet, she can go to school. If she goes to school, she can get a job. Enough girls join the workforce and a country gets uplifted. But Janah sees another benefit. “They want to say what happened to them, to tell their own stories,” she says.
She has collected so many of these stories--of the poor and the embattled and their search for basic human rights through employment--that she’s writing a book. “I just interviewed a security guard at a hotel in Freetown," Janah said. "He grew up as a rebel and child soldier in the conflict--think about that for a minute--then forced into the diamond mines. His life was so full of conflict, a constant struggle to access basic human resources, that it’s impossible to wrap my mind around.” Recalling young Kaiwo, “For someone like her, being able to tell her story and help other girls not become a victim is a very powerful thing.”
Samasource, which last month closed a $7.5 million round of philanthropic funding led by The MasterCard Foundation, has become the darling of the tech crowd for its deft use of the Internet to match an excess capacity of potential workers with the jobs they need to live in dignity. “But the scale of the problems can seem so great compared to the resources you have to address them,” says Janah--thus, the crowdsourcing project. Unlike her peers in the for-proft tech world, she is not going to be able to turn to her staffers with breathless reports of sky-high valuations, rounds of venture funding, or promises of equity upside.
And yet, Janah is convinced that dignified work can resurrect even the most damaged lives, and that her own business case is sound. “We’ve gotten the microwork model on the agenda of a lot of foundations and government entities. We just have to prove it can scale.”
Janah recalls with fondness the “aha” moment when she knew that Samasource could actually be a business. But the slog of ramping up to achieve a massive goal is largely free of lightbulb moments. “Not a day goes by when I don’t doubt myself or question something,” she says. So instead, she takes the power of microwork and puts it to work for herself and her team. “I had to manage my own psychology around this. So, I’ve trained myself to pause and celebrate each step.”
She rattles off a list of things that sound more startup than do-good: Realign your expectations, hit your goals, stay close to the customer, stay connected to your mission. She ends up sounding more like a Zen master than elevator pitcher. “It’s about looking down and doing what’s in front of you. Truly savor it. Then do the next thing. My job is to make sure we’re all going in the right direction, and at the end of the year, the sum of those steps adds up to something really great.”
Thursday, November 1, 2012
Mechanical Turk and the Limits of Big Data
Walter Frick, MIT Technology Review, November 1, 2012
The Internet is transforming how researchers perform experiments across the social sciences.
It’s telling that the most interesting presenter during MIT Technology Review’s EmTech session on big data last week was not really about big data at all. It was about Amazon’s Mechanical Turk, and the experiments it makes possible.
Like many other researchers, sociologist and Microsoft researcher Duncan Watts performs experiments using Mechanical Turk, an online marketplace that allows users to pay others to complete tasks. Used largely to fill in gaps in applications where human intelligence is required, social scientists are increasingly turning to the platform to test their hypotheses.
The point Watts made at EmTech was that, from his perspective, the data revolution has less to do with the amount of data available and more to do with the newly lowered cost of running online experiments.
Compare that to Facebook data scientists Eytan Bakshy and Andrew Fiore, who presented right before Watts. Facebook, of course, generates a massive amount of data, and the two spoke of the experiments they perform to inform the design of its products.
But what might have looked like two competing visions for the future of data and hypothesis testing are really two sides of the big data coin. That’s because data on its own isn’t enough. Even the kind of experiment Bakshy and Fiore discussed—essentially an elaborate A/B test—has its limits.
This is a point political forecaster and author Nate Silver discusses in his recent book The Signal and the Noise. After discussing economic forecasters who simply gather as much data as possible and then make inferences without respect for theory, he writes:
This kind of statement is becoming more common in the age of Big Data. Who needs theory when you have so much information? But this is categorically the wrong attitude to take toward forecasting, especially in a field like economics, where the data is so noisy. Statistical inferences are much stronger when backed up by theory or at least some deeper thinking about their root causes.
Bakshy and Fiore no doubt understand this, as they cited plenty of theory in their presentation. But Silver’s point is an important one. Data on its own won’t spit out answers; theory needs to progress as well. That’s where Watts’s work comes in.
The Internet is transforming how researchers think of the “lab” and enabling new kinds of experiments across the social sciences. Those experiments will be critical in helping us collectively make sense of the huge amounts of data we’re now generating. And those huge data sets will help inform the direction of Watts’s and others’ experiments.
The value of big data isn’t simply in the answers it provides, but rather in the questions it suggests that we ask.
Big Data and the Democratisation of Decisions (Economist Intelligence Unit
Report of the Economist Intelligence Unit, 2012
Too much important and relevant data – including new sources of Big Data – remains out of reach from those in your organization who can turn it into value.
Download Big data and the democratisation of decisions to learn:
· Why 77% of executives said more employees need access to Big Data;
· What the two biggest opportunities for Big Data to deliver business value are;
· What other kinds of data can provide better context for new sources of Big Data;
Subscribe to:
Posts (Atom)