Showing posts with label Election 2012. Show all posts
Showing posts with label Election 2012. Show all posts

Thursday, December 20, 2012

What's Next for Obama For America's Data and Technology?


Jim Pugh and Nathan Woodhull, The Huffington Post, December 18, 2012

If you've been following post-election news, you've no doubt heard about Barack Obama's "big data" advantage. The story has been the unprecedented investment in mining personal information for all kinds of wacky stuff. According to various articles, the campaign used data for everything from inviting donors to dinner with Sarah Jessica Parker to identifying potential voters through their visits to porn sites.

It's true that data was a game changer for the Obama campaign, but the reason is much less salacious than many reporters would have you believe. The campaign built an integrated database, which combined supporters' online activity, actions in the field, and public voting records into a single unified view of every American. On top of that data foundation, a large team of developers built community organizing software that empowered volunteers to become more deeply involved with the President's grassroots field operation.

The result was unprecedented efficiency in volunteer engagement and voter outreach. Supporters who signed up to help online received a personal call from their neighborhood leader the next day. Volunteers called only the most statistically persuadable potential supporters. Anyone who connected their Facebook account to the campaign was encouraged to send voting turnout messages to the specific friends who needed the extra push most. All of this led to more volunteers, more supporters, and increased turnout -- which all meant more Obama votes on Election Day.

But Election Day shouldn't be the end for these systems. We need to keep moving forward, continue the investment, and make them available to the whole movement.

Don't Abandon the Technology

Now that the election is over, funding will be tight. The donations that poured in during the election have dried up and hard choices need to be made.

One option is to mothball this infrastructure and plan to spin it up again for the next presidential election. This approach is attractive from a financial perspective, as no additional resources would be needed in the off cycle to make it happen.

But this would be a mistake. The campaign's advantage this cycle wasn't just bits and bytes, but the institutional experience and knowledge that had been building since the President's primary campaign in 2007. Instead of shelving all these systems, we should make a continued investment in maintaining and improving them.

Republicans are lagging behind Democrats right now on the data and technology front, but after the shellacking they experienced in the last election, there's little doubt that they'll be pouring in resources to catch up. Without sustained investment from our party, our advantage may be erased in the coming years. Not only would Republicans be moving ahead, but without staff to maintain institutional knowledge and adapt the systems to changing technology, Democrats could actually start the next election cycle behind where we're at right now. We need to keep moving forward if we want to keep our advantage on this front.

The technology industry never stops moving forward. Neither should we.

Make Tools Available to All Progressives

An investment in building on these systems offers another possibility as well: providing access to the rest of the progressive movement. Right now, only presidential campaigns have the resources to build systems of this sophistication. The data and technology infrastructure from the Obama campaign cost millions of dollars to build, and even the most well-funded senate campaigns couldn't afford anything close to that.

But with some additional work, the data and tech infrastructure from the Obama campaign could be adapted to offer the same functionality to other progressive candidates and groups, giving them the opportunity to use these systems with their own supporters and volunteers. For smaller campaigns that would have no chance of creating these systems on their own, this could be a game-changing step forward. And beyond the benefit to the Democratic Party and progressive movement, it could provide a path to fund the continued investment, via paid licensing from these outside campaigns and organizations.

A Sustained Tech Commitment for 21st-Century Politics

The roller-coaster of scaling up and scaling down that comes with elections has always been the standard for political organizations. If we want to stay competitive on the data and technology front, this approach needs to be changed. A 21st-century political movement must have a serious on-going commitment to staying at the forefront of technological advancement. With the infrastructure coming out of the Obama campaign, we've got a huge lead in this area.

Sadly, that is not what has happened so far. Since Election Day, the Democratic National Committee has laid off an unprecedented number of technology staff members, some of whom had been at the party for over ten years. The Obama Campaign's technology team is scattering to the winds and returning to industry. The window of opportunity to stay ahead of the technological curve is closing -- our party needs to change course now or risk being left behind.

Thursday, December 6, 2012

Ethan Roeder (Big Data Czar for Obama) in NYT: I Am Not Big Brother


Ethan Roeder, The New York Times, December 6, 2012

I’VE grown accustomed to reading inaccurate accounts of my day job. I’m in political data.

If I’m not spying on private citizens through the security cam in the parking garage, I’m probably sifting through their garbage for discarded pages from their diaries or deploying billions of spambots to crack into their e-mail. Reading what others muse about my profession is the opposite of my middle-school experience: people with only superficial information about me make a bunch of assumptions to fill in what’s missing and decide that I’m an all-knowing super-genius.

Sadly for me, this is a bunch of malarkey. You may chafe at how much the online world knows about you, but campaigns don’t know anything more about your online behavior than any retailer, news outlet or savvy blogger.

There are two categories of online data: information users provide explicitly, and stuff they communicate implicitly through their behavior. The explicit data includes e-mails and comments that users share directly. The implicit data comes from “click tracking,” which tells a campaign what buttons are getting pressed and how often. Combined, these two categories of data allow a campaign to put together an online experience that will resonate with as many people as possible, but also to customize the experience so that you are more likely to encounter content that’s relevant to you.

At times it might seem like sorcery to the recipient of a targeted e-mail, but it’s just a product of two simple factors: remembering who you are and remembering what you like.

In the offline world, which is my personal area of expertise, campaigns don’t know much more today than they did 44 years ago. In 1968, George Romney, then a candidate for the Republican presidential nomination, made headlines for using a newfangled “secret weapon” — a voter file. As The Times explained that winter, it is “an electronic data bank” that “contains the only really accurate, up-to-date roster of enrolled New Hampshire Republicans any candidate here has ever compiled, plus pertinent information about all of them.” Imagine how freaked out those New Hampshire Republicans must have been.

Virtually all of the offline data that people like me traffic in is boring, basic and publicly available. Want to know the year of birth for everyone who is registered to vote in Ohio? Just Google “Ohio voter file download.” There you go. I was born in 1976. Now we’re even.

How do we predict whether people are going to vote or not? We look at the voter file. It tells us how often a person votes, although not for whom. Not all strategists agree about how to interpret this information, but the source of the data is no secret.

What’s really new in politics today is not the data itself but how campaigns make sense of it. Cheaper and more plentiful computing power allows campaigns to process far more information than ever before to look for patterns, trends and correlations.

The science of modeling is a modern-day application of a practice that has been around for nearly 200 years: polling. Pollsters ask voters whom they support for president and how strongly. Campaigns then take demographic information about these voters into account in order to make assumptions about the entire population of a given state. The mechanics are exactly the same for public polls and internal campaign analyses. The difference is that the campaigns use statistical techniques to apply these assumptions to individual records in the voter file rather than stopping short and simply assuming that entire sections of the electorate will behave identically.

Contemporary data practice also frees campaigns from having to make assumptions about voters in the first place. In 2011 and 2012, the Obama campaign, with the help of more than two million volunteers, had more than 24 million conversations with voters. Online tools gave Obama supporters resources to help them play a crucial role in their neighborhoods, and a series of “share your story” pages on the campaign Web site provided a venue for voters to communicate directly with the campaign in long form.

All of this feedback doesn’t neatly boil down to a “yes” or “no” in a database — and why should it? Numerous avenues of listening, combined with the digital capacity to hold on to qualitative feedback, make campaigns aware of the differences among voters’ motivations, attitudes, protestations — not just their demographics and voting history. In a nation of over 200 million eligible voters, technology is allowing campaigns to finally see through the fog of the crowd and engage voters one by one.

In other words, there is no giant blue computer sitting on the 101st floor of a sleek skyscraper, surrounded by bubbling tubes of illuminated liquid, spitting out the manifest destiny of America’s voters. Campaigns are moving away from the meaningless labels of pollsters and newsweeklies — “Nascar dads” and “waitress moms” — and moving toward treating each voter as a separate person.

In 2012 you didn’t just have to be an African-American from Akron or a suburban married female age 45 to 54. More and more, the information age allows people to be complicated, contradictory and unique. New technologies and an abundance of data may rattle the senses, but they are also bringing a fresh appreciation of the value of the individual to American politics.

Ethan Roeder was the data director of Obama for America.

Thursday, November 29, 2012

Nate Silver: In Silicon Valley, Technology Talent Gap Threatens G.O.P. Campaigns


Nate Silver, The New Yorks Times Five Thirty Eight Blog, November 28, 2012

SAN FRANCISCO – I live in Brooklyn, where President Obama won 81 percent of the vote this month. It’s hard to find anywhere in the country that is more Democratic-leaning.

But San Francisco qualifies. Here, Mr. Obama won 84 percent of the vote, while Mitt Romney took just 13 percent. Even John McCain, who won 14 percent of the vote four years ago, performed slightly better than Mr. Romney did.

And unlike the New York metropolitan area, where Long Island, the borough of Staten Island and many suburbs in New York and New Jersey remain competitive in presidential elections, it is hard to find any significant pockets of support for Republican candidates in the nine counties that make up the San Francisco Bay Area.

Instead, Mr. Obama won the nine counties of the Bay Area by margins ranging from 25 percentage points (in Napa County) to 71 percentage points (in the city and county of San Francisco). In Santa Clara County, home to much of the Silicon Valley, the margin was 42 percentage points.

Over all, Mr. Obama won the election by 49 percentage points in the Bay Area, more than double his 22-point margin throughout California.

Although San Francisco, Oakland and Berkeley have long been liberal havens, the rest of the region has not always been so. In 1980, Ronald Reagan won the Bay Area vote over all, along with seven of its nine counties. George H.W. Bush won Napa County in 1988.

Republicans have lost every county in the region by a double-digit margin since then. But Democratic margins have become more and more emphatic. Mr. Obama’s 49-point margin throughout the Bay Area this year was considerably larger than Al Gore’s 34-point win in 2000, for example, or Bill Clinton’s 31-point win in 1992.

Even without the Bay Area’s vote, Democrats would still be favored to win California by solid margins. So why does any of this matter?


The reason is that Democrats’ strength in the region is hard to separate out from the growth of its core industry — information technology – and the advantage that having access to the most talented individuals working in the field could provide to Democratic campaigns.

Companies like Google and Apple do not have their own precincts on Election Day. However, it is possible to make some inferences about just how overwhelmingly Democratic are the employees at these companies, based on fund-raising data. (The Federal Election Commission requires that donors to presidential campaigns disclose their employer when they make a campaign contribution.)

Among employees who work for Google, Mr. Obama received about $720,000 in itemized contributions this year, compared with only $25,000 for Mr. Romney. That means that Mr. Obama collected almost 97 percent of the money between the two major candidates.

Apple employees gave 91 percent of their dollars to Mr. Obama. At eBay, Mr. Obama received 89 percent of the money from employees.

Over all, among the 10 American-based information technology companies on Fortune’s list of “most admired companies,” Mr. Obama raised 83 percent of the funds between the two major party candidates.

Mr. Obama’s popularity among the staff at these companies holds even for those which are not headquartered in California. About 81 percent of contributions at Microsoft, which is headquartered in Redmond, Wash., went to Mr. Obama. So did 77 percent of those at I.B.M., which is based in Armonk, N.Y.

It does not require an algorithm to deduce that the sort of employees who may be willing to donate substantial money to a political campaign may also be those who would consider working for it.

Since Democrats had the support of 80 percent or 90 percent of the best and brightest minds in the information technology field, it shouldn’t be surprising that Mr. Obama’s information technology infrastructure was viewed as state-of-the-art exemplary, whereas everyone from Republican volunteers to Silicon Valley journalists have criticized Mr. Romney’s systems. Mr. Romney’s get-out-the-vote application, Project Orca, is widely viewed as having failed on Election Day, perhaps contributing to a disappointing Republican turnout.

This is not intended to absolve Mr. Romney and his campaign entirely. There were undoubtedly many bright and talented information technology professionals who worked for Mr. Romney, and who might have fielded a better product given better management.

Even if only 10 percent or 20 percent of elite information technology professionals would consider working for a Republican like Mr. Romney, this is still a reasonably large talent pool to draw from.

But Democrats are drawing from a much larger group of potential staff and volunteers in Silicon Valley.

Perhaps a different type of Republican candidate, one whose views on social policy were more in line with the tolerant and multicultural values of the Bay Area, and the youthful cultures of the leading companies here, could gather more support among information technology professionals.

Ron Paul, the libertarian-leaning Republican, raised about $42,000 from Google employees, considerably more than Mr. Romney did.




Wednesday, November 21, 2012

Campaigns’ Use of Supporters’ Data Worries Privacy Advocates


Craig Timberg, The Washington Post, November 20, 2012


Shortly before Election Day, a Stanford graduate student reported that the campaign Web sites of both President Obama and Republican Mitt Romney were “leaking” personal information about their supporters through careless data handling.

Had it been Facebook and Google, a federal investigation might have ensued, and the companies could have suffered significant public relations setbacks and perhaps fines. But the Federal Trade Commission, the government agency most focused on personal privacy, has no jurisdiction over campaigns or political groups.
That is a small example of what privacy advocates say is a big problem with efforts to protect personal information in the United States: The politicians are not guarding the chicken coop. They are the foxes.

Obama’s sophisticated use of Big Data gave him a crucial edge in what, based on popular support alone, should have been a close election. Republicans are desperate to catch up. But it’s not clear who is positioned to protect the rights of voters at a time when politicians from both parties increasingly build their campaigns on the insights that commercial data brokers provide.

Washington has a community of professional privacy advocates at places such as the ACLU, the Electronic Privacy Information Center and the Center for Digital Democracy. Jeff Chester, executive director of the Center for Digital Democracy, said he approached lawmakers from both parties to express his concerns long before the election. But he got nowhere.

“Maybe we’re digital Don Quixotes,” Chester said. “There was a lack of interest, not surprisingly.”

People routinely tell pollsters that they’re concerned about online privacy, and Chester and his colleagues in the field count some allies on Capitol Hill and in the White House. The FTC under Chairman Jon Leibowitz and David Vladeck, head of its Bureau of Consumer Protection, have made the agency far more aggressive on consumer privacy generally — even if political campaigns are beyond their reach.

Yet overall the laws in the United States are much less strict than in Europe, where there are tight limits on what personal information can be collected and how long it can be kept. Companies caught crossing the line can provoke furious backlashes among their users.

The American political landscape, by comparison, is amorphous when it comes to privacy. There are widespread concerns on both the right and left but no single, coherent constituency demanding greater protections.

For all the talk in recent years about online privacy, data-hungry Google remains the most popular search engine and data-hungry Facebook the most popular social media site. Both worked closely with the campaigns and also have growing lobbying operations in Washington. Google’s Executive Chairman Eric Schmidt was a regular visitor at Obama’s Chicago campaign headquarters, say those who worked there, offering advice to the campaign’s data-savvy technologists.

Privacy advocates say the tide will eventually turn, when Americans truly understand the extent to which their information is collected and traded. A recent poll by the University of Pennsylvania’s Annenberg School for Communications found that nearly two-thirds of people would be less likely to support a candidate who bought data about voters’ online activities and used it to tailor political ads.

“People still don’t quite understand this stuff,” said Joseph Turow, the lead researcher on the Annenberg poll. He said politicians are “hoping people will, quote-unquote, get used to it.”

Monday, November 19, 2012

Crovitz: Obama's 'Big Data' Victory


Marketing politicians is now like selling drinks. It involves filtering policies and voters through algorithms.

L. Gordon Crovitz, The Wall Street Journal, November 18, 2012

When the Obama campaign emailed supporters to join a $40,000-a-ticket dinner in June at the New York home of actress Sarah Jessica Parker, journalists at ProPublica noticed something odd. They uncovered seven versions of the email solicitation for the fundraiser, some mentioning a second fundraiser that night, a concert by Mariah Carey, others that Ms. Parker is a mother, and still others that Vogue editor Anna Wintour would be at the dinner.


Who got which email depended on "big data"—information about each fundraising prospect and how different people react to different messages. In this year's election, it looks as if the Obama team's use of such data was one of its biggest edges over the Romney effort.

Some uses of big data were known before the election—for instance, the Obama website was even more assiduous than online retailers like Best Buy about dropping "cookies" on users' computers to gather information about their online habits. Reporting since the election makes clear just how important the role of data was in deciding the election.

Campaign manager Jim Messina pledged to "measure every single thing in this campaign" and built an analytics department five times the size of the 2008 effort. A Time magazine reporter got access to the data scientists in the campaign's Chicago headquarters on the condition that the reporter would keep mum until after the election. "What they revealed as they pulled back the curtain," Time recently reported, "was a massive data effort that helped Obama raise $1 billion, remade the process of targeting TV ads and created detailed models of swing-state voters that could be used to increase the effectiveness of everything from phone calls and door knocks to direct mailings and social media."

According to the magazine, the campaign created a "single massive system that could merge the information collected from pollsters, fundraisers, field workers and consumer databases as well as social-media and mobile contacts with the main Democratic voter files."

The campaign's "chief scientist," Rayid Ghani, had been at Accenture, where he co-wrote an academic paper describing work helping companies that "analyze large amounts of transactional data but are unable to systematically 'understand' their products." For example, Mr. Ghani helped grocers figure out why people bought orange juice by reducing the product to attributes that could be analyzed by algorithms—"Brand: Tropicana, Pulp: low, Fortified with: Vitamin-D, Size: 1 liter, Bottle type: plastic."

Marketing politicians is now like selling drinks. It involves filtering polices and voters through algorithms.

The Obama campaign focused on data showing the "persuadability" of voters. Multivariate tests identified issues and positions that could move undecided voters, ProPublica said: "The persuasion scores allowed the campaign to focus its outreach efforts—and their volunteer calls—on voters who might actually change their minds as the result. It also guided them in what policy messages individual voters should hear."

Big data give incumbents a big advantage, which seems to have surprised the Romney team. The Obama campaign has used cookies to track its supporters online since the 2008 election. It spent the past 18 months creating a new, unified database, factoring in some 80 pieces of information about each person, from age, race and sex to voting history. (The campaign denied reports that it tracked visits to pornography sites in its outreach algorithms.) The Romney campaign says it tried to match the Obama campaign's collection and analysis of data but had to start from scratch and had just seven months after the primaries.

What does this mean for you? Voters need to develop buyer-beware habits. The era of politicianssaying the same thing to all voters is over. Campaigns aim to tell voters exactly what each wants to hear: data-driven pandering.

Another consequence is that efforts by the Federal Trade Commission and other agencies to regulate data mining in the name of privacy are destined to collapse. Last month, Sen. Jay Rockefeller (D., W.Va.) sent a letter to the top "information broker" companies, accusing them of being "elusive" about what data they collect. Companies such as Acxiom and Experian replied that much of their information comes from government databases. They should also point out that political campaigns are among the most sophisticated users of the consumer data they collect.

The Obama campaign deserves credit for its big win through the sophisticated use of big data. As for regulators, they should understand that the information genie will not go back into the bottle—whether consumer information is used to sell orange juice or politicians.

A version of this article appeared November 19, 2012, on page A17 in the U.S. edition of The Wall Street Journal, with the headline: Obama's 'Big Data' Victory.

Sunday, November 18, 2012

Thaler: Applause for the Numbers Machine


Richard H. Thaler, The New York Times, November 18, 2012

THE biggest winners on Election Day weren’t politicians; they were numbers folks.

Computer scientists, behavioral scientists, statisticians and everyone who works with data should be proud. They told us who was going to win, but they also helped to make many of those victories happen.

Three groups of geeks deserve the love they rarely receive: people who run political polls, those who analyze the polls and those who figure out how to help campaigns connect with voters.

Many people doubted the accuracy of political polling this year. Part of the skepticism was based on the wide range of predictions, with some showing President Obama in the lead, and others Mitt Romney. But there were additional, structural reasons to worry whether pollsters would be able to find representative samples of voters.

One problem is that people are harder to reach on the telephone these days. About a third of voters no longer have a land line, and many of those who have them don’t pick up calls from strangers. So modern polling companies have to work harder to find voters willing to answer questions, then have to guess which of these respondents will actually show up and vote.

So it may come as a surprise that, collectively, polling companies did quite well during this election season. Although there was a small tendency for the pollsters to overestimate Mr. Romney’s share of the vote, a simple average of the polls in swing states produced a very accurate prediction of the Electoral College outcome. Notably, the most accurate polls tended to be done via the Internet, many by companies new to this field. That’s geek victory No. 1.

This relatively accurate polling data provided the raw material for the second group of election pioneers: poll analysts like Nate Silver, who writes the FiveThirtyEight blog for The New York Times, as well as Simon Jackman at Stanford, Sam Wang at Princeton and Drew Linzer at Emory University.

What do poll analysts do? They are like the meteorologists who forecast hurricanes. Data for meteorologists comes from satellites and other tracking stations; data for the poll analysts comes from polling companies. The analysts’ job is to take the often conflicting data from the polls and explain what it all means.

Worry about the reliability of the polling data led to widespread skepticism, or even outright hostility, toward poll analysts. The phrase “garbage in, garbage out” was one of the more polite criticisms bouncing around the Internet in the days before the election.

Because the polls were not, in fact, garbage, the first job of a poll analyst was quite easy: to average the results of the various polls, weighing more reliable and recent polls more heavily and correcting for known biases. (Some polls consistently project higher voter shares for one party or the other.)

A harder but more valuable task is to help readers translate the polling data into forecasts of the probability of victory. In Florida, where the final polls showed essentially a tie, according to Mr. Silver’s weighting method, it’s easy to see why he said the chance of either candidate winning the state was 50 percent. Ultimately, President Obama would very narrowly carry the state.

But what about North Carolina, where Mr. Silver projected that Mitt Romney would get 50.6 percent of the vote and President Obama, 48.9 percent? Looking at that very small difference, what probability would you have assigned to a Romney victory in that state?

Most people would guess something very close to 50-50. But not a good numbers guy. By looking back at previous elections with polling data this close, Mr. Silver estimated that Mr. Romney’s chances of winning North Carolina were 74 percent, a number that may seem surprisingly high. (Mr. Romney won the state.)

The slightly larger but still seemingly tiny lead that the president held in Ohio, another swing state, led poll analysts to predict that the chance of an Obama victory in Ohio was around 90 percent. And because Mr. Romney would have to win several such states with small Obama leads in order to prevail in the Electoral College, the analysts ended up with similarly high degrees of confidence in an overall Obama victory. They ended up predicting the Electoral College outcome almost exactly right, especially if you consider the final outcome in Florida to be a virtual tie, as they had projected.

Pundits making forecasts, some of whom had mocked the poll analysts, didn’t fare as well, and many failed miserably. George F. Will predicted that Mr. Romney would win 321 electoral votes, which turned out to be very close to President Obama’s actual total of 332. Jim Cramer from CNBC was nearly as wrong in the opposite direction, projecting that the president would win 440 electoral votes.

There is a lesson here. When it comes to assessing the chances of some complicated combination of events, gut feelings are pretty much useless. Pundits are no better at forecasting election outcomes than they would be at predicting the final path of a hurricane. Smart pundits should consider either abandoning this activity, or consulting with the geeks before rendering their guesses.

The third set of folks who deserve recognition in this election cycle were a group of young people working in a windowless room at Obama headquarters, affectionately known as the cave. They were part of the effort by the numbers-oriented campaign manager, Jim Messina, to maximize turnout.

THERE are two basic parts of an election campaign. The first comes under the category of messaging — deciding what a candidate should say and what ads to run. Most of the commentary we read about elections focuses on this component.

The second part is turnout, and in some ways is even more important. Here is a simple bit of math that you don’t have to be a geek to understand: It doesn’t matter which candidate a person prefers unless that person shows up and votes.

Pundits will debate for eternity which campaign did a better job of communicating its message, but there is no doubt which campaign won the turnout contest. Young, black and Hispanic voters all turned out in higher numbers than expected, and they often supported President Obama.

Much was made of the big Obama advantage in field offices in swing states. But those field offices would have been little good to the campaign without modern tools to find potential voters, have them register and encourage them to vote. In the weeks leading up to the election, the Obama canvassers had accurate lists of potential voters and field-tested scripts for their contacts with voters. This explains in part why Democrats were such heavy users of early voting.

By contrast, Project Orca, a get-out-the-vote computer program for the Romney campaign that wasn’t designed to be used until Election Day, reportedly had some bugs.

There should be something reassuring about this Obama campaign efficiency to all Americans, even those who supported Mr. Romney based on his success in business. When it came to the business of running a campaign, it was the former professor and community organizer who had the more technologically savvy organization and made more effective use of its resources, including geek power.

Richard H. Thaler is a professor of economics and behavioral science at the Booth School of Business at the University of Chicago. He was an informal adviser to the Obama campaign.

Saturday, November 17, 2012

Beware the Smart Campaign


Zeynep Tufekci, The New York Times, November 16, 2012


“I AM not a number. I am a free man!” was the famous cry of prisoner Number Six, who could never escape his Kafkaesque village on the 1960s television show “The Prisoner.” This is a prescient cry for an era when numbers follow us everywhere. Jim Messina, the victorious Obama campaign manager, probably agrees that you are not a number. That’s because you are four numbers.

The Obama campaign assigned all potential swing-state voters one number, on a scale of 1 to 100, that represented the likelihood that they would support Mr. Obama, and another number for the prospect that they would show up at the polls. A third metric evaluated the odds that an Obama supporter who was an inconsistent voter could be nudged to the polls, and a fourth score estimated how persuadable someone was by a conversation on a particular issue (which was, of course, also determined by crunching more numbers).
Mr. Messina is understandably proud of his team, which included an unprecedented number of data analysts and social scientists. As a social scientist and a former computer programmer, I enjoy the recognition my kind are getting. But I am nervous about what these powerful tools may mean for the health of our democracy, especially since we know so little about it all.

For all the bragging on the winning side — and an explicit coveting of these methods on the losing side — there are many unanswered questions. What data, exactly, do campaigns have on voters? How exactly do they use it? What rights, if any, do voters have over this data, which may detail their online browsing habits, consumer purchases and social media footprints?

How did Mr. Obama win? The message and the candidate matter, of course; it’s easier to persuade voters if your policies are more popular and your candidate more appealing. But a modern winning campaign requires more. As Mr. Messina explained, his campaign made an “unparalleled” $100 million investment in technology, demanded “data on everything,” “measured everything” and ran 66,000 computer simulations every day. In contrast, Mitt Romney’s campaign’s data operations were lagging, buggy and nowhere as sophisticated. A senior Romney aide described the shock he experienced in seeing the Obama campaign turn out “voters they never even knew existed.” And that kind of ability matters: while Mr. Obama did win decisively, the size of his lead in four states that determined the outcome, Florida, Ohio, Virginia and Colorado, was about 400,000 votes — or about 1.2 percent of the eligible voters.

The confluence of marketing and politics goes back a long way. A blizzard of direct mail engineered by political consultants is credited with defeating President Harry S. Truman’s national health care proposal after World War II. The new methods, however, are not just better direct mail. Noxious TV ads and slick mailers are like machetes compared with the scalpels of social-science-based big-data. The crude methods may still work to soften the ground and drown out other voices, but in the end they are still very big sticks. Sometimes they kill the patient — just ask swing-state voters about the TV ads they were bombarded with.

The scalpels, on the other hand, can be precise and effective in a quiet, un-public way. They take persuasion into a private, invisible realm. Misleading TV ads can be countered and fact-checked. A misleading message sent in just the kind of e-mail you will open or ad you will click on remains hidden from challenge by the other campaign or the media. Or someone who visits evangelical Web sites might be carefully shielded from messages about gay rights, and someone who has hostile views toward environmentalism may receive messages stroking that sentiment even if the broader campaign woos the green vote elsewhere.

What I really worry about, though, is that these new methods are more effective in manipulating people. Social scientists increasingly understand that much of our decision making is irrational and emotional. For example, the Obama campaign used pictures of the president’s family at every opportunity. This was no accident. The campaign field-tested this as early as 2007 through a rigorous randomized experiment, the kind used in clinical trials for medical drugs, and settled on the winning combination of image, message and button placement. I agree that his family is wonderful and his daughters are cute. But an increasing role of “likability” factors, which we now understand better how to manipulate, is not good for democracy.

These methods will also end up empowering better-financed campaigns. The databases are expensive, the algorithms are proprietary, the results of experiments by campaigns are secret, and the analytics require special expertise. The Democrats have an early advantage partly because academics and data analysts tend to be Democrats. Money will solve that problem. This will shift power in both parties even more toward the richer campaigns and may well be the final nail in the coffin of public financing for presidential campaigns.
What is to be done? Campaigns should make public every outreach message so we at least know what they are saying. These messages can be placed in a public database like campaign contributions so the other side can be aware of, and have the right to respond to, false claims. Political access to proprietary databases should be regulated to provide an even playing field.

I’m not claiming that the Obama campaign used these methods to mislead. However, the fact that the winning campaign’s “chief data scientist” was previously employed to “maximize the efficiency of supermarket sales promotions” does not thrill me. You should be worried even if your candidate is — for the moment — better at these methods. Democracy should not just be about how to persuade people to vote for one candidate over another by any means necessary.

Zeynep Tufekci is a fellow at the Center for Information Technology Policy at Princeton University.

Obama's Approach to Big Data: Do As I Say, Not As I Do


Politicians' Policy Decisions May Stymie Tools That Got Them Elected

Kate Kaye, Ad Age, November 16, 2012


One of the keys to success for President Barack Obama's reelection bid was its masterful use of data. But lost in the hype is this: The administration supports a browser-based do not track system that, if pervasive, would throw a wrench into the data-collection tactics that empowered the campaign.

Even today BarackObama.com features data-tracking cookies from several online ad and analytics firms.
The Mitt Romney and Obama campaigns spent hundreds of thousands of dollars in 2012 on data and related services to enhance their own voter contact information, inform their online and offline messaging and target ads. At the same time, Congress is inspecting the practices of firms that buy, sell and filter consumer data for corporate marketers.

"The Obama administration and the GOP should confront head-on the privacy issues raised by [their] far-reaching use of digital profiling and targeting data," argued privacy advocate Jeffrey Chester, founder of the Center for Digital Democracy. "It would be unfortunate for the administration's work to advance Do Not Track and other key safeguards if they failed to tackle the use of powerful data targeting technologies by political campaigns."

Industry and privacy wonks actually agree
It's a rare occurrence, but both Mr. Chester and the ad industry are in agreement on one thing: They both appreciate the attention the Obama data machine is getting. Privacy groups want to raise awareness of data collection and usage in the hopes of generating public support for curbing what they see as an increasingly infiltrative violation of personal privacy by marketers and the mushrooming data industry.


"Protecting the privacy of consumers and citizens should require policymakers from both sides to confront the civil liberties implications of what has been unleashed," added Mr. Chester, noting that the 2012 campaigns should divulge what data they collected, how they targeted ads and what will happen to the information now that the election is over.

Industry players, especially their Capitol Hill lobbyists, aim to convince legislators that the very data practices some of them criticize are helping them and their colleagues win races.

"Big data isn't going to help Todd Aken," said Mike Zaneis, general counsel of the Interactive Advertising Bureau, referring to the disgraced Congressman from Missouri who lost his Senate campaign after claiming women can ward off pregnancy resulting from "legitimate rape." Continued Mr. Zaneis, "But the Obama campaign used a lot of online data and a tremendous amount of offline data to go precinct-by-precinct to get-out-the-vote."

Third-party tags
More than a week after the election, BarackObama.com houses an array of third-party tags that track users for ad targeting and campaign and site analytics. Yesterday, around fifteen ad company tags were surfaced by Evidon's Ghostery software, including tags from BlueKai, which calls itself a "big data activation solution," and Appnexus, which among other things allows advertisers to use a variety of user behavioral data to target ads to those users on Facebook.


Both the Obama and Romney campaigns used social-media-widget and data provider ShareThis to target fundraising ads and identify issues and trends swing state voters were interested in, according to ShareThis CEO Kurt Abrahamson. The company tracks when people visit web pages and share them on Twitter, Facebook, LinkedIn or other popular social sites and allows advertisers to target ads using that anonymized information.

Clashing goals of campaigning and governing
Data tracking tools and techniques that have helped legislators on both sides of the aisle build supporter lists, generate donations and get out the vote could be stymied by a do-not-track browser standard or restrictive privacy legislation.


In February, the Federal Trade Commission and the ad industry announced they'd work together with browser companies to develop a DNT standard. At the same time, the U.S. Commerce Department introduced a consumer privacy bill of rights that guided companies to provide individual control over data collection, better data security measures, and transparency of data use, and also called for "a reasonable amount of data collection by companies." Secretary of Commerce John Bryson said at the time the department would work with Congress to implement the privacy bill of righs -- which some deem to be supportive of industry's self-regulatory approach -- through legislation.

The Digital Advertising Alliance, a large coalition of ad industry trade groups, has conducted an "ongoing dialogue with the FTC as recently as yesterday to figure out how to implement the [DNT] standard," said Stu Ingis, counsel to the DAA, on Wednesday. The DAA oversees the industry's Ad Choices program, which allows people to opt-out from online ad targeting through display ads that include the group's small triangular symbol. It's not entirely clear whether the FTC is confident that the DAA's self-regulatory program is enough to protect consumer privacy.

As reported by Politico earlier this month, FTC Chairman Jon Leibowitz said, "If by the end of the year or early next year, we haven't seen a real Do Not Track option for consumers, I suspect the commission will go back and think about whether we want to endorse legislation." Mr. Leibowitz is expected by beltway insiders to step down at the end of the year, and some believe his goal to finalize a DNT standard before he leaves is pressurizing the situation.

A free pass for political data?
Enter the Bipartisan Congressional Privacy Caucus. The group recently received responses to inquiries into several data firms that manage and analyze, and in some cases buy and sell, online and offline consumer data. Nine firms -- Acxiom, Epsilon, Equifax, Experian, Harte-Hanks, Intelius, Fair Isaac, Merkle, and Meredith Corp. -- submitted lengthy and often vague answers to a series of questions about their data businesses and practices.


"Many questions about how these data brokers operate have been left unanswered, particularly how they analyze personal information to categorize and rate consumers," said lawmakers in a joint statement regarding the companies' responses.

Absent from the list of data firms questioned were similar companies that deal mainly in voter file and political information that is often enhanced with consumer demographic, shopping and other data. For instance, NGP Van, the Democratic data powerhouse favored by the Obama team was not part of the inquiry. The Obama campaign and DNC spent hundreds of thousands of dollars with NGP Van this election cycle alone. The firm matches its voter data with data from TargetSmart, which offers "the richest set of consumer and interest data, allowing the most sophisticated targeting," according to the NGP Van site.

Other political data firms left out of the inquiry include Catalist, another Democratic data firm; Campaign Grid, which offers Republican data and online ad targeting; and Aristotle, a well-established non-partisan political data company. People involved with the congressional inquiry deny that political data firms were left off the list for any strategic reason.

In a press release about the data broker responses, the Privacy Caucus stated it "will push for whatever steps are necessary to make sure Americans know how this industry operates and are granted control over their own information."

Rep. Ed Markey, a Democrat from Massachusetts and Caucus co-chair, has sponsored a Do Not Track Kids Act and a mobile privacy bill.

Observers don't expect a privacy bill to be passed anytime soon; if that does happen, it may not apply to political campaigns or groups anyway. For instance, political messages are exempt from CAN-SPAM laws, and political organizations are not restricted by the Do Not Call Registry.

"Often when data laws are being proposed and put forward, the politicians exempt themselves," said Don Hinman, senior VP for data strategy at Epsilon, which gets some of its data from political advertisers but mainly is a purveyor of consumer information.

Mr. Ingis considers it exemption for political messages to be a first amendment issue. "It would be very hard for such a limitation on political messages to be restricted. . . . and I think that would have been true in the context of Do Not Call if they would have gone there," he said.


Tuesday, November 6, 2012

How Big Data Could Determine the Winner of Today's Election


Tarun Wadhwa, Forbes, November 6, 2012

If your favorite soda is Diet Dr. Pepper, the chances are that you’ll be supporting Mitt Romney.  Pepsi drinker? You’re most likely voting for Barack Obama. If you drink Mountain Dew, you probably don’t care either way.

These types of conclusions may seem simplistic and superficial, but both campaigns are betting that they will be the key to deciding who the next President of the United States is.

It’s more than what you drink, what you shop for, who your friends are, what websites you visit: all reveal clues to your political leanings. Campaigns have entered the era of “Big Data”—they target voters based on scraps of information they gather from unlikely places.

Thanks to the rise of mobile technology and social media, the number of records collected by data brokers on voter behavior has tripled—from 300 pieces in 2004 to more than 900 pieces today.

Campaigns care about your personal life
Voters used to be the ones obsessing over details of a candidate’s personal life.  Now the tables have turned. Campaigns research the personal lives of the voter.

Micro-targeting, a technique that delivers ads based on the personal traits of a voter, was once considered impossible.  But in 2004, it was recognized for helping George W. Bush defeat John Kerry.  Now it is used by almost every campaign.

Because of the intricacies of our electoral system, a relatively small group of people ends up deciding the outcome of elections.  In the 2000 Presidential campaign, hundreds of millions of dollars was spent on reaching just 7 percent of voters—fewer than 8 million people.  Even a small advantage in mobilizing potential voters in a swing state can determine the difference between a win and a loss.

In this election cycle, more than $3 billion dollars has been spent on broadcast-television advertising, which has remained the dominant form of political communication for the last fifty years.  But times are changing. Television purchases are no longer as effective as they used to be.  A study showed that 88% of voters with DVRs skip ads and that 45% use something other than live TV as their primary mode for viewing videos. These proportions are even higher in younger demographics.

The next frontier: digital behavioral advertising
Just as television advertising revolutionized the field in the 1960s, this election will likely mark digital-behavioral advertising as the next frontier in voter outreach.

As a nation, we are already divided along partisan lines.  We access different media, each with its own messaging and focus. Now we will receive different messages depending on who we are.  Zac Moffatt, digital director for Mitt Romney’s campaign, said to The New York Times that “two people in the same house could get different messages,” and that “not only would the message change, the type of content would change.”

In an article for Stanford Law Review, Daniel Kreiss, a journalism professor at University of North Carolina, Chapel Hill, explains how this can have negative long-term consequences for democratic participation.  With so much sensitive personal information in so many hands, there are risks of data breaches and unauthorized disclosure.

Citizens may hesitate to engage in political discussion on line for fear of being tagged and put into a marketing database.  And the high cost of political data and consulting activities may make it difficult for less affluent candidates to compete effectively.  Perhaps most worrying, campaigns may “redline” an electorate (by ignoring voters who won’t be sympathetic to their views because a model deems them unworthy of investment).

Political targeting – what’s next?
Sophisticated modeling and targeting will become commonplace at every step of the political process. NGOs, interest groups, and candidates for local office will be the next to adopt these methods.

United in Purpose, an evangelical Christian non-profit, is currently using such technology to assign points to voters based on whether they like NASCAR or fishing, and whether they are on anti-abortion or traditional marriage lists.  If these voters have a score of over 600 points, they are considered “serious about their faith”. They will be contacted if they have not registered to vote.

Many voters would be surprised to learn that their interactions with both campaigns are being recorded and analyzed using technology similar to what Target uses to determine whether  teenage girls are pregnant. When voters do learn what their candidates are doing, as many as 86 percent want this to stop. They regard it as an invasion of privacy.  Yet these types of activities are legally considered political speech, so there are hardly any restrictions in place.

What is most worrisome is that there is no easy way to opt out of these databases, or to limit what information is collected about you, or how it is used.  Sadly, we can’t “de-friend” or “unfollow” the politicians.



Wednesday, August 29, 2012

Facebook: The Real Presidential Swing State



David Talbot, Technology Review, July 20, 2012

The outcome of the 2012 campaign could have less to do with grand vision than with online data analytics and peer-to-peer voter targeting. 

Facebook and Internet campaign strategies grew up at the same time. In 2003 and early 2004, when Facebook was a new dorm-room plaything, Howard Dean's presidential campaign pioneered Internet fund-raising. By 2008, Facebook had crossed the 100-million-user mark and was coming to dominate online social networking; that year, Barack Obama's campaign wielded a custom social-networking site that helped win the White House (see "How Obama Really Did It"). A Facebook cofounder, Chris Hughes, helped build that site.

Now, in 2012, Facebook is central to the upcoming presidential election. Both Obama and his Republican opponent, Mitt Romney, are well aware that half or more of the electorate is on Facebook. Both campaigns' websites are entwined with Facebook pages; visitors are encouraged to log in with their Facebook accounts and then post messages supporting the candidates for their friends to see. What Facebook also gives the candidates is an arena for testing, analyzing, and distributing precisely targeted political advertising. Both campaigns can also use Facebook to urge their supporters to vote and, potentially, to lobby their undecided friends in swing states. That means this is where the 2012 election might be won or lost—even if far more money will be spent elsewhere, especially on TV ads.

Making use of social connections can lead to the ideal form of marketing: individual messages of persuasion delivered by trusted friends. You can see the president's campaign reaching for this goal with Obama 2012, an app that his supporters can use to integrate their Facebook accounts with the campaign's website. The app's avowed task is to give people a quick and easy way to access the volunteering and organizing functions that worked so well for Obama in 2008. But the permission screen that comes with the app makes clear that it has another purpose as well. When I installed the app, I noticed that it said it would grab information about my friends: their birthdates, locations, and "likes."

Facebook's policies require that such data be used in only the context of the app itself, but even so, the campaign should be able to create tools that prompt supporters to approach voting-age friends in swing states and craft personalized appeals based on what the campaign can infer about those friends' interests and views. Similar tools are coming from other quarters, too. In July NGP VAN, a company in Somerville, Massachusetts, that maintains a database on all registered U.S. voters and helps Democratic candidates access the data, released a Facebook app called Social Organizing. The app lets Democratic volunteers log in with Facebook and match their friends with voters in the database. Like the Obama app, NGP VAN's makes it possible for candidates to execute a peer-to-peer persuasion strategy using Facebook.

So don't be surprised—especially if you live in a state that is considered up for grabs, such as Ohio or Florida—if you hear from an old college friend with a political pitch based on what the campaign thinks is important to you, as suggested by your Facebook data. If you've "liked" a page blaming Obama for high gas prices, you might be reminded about his pro-drilling positions.

Don't be surprised if you hear from an old friend with a pitch based on what the campaign thinks is important to you. If you've "liked" a page blaming Obama for high gas prices, you might be reminded about his pro-drilling positions.

The Obama campaign didn't respond to requests for an interview about its plans, but Joe Trippi, the Democratic strategist who pioneered Internet fund-raising for Dean in 2003 and 2004, expects that the campaign will use sophisticated methods to determine how and when to encourage peer-to-peer appeals in the final weeks of the race. "What's most important in terms of being able to reach people is to know not only that the voter is undecided—and also what issues, what is holding them up from crossing the line—but who their friends are in the network that might be able to talk to them," he says. "And then get those friends the information that says, 'We need you to talk to your friend in Pennsylvania about these three issues that matter most to them.' This is a field organizer's dream." Certainly it is more than Trippi could have dreamed of as a $15-a-day campaign worker knocking on doors in Jones County, Iowa, for Senator Edward Kennedy in 1979, carrying shoeboxes of index cards indicating whether voters said they supported Kennedy for the next year's Democratic presidential nomination.

The Romney campaign's website also encourages supporters to log in using Facebook, but it requests permission only to view the individual user's information—not information about the user's friends "right now," says the Romney campaign's digital director, Zac Moffatt. The same is true for the Republican National Committee's Facebook app. This may change, though, because the Republicans share Trippi's view. "I think you will start to see, on our side, that app permissions will get changed," says a Republican digital strategist who spoke on condition of anonymity. "Republicans are working on apps that take advantage of all the things in the Facebook social graph."

How much information can the campaigns glean this way? Consider that the average friend count on Facebook is 190. As of early August, more than 150,000 people were using the Obama 2012 app. Multiply those numbers and you get more than 28 million people. Now, surely many friend lists overlap, and many of those people aren't even voters. And some users block the ability of apps like Obama's to gather information about them when their friends install the programs (a Consumer Reports study, however, found that only 37 percent of users touch app settings). But even if these factors make 90 percent of Obama supporters' friends useless to the campaign, the president's campaign app would still have intelligence on 2.8 million American voters who didn't necessarily take any explicit action to share it.

Persuading just a small percentage of those people could be crucial. In 2000, the contested election that put George W. Bush in office was determined largely by 537 votes in Florida, out of six million cast in that state. And in 2004 Bush beat John Kerry by fewer than 120,000 people out of 5.6 million who voted in Ohio. (Facebook is the virtual battleground within that battleground state. In 2012, just over five million account holders of voting age lived in Ohio—out of a total voting-age population of 8.8 million, according to Well & Lighthouse, a Democratic consulting firm.) Given math like that, the right peer-to-peer and message targeting strategies "could be the difference in swing states," Trippi says.

Fast and on target

In addition to any peer-to-peer strategies they might employ, the candidates are already waging online advertising campaigns that are more scientifically designed and demographically precise than the ones Obama and John McCain deployed in 2008. Political operatives can now rapidly test ad copy across multiple demographics, getting strategic insights within hours. They can even keep track of exactly which ads individual computer owners have clicked on.

These abilities were brought to bear in an ad campaign that rolled out in March of 2010, when President Obama signed the Patient Protection and Affordable Care Act—so-called Obamacare.

The midterm elections were just eight months away, and the president was concerned for a vulnerable ally, Harry Reid of Nevada, the Senate majority leader. On the health care issue alone, Reid's online strategist, Jon-David Schlough, developed 18 sets of targeted advertisements for people in different demographic groups. For example, the version geared to students pointed out that the legislation would let them keep their parents' insurance until age 26; the one for the elderly focused on what it would do to close a Medicaid benefit gap known as the doughnut hole.

Then, for each of the 18 campaigns, different versions were tested on Facebook. Schlough says the site gave him access to a wide range of demographic groups, made it possible to place small ads at low rates, and offered easy ways to experiment rapidly with different combinations of headline, image, and text. The versions that generated the most clicks would get wider distribution on multiple websites.

Eventually, the campaign could be sure that, say, an ad about being able to stay on parental insurance plans would be shown to a specific 25-year-old four times a day for two weeks. It's called nanotargeting, and "it's now a component of all campaigns," says Schlough, the founder of Well & Lighthouse. "Political types are used to large data-set analysis on things like polling data and turnout data. But the fact that so much more data is available, so much faster, is allowing us to innovate a lot quicker."

For Reid, such innovations might have been decisive. Consider that his opponent, Tea Party favorite Sharron Angle, spent about as much as Reid and was ahead in the polls in the weeks leading up to the election. In the end, Reid won by more than 5 percentage points.

David Talbot is Technology Review's chief correspondent.

This article was revised on August 15, 2012.

Thursday, August 2, 2012

Twitter's New Political Index Proves Big Data Knows What You're Thinking

Mat Honan,  Wired,  August 1, 2012
  
 
The new Twitter Political Index, or Twindex, will be updated daily throughout the election.
Twitter launched a new service on Wednesday called the Twitter Political Index, or Twindex. By applying highly tuned algorithms to Twitter's fire hose of data, the service offers a real-time look at voters' moods, and scores which presidential candidate is trending up (and who is trending down) day to day. 

Twindex is a joint effort between Twitter, Topsy, and two polling groups, the left-leaning Mellman Group and the more conservative NorthStar Opinion Research. The collective goal is to dive into Twitter's deep trove of data, and pull up insights faster than Gallup and other traditional polling companies. Expect to see Twindex results referenced in all political news and commentary as we head into the presidential election. 

Welcome to the age of big political data.

In 2008, Twitter co-founder Ev Williams walked into the then-tiny Twitter office's very small conference room, and saw something remarkable: a way for Twitter to track what people were saying about the upcoming presidential election in real-time. 

"If the dials are pointing in different directions, people are saying one thing to pollsters, and another in conversation." –Adam Sharp, Twitter's head of government news and social innovation.

‬The company had contracted Jeff Veen's Small Batch to build a site that could show how people were talking about the election. And on this day, Veen was in the office to show what he'd come up with, a subdomain on Twitter — election.twitter.com — that could track trending terms and follow message volumes about the various political candidates. 

When Veen's technology went live a few weeks later, it gave everyone a window into the vital discussions happening on Twitter. Williams was positively giddy. 

It was, Williams explained to Wired, a glimpse of what Twitter could be. This was in Twitter's salad days, literally, when the most common knock on Twitter was that it offered little more than people boasting about what they ate for lunch. "In the future, Twitter will be less personal," Williams explained. "Less about status, even. It will be more about what's happening with trends and events."

When election day rolled around in November 2008, Twitter had one of its biggest traffic days ever. Users posted some 1.8 million tweets. The mood at the company headquarters that night was ebullient. Sure, there were plenty of happy Obama supporters present, but mostly the team was excited because its servers stayed up under the load. As results came in, cheers went up as the team announced not who won the election, but tweet volumes.

Today, both the election site and the server load seem quaint. 1.8 million tweets? Twitter now does that every six minutes. And while that early election site was fun to look at and very interesting, it wasn't truly useful for drawing insight. Twitter's sample size was simply too small. But now, four years later, all of that has changed. 

Twitter is a big data company now. By its own reckoning, it has some 140 million active monthly users (outside estimates place it at 170 million) who tweet some 400 million times a day. And very, very many of them are talking politics. Now, with help from Topsy, Mellman and NorthStar, Twitter has found a way to extract voter sentiment from those conversations, measure it, and return a daily number. These results track very closely with the Gallup approval rating polling data. 

Here's how it works.
Topsy uses Twitter's high-volume fire hose of data to look at every tweet in the world, and establish a neutral baseline. Separately, it looks at all the tweets about Barack Obama and Mitt Romney, runs a sentiment analysis on them, and compares this analysis to the baseline. 
It looks at three days' worth of tweets each day, weighting the newer ones higher than then older ones. It then returns a numerical score for each candidate based on how tweets about the individual compare to all tweets as a whole. A completely neutral score would be 50. Anything above that is a net positive, while lower is a net negative. 

So, for example, if Obama has a score of 38, that would mean that tweets about him are more positive than 38 percent of all other messages on Twitter.

The project began when Twitter noticed that conversations about candidates on its own feeds accurately foreshadowed voter sentiments showing up in traditional polls. For example, during a FoxNews debate broadcast in which viewers were asked to rate candidates' responses as either "answer" or "dodge," Twitter saw a profound uptick in positive responses about Newt Gingrich. A few days days later, Gingrich was indeed moving up in the polls, but Twitter could see this shift in real-time, much, much earlier, during the debate. 

Similarly, in the run-up to the Michigan and Arizona primaries, Twitter saw Mitt Romney's follower count surge, while Rick Santorum's sputtered out. When the election results came in, they confirmed what Twitter was seeing internally: Its own social media provided an inside line on what voters were thinking.
Twitter's index tracks very closely with Gallup polling results, but it's where the results diverge that things get interesting.

So Twitter began working with polling groups and Topsy to look into the political data buried in the din of constant online chatter — they wanted a better way to measure the sentiment voters were expressing in real-time. Topsy would look at every single tweet sent in the world, every day, and create a three-day average baseline. It created an algorithm to understand which tweets skewed positive and which were negative. Together, Twitter and Topsy built a keyword engine, and via repetitive, ongoing spot checks by human observers, they found their algorithm would generate voter-accurate results 90 percent of the time. 

And that was just the beginning of a refinement process. Every time they ran the data set against human curators and found differences, they were able to improve the algorithm. What Twitter eventually built was the Twindex. It didn't rely on questions, and could be generated in real-time. And when Twitter compared the Twindex for Obama with Gallup's approval rating, the graph was remarkable. 

"We pulled this up and said 'Oh, I think we're onto something,'" says Adam Sharp, Twitter's head of government news and social innovation. "At first glance, you can readily see some parallels in the data."

As it continued to refine its methods, Twitter found that it had an increasingly strong correlation with Gallup polling data. But more interesting, obviously, is where the numbers diverge. 

"If the dials are pointing in different directions, people are saying one thing to pollsters, and another in conversation," explains Sharp. "That is where the Twitter index is providing a real service to journalists, because it's where we are saying we don't have a complete picture, and need to be asking better questions." 

Twitter attributes some of this to the differences between ongoing conversations (Twitter) and specific responses to specific questions (traditional polling). For example, in the weeks after Osama Bin Laden was killed, there was a discrepancy in what Twitter and Gallup found. 

A possible explanation of this is that voters might have answered approval rating poll questions very positively in the weeks after the raid, but in ongoing conversations with each other on Twitter, sentiment focused more on normal, day-to-day concerns about the economy. 

Twitter hopes to apply the Twindex to other issues — including, of course, analyzing sentiment around brands. But it's also hopeful that others will take its findings and run with them. 

"One of the reasons why we partnered with Topsy was because a secondary goal was to boost the ecosystem around big Twitter data," says Sharp. "To demonstrate the data was big enough, and show that it was available via existing entirely publicly available data."