Friday, December 21, 2012

The Best Open Data Releases of 2012


Picture (Device Independent Bitmap)Emily Badger, The Atlantic, December 19, 2012

Last year, Cities named ten of its favorite metro datasets of 2011 from cities across North America, illustrating the breadth of what we might learn (regarding mosquito traps! misplaced vehicles! energy consumption!) in the still relatively young field of urban open data. For this year's installment, we're going one step further. Sure, raw data is great. But useful tools, maps and data visualizations built with said data are even better.

Below, you'll find our picks for 2012's best open data releases from municipal vaults, with an emphasis on tools that can be used by anyone, not just developers and data geeks. If we missed your favorite, please add it in the comments.

1. Crime in Philadelphia. Philadelphia snuck onto our 2012 list just under the wire, publishing last week a big data set on all major crimes in the city dating back to January 1, 2006. The data is now updated daily, covering incidents of homicide, rape, robbery, assault and theft. Philadelphia now joins Chicago, which released 10 years of crime data last year. Baltimore has a similar dataset. For Philadelphians more interested in live trends than historic ones, the city is now also mapping recent crimes (see the above graphic). A smart bonus feature: when you click on an individual incident, the map gives you an opportunity to "submit a tip" to the Philadelphia Police Department.

2. Bikeshare rides in Boston. Boston’s Hubway bikeshare system published a massive file of historic trip data earlier this year, then invited riders and developers to turn the information into something useful with a data visualization challenge. This map comes from one of the winners, Ari Ofsevit, showing the average speeds across different routes between bikeshare stations.

Capital Bikeshare in Washington also publishes trip data. Nice Ride in Minneapolis did so earlier this year as well, although that release ran into privacy complications when it turned out the anonymoized data wasn’t so anonymous after all.

3. Public transit in Atlanta. Earlier this year, we wrote about the handful of large metros in the U.S. that were still not opening up their GTFS files of public transit data to anyone other than Google. Atlanta was one of the notable holdouts. In October, however, the Metropolitan Atlanta Rapid Transit Authority finally published this feed, making it possible for developers anywhere – and not just Google Maps – to produce apps, maps, trip planners or other tools with the city’s real-time transit data.

4. Blight in New Orleans. Code for America helped the city build a web tool this year mapping and tracking blighted properties all over town to help neighbors and community groups keep track of the status of abandoned and code-violating properties. With Blight Status, residents can for the first time follow blighted properties through the process of inspections, legal hearings, judgment and resolution. As a result, the city is able to convey that it’s actually working on the problem, while residents are given some reassurance of that progress.

5. Pedestrian injuries in San Francisco. San Francisco’Department of Public Health publishes a slew of data and maps (see their great Sustainable Communities Indicators website tracking everything from air quality to food access). One project in particular has had a significant impact. San Francisco’s High-Injury Corridors map tracks data on pedestrian injuries across the city. But instead of mapping individual collisions, the tool weights pedestrian deaths more heavily than other injuries and highlights injury-prone corridors instead of intersections. “The story that this map tells is that 5 percent of the city streets account for 55 percent of the injuries weighted for severity,” says Rajiv Bhatia, director of environmental health for the city’s Department of Public Health. “That was a transformative map to both the pedestrian safety stakeholders, to the police department and the transportation agency.” This is the map he’s talking about:
Picture (Device Independent Bitmap)
San Francisco High-Injury Corridors map.

“The transportation agency realized that if you overlay this map with a map of where the city of San Francisco has made their traffic calming and traffic safety investments,” Bhatia adds, “you’ll see almost no overlap.”
6. Green roofs in Chicago. The city has identified 359 vegetated roofs across town, with a total surface area of more than 5 million square feet. The city’s data portal now publishes data complete with location, dimensions and satellite imagery for all of them, as well as a map of their locations. Perhaps the most well-known green roof in town? This one, above Chicago’s City Hall:

Chicago City Hall, via the city's green roofs map.

7. Rat sightings in New York City. No, New York doesn't keep a special call-in line or data log just for rat spottings across the city. This data comes instead from the broader database of 311 service requests to city hall over the past three years. The good news? "This information is automatically updated daily." And you can use the New York data portal to map the results. But we're holding out for someone else to do this with a bunch of tiny rat icons.
8. Tsunami sirens in Honolulu. Honolulu published data on the location of dozens of warning sirens around the island, and Code for America helped to build an app on top of the data allowing local citizens to "adopt" a siren in the same way that other communities have created adopt-a-hydrant programs. In this case, instead of volunteering to dig out hydrants during snow storms, Honolulu residents can take responsibility for listening to siren tests and reporting problems ahead of any Tsunami. As this map shows, about half of the sirens (those in green) have already been adopted:
Picture (Device Independent Bitmap)
Honolulu Adopt-a-Siren.

9. Dangerous dogs in Austin. "Declared Dangerous Dogs" in Austin are court-ordered to be restrained at all times and are required to wear large tags identifying them as such. "They have attacked in the past," warns the city. "If they attack again the court could order them put to sleep." Want to know where they live? The city has now mapped them, complete with useful dog descriptions. Watch out, for instance, on Daleview Drive for "Nibbles," a female red-and-tan Golden Retriever/Chow mix.

10. Fixed speed cameras in Baltimore. Baltimore has dozens of these things around town waiting to snap photos of aggressive drivers. The penalty? A $40 mailed citation for going more than 12 miles per hour over the speed limit. But at least the city is up-front about where these cameras are located. This map and dataset on Baltimore's open-data portal identifies the intersection, coordinates and even driving direction (southbound, eastbound, etc.) for all of them. We can imagine such data might come in handy in any number of navigation apps.

Thursday, December 20, 2012

Tim Wu n NYT on Peer Economy: Apps to Regulate Apps


Tim Wu, The New York Times, December 19, 2012

NOBODY ever said that big cities make for easy living. The apps of the moment, Uber and Airbnb, try to mitigate matters by letting you book a car ride or rent someone’s apartment using your smartphone or computer. They are beloved by those contemplating scarce taxis or $500 hotel rooms. But they’re considerably less popular among city regulators, whose reactions recall Ned Ludd’s response to the automated loom.

Last month, Uber was effectively outlawed by Vancouver, British Columbia (by setting a minimum fare so high it discouraged users), and there are proposals to ban it in New York and other cities. Airbnb is already illegal in cities like San Francisco and New York, where unpredictable enforcement can result in enormous fines for its users.

To be fair, the cities think they’re protecting consumers, and the apps, particularly Airbnb, do raise some real concerns. You might be less of an Airbnb fan if your neighbor ran a hostel for the world’s misbehaving youth. As for Uber, its rates aren’t governed by city agencies, and it raised them in the aftermath of Hurricane Sandy in New York City — which critics labeled price gouging.

Yet many of the complaints are anecdotal, and too many have the odor of industry protectionism. Banning Airbnb helps hotels more than homeowners; banning Uber helps taxi companies more than passengers. Boston, in one egregious example, tried to ban Uber simply because it used GPS to measure fares, instead of an old-fashioned meter. (The ban was later reversed.)

Regulators can do their job and protect consumers against harm without being so heavy-handed. The current approach recalls Prohibition: total bans that are widely violated, with semi-random enforcement and huge fines for unlucky individuals (one Airbnb host in New York recently faced more than $40,000 in potential penalties, before the case was dismissed). It’s a clumsy approach that turns ordinary citizens into scofflaws.
This isn’t to say that the apps should have some kind of special legal immunity. It’s just that there are much smarter and more effective ways to protect consumers against potential harms. The trick is using the same techniques of real-time information access that the apps employ.

Here’s how it might work for Airbnb. Cities could require the company to provide co-op boards and landlords, at their request, with an app that listed any advertised rentals at their addresses. That way, instead of one uniform rule for the city, landlords and boards could decide for themselves how they wanted to handle Airbnb rentals. They could take a laissez-faire approach, or ban tenants from advertising on Airbnb entirely — in which case the app would help them detect and fine any violators. Others might choose to include Airbnb guidelines in rental contracts, take a part of the proceeds or limit usage to certain levels. Over all, getting better information about how Airbnb is actually being used would yield better solutions for all.

Another info-access app could help cities regulate Uber’s rates, if they really were afraid of price gouging. Regulators could simply require Uber to disclose the prices it charged and where its cars were going. If cities wanted to ban rate hikes during emergencies, they could watch to see that the law was obeyed.

This kind of precise, data-driven regulation could protect consumers while also protecting their right to pay for a valuable service. No one can deny that these apps are responding to real demands and helping cities become easier to live in and visit. There are also basic rights at stake. Using Uber to hire a car is basically an exercise in freedom to contract. Taking away the right to rent out your room diminishes the value of your property and drives up costs for visitors.

Change isn’t always pretty, but a healthy city is one where old systems — even the hallowed taxi medallion — stand to be challenged by the winds of creative destruction. Uber and Airbnb are just the first examples of a wave of services trying to match willing buyers and sellers in unexpected ways. That’s why it is so important that regulators get this right, lest they discourage those who are trying to follow their lead. The challenge for regulators is to simultaneously allow change while protecting us from the worst effects of it. It is, in short, a time to think carefully, rather than banning first and asking questions later.

Tim Wu, a professor of law at Columbia, is the author of “The Master Switch: The Rise and Fall of Information Empires.”

What's Next for Obama For America's Data and Technology?


Jim Pugh and Nathan Woodhull, The Huffington Post, December 18, 2012

If you've been following post-election news, you've no doubt heard about Barack Obama's "big data" advantage. The story has been the unprecedented investment in mining personal information for all kinds of wacky stuff. According to various articles, the campaign used data for everything from inviting donors to dinner with Sarah Jessica Parker to identifying potential voters through their visits to porn sites.

It's true that data was a game changer for the Obama campaign, but the reason is much less salacious than many reporters would have you believe. The campaign built an integrated database, which combined supporters' online activity, actions in the field, and public voting records into a single unified view of every American. On top of that data foundation, a large team of developers built community organizing software that empowered volunteers to become more deeply involved with the President's grassroots field operation.

The result was unprecedented efficiency in volunteer engagement and voter outreach. Supporters who signed up to help online received a personal call from their neighborhood leader the next day. Volunteers called only the most statistically persuadable potential supporters. Anyone who connected their Facebook account to the campaign was encouraged to send voting turnout messages to the specific friends who needed the extra push most. All of this led to more volunteers, more supporters, and increased turnout -- which all meant more Obama votes on Election Day.

But Election Day shouldn't be the end for these systems. We need to keep moving forward, continue the investment, and make them available to the whole movement.

Don't Abandon the Technology

Now that the election is over, funding will be tight. The donations that poured in during the election have dried up and hard choices need to be made.

One option is to mothball this infrastructure and plan to spin it up again for the next presidential election. This approach is attractive from a financial perspective, as no additional resources would be needed in the off cycle to make it happen.

But this would be a mistake. The campaign's advantage this cycle wasn't just bits and bytes, but the institutional experience and knowledge that had been building since the President's primary campaign in 2007. Instead of shelving all these systems, we should make a continued investment in maintaining and improving them.

Republicans are lagging behind Democrats right now on the data and technology front, but after the shellacking they experienced in the last election, there's little doubt that they'll be pouring in resources to catch up. Without sustained investment from our party, our advantage may be erased in the coming years. Not only would Republicans be moving ahead, but without staff to maintain institutional knowledge and adapt the systems to changing technology, Democrats could actually start the next election cycle behind where we're at right now. We need to keep moving forward if we want to keep our advantage on this front.

The technology industry never stops moving forward. Neither should we.

Make Tools Available to All Progressives

An investment in building on these systems offers another possibility as well: providing access to the rest of the progressive movement. Right now, only presidential campaigns have the resources to build systems of this sophistication. The data and technology infrastructure from the Obama campaign cost millions of dollars to build, and even the most well-funded senate campaigns couldn't afford anything close to that.

But with some additional work, the data and tech infrastructure from the Obama campaign could be adapted to offer the same functionality to other progressive candidates and groups, giving them the opportunity to use these systems with their own supporters and volunteers. For smaller campaigns that would have no chance of creating these systems on their own, this could be a game-changing step forward. And beyond the benefit to the Democratic Party and progressive movement, it could provide a path to fund the continued investment, via paid licensing from these outside campaigns and organizations.

A Sustained Tech Commitment for 21st-Century Politics

The roller-coaster of scaling up and scaling down that comes with elections has always been the standard for political organizations. If we want to stay competitive on the data and technology front, this approach needs to be changed. A 21st-century political movement must have a serious on-going commitment to staying at the forefront of technological advancement. With the infrastructure coming out of the Obama campaign, we've got a huge lead in this area.

Sadly, that is not what has happened so far. Since Election Day, the Democratic National Committee has laid off an unprecedented number of technology staff members, some of whom had been at the party for over ten years. The Obama Campaign's technology team is scattering to the winds and returning to industry. The window of opportunity to stay ahead of the technological curve is closing -- our party needs to change course now or risk being left behind.

Tim O'Reilly: Three Lessons for the Industrial Internet



Simplicity, generativity and robustness shaped the Internet. Tim O'Reilly explains how they can also define the industrial Internet.

Mac Slocum, O'Reilly Radar, December 19, 2012

The map of the industrial Internet is still being drawn, which means the decisions we’re making about it now will determine the extent to which it shapes our world.

With that as a backdrop, Tim O’Reilly (@timoreilly) used his presentation at the recent Minds + Machines event to urge the industrial Internet’s architects to apply three key lessons from the Internet’s evolution. These three characteristics gave the Internet its ability to be open, to scale and to adapt — and if these same attributes are applied to the industrial Internet, O’Reilly believes this growing domain has the ability to “change who we are.”

Full video and slides from O’Reilly’s talk are embedded at the end of this piece. You’ll find a handful of insights from the presentation outlined below.

Lesson 1: Simplicity
“Standardize as little as possible, but as much as is needed so the system is able to evolve,” O’Reilly said.

To illustrate this point, O’Reilly drew a line between the simplicity and openness of TCP/IP, the creation and growth of the World Wide Web, and the emergence of Google.

“The Internet is fundamentally permission-less,” O’Reilly said. “Those of us who were early pioneers on the web, all we had to do was download the software and start playing. That’s how the web grew organically. So much more came from that.”

A nice side-effect of this model is that you can easily determine its success. “A new platform can be said to succeed when your customers and partners build new features before you do,” O’Reilly noted.

“The IP protocol did the smallest, necessary thing: it specified the format of the data that would be exchanged between machines. Everything else could vary, from the transport protocols and transport medium all the way to the kinds of applications and services that were exchanging that data. Jonathan Zittrain refers to this as the ‘hourglass architecture’ of the Internet.”
— Tim O’Reilly, “Lessons for the Industrial Internet,” slide 6
(The “simplicity” segment begins at the 1:43 mark in the accompanying video.)

Lesson 2: Generativity
“Create an architecture of participation that leads to unexpected innovations and discoveries, and builds a new ecosystem of companies that add value to the network,” O’Reilly explained.

O’Reilly pointed to two examples relevant to this lesson:
1.      Google Maps — Online mapping services were already common when Google launched its Maps service in 2005. So how did Google push to the front of the line? When hackers started mashing up Maps data with external sources, Google embraced those efforts and launched APIs. The result is a mapping platform that’s achieved ubiquity through accessibility. It’s unlikely anyone within Google could have anticipated the innovations that would come once the doors were thrown open.

2.      Apple’s App Store / the rise of Android — The first iPhone didn’t launch with an App Store. Third-party development at that point was limited to web apps. But Apple plotted a new course when it saw jailbreaking grow, and the company now oversees a marketplace with more than 700,000 applications. However, O’Reilly pointed to the rise of Android as evidence that Apple didn’t completely embrace a participatory architecture. “Apple didn’t learn the lesson well enough,” he said.

O’Reilly also noted that the industrial Internet’s considerable upside makes the need for participatory architecture vital. “I think this industrial Internet idea is so powerful and so right, that everybody is going to want to get on board. It’s going to be really important to figure out how you create an open systems approach to this. That doesn’t mean you can’t create enormous value for yourselves, but it’s super important to think about that aspect of it.”

“When Paul Rademacher reverse-engineered the format of Google’s new mapping app to create the first map mashup, HousingMaps.com, Google could have branded him a ‘hacker’ and tried to shut him down. Instead, they responded by opening up free APIs for developers. Other, more closed platforms were left in the dust, and Google Maps became the preferred mapping platform for the web.”
— Tim O’Reilly, “Lessons for the Industrial Internet,” slide 11

(This “generativity” segment begins at the 3:01 mark in the accompanying video.)

Lesson 3: Robustness
Build “the ability to tolerate failure and degrade gracefully rather than catastrophically,” O’Reilly said.

“Graceful failure” may sound like an excuse a parent uses to soothe a child’s bruised ego, but it’s actually a fundamental building block of the Internet.

As an example, O’Reilly said that one of the key innovations Tim Berners-Lee constructed when he developed the World Wide Web was the ability for anyone to create a hypertext link that didn’t resolve. A user would bump up against a 404 message if a link hit a dead end, yet everything would still function and the user could flow around the obstacle without the entire system crumbling.

Now, you may think an innocuous 404 failure is a far cry from the failure of a massive chunk of industrial machinery. That’s not necessarily the case. O’Reilly told the story of a Boeing engineer who addressed catastrophic metal fatigue in airplanes through a form of graceful failure. “The right answer wasn’t to eliminate all the cracks,” O’Reilly said. “You had to figure out how to live with them.”

Bottom line: Whether we’re talking about hypertext or airliner materials, an embrace of robustness and graceful failure expands a domain’s possibilities. “One of the big lessons from the Internet is if you don’t know how to fail, you’ll never be able to scale,” O’Reilly said.

“While it may seem that this philosophy of the Internet is inappropriate for the highly engineered systems of the industrial Internet, I’ll remind you of the failure of de Havilland’s Comet in 1954 and the rise of Boeing as the dominant provider of commercial aircraft. Over the course of three years, three Comets fell out of the sky for initially unexplained reasons. It eventually became clear that the problem was metal fatigue. De Havilland tried to eliminate all cracks; Boeing learned to live with them.”
— Tim O’Reilly, “Lessons for the Industrial Internet,” slide 18

(This “robustness” segment begins at the 6:40 mark in the accompanying video.)

Full video: “Minds + Machines 2012: ‘Closing the Loop’ – Lessons of Data for the Industrial Internet”
Additional insights and discussion are contained in this video. The first 13 minutes features O’Reilly’s presentation, which is then followed by a panel discussion between O’Reilly, Paul Maritz of EMC, DJ Patil of Greylock Partners, Hilary Mason of bitly, and Matt Reilly of Accenture Management Consulting.

Slides: “Lessons for the Industrial Internet”
Lessons for the Industrial Internet (pdf with notes) from Tim O’Reilly

This is a post in our industrial Internet series, an ongoing exploration of big machines and big data. The series is produced as part of a collaboration between O’Reilly and GE.

Tuesday, December 18, 2012

Crowds Are Not People, My Friend

December 18, 2012

Crowds Are Not People, My Friend

By MAGGIE KOERTH-BAKER

http://www.nytimes.com/2012/12/23/magazine/crowds-are-not-people-my-friend.html?pagewanted=1&_r=0&pagewanted=print#h[]

We have all been to a church or a concert without merging with the rest of the audience into some sort of hive mind. Likewise, we all know from experience that Internet message boards aren’t Borg-like mind-melding machines that multiply users’ brain power. So why, then, do we insist on treating crowds, real or virtual, like sentient beings? We’ve long believed that physical crowds are emotional, irrational and prone to violence. Over the last decade, we’ve come to think of virtual crowds as sources of wisdom that can’t be found in individuals. Both these ideas treat crowds as entities, rather than groups of people — an idea that has its origins in 19th-century sociology, which, according to scientists studying crowd behavior today, is deeply flawed.

Clark McPhail, an emeritus professor of sociology at the University of Illinois at Urbana-Champaign, and one of the first people to actually document and study how people behave when they come together in large gatherings, doesn’t even like to use the word “crowd.” It’s too weighed down by inaccurate stereotypes. For years, sociologists thought a crowd behaved like a herd of animals: at some point, it reaches a critical mass and the will of the crowd overrides individual intelligence and individual decision making.

But that’s not what happens. Groups of people are still made up of people. They can behave in helpful and intelligent ways, or they can behave in dumb and dangerous ways. But in either case, a crowd’s behavior depends on what individuals are thinking and how they interact with one another — not some overpowering collective consciousness. “Crowds don’t have central nervous systems,” McPhail said. And that is true whether the crowds you’re talking about are physical or virtual.

Gustave Le Bon was one of the first people to write about crowds as entities separate from the people in them. His 1895 book, “The Crowd: A Study of the Popular Mind,” shaped academic discussions of human gatherings for half a century and encouraged 20th-century fascist dictators, including Benito Mussolini, to treat crowds as emotional organisms — something to be manipulated and controlled. (Perhaps a Le Bonian understanding of crowds makes us feel more comfortable about the atrocities of the 20th century.) But “The Crowd” was more a work of philosophy than of science, McPhail told me. Le Bon’s ideas were based on armchair analysis of past events, not on carefully documented studies of crowds in action. In the 1960s, sociologists began to study protests and public gatherings, and they realized that the things they believed about crowd behavior didn’t align with what took place in the real world.

Take, for example, the effect fear has on a crowd. Common sense — which is to say, the Le Bon-influenced myths you’ve been steeped in since high school — would suggest that a panicked crowd loses all semblance of rationality, charging madly and trampling anyone who doesn’t keep up. But despite individual instances that come to mind — the tragic Who concert in Cincinnati in 1979, say — studies since the early 1980s have shown that groups of people generally don’t move as a collective front, and they aren’t all crazed with terror, even in terrifying situations. On Sept. 11, for instance, large numbers of people organized themselves into a quick, careful and efficient evacuation of the World Trade Center towers. They knew one another, so they discussed plans, they made decisions, they behaved rationally and independently.

In 1997, McPhail and a team of researchers documented the behavior of individuals and small groups that made up the crowd of 500,000 at a Promise Keepers rally in Washington. The Promise Keepers are an evangelical Christian men’s organization, and as such, the event was highly structured, with performers and preachers explicitly asking the audience to do certain things: pray, sing, etc. But at no point during the entire rally were more than 80 percent of the participants doing the same thing simultaneously. Most of the time when the audience acted in unison, less than 55 percent participated.

Scientists who focus on virtual groups see similar patterns. Conor Mayo-Wilson is a researcher of mathematical philosophy at Carnegie Mellon University who studies how people learn and solve problems by sharing information — how scientists in a given field reach a consensus, for instance, or even how African farmers choose which crops to plant. These groups, however different, take advantage of a diverse range of experiences and knowledge, so it’s reasonable to think that collective intelligence might come to a more accurate conclusion than any one individual. But research done by Mayo-Wilson and others shows that this isn’t exactly the case.

For instance, we know today that stomach ulcers are caused by bacteria, Mayo-Wilson said. Scientists were making connections between bacteria and stomach ulcers as early as 1889. But in 1954, Edward Palmer published a paper that claimed to find no bacteria whatsoever in 1,000 human stomachs. Palmer’s study was flawed, but knowledge of that paper spread faster and more widely than the earlier work. Soon everybody knew that bacteria couldn’t live in the stomach, but what they knew was completely false. Linking people into virtual groups enables the sharing of knowledge, but when that information isn’t accurate, it can lead the group consensus astray. “When information comes from a common source, that can cause problems with individual decision making, because it can eradicate minority viewpoints,” Mayo-Wilson said.

This assumption that crowds have some non-fragmented consciousness leads us to the false dichotomy we draw between physical and virtual crowds: one is dumb, the other is smart. But in both cases, we’re placing too much emphasis on the crowd as distinct from the people involved in it. “The thing we’re trying to emphasize is that it’s the individuals and how they interact with one another,” Mayo-Wilson told me. “How those people receive information can influence whether or not they make a good decision.”

This has real-world consequences. When police officers show up at a protest or political rally, they tend to think of the crowd in Le Bonian terms, McPhail told me. That can be dangerous. If the police assume the crowd is acting as one, it becomes easier for a handful of people to provoke a violent reaction from law enforcement — and vice versa. McPhail uses what he has learned from 40 years of studying groups of people to advise law enforcement on better, safer ways to deal with crowds. He told me that 150 years of records from Europe and the United States show violence happens at less than 15 percent of political gatherings. So he instructs officers to never respond categorically to a crowd. If one person is breaking the law, address that person in an unobtrusive way. “If you are blatant and violent, you affect people who weren’t doing anything, and that . . . turns them against you,” he said.

At the same time, knowing that virtual crowds are merely human helps us better predict when one is likely to be smart and when it’s likely to be stupid. Reddit can help someone understand a medical diagnosis just as easily as it can foster a men’s rights movement. Scientists, working as a virtual group, are capable of sharing diverse research to reach a consensus on climate change, but they’re also capable of passing down the received wisdom that crowds have minds. The group itself isn’t what matters. What matters is who they are, what they know and how they interact.

Maggie Koerth-Baker is science editor at BoingBoing.net and author of “Before the Lights Go Out,” on the future of energy production and consumption.

How Main Street will fight big business with 'big data'

How Main Street will fight big business with ‘big data’

By Christina Farr | VentureBeat.com, Published: December 17

http://www.washingtonpost.com/business/technology/how-main-street-will-fight-big-business-with-big-data/2012/12/17/cabe2936-461f-11e2-8c8f-fbebf7ccab4e_print.html

When we consider the “big data” trend, it’s most often in terms of how large corporations like Macy’s and Starbucks will use vast volumes of consumer data to grow their business.

But how about the little guy — the coffee shop on Main Street that is struggling to compete with Starbucks to keep its doors open?

Intuit, the maker of financial management products for small businesses, currently has hundreds of engineers tasked with bringing the benefits of big data to Main Street. The company envisions a future where small businesses will be armed with data to help them make strategic business decisions and drive consumer sales. It’s turning its database of small business finances into tools to help its customers save time and money.

In an intimate meeting at the company’s San Francisco office, CEO Brad Smith told reporters that Intuit is stressing three things: privacy, innovation, and the accessibility of data. For Smith, that big data can be used to benefit the little guy by 2020 is a no brainer, given the sheer quantity of relevant data in Intuit’s grasp (over the next decade, Emergent Research posits that it will increase more than 40-fold.) With its 60 million global customers, Intuit is currently sitting on a “treasure trove of data,” in Smith’s opinion. He is fond of saying that 20 percent of U.S. GDP flows through Quickbooks, Intuit’s small business accounting software.

One of the most prominent examples is Intuit’s Small Business Index, which pulls together sales, profit and employment data from a statistical sample of 70,000 small businesses (with fewer than 20 employees) that use Quickbooks and its online payroll software. Small business owners can use this data to determine when it’s the right time to cut back on expenses, or hire some additional help. This effort was spearheaded by Nora Denzel, senior vice president of Big Data, Social Design and Marketing, and her 100 person-strong team of researchers.

However, consumers are concerned that their personal information is being used by businesses for targeted advertising.

Surrounded by members of the media, Smith is well-prepared to deflect criticism on the topic of consumer protections and privacy — chief privacy officer Barb Lawler sits front and center. Smith frequently repeats the company mantra, ”Our view is that this is not our data, this is our company’s data.” He said the company has been working with “every agency there is” to ensure that your data will be secure, and that it won’t be used without permission.

Smith hopes that most consumers and small business owners will voluntarily give up their data — “[we] can improve your life with your permission,” he stresses.

One of his favorite examples is how employees noted that two-thirds of Intuit’s customers using the Quickbooks product, a small business accounting software, were recently denied a loan. ”Banks were afraid to take the risk and we are sitting in the middle of the Quickbooks data [and] can qualify that individual,” Smith explained. The company analyzed whether applicants were paying their bills on time, and the banks were notified about those that were least likely to default on a loan. Smith claims that $10 million in loans were subsequently provided to Intuit’s customers.

I asked whether the company will consider social data  – the number of “likes” on a brand’s Facebook page — as an indication of an individual’s ability to repay a loan on time. Social finance startups like LendUp are beginning to use this information to get a better picture of the borrower and their likelihood of repaying a loan on time. Smith said the viability of leveraging data from social networking sites is “yet to be determined.”

Inspired by Google, Intuit’s next-door neighbor at its Mountain View, Calif. headquarters, it has been company policy for years to let engineers devote 10 percent of their time to innovative projects. According to Smith, this yielded Intuit $100 million in revenue in 2011.

The top execs are keen advocates of Eric Ries’s ”Lean Startup” method — the company recently invited the author and serial entrepreneur in to its offices to help them internalize a culture of experimentation, and rapid product cycles. Ries worked with the company’s most talented engineers, tasked with developing new technologies to arm small businesses.

One of the top engineers is Michael J. Radwin, a data nerd who has introduced text analytics, recommendation services, and data-driven algorithms in his two-year tenure at Intuit.

Shortly after the meeting, Radwin pulls me aside to demo one of the projects he’s most proud of — a local business recommendations service intended to provide more detail than Yelp. The idea was conceived by an engineer, and a team formed to bring it to fruition. It’s in beta, and is primarily used today to give Intuit employees, analysts, and the press an idea of how data might be used in the future.

Radwin views Yelp data as highly deceptive, and skewed to favor large franchises with digitally-savvy customers. By integrating information from a variety of external sources including Mint.com (the personal finance company Intuit acquired a few years ago), his team created a detailed map of businesses in any area. Customers can perform a simple search to uncover the hair salons or coffee shops with the highest customer retention rates and most competitive prices — more often than not, its the family-owned businesses that emerge on top.

If Radwin has his way, this prototype will likely be under development in the coming years. It’s just one example of how data can inform consumers, and be used a weapon for the little guy on Main Street.

Monday, December 17, 2012

Nate Silver, FiveThirtyEight Prove Predictive Analytics Getting Real

·       December 16, 2012, 12:03 PM ET

Nate Silver, FiveThirtyEight Prove Predictive Analytics Getting Real

http://blogs.wsj.com/cio/2012/12/16/nate-silver-fivethirtyeight-prove-predictive-analytics-getting-real/?KEYWORDS=big+data

Picture (Device Independent Bitmap)       
Irving Wladawsky-Berger

Guest Contributor

Information analytics is emerging as one of the most powerful tools to help companies make better decisions, plans and forecasts.  But, like any such IT-based tools, it’s most effective in the hands of people who understand its pitfalls and limitations as well as its potential. It’s one of the key areas where CIOs can help their companies leverage advances in technology to improve their overall competitiveness.

Our recently concluded 2012 presidential campaign is an interesting  case study in the effective use of information. Up to the very end, many pollsters and political pundits were saying that the election was too close to call. But, whenever friends and colleagues asked my opinion, I invariably told them to go look at FiveThirtyEight.com, the political polling website and blog created by Nate Silver.

Silver launched the FiveThirtyEight in March of 2008. He gained national attention when he correctly predicted the results of the 2008 Democratic Party presidential primaries. His final forecasts for the 2008 presidential elections predicted the winner in 49 of the 50 states, as well as the winner of every race in the Senate. The FiveThirtyEight has been affiliated with the NY Times since August of 2010.

Political forecasting attracts a variety of people. Many pollsters seem to still be using methodologies that worked well a decade or two ago, but look woefully behind the times in the emerging world of big data and advanced analytics. Most political pundits seem to inhabit a kind-of magical realism world of their own: if you say what you want to be true often enough and loud enough, it will eventually come to pass.

Silver, on the another hand, views information-based predictions, including political forecasting, as a scientific discipline. You use all available information; you analyze and extract insights out of all that information using sophisticated models and algorithms; you apply human judgment to make predictions based on those insights; and you keep evaluating and adjusting your models and predictions based on how well they perform in the real world.

But equally important, you must be aware not only of the possibilities but also the limitations and pitfalls inherent in any such predictions. In political elections, as is the case when dealing with any highly complex, fast changing, chaotic and unpredictable system, predictions can only be expressed in terms of probability distributions, and they are generally quite volatile as new information and unanticipated events are factored in. Moreover, predictions are based on models reflecting your views of how the future is likely to evolve. Different models applied to the same data can lead to widely different predictions.

Silver started tracking the 2012 election in June. In his initial forecast he estimated that Obama would win the election with 291.3 electoral votes, compared to 246.7 for Mitt Romney, which gave the President a 61.8% of reelection. As he explained in his blog, there are wide error margins inherent in such early forecasts, so they should only be taken as a pretty rough guide, not unlike forecasting the path of a hurricane five to seven days later. You can give an estimate, but the cone of uncertainty is very wide.

As it turned out, Silver correctly predicted the winner in all 50 states, including all nine highly contested swing states. He also correctly predicted the winner in 31 of the 33 Senate races.

How does he do it? The answer, I believe, is that Nate Silver brings a scientific point of view to the Wild West world of forecasting political elections. This is evident throughout his recently published book, The Signal and the Noise: Why Most Predictions Fail but Some Don’t. In the book, Silver explains not only his own particular approach to information-based predictions, but examines the growing field of predictions and why so many fail in spite of–or perhaps because of–the vast quantities of information we now have available. He writes:

“The exponential growth in information is sometime seen as a cure-all, as computers were in the 1970s. Chris Anderson, the editor of Wired magazine, wrote in 2008 that the sheer volume of data would obviate the need for theory, and even the scientific method. This is an emphatically pro-science and pro-technology book, and I think of it as a very optimistic one. But it argues that these views are badly mistaken. The numbers have no way of speaking for themselves. We speak for them. We imbue them with meaning. . . Data-driven predictions can succeed – but they can fail. It is when we deny our role in the process that the odds of failure rise. Before we demand more of our data, we need to demand more of ourselves.”

Silver learned his craft in the new field of sabermetrics–the use of statistics in baseball to project a player’s performance and career. Sabermetrics was popularized by Michael Lewis in Moneyball, his bestseller book–later turned into a film–about Billy Beane, the Oakland Athletics general manager who used such statistical techniques to make his small-market team highly competitive against teams with much larger budgets.

Sabermetrics combined Silver’s love of statistics and baseball. He developed a sophisticated statistical system,PECOTA, for forecasting the future performance of baseball players by comparing the characteristics of the player being evaluated to the characteristics of all past and present players, and then looking at the actual career paths of those players with the most similar characteristics. This enabled the system to assign a probability to different career trajectories for the player being evaluated.

Sabermetrics is now widely used by every team in baseball. But, Silver asks in his book, can statistics alone tell you everything you want to know about a player? When Moneyball first came out, many viewed it as a story about the conflict between the traditional approach of the scouts versus the new approaches being introduced by the statheads. Is there still a valuable role for the scouts, the professional talent evaluators who personally travel to learn about the players first hand by actually watching them play and meeting them in person?

Billy Bean attributes the success of the Oakland A’s not just to their statistical aptitude but to their careful scouting of amateur players. In fact, their scouting budget is now much higher than it has ever been as a way of complementing their statistical analysis. Scouting is particularly valuable when evaluating young amateur players, while they are still learning the game. This is when a good scout can spot certain intangibles that may make a player stand out, such as their mental makeup and overall attitude toward the game.

“The key to making a good forecast,” writes Silver, “is not in limiting yourself to quantitative information. Rather, it’s having a good process for weighing the information appropriately. This is the essence of Beane’s philosophy, collect as much information as possible, but then be as rigorous and disciplined as possible when analyzing it.”

The Signal and the Noise describes a number of areas where information-based predictions have been successful, including baseball, political elections and the forecasting of hurricanes. Good hurricane forecasting has life-and-death consequences. Twenty five years ago, for example, you could only predict a hurricane’s landfall 72 hours in advance to within 350 miles, an area way too large to evacuate if necessary. You can now come within a much more manageable 100 miles, and the forecasts are getting better every year. As we just saw with Hurricane Sandy, such accurate predictions truly help save lives.

“But,” writes Silver, “these cases of progress in forecasting must be weighted against a series of failures.”

As one reviewer observed, given his accomplishments and notoriety, it’s slightly heartbreaking that instead of a “geek-conquers-world” book about “his rise to statistical godliness”, Silver wrote a book examining the state of predictions in a variety of fields, with a handful of successes amidst a large number of failures. “As science, this investigation is wholly satisfying. As a literary proposition, it’s a bit disappointing. It’s always more gripping to read about how we might achieve the improbable than why we can’t.”

But, good scientists and engineers are fatalistic by nature, always worrying about what can go wrong with their predictions and designs. We have gotten pretty good at predictions and designs in the more mature fields like physics and civil engineering, but we are just learning how to do so in sociotechnical systems, that is, systems involving people and organizations. Such systems generally exhibit a level of complexity that is often beyond our ability to understand and control. Not only do we have to have to deal with very tough mathematical problems, but with the even more complex issues involved in human and organizational behaviors. Our instincts will often lead us to see patterns where there might be none. We have to work very hard to become aware of and try to overcome our biases.

In business, for example, we are surprised when our strategies did not work in spite of being based on a careful analysis of market conditions.  Often, the reason is that we use our analysis to confirm our biases and prejudices, – what we want to be true, – even without realizing we are doing so.

“Prediction is difficult for us for the same reason that it is so important: it is where objective and subjective reality intersect,” writes Nate Silver in the concluding paragraphs of his book. “Distinguishing the signal from the noise requires both scientific knowledge and self knowledge: the serenity to accept the things we cannot predict, the courage to predict the things we can, and the wisdom to know the difference.”

Irving Wladawsky-Berger is a former vice-president of technical strategy and innovation at IBM. He is a strategic advisor to Citigroup and is a regular contributor to CIO Journal.

IBM Looks Ahead to a Sensor Revolution and Cognitive Computers

IBM Looks Ahead to a Sensor Revolution and Cognitive Computers

By STEVE LOHR

http://bits.blogs.nytimes.com/2012/12/17/ibm-looks-ahead-to-a-sensor-revolution-and-cognitive-computers/

The year-end prediction lists from technology companies and research firms are — let’s be honest — in good part thinly disguised marketing pitches: These are the big trends for next year, and — surprise — our products are tailored-made to help you turn those trends into moneymakers.

But I.B.M. has a different spin on this year-end ritual. It taps its top researchers worldwide to come up with a list of five technologies that are likely to advance remarkably over the next five years. The company calls the list Five in Five, with the latest released on Monday. And this year’s nominees are innovations in computing sensors for touch, sight, hearing, taste and smell.

Touch technologies may mean that tomorrow’s smartphones and tablets will be gateways to a tactile world. Haptics feedback techniques, infrared and pressure-sensitive technologies, I.B.M. researchers predict, will enable a user to brush a finger over the screen and feel the simulated touch of a fabric, its texture and weave. The feel of objects can be translated into unique vibration patterns, as if the tactile version of fingerprints or voice patterns. The resulting vibration patterns will simulate a different feel, for example, of fabrics like wool, cotton or silk.

The coming sensor innovations, said Bernard Meyerson, an I.B.M. scientist and vice president of innovation, are vital ingredients in what is called cognitive computing. The idea is that in the future computers will be increasingly able to sense, adapt and learn, in their way.

That vision, of course, has been around for a long time — a pursuit of artificial intelligence researchers for decades. But there seem to be two reasons that cognitive computing is something I.B.M., and others, are taking seriously these days. The first is that the vision is becoming increasingly possible to achieve, though formidable obstacles remain. I wrote an article in the Science section last year on I.B.M.’s cognitive computing project.

The other reason is a looming necessity. When I asked Dr. Meyerson why the five-year prediction exercise was a worthwhile use of researchers’ time, he replied that it helped focus thinking. Actually, his initial reply was a techie epigram. “In a nutshell,” he said, “seven nanometers.”

Dr. Meyerson, who has a Ph.D. in solid-state physics, was talking about the physical limits on the width of semiconductor circuits, when they can’t be shrunk any further. (The width of a human hair is roughly 80,000 nanometers.) Today, the most advanced chips have circuits 22 nanometers in width. Next comes 14 nanometers, then 10 and then 7, Dr. Meyerson said.

“We have three more cycles, and then the biggest knobs for improving performance in silicon are gone.” he said. “You have to change the architecture, use a different approach.”

“With a cognitive computer, you train it rather than program it,” Dr. Meyerson said.

The cognitive path, if successful, would raise a machine’s level of recognition of the world. Today, computers mimic human intelligence with brute force, collecting a vast amount of data and then sifting for statistical patterns that identify specific words, images, biological or chemical compounds.

But a cognitive computer, Dr. Meyerson said, would “not have to go for all the fine detail, but go up and see the interesting thing. This is about moving computing way, way up from where it is today.”

He provides a sensor-based example. The computer is presented with a white powder in two piles. One is salt and the other is sugar. “It could taste the difference, without having to do a detailed chemical analysis,” Dr. Meyerson explained.

Cognitive computers, by learning some tricks from the way human brains compute, could in theory deliver big energy savings. I.B.M.’s “Jeopardy”-winning computer is a clever machine, defeating its human rivals last year. But the computer, called Watson, runs on 85,000 watts of electricity. The human brain hums along on 20 watts.

“We need to make the machines much, much more efficient,” Dr. Meyerson said.

Elections and Big Data

Elections and Big Data

How President Obama's campaign used big data to rally individual voters, part 1 and 2.

By Sasha Issenberg on December 17, 2012

Sasha Issenberg is the author of The Victory Lab: The Secret Science of Winning Campaigns.

http://www.technologyreview.com/featuredstory/508851/how-obama-used-big-data-to-rally-voters-part-2/

Two years after Barack Obama's election as president, Democrats suffered their worst defeat in decades. The congressional majorities that had given Obama his legislative successes, reforming the health-insurance and financial markets, were swept away in the midterm elections; control of the House flipped and the Democrats' lead in the Senate shrank to an ungovernably slim margin. Pundits struggled to explain the rise of the Tea Party. Voters' disappointment with the Obama agenda was evident as independents broke right and Democrats stayed home. In 2010, the Democratic National Committee failed its first test of the Obama era: it had not kept the Obama coalition together.

But for Democrats, there was bleak consolation in all this: Dan Wagner had seen it coming. When Wagner was hired as the DNC's targeting director, in January of 2009, he became responsible for collecting voter information and analyzing it to help the committee approach individual voters by direct mail and phone. But he appreciated that the raw material he was feeding into his statistical models amounted to a series of surveys on voters' attitudes and preferences. He asked the DNC's technology department to develop software that could turn that information into tables, and he called the result Survey Manager.

That fall, when a special election was held to fill an open congressional seat in upstate New York, Wagner successfully predicted the final margin within 150 votes—well before Election Day. Months later, pollsters projected that Martha Coakley was certain to win another special election, to fill the Massachusetts Senate seat left empty by the death of Ted Kennedy. But Wagner's Survey Manager correctly predicted that the Republican Scott Brown was likely to prevail in the strongly Democratic state. "It's one thing to be right when you're going to win," says Jeremy Bird, who served as national deputy director of Organizing for America, the Obama campaign in abeyance, housed at the DNC. "It's another thing to be right when you're going to lose."

It is yet another thing to be right five months before you're going to lose. As the 2010 midterms approached, Wagner built statistical models for selected Senate races and 74 congressional districts. Starting in June, he began predicting the elections' outcomes, forecasting the margins of victory with what turned out to be improbable accuracy. But he hadn't gotten there with traditional polls. He had counted votes one by one. His first clue that the party was in trouble came from thousands of individual survey calls matched to rich statistical profiles in the DNC's databases. Core Democratic voters were telling the DNC's callers that they were much less likely to vote than statistical probability suggested. Wagner could also calculate how much the Democrats' mobilization programs would do to increase turnout among supporters, and in most races he knew it wouldn't be enough to cover the gap revealing itself in Survey Manager's tables.

His congressional predictions were off by an average of only 2.5 percent. "That was a proof point for a lot of people who don't understand the math behind it but understand the value of what that math produces," says Mitch Stewart, Organizing for America's director. "Once that first special [election] happened, his word was the gold standard at the DNC."

The significance of Wagner's achievement went far beyond his ability to declare winners months before Election Day. His approach amounted to a decisive break with 20th-century tools for tracking public opinion, which revolved around quarantining small samples that could be treated as representative of the whole. Wagner had emerged from a cadre of analysts who thought of voters as individuals and worked to aggregate projections about their opinions and behavior until they revealed a composite picture of everyone. His techniques marked the fulfillment of a new way of thinking, a decade in the making, in which voters were no longer trapped in old political geographies or tethered to traditional demographic categories, such as age or gender, depending on which attributes pollsters asked about or how consumer marketers classified them for commercial purposes. Instead, the electorate could be seen as a collection of individual citizens who could each be measured and assessed on their own terms. Now it was up to a candidate who wanted to lead those people to build a campaign that would interact with them the same way.

ole0

Dan Wagner, the chief analytics officer for Obama 2012, led the campaign's "Cave" of data scientists.

After the voters returned Obama to office for a second term, his campaign became celebrated for its use of technology—much of it developed by an unusual team of coders and engineers—that redefined how individuals could use the Web, social media, and smartphones to participate in the political process. A mobile app allowed a canvasser to download and return walk sheets without ever entering a campaign office; a Web platform called Dashboard gamified volunteer activity by ranking the most active supporters; and "targeted sharing" protocols mined an Obama backer's Facebook network in search of friends the campaign wanted to register, mobilize, or persuade.

But underneath all that were scores describing particular voters: a new political currency that predicted the behavior of individual humans. The campaign didn't just know who you were; it knew exactly how it could turn you into the type of person it wanted you to be.

The Scores

Four years earlier, Dan Wagner had been working at a Chicago economic consultancy, using forecasting skills developed studying econometrics at the University of Chicago, when he fell for Barack Obama and decided he wanted to work on his home-state senator's 2008 presidential campaign. Wagner, then 24, was soon in Des Moines, handling data entry for the state voter file that guided Obama to his crucial victory in the Iowa caucuses. He bounced from state to state through the long primary calendar, growing familiar with voter data and the ways of using statistical models to intelligently sort the electorate. For the general election, he was named lead targeter for the Great Lakes/Ohio River Valley region, the most intense battleground in the country.

After Obama's victory, many of his top advisors decamped to Washington to make preparations for governing. Wagner was told to stay behind and serve on a post-election task force that would review a campaign that had looked, to the outside world, technically flawless.

In the 2008 presidential election, Obama's targeters had assigned every voter in the country a pair of scores based on the probability that the individual would perform two distinct actions that mattered to the campaign: casting a ballot and supporting Obama. These scores were derived from an unprecedented volume of ongoing survey work. For each battleground state every week, the campaign's call centers conducted 5,000 to 10,000 so-called short-form interviews that quickly gauged a voter's preferences, and 1,000 interviews in a long-form version that was more like a traditional poll. To derive individual-level predictions, algorithms trawled for patterns between these opinions and the data points the campaign had assembled for every voter—as many as one thousand variables each, drawn from voter registration records, consumer data warehouses, and past campaign contacts.

ole1

This innovation was most valued in the field. There, an almost perfect cycle of microtargeting models directed volunteers to scripted conversations with specific voters at the door or over the phone. Each of those interactions produced data that streamed back into Obama's servers to refine the models pointing volunteers toward the next door worth a knock. The efficiency and scale of that process put the Democrats well ahead when it came to profiling voters. John McCain's campaign had, in most states, run its statistical model just once, assigning each voter to one of its microtargeting segments in the summer. McCain's advisors were unable to recalculate the probability that those voters would support their candidate as the dynamics of the race changed. Obama's scores, on the other hand, adjusted weekly, responding to new events like Sarah Palin's vice-presidential nomination or the collapse of Lehman Brothers.

Within the campaign, however, the Obama data operations were understood to have shortcomings. As was typical in political information infrastructure, knowledge about people was stored separately from data about the campaign's interactions with them, mostly because the databases built for those purposes had been developed by different consultants who had no interest in making their systems work together.

But the task force knew the next campaign wasn't stuck with that situation. Obama would run his final race not as an insurgent against a party establishment, but as the establishment itself. For four years, the task force members knew, their team would control the Democratic Party's apparatus. Their demands, not the offerings of consultants and vendors, would shape the marketplace. Their report recommended developing a "constituent relationship management system" that would allow staff across the campaign to look up individuals not just as voters or volunteers or donors or website users but as citizens in full. "We realized there was a problem with how our data and infrastructure interacted with the rest of the campaign, and we ought to be able to offer it to all parts of the campaign," says Chris Wegrzyn, a database applications developer who served on the task force.

Wegrzyn became the DNC's lead targeting developer and oversaw a series of costly acquisitions, all intended to free the party from the traditional dependence on outside vendors. The committee installed a Siemens Enterprise System phone-dialing unit that could put out 1.2 million calls a day to survey voters' opinions. Later, party leaders signed off on a $280,000 license to use Vertica software from Hewlett-Packard that allowed their servers to access not only the party's 180-million-person voter file but all the data about volunteers, donors, and those who had interacted with Obama online.

Many of those who went to Washington after the 2008 election in order to further the president's political agenda returned to Chicago in the spring of 2011 to work on his reëlection. The chastening losses they had experienced in Washington separated them from those who had known only the ecstasies of 2008. "People who did '08, but didn't do '10, and came back in '11 or '12—they had the hardest culture clash," says Jeremy Bird, who became national field director on the reëlection campaign. But those who went to Washington and returned to Chicago developed a particular appreciation for Wagner's methods of working with the electorate at an atomic level. It was a way of thinking that perfectly aligned with their simple theory of what it would take to win the president reëlection: get everyone who had voted for him in 2008 to do it again. At the same time, they knew they would need to succeed at registering and mobilizing new voters, especially in some of the fastest-growing demographic categories, to make up for any 2008 voters who did defect.

Obama's campaign began the election year confident it knew the name of every one of the 69,456,897 Americans whose votes had put him in the White House. They may have cast those votes by secret ballot, but Obama's analysts could look at the Democrats' vote totals in each precinct and identify the people most likely to have backed him. Pundits talked in the abstract about reassembling Obama's 2008 coalition. But within the campaign, the goal was literal. They would reassemble the coalition, one by one, through personal contacts

The Experiments

When Jim Messina arrived in Chicago as Obama's newly minted campaign manager in January of 2011, he imposed a mandate on his recruits: they were to make decisions based on measurable data. But that didn't mean quite what it had four years before. The 2008 campaign had been "data-driven," as people liked to say. This reflected a principled imperative to challenge the political establishment with an empirical approach to electioneering, and it was greatly influenced by David Plouffe, the 2008 campaign manager, who loved metrics, spreadsheets, and performance reports. Plouffe wanted to know: How many of a field office's volunteer shifts had been filled last weekend? How much money did that ad campaign bring in?

But for all its reliance on data, the 2008 Obama campaign had remained insulated from the most important methodological innovation in 21st-century politics. In 1998, Yale professors Don Green and Alan Gerber conducted the first randomized controlled trial in modern political science, assigning New Haven voters to receive nonpartisan election reminders by mail, phone, or in-person visit from a canvasser and measuring which group saw the greatest increase in turnout. The subsequent wave of field experiments by Green, Gerber, and their followers focused on mobilization, testing competing modes of contact and get-out-the-vote language to see which were most successful.

The first Obama campaign used the findings of such tests to tweak call scripts and canvassing protocols, but it never fully embraced the experimental revolution itself. After Dan Wagner moved to the DNC, the party decided it would start conducting its own experiments. He hoped the committee could become "a driver of research for the Democratic Party."

To that end, he hired the Analyst Institute, a Washington-based consortium founded under the AFL-CIO's leadership in 2006 to coördinate field research projects across the electioneering left and distribute the findings among allies. Much of the experimental world's research had focused on voter registration, because that was easy to measure. The breakthrough was that registration no longer had to be approached passively; organizers did not have to simply wait for the unenrolled to emerge from anonymity, sign a form, and, they hoped, vote. New techniques made it possible to intelligently profile nonvoters: commercial data warehouses sold lists of all voting-age adults, and comparing those lists with registration rolls revealed eligible candidates, each attached to a home address to which an application could be mailed. Applying microtargeting models identified which nonregistrants were most likely to be Democrats and which ones Republicans.

The Obama campaign embedded social scientists from the Analyst Institute among its staff. Party officials knew that adding new Democratic voters to the registration rolls was a crucial element in their strategy for 2012. But already the campaign had ambitions beyond merely modifying nonparticipating citizens' behavior through registration and mobilization. It wanted to take on the most vexing problem in politics: changing voters' minds.

The expansion of individual-level data had made possible the kind of testing that could help do that. Experimenters had typically calculated the average effect of their interventions across the entire population. But as campaigns developed deep portraits of the voters in their databases, it became possible to measure the attributes of the people who were actually moved by an experiment's impact. A series of tests in 2006 by the women's group Emily's List had illustrated the potential of conducting controlled trials with microtargeting databases. When the group sent direct mail in favor of Democratic gubernatorial candidates, it barely budged those whose scores placed them in the middle of the partisan spectrum; it had a far greater impact upon those who had been profiled as soft (or nonideological) Republicans.

That test, and others that followed, demonstrated the limitations of traditional targeting. Such techniques rested on a series of long-standing assumptions—for instance, that middle-of-the-roaders were the most persuadable and that infrequent voters were the likeliest to be captured in a get-out-the-vote drive. But the experiments introduced new uncertainty. People who were identified as having a 50 percent likelihood of voting for a Democrat might in fact be torn between the two parties, or they might look like centrists only because no data attached to their records pushed a partisan prediction in one direction or another. "The scores in the middle are the people we know less about," says Chris Wyant, a 2008 field organizer who became the campaign's general election director in Ohio four years later. "The extent to which we were guessing about persuasion was not lost on any of us."

ole2

One way the campaign sought to identify the ripest targets was through a series of what the Analyst Institute called "experiment-informed programs," or EIPs, designed to measure how effective different types of messages were at moving public opinion.

The traditional way of doing this had been to audition themes and language in focus groups and then test the winning material in polls to see which categories of voters responded positively to each approach. Any insights were distorted by the artificial settings and by the tiny samples of demographic subgroups in traditional polls. "You're making significant resource decisions based on 160 people?" asks Mitch Stewart, director of the Democratic campaign group Organizing for America. "Isn't that nuts? And people have been doing that for decades!"

An experimental program would use those steps to develop a range of prospective messages that could be subjected to empirical testing in the real world. Experimenters would randomly assign voters to receive varied sequences of direct mail—four pieces on the same policy theme, each making a slightly different case for Obama—and then use ongoing survey calls to isolate the attributes of those whose opinions changed as a result.

In March, the campaign used this technique to test various ways of promoting the administration's health-care policies. One series of mailers described Obama's regulatory reforms; another advised voters that they were now entitled to free regular check-ups and ought to schedule one. The experiment revealed how much voter response differed by age, especially among women. Older women thought more highly of the policies when they received reminders about preventive care; younger women liked them more when they were told about contraceptive coverage and new rules that prohibited insurance companies from charging women more.

When Paul Ryan was named to the Republican ticket in August, Obama's advisors rushed out an EIP that compared different lines of attack about Medicare. The results were surprising. "The electorate [had seemed] very inelastic," says Terry Walsh, who coördinated the campaign's polling and paid-media spending. "In fact, when we did the Medicare EIPs, we got positive movement that was very heartening, because it was at a time when we were not seeing a lot of movement in the electorate." But that movement came from quarters where a traditional campaign would never have gone hunting for minds it could change. The Obama team found that voters between 45 and 65 were more likely to change their views about the candidates after hearing Obama's Medicare arguments than those over 65, who were currently eligible for the program.

A similar strategy of targeting an unexpected population emerged from a July EIP testing Obama's messages aimed at women. The voters most responsive to the campaign's arguments about equal-pay measures and women's health, it found, were those whose likelihood of supporting the president was scored at merely 20 and 40 percent. Those scores suggested that they probably shared Republican attitudes; but here was one thing that could pull them to Obama. As a result, when Obama unveiled a direct-mail track addressing only women's issues, it wasn't to shore up interest among core parts of the Democratic coalition, but to reach over for conservatives who were at odds with their party on gender concerns. "The whole goal of the women's track was to pick off votes for Romney," says Walsh. "We were able to persuade people who fell low on candidate support scores if we gave them a specific message."

At the same time, Obama's campaign was pursuing a second, even more audacious adventure in persuasion: one-on-one interaction. Traditionally, campaigns have restricted their persuasion efforts to channels like mass media or direct mail, where they can control presentation, language, and targeting. Sending volunteers to persuade voters would mean forcing them to interact with opponents, or with voters who were undecided because they were alienated from politics on delicate issues like abortion. Campaigns have typically resisted relinquishing control of ground-level interactions with voters to risk such potentially combustible situations; they felt they didn't know enough about their supporters or volunteers. "You can have a negative impact," says Jeremy Bird, who served as national deputy director of Organizing for America. "You can hurt your candidate."

In February, however, Obama volunteers attempted 500,000 conversations with the goal of winning new supporters. Voters who'd been randomly selected from a group identified as persuadable were polled after a phone conversation that began with a volunteer reading from a script. "We definitely find certain people moved more than other people," says Bird. Analysts identified their attributes and made them the core of a persuasion model that predicted, on a scale of 0 to 10, the likelihood that a voter could be pulled in Obama's direction after a single volunteer interaction. The experiment also taught Obama's field department about its volunteers. Those in California, which had always had an exceptionally mature volunteer organization for a non-battleground state, turned out to be especially persuasive: voters called by Californians, no matter what state they were in themselves, were more likely to become Obama supporters.

ole3

Alex Lundry created Mitt Romney's data science unit. It was less than one-tenth the size of Obama's analytics team.

With these findings in hand, Obama's strategists grew confident that they were no longer restricted to advertising as a channel for persuasion. They began sending trained volunteers to knock on doors or make phone calls with the objective of changing minds.

That dramatic shift in the culture of electioneering was felt on the streets, but it was possible only because of advances in analytics. Chris Wegrzyn, a database applications developer, developed a program code-named Airwolf that matched county and state lists of people who had requested mail ballots with the campaign's list of e-mail addresses. Likely Obama supporters would get regular reminders from their local field organizers, asking them to return their ballots, and, once they had, a message thanking them and proposing other ways to be involved in the campaign. The local organizer would receive daily lists of the voters on his or her turf who had outstanding ballots so that the campaign could follow up with personal contact by phone or at the doorstep. "It is a fundamental way of tying together the online and offline worlds," says Wagner.

Wagner, however, was turning his attention beyond the field. By June of 2011, he was chief analytics officer for the campaign and had begun making the rounds of the other units at headquarters, from fund-raising to communications, offering to help "solve their problems with data." He imagined the analytics department—now a 54-person staff, housed in a windowless office known as the Cave—as an "in-house consultancy" with other parts of the campaign as its clients. "There's a process of helping people learn about the tools so they can be a participant in the process," he says. "We essentially built products for each of those various departments that were paired up with a massive database we had."

The Flow

As job notices seeking specialists in text analytics, computational advertising, and online experiments came out of the incumbent's campaign, Mitt Romney's advisors at the Republicans' headquarters in Boston's North End watched with a combination of awe and perplexity. Throughout the primaries, Romney had appeared to be the only Republican running a 21st-century campaign, methodically banking early votes in states like Florida and Ohio before his disorganized opponents could establish operations there.

But the Republican winner's relative sophistication in the primaries belied a poverty of expertise compared with the Obama campaign. Since his first campaign for governor of Massachusetts, in 2002, Romney had relied upon TargetPoint Consulting, a Virginia firm that was then a pioneer in linking information from consumer data warehouses to voter registration records and using it to develop individual-level predictive models. It was TargetPoint's CEO, Alexander Gage, who had coined the term "microtargeting" to describe the process, which he modeled on the corporate world's approach to customer relationship management.

Such techniques had offered George W. Bush's reëlection campaign a significant edge in targeting, but Republicans had done little to institutionalize that advantage in the years since. By 2006, Democrats had not only matched Republicans in adopting commercial marketing techniques; they had moved ahead by integrating methods developed in the social sciences.

Romney's advisors knew that Obama was building innovative internal data analytics departments, but they didn't feel a need to match those activities. "I don't think we thought, relative to the marketplace, we could be the best at data in-house all the time," Romney's digital director, Zac Moffatt, said in July. "Our idea is to find the best firms to work with us." As a result, Romney remained dependent on TargetPoint to develop voter segments, often just once, and then deliver them to the campaign's databases. That was the structure Obama had abandoned after winning the nomination in 2008.

In May a TargetPoint vice president, Alex Lundry, took leave from his post at the firm to assemble a data science unit within Romney's headquarters. To round out his team, Lundry brought in Tom Wood, a University of Chicago postdoctoral student in political science, and Brent McGoldrick, a veteran of Bush's 2004 campaign who had left politics for the consulting firm Financial Dynamics (later FTI Consulting), where he helped financial-services, health-care, and energy companies communicate better. But Romney's data science team was less than one-tenth the size of Obama's analytics department. Without a large in-house staff to handle the massive national data sets that made it possible to test and track citizens, Romney's data scientists never tried to deepen their understanding of individual behavior. Instead, they fixated on trying to unlock one big, persistent mystery, which Lundry framed this way: "How can we get a sense of whether this advertising is working?"

"You usually get GRPs and tracking polls," he says, referring to the gross ratings points that are the basic unit of measuring television buys. "There's a very large causal leap you have to make from one to the other."

Lundry decided to focus on more manageable ways of measuring what he called the information flow. His team converted topics of political communication into discrete units they called "entities." They initially classified 200 of them, including issues like the auto industry bailout, controversies like the one surrounding federal funding for the solar-power company Solyndra, and catchphrases like "the war on women." When a new concept (such as Obama's offhand remark, during a speech about our common dependence on infrastructure, that "you didn't build that") emerged as part of the election-year lexicon, the analysts added it to the list. They tracked each entity on the National Dialogue Monitor, TargetPoint's system for measuring the frequency and tone with which certain topics are mentioned across all media. TargetPoint also integrated content collected from newspaper websites and closed-caption transcripts of broadcast programs. Lundry's team aimed to examine how every entity fared over time in each of two categories: the informal sphere of social media, especially Twitter, and the journalistic product that campaigns call earned press coverage.

Ultimately, Lundry wanted to assess the impact that each type of public attention had on what mattered most to them: Romney's position in the horse race. He turned to vector autoregression models, which equities traders use to isolate the influence of single variables on market movements. In this case, Lundry's team looked for patterns in the relationship between the National Dialogue Monitor's data and Romney's numbers in Gallup's daily tracking polls. By the end of July, they thought they had identified a three-step process they called "Wood's Triangle."

Within three or four days of a new entity's entry into the conversation, either through paid ads or through the news cycle, it was possible to make a well-informed hypothesis about whether the topic was likely to win media attention by tracking whether it generated Twitter chatter. That informal conversation among political-class elites typically led to traditional print or broadcast press coverage one to two days later, and that, in turn, might have an impact on the horse race. "We saw this process over and over again," says Lundry.

They began to think of ads as a "shock to the system"—a way to either introduce a new topic or restore focus on an area in which elite interest had faded. If an entity didn't gain its own energy—as when the Republicans charged over the summer that the White House had waived the work requirements in the federal welfare rules—Lundry would propose a "re-shock to the system" with another ad on the subject five to seven days later. After 12 to 14 days, Lundry found, an entity had moved through the system and exhausted its ability to move public opinion—so he would recommend to the campaign's communications staff that they move on to something new.

Those insights offered campaign officials a theory of information flows, but they provided no guidance in how to allocate campaign resources in order to win the Electoral College. Assuming that Obama had superior ground-level data and analytics, Romney's campaign tried to leverage its rivals' strategy to shape its own; if Democrats thought a state or media market was competitive, maybe that was evidence that Republicans should think so too. "We were necessarily reactive, because we were putting together the plane as it took off," Lundry says. "They had an enormous head start on us."

Romney's political department began holding regular meetings to look at where in the country the Obama campaign was focusing resources like ad dollars and the president's time. The goal was to try to divine the calculations behind those decisions. It was, in essence, the way Microsoft's Bing approached Google: trying to reverse-engineer the market leader's code by studying the visible output. "We watch where the president goes," Dan Centinello, the Romney deputy political director who oversaw the meetings, said over the summer.

Obama's media-buying strategy proved particularly hard to decipher. In early September, as part of his standard review, Lundry noticed that the week after the Democratic convention, Obama had aired 68 ads in Dothan, Alabama, a town near the Florida border. Dothan was one of the country's smallest media markets, and Alabama one of the safest Republican states. Even though the area was known to savvy ad buyers as one of the places where a media market crosses state lines, Dothan TV stations reached only about 9,000 Florida voters, and around 7,000 of them had voted for John McCain in 2008. "This is a hard-core Republican media market," Lundry says. "It's incredibly tiny. But they were advertising there."

Romney's advisors might have formed a theory about the broader media environment, but whatever was sending Obama hunting for a small pocket of votes was beyond their measurement. "We could tell," says McGoldrick, "that there was something in the algorithms that was telling them what to run."

Tomorrow: Part 3—The Community