Data brokers, workplace sensor studies, unreported drug side effects revealed in search data, and the dark side of big data.
ProPublica's Lois Beckett takes a look this week at data brokers. She says that though Congress is making moves to make such companies give consumers more control over their data and what happens to it, many people not only don't know these data brokers exist, but they also don't know the extent of the data gathered and how it's used.
Beckett takes a step-by-step look at who these companies are, how much they know, how they get our data and from where, what kind of data they're allowed to collect, how they use the data, and much, much more. She says most of the time, consumers have no idea their data has been purchased. For instance, "When you're checking out at a store and a cashier asks you for your Zip code, the store isn't just getting that single piece of information," she writes. "Acxiom and other data companies offer services that allow stores to use your Zip code and the name on your credit card to pinpoint your home address — without asking you for it directly."
It's possible but very, very difficult for consumers to stop the companies from collecting and sharing their data, Beckett notes. Though most data brokers have an "opt-out" policy, consumers would "need to know about all the different data brokers and where to find their opt-outs" — information that most consumers don't have and don't know how to find. You can find Beckett's full report at ProPublica — it's this week's recommended read.
In related news, Rachel Emma Silverman at the Wall Street Journal takes a look at the use of sensors and data gathering practices in the workplace. "As big data becomes a fixture of office life, companies are turning to tracking devices to gather real-time information on how teams of employees work and interact," she writes. "Sensors, worn on lanyards or placed on office furniture, record how often staffers get up from their desks, consult other teams and hold meetings."
Though there's a fine line between big data and big brother, Silverman says, "[s]ensor proponents … argue that smartphones and corporate ID badges already can transmit their owner's location" and most companies will allow workers to opt out of the sensor studies. Silverman reviews a few real-world sensor study cases, along with the results and insights gleaned. You can read her full report at the Wall Street Journal.
Study shows search data reveals unreported drug side effects
John Markoff reports at the New York Times on a study published this week that shows by data mining Internet search data, scientists from Microsoft, Stanford and Columbia University have been able to discover unreported prescription drug side effects before the FDA's warning system flagged them. Markoff writes:
"Using automated software tools to examine queries by six million Internet users taken from web search logs in 2010, the researchers looked for searches relating to an antidepressant, paroxetine, and a cholesterol lowering drug, pravastatin. They were able to find evidence that the combination of the two drugs caused high blood sugar."
Users who opted to participate in the study installed a browser toolbar that gathered anonymized data, Markoff reports. Using data from 82 million drug-, symptom- and condition-related searches in 2010, researchers were able to cross-reference searches for "paroxetine" and "pravastatin" with the number of times users would also search for "hyperglycemia" or any of its 80 or so symptoms.
"They determined that people who searched for both drugs during the 12-month period were significantly more likely to search for terms related to hyperglycemia than were those who searched for just one of the drugs," Markoff reports. He also notes that the searches for symptoms relating to both drugs occurred within a short period of time — 30% the same day, 40% the same week, and 50% the same month. You can read Markoff's full report at The New York Times.
Keeping an eye on big data's dark side
Viktor Mayer-Schönberger's and Kenneth Cukier's new book Big Data: A Revolution That Will Transform How We Live, Work, and Think was released this week. The duo addressed one of the topics in their book — predicting and punishing crime before it happens, ala Minority Report — in a post at PopSci. They warn that along with all the benefits we're reaping from big data, we need to be conscious of big data's dark side, too:
"Already we see the seedlings of Minority Report-style predictions penalizing people. Parole boards in more than half of all U.S. states use predictions founded on data analysis as a factor in deciding whether to release somebody from prison or to keep him incarcerated. A growing number of places in the United States — from precincts in Los Angeles to cities like Richmond, Virginia — employ 'predictive policing': using big-data analysis to select what streets, groups, and individuals to subject to extra scrutiny, simply because an algorithm pointed to them as more likely to commit crime."
They warn further that it won't stop there — law enforcement will eventually attempt to predict crime on individual levels and, ultimately, to use big data to prevent the crime in the first place. You can read more from Mayer-Schönberger's and Cukier's piece at PopSci.
Cukier also sat down with O'Reilly Radar online managing editor Mac Slocum at the recent Strata Conference in Santa Clara to talk about government's use of big data, regulations and restrictions, and what needs to be done to keep data open. You can watch their interview in the following video:
Steve Loh, The New York Times, February 28, 2013
Personal data is a valuable asset that ought to be put to work.
Fluid data markets will benefit economies, societies and individuals.
Privacy rules should focus on how data is used rather than on the widespread collection of personal data.
That is the gist of a new report from World Economic Forum's Personal Data project, "Unlocking the Value of Personal Data: From Collection to Usage."
The modern digital world, with its explosion of data, has made the traditional approach to privacy based on "notice and consent" typically between two parties — a marketer and a consumer — obsolete, in the view of the report's authors.
"The technology has overrun the classical model," said Craig Mundie, a senior adviser to Microsoft's chief executive, Steven A. Ballmer.
Mr. Mundie was on the five-member steering board for the report. All five people represent corporations that stand to gain from tapping personal data.
Privacy advocates and regulators in Europe and the United States have been reluctant to give up on efforts to control the collection of data. Their concern is that once personal data is collected, its use is very difficult to monitor and control. Information brokers that consumers never see — and few know about — market personal data to advertisers, retailers, financial institutions and others. That problem prompted the Federal Trade Commission in a report last year to recommend that Congress enact legislation "to provide greater transparency for, and control over, the practices of information brokers."
But while recognizing the privacy challenges, the companies participating in the World Economic Forum project say what was needed was a careful balance. In a blog post on Wednesday, Raymond J. Baxter, a senior vice president of Kaiser Permanente, a major health care provider and insurer, emphasized the value of personal data, when used properly. He cited Kaiser's use of personal medical data for research.
For example, mining family data and outcomes over years, Kaiser scientists found that the children of women who took anti-depressant drugs while pregnant had more than twice the risk of developing autism disorders. "By discovering this correlation and leveraging this data in new ways, lives are improved," Mr. Baxter wrote.
According to Mr. Mundie of Microsoft, technology can help strike the right balance between individuals' concerns about privacy and the benefit of a fluid market in personal data. He said independent organizations, most likely nonprofits, would develop automated privacy preference services that individuals could subscribe to. A person would check off what he or she wanted his data to be used for and not. Those preferences, he explained, would then be encoded as software tags that traveled with the person's data.
Those preferences, Mr. Mundie added, could vary depending on context. For example, a person might say he or she did not want personal medical data shared beyond a family doctor and one or two specialists — unless the person was taken to an emergency ward.
"You can intelligently use computing technology to provide the benefits and curtail abuse," Mr. Mundie said.
A small group of academics, business executives and journalists gathered at the M.I.T. Media Lab last Thursday, and the purpose was to toss out ideas and discuss the concept of "Data-Driven Societies." A daunting topic, ambitious and vague at once, it seems.
Up to now, the focus on the power and implications of Big Data technology has been involved social media, business decision-making and online privacy. Those are big subjects in their own right. So it's not surprising that the notion of a data-driven society has not been much considered.
But someone who has was host of the meeting: Alex Pentland, a computational social scientist at the Media Lab. He put his intellectual stake in the ground last year in a presentation posted on Edge.org, "Reinventing Society in the Wake of Big Data."
Mr. Pentland's starting point is that the most important data that is becoming available on a vast new scale is information about people's behavior. For example, he cites location data from cellphones and evermore consumption data as people increasingly use credit cards for even the smallest purchases. He distinguishes this behavioral data from less-telling data — about people's beliefs like Facebook communications or Google searches.
The fine-grained behavioral data, according to Mr. Pentland, opens the way to changing how we think about society and how a society is governed. Adam Smith and Karl Marx, he explains, thought about markets and classes, respectively. "But those are aggregates," he said. "They're averages."
Yet now, Mr. Pentland says, it becomes possible to track social phenomena down to the individual level and the social and economic connections among individuals. The ability to monitor these "micro-patterns," Mr. Pentland said, means "we're entering a new era of social physics."
What might that mean in practice? Reed Hundt, the chairman of the Federal Communications Commission in the Clinton administration, observed at the meeting that Big Data played a major role in the last election — a reference to the Obama campaign's deft use of data analysis to identify potential Obama voters and encourage them to cast their ballots.
"You get elected with Big Data, but you govern without it," Mr. Hundt said. "How much sense does that make?"
Mr. Hundt, chief executive for the Coalition for Green Capital, a nonprofit organization, pointed to the waste in a range of government incentive and benefit programs, from tax credits for solar panels to Social Security, that results from the across-the-board approach — or policy by averages, as Mr. Pentland might put it.
Instead, a data-driven approach to solar-energy incentives would concentrate government incentives to where the payoff is greatest in terms of efficiently generating alternative energy — larger buildings with a lot of roof space instead of small houses, Mr. Hundt said. A by-the-data model for benefits programs, he added, would suggest means-testing Social Security payments as well as adjusting payments locally for differences in costs of living.
"So all men are created equal, but are subsidized individually," Mr. Hundt quipped.
Intriguing, and perhaps wise policy, but it would also seem to be a redefinition of fairness as it applies to broad benefit programs, like Social Security and Medicare, which typically make standard payments and avoid means testing.
What are the chances such a data-driven course would be politically acceptable? If the data points the way to greater efficiency, why not, Mr. Hundt replied. After all, he said, a major role of government is to transfer income to people who would benefit most — better data, closely analyzed, means government can perform that role more effectively, Mr. Hundt said.
An underlying assumption of tilting toward a data-driven society is that, as one participant put it, "information over time wins out." That is, data will change attitudes and policy, combating bias and causing policy-making to be more of a science. To data optimists, then, the endless political squabbling and stalemate in Washington points to all the room there is for improvement.
In a Big Data world, the data-mining for patterns and insights to guide policy will be done automatically — by software algorithms. Of course, algorithms are created by people and they contain inferences and assumptions coded in. Those coded-in values shape the output — computer-generated predictions, recommendations and simulations.
That raises question of the human design and control of the computerized helpers in policy-making, as in other realms of decision-making. "At some point, you're in the hands of the algorithm," observed John Henry Clippinger, chief executive of the Institute for Data Driven Design, a nonprofit research and educational organization. "You're whistling in the dark if you don't think that day is coming."
Ms. Smith, Network World, December 2, 2012
There are several Internet security experts who agree with Steve Rambam's claim [1] that "Privacy is dead - get over it." Yet other privacy and security experts such as Bruce Schneier completely disagree. In The Value of Privacy [2] Schneier wrote, "Privacy protects us from abuses by those in power, even if we're doing nothing wrong at the time of surveillance." When it comes to data protection and protecting people's privacy in the digital age, Europe is far more advanced than America. [3]
In fact, the head of France's data protection agency, Isabelle Falque-Pierrotin, did an excellent job summing it up as: "In Europe, we consider privacy a fundamental right. That doesn't mean it is exclusive of other rights, but economic rights are not superior to privacy." The New York Times also reported [4] that she said in the United States, "personal data are seen as raw material for business."
In November, Microsoft's Chief Privacy Officer Brendon Lynch said [5] of the IAPP European Data Protection Congress 2012 [6], "One area of strong consensus was the tremendous potential the digital economy holds for companies on both sides of the pond. Accordingly, it's important to strike the right balance between data protection with business growth through interoperability between privacy regulation in the EU, U.S. and elsewhere."
Many privacy advocates cringe when hearing the word "balance," such as striking a balance between security and privacy. Hopefully people won't come to cringe when they hear the word balance applied to big data security protections and privacy. As Bruce Schneier wrote [2] way back in 2006:
Too many wrongly characterize the debate as "security versus privacy." The real choice is liberty versus control. Tyranny, whether it arises under threat of foreign physical attack or under constant domestic authoritative scrutiny, is still tyranny. Liberty requires security without intrusion, security plus privacy. Widespread police surveillance is the very definition of a police state. And that's why we should champion privacy even when we have nothing to hide.
Whether people realize it or not, big data is not privacy-friendly even when it is supposedly anonymized or contains obfuscated PII (Personally Identifiable Information) data. Researchers have shown that "linkability threats" can re-identity individuals. Since it boils down to the fact that you are not anonymous when it comes to big data [7], Microsoft has developed "Differential Privacy for everyone" [download PDF [8]].
In the IAPP keynote address [download PDF [9]], Lynch made some excellent and thought-provoking privacy points regarding big data. He said:
Data is the fuel that drives all of these powerful technologies, but what can be done with the data today can at times seem enormously helpful or enormously threatening. Consider two scenarios shown here. In the first case, I am using my phone in a grocery store to find out more about the items on the shelves and it is mashing up that with my private data to personalize my experience. So here I downloaded a recipe and customized it for my dietary needs. If it's a trusted system, that's a great experience. On the other hand, consider the US company, Target, which recently generated a lot of press about its pregnancy prediction score. This was based on what people were purchasing in Target stores, they are able to indicate a shopper that appeared to be pregnant. The concern about how Target can figure out such details about customers shopping in its stores, who are not explicitly sharing that information, is the concern. And what does it do with those insights? In this particular case, they sent some mailers to the individual involved - it was a teenage girl and her father was very offended that they were wrongly marketing to her, but it eventually did come out that she was in fact pregnant. Target knew a lot more than her father knew.
Peter Cullen, Microsoft's Chief Privacy Strategist, wrote [10] about "notice and consent" as a means of privacy protection and how data privacy frameworks need to "focus on the 'harms' or 'impacts' of data use, which should not only include physical and financial injury, but also broader concepts such as reputational or social harm."
Yet after showing a video that highlighted data transfers in today's world at the IAPP conference, Lynch said, "How could there possibly be meaningful notice and consent mechanisms in place for every transfer of data that was involved?" He added, "It would seem that advances in technology and the rise of big data can create amazing societal benefits but they can also strain traditional notions of secrecy and the notice and consent approach to privacy protection."
In his keynote, Lynch said:
Some technology and internet companies today take the position that privacy is dead, or at least that privacy is an outdated concept that people need to get over so technology companies can help them reap the benefits of sharing as much information as possible. But we disagree that privacy is not relevant or desirable, in this sensor-driven, social everywhere, big data world that we are heading towards. People today expect strong privacy protections because they are increasingly aware of, and concerned about, the digital trails they leave behind online and indeed there's plenty of evidence that people still care deeply about privacy.
Of course people care about privacy. Europe continues to illustrate this to the world by taking a hard stance when data is used without "informed consent" and when users cannot "opt out." Lynch believes we need to not only protect privacy in regards to big data, but also that people "need an updated notion of privacy and data protection principles, one that shifts from a focus on secrecy to a more nuanced approach, based on reasonable consumer expectations, context and a greater emphasis on how personal information is used."
Big data definitely represents significant threats to personal privacy. Let's hope this "shift" and "updated notion of privacy" won't include the word "balance" that puts individuals on the losing end as it generally has when the government talks of striking a balance between security and privacy.
Craig Timberg, The Washington Post, November 20, 2012
Shortly before Election Day, a Stanford graduate student reported that the campaign Web sites of both President Obama and Republican Mitt Romney were “leaking” personal information about their supporters through careless data handling.
Had it been Facebook and Google, a federal investigation might have ensued, and the companies could have suffered significant public relations setbacks and perhaps fines. But the Federal Trade Commission, the government agency most focused on personal privacy, has no jurisdiction over campaigns or political groups.
That is a small example of what privacy advocates say is a big problem with efforts to protect personal information in the United States: The politicians are not guarding the chicken coop. They are the foxes.
Obama’s sophisticated use of Big Data gave him a crucial edge in what, based on popular support alone, should have been a close election. Republicans are desperate to catch up. But it’s not clear who is positioned to protect the rights of voters at a time when politicians from both parties increasingly build their campaigns on the insights that commercial data brokers provide.
Washington has a community of professional privacy advocates at places such as the ACLU, the Electronic Privacy Information Center and the Center for Digital Democracy. Jeff Chester, executive director of the Center for Digital Democracy, said he approached lawmakers from both parties to express his concerns long before the election. But he got nowhere.
“Maybe we’re digital Don Quixotes,” Chester said. “There was a lack of interest, not surprisingly.”
People routinely tell pollsters that they’re concerned about online privacy, and Chester and his colleagues in the field count some allies on Capitol Hill and in the White House. The FTC under Chairman Jon Leibowitz and David Vladeck, head of its Bureau of Consumer Protection, have made the agency far more aggressive on consumer privacy generally — even if political campaigns are beyond their reach.
Yet overall the laws in the United States are much less strict than in Europe, where there are tight limits on what personal information can be collected and how long it can be kept. Companies caught crossing the line can provoke furious backlashes among their users.
The American political landscape, by comparison, is amorphous when it comes to privacy. There are widespread concerns on both the right and left but no single, coherent constituency demanding greater protections.
For all the talk in recent years about online privacy, data-hungry Google remains the most popular search engine and data-hungry Facebook the most popular social media site. Both worked closely with the campaigns and also have growing lobbying operations in Washington. Google’s Executive Chairman Eric Schmidt was a regular visitor at Obama’s Chicago campaign headquarters, say those who worked there, offering advice to the campaign’s data-savvy technologists.
Privacy advocates say the tide will eventually turn, when Americans truly understand the extent to which their information is collected and traded. A recent poll by the University of Pennsylvania’s Annenberg School for Communications found that nearly two-thirds of people would be less likely to support a candidate who bought data about voters’ online activities and used it to tailor political ads.
“People still don’t quite understand this stuff,” said Joseph Turow, the lead researcher on the Annenberg poll. He said politicians are “hoping people will, quote-unquote, get used to it.”
Marketing politicians is now like selling drinks. It involves filtering policies and voters through algorithms.
L. Gordon Crovitz, The Wall Street Journal, November 18, 2012
When the Obama campaign emailed supporters to join a $40,000-a-ticket dinner in June at the New York home of actress Sarah Jessica Parker, journalists at ProPublica noticed something odd. They uncovered seven versions of the email solicitation for the fundraiser, some mentioning a second fundraiser that night, a concert by Mariah Carey, others that Ms. Parker is a mother, and still others that Vogue editor Anna Wintour would be at the dinner.
Who got which email depended on "big data"—information about each fundraising prospect and how different people react to different messages. In this year's election, it looks as if the Obama team's use of such data was one of its biggest edges over the Romney effort.
Some uses of big data were known before the election—for instance, the Obama website was even more assiduous than online retailers like Best Buy about dropping "cookies" on users' computers to gather information about their online habits. Reporting since the election makes clear just how important the role of data was in deciding the election.
Campaign manager Jim Messina pledged to "measure every single thing in this campaign" and built an analytics department five times the size of the 2008 effort. A Time magazine reporter got access to the data scientists in the campaign's Chicago headquarters on the condition that the reporter would keep mum until after the election. "What they revealed as they pulled back the curtain," Time recently reported, "was a massive data effort that helped Obama raise $1 billion, remade the process of targeting TV ads and created detailed models of swing-state voters that could be used to increase the effectiveness of everything from phone calls and door knocks to direct mailings and social media."
According to the magazine, the campaign created a "single massive system that could merge the information collected from pollsters, fundraisers, field workers and consumer databases as well as social-media and mobile contacts with the main Democratic voter files."
The campaign's "chief scientist," Rayid Ghani, had been at Accenture, where he co-wrote an academic paper describing work helping companies that "analyze large amounts of transactional data but are unable to systematically 'understand' their products." For example, Mr. Ghani helped grocers figure out why people bought orange juice by reducing the product to attributes that could be analyzed by algorithms—"Brand: Tropicana, Pulp: low, Fortified with: Vitamin-D, Size: 1 liter, Bottle type: plastic."
Marketing politicians is now like selling drinks. It involves filtering polices and voters through algorithms.
The Obama campaign focused on data showing the "persuadability" of voters. Multivariate tests identified issues and positions that could move undecided voters, ProPublica said: "The persuasion scores allowed the campaign to focus its outreach efforts—and their volunteer calls—on voters who might actually change their minds as the result. It also guided them in what policy messages individual voters should hear."
Big data give incumbents a big advantage, which seems to have surprised the Romney team. The Obama campaign has used cookies to track its supporters online since the 2008 election. It spent the past 18 months creating a new, unified database, factoring in some 80 pieces of information about each person, from age, race and sex to voting history. (The campaign denied reports that it tracked visits to pornography sites in its outreach algorithms.) The Romney campaign says it tried to match the Obama campaign's collection and analysis of data but had to start from scratch and had just seven months after the primaries.
What does this mean for you? Voters need to develop buyer-beware habits. The era of politicianssaying the same thing to all voters is over. Campaigns aim to tell voters exactly what each wants to hear: data-driven pandering.
Another consequence is that efforts by the Federal Trade Commission and other agencies to regulate data mining in the name of privacy are destined to collapse. Last month, Sen. Jay Rockefeller (D., W.Va.) sent a letter to the top "information broker" companies, accusing them of being "elusive" about what data they collect. Companies such as Acxiom and Experian replied that much of their information comes from government databases. They should also point out that political campaigns are among the most sophisticated users of the consumer data they collect.
The Obama campaign deserves credit for its big win through the sophisticated use of big data. As for regulators, they should understand that the information genie will not go back into the bottle—whether consumer information is used to sell orange juice or politicians.
A version of this article appeared November 19, 2012, on page A17 in the U.S. edition of The Wall Street Journal, with the headline: Obama's 'Big Data' Victory.
Big Data: Opportunities and Risks Co-Chairs
Tim Armstrong, Chairman and CEO, AOL Inc.
Dominic Barton, Global Managing Director, McKinsey & Co.
R. Marcelo Claure, Chairman, President and CEO, Brightstar Corp.
Subject Expert
David J. Rothkopf, President and CEO, Garten Rothkopf
Big Data: The Top Four Recommendations
1. Big Data Is Opportunity
Industry already recognizes that the advent of Big Data is a potential engine of significant economic growth. Policy makers must address issues such as security and privacy but not restrict this opportunity. Consumers, government and companies all have a role to play in defining policy. Countries need to recognize that the treatment of data is ultimately an issue of national competitiveness.
2. Create Rules of the Road
The government should define which activities concerning data are legal or illegal. For activities that are legal, consumers should then define what is public or private in their own cases through opting in or out in terms of what they choose to reveal about themselves.
3. Enact Cybersecurity Standards
The private sector, government, the military and consumers should jointly develop detailed standards and incidence-reporting practices for cybersecurity, building on existing industry best practices.
4. Create a Global Data Treaty
The U.S. should play a leadership role in coordinating with international bodies to pass a treaty to establish clear, universal standards on data privacy and ownership.
It's a phenomenon commonly referred to as Big Data, and it has generated widespread debate over a host of issues. What should companies do with such data? How can it be used profitably? How should it be treated? Should there be laws governing its use? Internationally recognized safeguards for consumer privacy? And how do U.S. companies stay competitive as more countries learn how to process and interpret such data?
The Wall Street Journal's John Bussey moderated a task-force discussion about just such issues. Here are edited excerpts of their presentations to the CEO Council:
Opportunity First
JOHN BUSSEY: Our group ended up with essentially two principles and two action items. Dominic, if you could take our first, please?
DOMINIC BARTON: Our group felt very strongly that this is a huge opportunity, and we shouldn't focus on how to protect or regulate Big Data before we recognize how important it is. Companies that use Big Data effectively get about a 6% productivity improvement versus others.
We're at the early stages. It's in every sector. And we all felt it can be the next wave of productivity growth if we use this data effectively.
Depending on how you protect or use data, it can actually lead to a country's competitive advantage. We've got small city-states like Singapore and Abu Dhabi that are allowing, for example, consumer medical information to be publicly available to get innovation.
So let's not lose sight of how important it can be for productivity improvement as we think about the protection.
MR. BUSSEY: Marcelo, our second principle.
MARCELO CLAURE: First of all, we had a fascinating group. We had a tremendous amount of interaction. And we quickly realized that all of us live in a hyper-connected world. By that I mean we're passing a huge amount of data every day through a connected car, a connected cellphone, a connected tablet, a connected home.
Big Data: The Top Four Recommendations
1. Big Data Is Opportunity
Industry already recognizes that the advent of Big Data is a potential engine of significant economic growth. Policy makers must address issues such as security and privacy but not restrict this opportunity. Consumers, government and companies all have a role to play in defining policy. Countries need to recognize that the treatment of data is ultimately an issue of national competitiveness.
2. Create Rules of the Road
The government should define which activities concerning data are legal or illegal. For activities that are legal, consumers should then define what is public or private in their own cases through opting in or out in terms of what they choose to reveal about themselves.
3. Enact Cybersecurity Standards
The private sector, government, the military and consumers should jointly develop detailed standards and incidence-reporting practices for cybersecurity, building on existing industry best practices.
4. Create a Global Data Treaty
The U.S. should play a leadership role in coordinating with international bodies to pass a treaty to establish clear, universal standards on data privacy and ownership.
Companies around the world increasingly collect and process vast amounts of customer data, particularly data tracking consumer behavior on the Web.
We see it as a tremendous opportunity, but what is the government role?
Some members of our group argued that government shouldn't be allowed to regulate or do anything related to Big Data. But the rest of the group agreed that we've got to define the basics of what the government is going to do.
And where we came out at the end was, government should have the role of defining what is legal and what is illegal in the use of Big Data. And that's where the government should stop. We think it should be left up to the consumer to define what he or she wants to share and doesn't want to share, what is public and what is private.
Limiting government control to what's legal and illegal will allow the consumer to make more choices. As a consumer, you should be able to determine how much of your Facebook profile you want to share, or whether you want to share information about where you shopped last.
Security and Treaty
MR. BUSSEY: And Tim has our two action items.
TIM ARMSTRONG: The first one is probably more important than it seems on the surface, especially for private companies. I think we started this as kind of a government conversation, but mostly it's private companies that actually own the infrastructure that makes the company or the country work.
So cybersecurity is something that really needs to be dealt with. And specifically, having standards around this is important in two different directions.
One is standards of what happens when something gets attacked. How do you report it, and what is the infrastructure to help companies with that?
The second piece which is important and which came up in our discussion is when your data lives outside the country, or when you're going to do things with data in other countries. There are groups like CFIUS [the U.S. government's Committee on Foreign Investment in the U.S., which reviews foreign investment deemed to have an effect on national security] which are inside the government, which, if you haven't dealt with them, you will, over these type of data assets. Cybersecurity is really important.
The last action item is coming up with a global data treaty. The U.S. is probably the largest economy involved in Big Data right now, and we think it's important for the U.S. to take a leading role in defining what some of the Big Data policies and standards should be in such a treaty. We would hope that we, the U.S., could define a basic treaty to start off, and then other countries could either adopt or augment the treaty over time.
Politicians' Policy Decisions May Stymie Tools That Got Them Elected
Kate Kaye, Ad Age, November 16, 2012
One of the keys to success for President Barack Obama's reelection bid was its masterful use of data. But lost in the hype is this: The administration supports a browser-based do not track system that, if pervasive, would throw a wrench into the data-collection tactics that empowered the campaign.
Even today BarackObama.com features data-tracking cookies from several online ad and analytics firms.
The Mitt Romney and Obama campaigns spent hundreds of thousands of dollars in 2012 on data and related services to enhance their own voter contact information, inform their online and offline messaging and target ads. At the same time, Congress is inspecting the practices of firms that buy, sell and filter consumer data for corporate marketers.
"The Obama administration and the GOP should confront head-on the privacy issues raised by [their] far-reaching use of digital profiling and targeting data," argued privacy advocate Jeffrey Chester, founder of the Center for Digital Democracy. "It would be unfortunate for the administration's work to advance Do Not Track and other key safeguards if they failed to tackle the use of powerful data targeting technologies by political campaigns."
Industry and privacy wonks actually agree
It's a rare occurrence, but both Mr. Chester and the ad industry are in agreement on one thing: They both appreciate the attention the Obama data machine is getting. Privacy groups want to raise awareness of data collection and usage in the hopes of generating public support for curbing what they see as an increasingly infiltrative violation of personal privacy by marketers and the mushrooming data industry.
"Protecting the privacy of consumers and citizens should require policymakers from both sides to confront the civil liberties implications of what has been unleashed," added Mr. Chester, noting that the 2012 campaigns should divulge what data they collected, how they targeted ads and what will happen to the information now that the election is over.
Industry players, especially their Capitol Hill lobbyists, aim to convince legislators that the very data practices some of them criticize are helping them and their colleagues win races.
"Big data isn't going to help Todd Aken," said Mike Zaneis, general counsel of the Interactive Advertising Bureau, referring to the disgraced Congressman from Missouri who lost his Senate campaign after claiming women can ward off pregnancy resulting from "legitimate rape." Continued Mr. Zaneis, "But the Obama campaign used a lot of online data and a tremendous amount of offline data to go precinct-by-precinct to get-out-the-vote."
Third-party tags
More than a week after the election, BarackObama.com houses an array of third-party tags that track users for ad targeting and campaign and site analytics. Yesterday, around fifteen ad company tags were surfaced by Evidon's Ghostery software, including tags from BlueKai, which calls itself a "big data activation solution," and Appnexus, which among other things allows advertisers to use a variety of user behavioral data to target ads to those users on Facebook.
Both the Obama and Romney campaigns used social-media-widget and data provider ShareThis to target fundraising ads and identify issues and trends swing state voters were interested in, according to ShareThis CEO Kurt Abrahamson. The company tracks when people visit web pages and share them on Twitter, Facebook, LinkedIn or other popular social sites and allows advertisers to target ads using that anonymized information.
Clashing goals of campaigning and governing
Data tracking tools and techniques that have helped legislators on both sides of the aisle build supporter lists, generate donations and get out the vote could be stymied by a do-not-track browser standard or restrictive privacy legislation.
In February, the Federal Trade Commission and the ad industry announced they'd work together with browser companies to develop a DNT standard. At the same time, the U.S. Commerce Department introduced a consumer privacy bill of rights that guided companies to provide individual control over data collection, better data security measures, and transparency of data use, and also called for "a reasonable amount of data collection by companies." Secretary of Commerce John Bryson said at the time the department would work with Congress to implement the privacy bill of righs -- which some deem to be supportive of industry's self-regulatory approach -- through legislation.
The Digital Advertising Alliance, a large coalition of ad industry trade groups, has conducted an "ongoing dialogue with the FTC as recently as yesterday to figure out how to implement the [DNT] standard," said Stu Ingis, counsel to the DAA, on Wednesday. The DAA oversees the industry's Ad Choices program, which allows people to opt-out from online ad targeting through display ads that include the group's small triangular symbol. It's not entirely clear whether the FTC is confident that the DAA's self-regulatory program is enough to protect consumer privacy.
As reported by Politico earlier this month, FTC Chairman Jon Leibowitz said, "If by the end of the year or early next year, we haven't seen a real Do Not Track option for consumers, I suspect the commission will go back and think about whether we want to endorse legislation." Mr. Leibowitz is expected by beltway insiders to step down at the end of the year, and some believe his goal to finalize a DNT standard before he leaves is pressurizing the situation.
A free pass for political data?
Enter the Bipartisan Congressional Privacy Caucus. The group recently received responses to inquiries into several data firms that manage and analyze, and in some cases buy and sell, online and offline consumer data. Nine firms -- Acxiom, Epsilon, Equifax, Experian, Harte-Hanks, Intelius, Fair Isaac, Merkle, and Meredith Corp. -- submitted lengthy and often vague answers to a series of questions about their data businesses and practices.
"Many questions about how these data brokers operate have been left unanswered, particularly how they analyze personal information to categorize and rate consumers," said lawmakers in a joint statement regarding the companies' responses.
Absent from the list of data firms questioned were similar companies that deal mainly in voter file and political information that is often enhanced with consumer demographic, shopping and other data. For instance, NGP Van, the Democratic data powerhouse favored by the Obama team was not part of the inquiry. The Obama campaign and DNC spent hundreds of thousands of dollars with NGP Van this election cycle alone. The firm matches its voter data with data from TargetSmart, which offers "the richest set of consumer and interest data, allowing the most sophisticated targeting," according to the NGP Van site.
Other political data firms left out of the inquiry include Catalist, another Democratic data firm; Campaign Grid, which offers Republican data and online ad targeting; and Aristotle, a well-established non-partisan political data company. People involved with the congressional inquiry deny that political data firms were left off the list for any strategic reason.
In a press release about the data broker responses, the Privacy Caucus stated it "will push for whatever steps are necessary to make sure Americans know how this industry operates and are granted control over their own information."
Rep. Ed Markey, a Democrat from Massachusetts and Caucus co-chair, has sponsored a Do Not Track Kids Act and a mobile privacy bill.
Observers don't expect a privacy bill to be passed anytime soon; if that does happen, it may not apply to political campaigns or groups anyway. For instance, political messages are exempt from CAN-SPAM laws, and political organizations are not restricted by the Do Not Call Registry.
"Often when data laws are being proposed and put forward, the politicians exempt themselves," said Don Hinman, senior VP for data strategy at Epsilon, which gets some of its data from political advertisers but mainly is a purveyor of consumer information.
Mr. Ingis considers it exemption for political messages to be a first amendment issue. "It would be very hard for such a limitation on political messages to be restricted. . . . and I think that would have been true in the context of Do Not Call if they would have gone there," he said.
Tarun Wadhwa, Forbes, November 6, 2012
If your favorite soda is Diet Dr. Pepper, the chances are that you’ll be supporting Mitt Romney. Pepsi drinker? You’re most likely voting for Barack Obama. If you drink Mountain Dew, you probably don’t care either way.
These types of conclusions may seem simplistic and superficial, but both campaigns are betting that they will be the key to deciding who the next President of the United States is.
It’s more than what you drink, what you shop for, who your friends are, what websites you visit: all reveal clues to your political leanings. Campaigns have entered the era of “Big Data”—they target voters based on scraps of information they gather from unlikely places.
Thanks to the rise of mobile technology and social media, the number of records collected by data brokers on voter behavior has tripled—from 300 pieces in 2004 to more than 900 pieces today.
Campaigns care about your personal life
Voters used to be the ones obsessing over details of a candidate’s personal life. Now the tables have turned. Campaigns research the personal lives of the voter.
Micro-targeting, a technique that delivers ads based on the personal traits of a voter, was once considered impossible. But in 2004, it was recognized for helping George W. Bush defeat John Kerry. Now it is used by almost every campaign.
Because of the intricacies of our electoral system, a relatively small group of people ends up deciding the outcome of elections. In the 2000 Presidential campaign, hundreds of millions of dollars was spent on reaching just 7 percent of voters—fewer than 8 million people. Even a small advantage in mobilizing potential voters in a swing state can determine the difference between a win and a loss.
In this election cycle, more than $3 billion dollars has been spent on broadcast-television advertising, which has remained the dominant form of political communication for the last fifty years. But times are changing. Television purchases are no longer as effective as they used to be. A study showed that 88% of voters with DVRs skip ads and that 45% use something other than live TV as their primary mode for viewing videos. These proportions are even higher in younger demographics.
The next frontier: digital behavioral advertising
Just as television advertising revolutionized the field in the 1960s, this election will likely mark digital-behavioral advertising as the next frontier in voter outreach.
As a nation, we are already divided along partisan lines. We access different media, each with its own messaging and focus. Now we will receive different messages depending on who we are. Zac Moffatt, digital director for Mitt Romney’s campaign, said to The New York Times that “two people in the same house could get different messages,” and that “not only would the message change, the type of content would change.”
In an article for Stanford Law Review, Daniel Kreiss, a journalism professor at University of North Carolina, Chapel Hill, explains how this can have negative long-term consequences for democratic participation. With so much sensitive personal information in so many hands, there are risks of data breaches and unauthorized disclosure.
Citizens may hesitate to engage in political discussion on line for fear of being tagged and put into a marketing database. And the high cost of political data and consulting activities may make it difficult for less affluent candidates to compete effectively. Perhaps most worrying, campaigns may “redline” an electorate (by ignoring voters who won’t be sympathetic to their views because a model deems them unworthy of investment).
Political targeting – what’s next?
Sophisticated modeling and targeting will become commonplace at every step of the political process. NGOs, interest groups, and candidates for local office will be the next to adopt these methods.
United in Purpose, an evangelical Christian non-profit, is currently using such technology to assign points to voters based on whether they like NASCAR or fishing, and whether they are on anti-abortion or traditional marriage lists. If these voters have a score of over 600 points, they are considered “serious about their faith”. They will be contacted if they have not registered to vote.
Many voters would be surprised to learn that their interactions with both campaigns are being recorded and analyzed using technology similar to what Target uses to determine whether teenage girls are pregnant. When voters do learn what their candidates are doing, as many as 86 percent want this to stop. They regard it as an invasion of privacy. Yet these types of activities are legally considered political speech, so there are hardly any restrictions in place.
What is most worrisome is that there is no easy way to opt out of these databases, or to limit what information is collected about you, or how it is used. Sadly, we can’t “de-friend” or “unfollow” the politicians.
Tim O'Reilly, O'Reilly Radar, October 16, 2012
I’m convinced that there’s a wave of innovation coming in healthcare, driven by new kinds of data, new ways of extracting meaning from that data, and new business models that data can enable. That’s one of the reasons why we launched our StrataRx Conference, which focuses on the importance of data science to the future of health care.
Unfortunately, much of the data that will enable an entrepreneurial explosion is still locked up — in paper records, in proprietary data formats, and by well-intentioned but conflicting privacy regulations.
We’re making progress towards open data in healthcare, but there are still so many obstacles! Ann Waldo recently introduced me to one of these.
A 2009 law modernized patient access rights by allowing individuals to get copies of their medical records in electronic format. Unfortunately, however, these patients’ access rights surprisingly do not include lab test results – one of the types of medical records that people are most likely to find urgent and useful. Due to the interaction of HIPAA (the Federal medical privacy law), CLIA (a Federal laboratory regulatory law), and state laws, patients can only get direct access to their their test results from labs in a handful of states.
A recent New York Times story highlighted just how much pain and suffering can be caused by this inability to get access to your own lab results.
In 2011, the Department of Health and Human Services put forward a proposed Rule that would give patients the right to get their test results directly from laboratories. This Rule is still waiting to be finalized. In hopes of breaking the logjam, O’Reilly Media and a variety of other players have written a consensus letter that voices our whole-hearted support for that proposed Rule and encourages the Federal government to finalize it promptly.
We’d love to invite you to join us in signing this letter.
Patients’ rights should include direct access to their lab results, just like all their other medical records!
Alexis Madrigal, The Atlantic, October 2012
Here's a pocket history of the web, according to many people. In the early days, the web was just pages of information linked to each other. Then along came web crawlers that helped you find what you wanted among all that information. Some time around 2003 or maybe 2004, the social web really kicked into gear, and thereafter the web's users began to connect with each other more and more often. Hence Web 2.0, Wikipedia, MySpace, Facebook, Twitter, etc. I'm not strawmanning here. This is the dominant history of the web as seen, for example, in this Wikipedia entry on the 'Social Web.'
tl;dr version
1. The sharing you see on sites like Facebook and Twitter is the tip of the 'social' iceberg. We are impressed by its scale because it's easy to measure.
2. But most sharing is done via dark social means like email and IM that are difficult to measure.
3. According to new data on many media sites, 69% of social referrals came from dark social. 20% came from Facebook.
4. Facebook and Twitter do shift the paradigm from private sharing to public publishing. They structure, archive, and monetize your publications.
But it's never felt quite right to me. For one, I spent most of the 90s as a teenager in rural Washington and my web was highly, highly social. We had instant messenger and chat rooms and ICQ and USENET forums and email. My whole Internet life involved sharing links with local and Internet friends. How was I supposed to believe that somehow Friendster and Facebook created a social web out of what was previously a lonely journey in cyberspace when I knew that this has not been my experience? True, my web social life used tools that ran parallel to, not on, the web, but it existed nonetheless.
To be honest, this was a very difficult thing to measure. One dirty secret of web analytics is that the information we get is limited. If you want to see how someone came to your site, it's usually pretty easy. When you follow a link from Facebook to The Atlantic, a little piece of metadata hitches a ride that tells our servers, "Yo, I'm here from Facebook.com." We can then aggregate those numbers and say, "Whoa, a million people came here from Facebook last month," or whatever.
There are circumstances, however, when there is no referrer data. You show up at our doorstep and we have no idea how you got here. The main situations in which this happens are email programs, instant messages, some mobile applications*, and whenever someone is moving from a secure site ("https://mail.google.com/blahblahblah") to a non-secure site (http://www.theatlantic.com).
This means that this vast trove of social traffic is essentially invisible to most analytics programs. I call it DARK SOCIAL. It shows up variously in programs as "direct" or "typed/bookmarked" traffic, which implies to many site owners that you actually have a bookmark or typed in www.theatlantic.com into your browser. But that's not actually what's happening a lot of the time. Most of the time, someone Gchatted someone a link, or it came in on a big email distribution list, or your dad sent it to you.
Nonetheless, the idea that "social networks" and "social media" sites created a social web is pervasive. Everyone behaves as if the traffic your stories receive from the social networks (Facebook, Reddit, Twitter, StumbleUpon) is the same as all of your social traffic. I began to wonder if I was wrong. Or at least that what I had experienced was a niche phenomenon and most people's web time was not filled with Gchatted and emailed links. I began to think that perhaps Facebook and Twitter has dramatically expanded the volume of -- at the very least -- linksharing that takes place.
Everyone else had data to back them up. I had my experience as a teenage nerd in the 1990s. I was not about to shake social media marketing firms with my tales of ICQ friends and the analogy of dark social to dark energy. ("You can't see it, dude, but it's what keeps the universe expanding. No dark social, no Internet universe, man! Just a big crunch.")
And then one day, we had a meeting with the real-time web analytics firm, Chartbeat. Like many media nerds, I love Chartbeat. It lets you know exactly what's happening with your stories, most especially where your readers are coming from. Recently, they made an accounting change that they showed to us. They took visitors who showed up without referrer data and split them into two categories. The first was people who were going to a homepage (theatlantic.com) or a subject landing page (theatlantic.com/politics). The second were people going to any other page, that is to say, all of our articles. These people, they figured, were following some sort of link because no one actually types "http://www.theatlantic.com/technology/archive/2012/10/atlast-the-gargantuan-telescope-designed-to-find-life-on-other-planets/263409/." They started counting these people as what they call direct social.
The second I saw this measure, my heart actually leapt (yes, I am that much of a data nerd). This was it! They'd found a way to quantify dark social, even if they'd given it a lamer name!
On the first day I saw it, this is how big of an impact dark social was having on The Atlantic.
Just look at that graph. On the one hand, you have all the social networks that you know. They're about 43.5 percent of our social traffic. On the other, you have this previously unmeasured darknet that's delivering 56.5 percent of people to individual stories. This is not a niche phenomenon! It's more than 2.5x Facebook's impact on the site.
Day after day, this continues to be true, though the individual numbers vary a lot, say, during a Reddit spike or if one of our stories gets sent out on a very big email list or what have you. Day after day, though, dark social is nearly always our top referral source.
Perhaps, though, it was only The Atlantic for whatever reason. We do really well in the social world, so maybe we were outliers. So, I went back to Chartbeat and asked them to run aggregate numbers across their media sites.
Get this. Dark social is even more important across this broader set of sites. Almost 69 percent of social referrals were dark! Facebook came in second at 20 percent. Twitter was down at 6 percent.
All in all, direct/dark social was 17.5 percent of total referrals; only search at 21.5 percent drove more visitors to this basket of sites. (FWIW, at The Atlantic, social referrers far outstrip search. I'd guess the same is true at all the more magaziney sites.)
There are a couple of really interesting ramifications of this data. First, on the operational side, if you think optimizing your Facebook page and Tweets is "optimizing for social," you're only halfway (or maybe 30 percent) correct. The only real way to optimize for social spread is in the nature of the content itself. There's no way to game email or people's instant messages. There's no power users you can contact. There's no algorithms to understand. This is pure social, uncut.
Second, the social sites that arrived in the 2000s did not create the social web, but they did structure it. This is really, really significant. In large part, they made sharing on the Internet an act of publishing (!), with all the attendant changes that come with that switch. Publishing social interactions makes them more visible, searchable, and adds a lot of metadata to your simple link or photo post. There are some great things about this, but social networks also give a novel, permanent identity to your online persona. Your taste can be monetized, by you or (much more likely) the service itself.
Third, I think there are some philosophical changes that we should consider in light of this new data. While it's true that sharing came to the web's technical infrastructure in the 2000s, the behaviors that we're now all familiar with on the large social networks was present long before they existed, and persists despite Facebook's eight years on the web. The history of the web, as we generally conceive it, needs to consider technologies that were outside the technical envelope of "webness."
People layered communication technologies easily and built functioning social networks with most of the capabilities of the web 2.0 sites in semi-private and without the structure of the current sites.
If what I'm saying is true, then the tradeoffs we make on social networks is not the one that we're told we're making. We're not giving our personal data in exchange for the ability to share links with friends. Massive numbers of people -- a larger set than exists on any social network -- already do that outside the social networks. Rather, we're exchanging our personal data in exchange for the ability to publish and archive a record of our sharing. That may be a transaction you want to make, but it might not be the one you've been told you made.
* Chartbeat datawiz Josh Schwartz said it was unlikely that the mobile referral data was throwing off our numbers here. "Only about four percent of total traffic is on mobile at all, so, at least as a percentage of total referrals, app referrals must be a tiny percentage," Schwartz wrote to me in an email. "To put some more context there, only 0.3 percent of total traffic has the Facebook mobile site as a referrer and less than 0.1 percent has the Facebook mobile app."