Showing posts with label Data Mining. Show all posts
Showing posts with label Data Mining. Show all posts

Saturday, October 6, 2012

Predicting the Tuture through Online Data Mining


Santiago Zabala, Al Jazeera, October 5, 2012

It is often said philosophers are either late when it comes to comment upon new technological innovations or in advance, that is, so early that they actually seem to predict them. When they are late, it's usually because they prefer to carefully examine the new innovations in order to achieve insightful analysis, and when they foresee such discoveries, it arises from an ethical concern over the direction the world is taking.

In other words, their insight is not expressed by envisioning the day a software company manages to predict the future by scanning information from the internet and therefore framing our freedom, but rather by working through the existential consequences these innovations might have upon our life. 

This is probably why among Martin Heidegger's greatest concerns when it came to technological innovations was the formation of existential conditions where, as he said, the "lack of emergency is the only emergency". In this condition human beings would be completely "uprooted" from the earth, that is, "framed" ("Ge-stell") by a technological power they are no longer able to control.

As it turns out, the software company Recorded Future (which has recently been praised by Wired , the MIT Technology Review and other media outlets, after the CIA and Google invested millions in their services) seems to be offering its clients something similar: a world where emergencies, that is, future events, can be calculated in advance. But how does this start-up actually function, and why are the German philosopher’s concerns relevant to its services? 

Recorded Future
Recorded Future is based in Gothenburg and has offices in London, Boston, Arlington and New York. A team of 20 computer scientists, statisticians and experts in linguistics "calculate" the future. While Yahoo, Google and Bing use links to connect and rank different web pages, Recorded Future goes further by scouring (in real time) thousands of available information sources such as blogs, websites and Twitter comments in order to find "invisible links", that is, relationships among actions, people and institutions that refer to related events in the future. 

Even though this might not seem particularly relevant, considering that we can also predict next week's weather by searching through different weather stations, if we look at the amount of data this company is capable of analysing and relating in just a few hours, it becomes clear it can obtain better data than public internet users have access to.

The information we need to predict whether tomorrow it will rain is limited by the number of weather sites available, but sites that might refer to upcoming anti-American demonstrations in the Middle East are infinitely more numerous given the political, economic and military aspects of these sorts of events.

After mining from the web all the related people ("Bashar al-Assad"), places ("Syria") and activities ("military interventions") that refer to a possible demonstration, Recorded Future uses algorithms to predict when and where a demonstration will occur. 

An example of a predicted demonstration is available in a video on the company website which illustrates how its powerful engines monitor these protests not only in the Middle East, but also in South America and North Africa. The fact that Google and the CIA have already invested millions in this company is an indication that it will be used to conserve certain interest against others as the example above indicates.
The different fee levels for the customers of Recorded Future are probably related to the quality and quantity of information they wish to purchase, making this, and similar companies, at the service of the wealthiest and most powerful. 

Lack of emergencies
From a philosophical point of view, the most interesting feature of all this is not that these demonstrations can be predicted, but rather how technology has finally uprooted and dislodged man from the world, that is, has given human existence to a power beyond human control. The secured, comfortable and calculated environment that allowed the creation of society has now become so functional and rationalised that we cannot help but become victims by existing in it.

This existential dilemma does not arise from the fact that it’s finally possible to organise all the things the web already knows about the future, which could certainly become useful to prevent diseases or famine, but rather that a private company now means to know everything, that is, all human projects. 

We have entered an age where only those framed within the approved interests of Recorded Future clients will be able to live freely, that is, without being predicted. But how free is an existence that is completely revealed to the modern "lack of emergencies"?

As Heidegger explained , emergencies do not arise when something doesn't function correctly, but rather when "everything functions … and propels everything more and more toward further functioning". It's within this logic that as soon as something critical to the interests of those who can afford it fails to function, Recorded Future will alert its customers, who will then take the appropriate measures to conserve the previous condition.

In sum, Heidegger's concerns over a world lacking "emergencies" more than 50 years ago was meant to point out how technologies such as that employed by Recorded Future (and similar companies ) aim to avoid the future, that is, to change the world.

Santiago Zabala is ICREA Research Professor of Philosophy at the University of Barcelona. His books include The Hermeneutic Nature of Analytic Philosophy (2008), The Remains of Being (2009) and most recently, Hermeneutic Communism (2011, co-authored with G Vattimo), all published by Columbia University Press. 

Thursday, September 20, 2012

Big Data for All


Omer Tene, Concurring Opinions, September 20, 2012

Much has been written over the past couple of years about “big data” (See, for example, here and here and here). In a new article, Big Data for All: Privacy and User Control in the Age of Analytics, which will be published in the Northwestern Journal of Technology and Intellectual Property, Jules Polonetsky and I try to reconcile the inherent tension between big data business models and individual privacy rights. We argue that going forward, organizations should provide individuals with practical, easy to use access to their information, so they can become active participants in the data economy. In addition, organizations should be required to be transparent about the decisional criteria underlying their data processing activities.

The term “big data” refers to advances in data mining and the massive increase in computing power and data storage capacity, which have expanded by orders of magnitude the scope of information available for organizations. Data are now available for analysis in raw form, escaping the confines of structured databases and enhancing researchers’ abilities to identify correlations and conceive of new, unanticipated uses for existing information. In addition, the increasing number of people, devices, and sensors that are now connected by digital networks has revolutionized the ability to generate, communicate, share, and access data.

Data creates enormous value for the world economy, driving innovation, productivity, efficiency and growth. In the article, we flesh out some compelling use cases for big data analysis. Consider, for example, a group of medical researchers who were able to parse out a harmful side effect of a combination of medications, which were used daily by millions of Americans, by analyzing massive amounts of online search queries. Or scientists who analyze mobile phone communications to better understand the needs of people who live in settlements or slums in developing countries.

At the same time, the “data deluge” presents formidable privacy concerns. Protecting privacy become harder as information is multiplied and shared ever more widely among multiple parties around the world. As more information regarding individuals’ health, financials, location, electricity use and online activity percolates, concerns arise about profiling, tracking, discrimination, exclusion, government surveillance and loss of control. From a more technical legal angle, big data challenges some of the most fundamental concepts of privacy law, including the definition of “personally identifiable information”, the role of individual control, and the principles of data minimization and purpose limitation.

In our article, we make the case for providing individuals with usable access to their data. The call for transparency is not new, of course. Rather the emphasis is on access to data in usable format, which can work to create value to individuals. Transparency and access alone have not emerged as potent tools because individuals do not care for, and cannot afford to indulge in transparency and access for their own sake (see one oft-cited counterexample here). The enabler of transparency and access is the ability to use the information and benefit from it in a tangible way. This will be achieved through “featurization” or “app-ification” of privacy. Organizations should build as many dials and levers as needed for individuals to engage with their data.

We expect that “featurization” of big data, harnessing its immense force for not only organizational but also individual benefit, will unleash a wave of innovation and create a market for personal data applications. The technological groundwork has already been completed with mash-ups and real-time APIs making it easier for organizations to combine information from different sources and services into a single user experience. Regardless of lingering questions concerning who – if anyone – “owns” the information, we think that fairness dictates that individuals enjoy beneficial use of the data about them.

Our second proposal would require organizations to disclose the decisional criteria underpinning their data analytics machinery. In a big data world, it is often not the data but rather the inferences drawn from them that give cause for concern. Inaccurate, manipulative or discriminatory conclusions may be drawn from perfectly innocuous, accurate data. Much like in quantum physics, the observer in big data analysis can affect the results of her research by defining the data set, proposing a hypothesis or writing an algorithm. At the end of the day, big data analysis is an interpretative process, in which one’s identity and perspective informs one’s results. Like any interpretative process, it is subject to error, inaccuracy and bias. Louis Brandeis, who together with Samuel Warren “invented” the legal right to privacy in 1890, has also written that “[s]unlight is said to be the best of disinfectants”. We trust if the existence and uses of databases were visible to the public, organizations would be more likely to avoid unethical or socially unacceptable uses of data.

Wednesday, September 5, 2012

Brookings Report: Big Data for Education: Data Mining, Data Analytics, and Web Dashboards



Darrell M. West, The Brookings Institution, September 4, 2012

Imagine this scenario: twelve-year-old Susan took a course designed to improve her reading skills. She read short stories and the teacher would give her and her fellow students a written test every other week measuring vocabulary and reading comprehension. A few days later, Susan’s instructor graded the paper and returned her exam. The test showed that she did well on vocabulary, but needed to work on retaining key concepts.

In the future, her younger brother Richard is likely to learn reading through a computerized software program. As he goes through each story, the computer will collect data on how long it takes him to master the material. After each assignment, a quiz will pop up on his screen and ask questions concerning vocabulary and reading comprehension. As he answers each item, Richard will get instant feedback showing whether his answer is correct and how his performance compares to classmates and students across the country. For items that are difficult, the computer will send him links to websites that explain words and concepts in greater detail. At the end of the session, his teacher will receive an automated readout on Richard and the other students in the class summarizing their reading time, vocabulary knowledge, reading comprehension, and use of supplemental electronic resources.

In comparing these two learning environments, it is apparent that current school evaluations suffer from several limitations. Many of the typical pedagogies provide little immediate feedback to students, require teachers to spend hours grading routine assignments, aren’t very proactive about showing students how to improve comprehension, and fail to take advantage of digital resources that can improve the learning process. This is unfortunate because data-driven approaches make it possible to study learning in real-time and offer systematic feedback to students and teachers.

In this report, I examine the potential for improved research, evaluation, and accountability through data mining, data analytics, and web dashboards. So-called “big data” make it possible to mine learning information for insights regarding student performance and learning approaches.[1] Rather than rely on periodic test performance, instructors can analyze what students know and what techniques are most effective for each pupil. By focusing on data analytics, teachers can study learning in far more nuanced ways.[2] Online tools enable evaluation of a much wider range of student actions, such as how long they devote to readings, where they get electronic resources, and how quickly they master key concepts.


[1] James Manyika, Michael Chui, Brad Brown, Jacques Bughin, Richard Dobbs, Charles Roxburgh, and Angela Byers, “Big Data: The Next Frontier for Innovation, Competition, and Productivity,” McKinsey Global Institute, May, 2011.
[2] Felix Castro, Alfredo Vellido, Angela Nebot, and Francisco Mugica, “Applying Data Mining Techniques to e-Learning Problems,” Studies in Computational Intelligence, Volume 62, 2007, pp. 183-221.

Tuesday, August 21, 2012

Three kinds of big data

Looking ahead at big data's role in enterprise business intelligence, civil engineering, and customer relationship optimization.

Alistair Croll,  O'Reilly Radar,  August 21, 2012 
 
In the past couple of years, marketers and pundits have spent a lot of time labeling everything ”big data.” The reasoning goes something like this:

· Everything is on the Internet.

· The Internet has a lot of data.

· Therefore, everything is big data.

When you have a hammer, everything looks like a nail. When you have a Hadoop deployment, everything looks like big data. And if you’re trying to cloak your company in the mantle of a burgeoning industry, big data will do just fine. But seeing big data everywhere is a sure way to hasten the inevitable fall from the peak of high expectations to the trough of disillusionment.

We saw this with cloud computing. From early idealists saying everything would live in a magical, limitless, free data center to today’s pragmatism about virtualization and infrastructure, we soon took off our rose-colored glasses and put on welding goggles so we could actually build stuff.

So where will big data go to grow up?

Once we get over ourselves and start rolling up our sleeves, I think big data will fall into three major buckets: Enterprise BI, Civil Engineering, and Customer Relationship Optimization. This is where we’ll see most IT spending, most government oversight, and most early adoption in the next few years.

Enterprise BI 2.0

For decades, analysts have relied on business intelligence (BI) products like Hyperion, Microstrategy and Cognos to crunch large amounts of information and generate reports. Data warehouses and BI tools are great at answering the same question — such as “what were Mary’s sales this quarter?” — over and over again. But they’ve been less good at the exploratory, what-if, unpredictable questions that matter for planning and decision making because that kind of fast exploration of unstructured data is traditionally hard to do and therefore expensive.

Most “legacy” BI tools are constrained in two ways:

· First, they’ve been schema-then-capture tools in which the analyst decides what to collect, then later capture that data for analysis.

· Second, they’ve typically focused on reporting what Avinash Kaushik (channeling Donald Rumsfeld) refers to as “known unknowns” — things we know we don’t know, and generate reports for.

These tools are used for reporting and operational purposes, usually focused on controlling costs, executing against an existing plan, and reporting on how things are going.

As my Strata co-chair Edd Dumbill pointed out when I asked for thoughts on this piece:

“The predominant functional application of big data technologies today is in ETL (Extract, Transform, and Load). I’ve heard the figure that it’s about 80% of Hadoop applications. Just the real grunt work of log file or sensor processing before loading into an analytic database like Vertica.”

The availability of cheap, fast computers and storage, as well as open source tools, have made it okay to capture first and ask questions later. That changes how we use data because it makes it okay to speculate beyond the initial question that triggered the collection of data.

What’s more, the speed with which we can get results — sometimes as fast as a human can ask them — makes data easier to explore interactively. This combination of interactivity and speculation takes BI into the realm of “unknown unknowns,” the insights that can produce a competitive advantage or an out-of-the-box differentiator.

We saw this shift in cloud computing: first, big public clouds wooed green-field startups. Then, in a few years, incumbent IT vendors introduced their private cloud offerings. Private clouds included only a fraction of the benefits of public clouds, but were nevertheless a sufficient blend of smoke, mirrors, and features to delay the inevitable move to public resources by a few years and appease the business. For better or worse, that’s where most of IT budgets are being spent today according to IDC, Gartner, and others.

In the next few years, then, look for acquisitions and product introductions — and not a little vaporware — as BI vendors that enterprises trust bring them “big data lite”: enough to satisfy their CEO’s golf buddies, but not so much that their jobs are threatened. This, after all, is how change comes to big organizations.

Ultimately, we’ll see traditional “known unknowns” BI reporting living alongside big-data-powered data import and cleanup, and fast, exploratory data “unknown unknown” interactivity.

Civil Engineering

The second use of big data is in society and government. Already, data mining can be used to predict disease outbreaks, understand traffic patterns, and improve education.

Cities are facing budget crunches, infrastructure problems, and a crowding from rural citizens. Solving these problems is urgent, and cities are perfect labs for big data initiatives. Take a metropolis like New York: hackathons; open feeds of public data; and a population that generates a flood of information as it shops, commutes, gets sick, eats, and just goes about its daily life.



I think municipal data is one of the big three for several reasons: it’s a good tie breaker for partisanship, we have new interfaces everyone can understand, and we finally have a mostly-connected citizenry.

In an era of partisan bickering, hard numbers can settle the debate. So, they’re not just good government; they’re good politics. Expect to see big data applied to social issues, helping us to make funding more effective and scarce government resources more efficient (perhaps to the chagrin of some public servants and lobbyists). As this works in the world’s biggest cities, it’ll spread to smaller ones, to states, and to municipalities.

Making data accessible to citizens is possible, too: Siri and Google Now show the potential for personalized agents; Narrative Science takes complex data and turns it into words the masses can consume easily; Watson and Wolfram Alpha can give smart answers, either through curated reasoning or making smart guesses.

For the first time, we have a connected citizenry armed (for the most part) with smartphones. Nielsen estimated that smartphones would overtake feature phones in 2011, and that concentration is high in urban cores. The App Store is full of apps for bus schedules, commuters, local events, and other tools that can quickly become how governments connect with their citizens and manage their bureaucracies.

The consequence of all this, of course, is more data. Once governments go digital, their interactions with citizens can be easily instrumented and analyzed for waste or efficiency. That’s sure to provoke resistance from those who don’t like the scrutiny or accountability, but it’s a side effect of digitization: every industry that goes digital gets analyzed and optimized, whether it likes it or not.

Customer Relationship Optimization

The final home of applied big data is marketing. More specifically, it’s improving the relationship with consumers so companies can, as Sergio Zyman once said, sell them more stuff, more often, for more money, more efficiently.

The biggest data systems today are focused on web analytics, ad optimization, and the like. Many of today’s most popular architectures were weaned on ads and marketing, and have their ancestry in direct marketing plans. They’re just more focused than the comparatively blunt instruments with which direct marketers used to work.

The number of contact points in a company has multiplied significantly. Where once there was a phone number and a mailing address, today there are web pages, social media accounts, and more. Tracking users across all these channels — and turning every click, like, share, friend, or retweet into the start of a long funnel that leads, inexorably, to revenue is a big challenge. It’s also one that companies like Salesforce understand, with its investments in chat, social media monitoring, co-browsing, and more.

This is what’s lately been referred to as the “360-degree customer view” (though it’s not clear that companies will actually act on customer data if they have it, or whether doing so will become a compliance minefield). Big data is already intricately linked to online marketing, but it will branch out in two ways.

First, it’ll go from online to offline. Near-field-equipped smartphones with ambient check-in are a marketer’s wet dream, and they’re coming to pockets everywhere. It’ll be possible to track queue lengths, store traffic, and more, giving retailers fresh insights into their brick-and-mortar sales. Ultimately, companies will bring the optimization that online retail has enjoyed to an offline world as consumers become trackable.

Second, it’ll go from Wall Street (or maybe that’s Madison Avenue and Middlefield Road) to Main Street. Tools will get easier to use, and while small businesses might not have a BI platform, they’ll have a tablet or a smartphone that they can bring to their places of business. Mobile payment players like Square are already making them reconsider the checkout process. Adding portable customer intelligence to the tool suite of local companies will broaden how we use marketing tools.

Headlong into the trough

That’s my bet for the next three years, given the molasses of market confusion, vendor promises, and unrealistic expectations we’re about to contend with. Will big data change the world? Absolutely. Will it be able to defy the usual cycle of earnest adoption, crushing disappointment, and eventual rebirth all technologies must travel? Certainly not.


O'Reilly Radar (http://s.tt/1lhOc)

Monday, July 16, 2012

The Way the Digital Cookie Crumbles

If regulators and lawyers limit the use of data, advertising online will become less efficient. 

 

L. Gordon Crovitz, The Wall Street Journal, July 15, 2012.

For a measure of how technology is changing human expectations, consider the "cookies" on your computers. These invisible text files are how websites track activity, delivering to marketers detailed information about individual behavior and preferences. In exchange for data, we get highly personalized online services. 

This use of cookies fuels the economics of the Web, but it has also caused anxiety as people have had to reconsider analog-era expectations of privacy to embrace digital-era benefits of sharing data. A Wall Street Journal report last week caused some consternation when it revealed how the travel website Orbitz uses data to give different offers to people who use Apple computers and those using Windows-based machines. 

Data analysts at Orbitz detected patterns showing that Apple users spend up to 30% more a night on hotels and are likelier to book four- or five-star lodgings than PC users. Apple users buy more expensive computers, and the average household income for adult owners of Mac computers is almost $100,000, compared with about $75,000 for PC owners, according to Forrester. 

When Orbitz used these data to feature higher-priced hotels more prominently in Apple users' search results, privacy lobbyists claimed outrage. But even in the analog era, readers of this newspaper saw advertisements for different products and services than readers of less high-end papers. 

It wouldn't be surprising if data showed less price sensitivity among Apple users, so they could be offered higher prices, but at least for now Orbitz shows the same prices for the same rooms regardless of how users access the site. Price-customization software is being used by many retailers so that online buyers who click directly to pay for products do not get special offers shown to less-persuaded shoppers.

These uses of personal data can seem a bit creepy, but the evidence also shows how quickly consumers have gotten used to being tracked. When given the choice, few consumers opt out of cookies. People accept the benefits of more relevant ads and more personalized websites in exchange for letting marketers track their interests. 

There is an enormous industry in predictive analytics and "big data." Consumers are loyal to Amazon in part because of its recommendation tools—if you liked that book, you may like this one—which mine user data to determine relevancy. Facebook says it will deliver targeted ads based on what other websites and apps users access. Apple tells users it will target ads to them based on apps they download. Google delivers advertising based on how people use its various services, including what they write in messages sent via Gmail.

Left alone, people would continue to make their own evolving judgments about how much data to share. Instead, regulators issue edicts. The Federal Trade Commission has extracted 20-year consent decrees from Google, Facebook, Twitter and Myspace, giving regulators broad review over their privacy and data practices. This would be fine if the purpose were to ensure that companies comply with disclosures about how they use data, but the FTC wants to define privacy standards. 

One result of FTC meddling is that plaintiff lawyers have open invitations to file nuisance suits on behalf of supposed privacy victims. A federal judge is considering a $20 million settlement offer by Facebook, which has agreed to make its disclosures clearer that when users click "Like" to promote a product on Facebook, their names and photos can be used. 

The $20 million would be divided equally between plaintiff lawyers and privacy interest groups such as the Electronic Frontier Foundation. Nothing would go to the allegedly harmed 900 million Facebook users.

"The plaintiff's lawyers get rich, class members get little and nonprofit groups often reap millions by urging judges to approve the deal regardless of its merits," Wired magazine reported last week on the Facebook settlement. This case "provides a glimpse into the dark side of large class-action settlements."

If regulators and lawyers push too hard to limit the use of cookie data, advertising online will become less efficient. This in turn will reduce the amount of free, advertising-supported services enjoyed by consumers, such as social media, entertainment and email. 

Consumers seem to understand there's no such thing as a free lunch, even online: If they are not paying for a product, then for better or worse, they are the product. Each consumer should be able to decide how to make this trade-off between sharing data and getting advertising-supported services.

The privacy debate shows how naive Silicon Valley firms were to sign 20-year agreements granting Washington regulators broad authority over how they operate. Digital entrepreneurs should be allowed to innovate freely, with consumers also free to choose their individual trade-off between how their data are used and the benefits they get in return. Overregulation is the way the digital cookie crumbles. 

A version of this article appeared July 16, 2012, on page A11 in the U.S. edition of The Wall Street Journal, with the headline: The Way the Digital Cookie Crumbles.

Monday, June 18, 2012

Privacy and Big Data

You for Sale: Mapping, and Sharing, the Consumer Genome

Natasha Singer, The New York Times, June 17, 2012

IT knows who you are. It knows where you live. It knows what you do.
It peers deeper into American life than the F.B.I. or the I.R.S., or those prying digital eyes at Facebook and Google. If you are an American adult, the odds are that it knows things like your age, race, sex, weight, height, marital status, education level, politics, buying habits, household health worries, vacation dreams — and on and on.

Right now in Conway, Ark., north of Little Rock, more than 23,000 computer servers are collecting, collating and analyzing consumer data for a company that, unlike Silicon Valley’s marquee names, rarely makes headlines. It’s called the Acxiom Corporation, and it’s the quiet giant of a multibillion-dollar industry known as database marketing.

Few consumers have ever heard of Acxiom. But analysts say it has amassed the world’s largest commercial database on consumers — and that it wants to know much, much more. Its servers process more than 50 trillion data “transactions” a year. Company executives have said its database contains information about 500 million active consumers worldwide, with about 1,500 data points per person. That includes a majority of adults in the United States.

Such large-scale data mining and analytics — based on information available in public records, consumer surveys and the like — are perfectly legal. Acxiom’s customers have included big banks like Wells Fargo and HSBC, investment services like E*Trade, automakers like Toyota and Ford, department stores like Macy’s — just about any major company looking for insight into its customers.

For Acxiom, based in Little Rock, the setup is lucrative. It posted profit of $77.26 million in its latest fiscal year, on sales of $1.13 billion.

But such profits carry a cost for consumers. Federal authorities say current laws may not be equipped to handle the rapid expansion of an industry whose players often collect and sell sensitive financial and health information yet are nearly invisible to the public. In essence, it’s as if the ore of our data-driven lives were being mined, refined and sold to the highest bidder, usually without our knowledge — by companies that most people rarely even know exist.

Julie Brill, a member of the Federal Trade Commission, says she would like data brokers in general to tell the public about the data they collect, how they collect it, whom they share it with and how it is used. “If someone is listed as diabetic or pregnant, what is happening with this information? Where is the information going?” she asks. “We need to figure out what the rules should be as a society.”
 Although Acxiom employs a chief privacy officer, Jennifer Barrett Glasgow, she and other executives declined requests to be interviewed for this article, said Ines Rodriguez Gutzmer, director of corporate communications.

In March,  however, Ms. Barrett Glasgow  endorsed increased industry openness. “It’s not an unreasonable request to have more transparency among data brokers,” she said in an interview with The New York Times.  In marketing materials, Acxiom promotes itself as “a global thought leader in addressing consumer privacy issues and earning the public trust.”

But, in interviews, security experts and consumer advocates paint a portrait of a company with practices that privilege corporate clients’ interests over those of consumers and contradict the company’s stance on transparency. Acxiom’s marketing materials, for example, promote a special security system for clients and associates to encrypt the data they send. Yet cybersecurity experts who examined Acxiom’s Web site for The Times found basic security lapses on an online form for consumers seeking access to their own profiles. (Acxiom says it has fixed the broken link that caused the problem.)

In a fast-changing digital economy, Acxiom is developing even more advanced techniques to mine and refine data. It has recruited talent from Microsoft, Google, Amazon.com and Myspace and is using a powerful, multiplatform approach to predicting consumer behavior that could raise its standing among investors and clients.

 Of course, digital marketers already customize pitches to users, based on their past activities. Just think of “cookies,” bits of computer code placed on browsers to keep track of online activity. But Acxiom, analysts say, is pursuing far more comprehensive techniques in an effort to influence consumer decisions. It is integrating what it knows about our offline, online and even mobile selves, creating in-depth behavior portraits in pixilated detail. Its executives have called this approach a “360-degree view” on consumers.

 “There’s a lot of players in the digital space trying the same thing,” says Mark Zgutowicz, a Piper Jaffray analyst. “But Acxiom’s advantage is they have a database of offline information that they have been collecting for 40 years and can leverage that expertise in the digital world.”

 Yet some prominent privacy advocates worry that such techniques could lead to a new era of consumer profiling.

Jeffrey Chester, executive director of the Center for Digital Democracy, a nonprofit group in Washington, says: “It is Big Brother in Arkansas.”

 SCOTT HUGHES, an up-and-coming small-business owner and Facebook denizen, is Acxiom’s ideal consumer. Indeed, it created him.

 Mr. Hughes is a fictional character who appeared in an Acxiom investor presentation in 2010. A frequent shopper, he was designed to show the power of Acxiom’s multichannel approach.
 In the presentation, he logs on to Facebook and sees that his friend Ella has just become a fan of Bryce Computers, an imaginary electronics retailer and Acxiom client. Ella’s update prompts Mr. Hughes to check out Bryce’s fan page and do some digital window-shopping for a fast inkjet printer.
Such browsing seems innocuous — hardly data mining. But it cues an Acxiom system designed to recognize consumers, remember their actions, classify their behaviors and influence them with tailored marketing.

When Mr. Hughes follows a link to Bryce’s retail site, for example, the system recognizes him from his Facebook activity and shows him a printer to match his interest. He registers on the site, but doesn’t buy the printer right away, so the system tracks him online. Lo and behold, the next morning, while he scans baseball news on ESPN.com, an ad for the printer pops up again.

That evening, he returns to the Bryce site where, the presentation says, “he is instantly recognized” as having registered. It then offers a sweeter deal: a $10 rebate and free shipping.

 It’s not a random offer. Acxiom has its own classification system, PersonicX, which assigns consumers to one of 70 detailed socioeconomic clusters and markets to them accordingly. In this situation, it pegs Mr. Hughes as a “savvy single” — meaning he’s in a cluster of mobile, upper-middle-class people who do their banking online, attend pro sports events, are sensitive to prices — and respond to free-shipping offers.

 Correctly typecast, Mr. Hughes buys the printer.

But the multichannel system of Acxiom and its online partners is just revving up. Later, it sends him coupons for ink and paper, to be redeemed via his cellphone, and a personalized snail-mail postcard suggesting that he donate his old printer to a nearby school.

 Analysts say companies design these sophisticated ecosystems to prompt consumers to volunteer enough personal data — like their names, e-mail addresses and mobile numbers — so that marketers can offer them customized appeals any time, anywhere.

Still, there is a fine line between customization and stalking. While many people welcome the convenience of personalized offers, others may see the surveillance engines behind them as intrusive or even manipulative.

“If you look at it in cold terms, it seems like they are really out to trick the customer,” says Dave Frankland, the research director for customer intelligence at Forrester Research. “But they are actually in the business of helping marketers make sure that the right people are getting offers they are interested in and therefore establish a relationship with the company.”

DECADES before the Internet as we know it, a businessman named Charles Ward planted the seeds of Acxiom. It was 1969, and Mr. Ward started a data processing company in Conway called Demographics Inc., in part to help the Democratic Party reach voters. In a time when Madison Avenue was deploying one-size-fits-all national ad campaigns, Demographics and its lone computer used public phone books to compile lists for direct mailing of campaign material.

Today, Acxiom maintains its own database on about 190 million individuals and 126 million households in the United States. Separately, it manages customer databases for or works with 47 of the Fortune 100 companies. It also worked with the government after the September 2001 terrorist attacks, providing information about 11 of the 19 hijackers.

To beef up its digital services, Acxiom recently mounted an aggressive hiring campaign. Last July, it named Scott E. Howe, a former corporate vice president for Microsoft’s advertising business group, as C.E.O. Last month, it hired Phil Mui, formerly group product manager for Google Analytics, as its chief product and engineering officer.

In interviews, Mr. Howe has laid out a vision of Acxiom as a new-millennium “data refinery” rather than a data miner. That description posits Acxiom as a nimble provider of customer analytics services, able to compete with Facebook and Google, rather than as a stealth engine of consumer espionage.
Still, the more that information brokers mine powerful consumer data, the more they become attractive targets for hackers — and draw scrutiny from consumer advocates.

This year, Advertising Age ranked Epsilon, another database marketing firm, as the biggest advertising agency in the United States, with Acxiom second. Most people know Epsilon, if they know it at all, because it experienced a major security breach last year, exposing the e-mail addresses of millions of customers of Citibank, JPMorgan Chase, Target, Walgreens and others. In 2003, Acxiom had its own security breaches.

 But privacy advocates say they are more troubled by data brokers’ ranking systems, which classify some people as high-value prospects, to be offered marketing deals and discounts regularly, while dismissing others as low-value — known in industry slang as “waste.”

 Exclusion from a vacation offer may not matter much, says Pam Dixon, the executive director of the World Privacy Forum, a nonprofit group in San Diego, but if marketing algorithms judge certain people as not worthy of receiving promotions for higher education or health services, they could have a serious impact.

 "Over time, that can really turn into a mountain of pathways not offered, not seen and not known about,” Ms. Dixon says.

Until now, database marketers operated largely out of the public eye. Unlike consumer reporting agencies that sell sensitive financial information about people for credit or employment purposes, database marketers aren’t required by law to show consumers their own reports and allow them to correct errors. That may be about to change. This year, the F.T.C. published a report calling for greater transparency among data brokers and asking Congress to give consumers the right to access information these firms hold about them.

ACXIOM’S Consumer Data Products Catalog offers hundreds of details — called “elements” — that corporate clients can buy about individuals or households, to augment their own marketing databases. Companies can buy data to pinpoint households that are concerned, say, about allergies, diabetes or “senior needs.” Also for sale is information on sizes of home loans and household incomes.
 Clients generally buy this data because they want to hold on to their best customers or find new ones — or both.

 A bank that wants to sell its best customers additional services, for example, might buy details about those customers’ social media, Web and mobile habits to identify more efficient ways to market to them. Or, says Mr. Frankland at Forrester, a sporting goods chain whose best customers are 25- to 34-year-old men living near mountains or beaches could buy a list of a million other people with the same characteristics. The retailer could hire Acxiom, he says, to manage a campaign aimed at that new group, testing how factors like consumers’ locations or sports preferences affect responses.

But the catalog also offers delicate information that has set off alarm bells among some privacy advocates, who worry about the potential for misuse by third parties that could take aim at vulnerable groups. Such information includes consumers’ interests — derived, the catalog says, “from actual purchases and self-reported surveys” — like “Christian families,” “Dieting/Weight Loss,” “Gaming-Casino,” “Money Seekers” and “Smoking/Tobacco.” Acxiom also sells data about an individual’s race, ethnicity and country of origin. “Our Race model,” the catalog says, “provides information on the major racial category: Caucasians, Hispanics, African-Americans, or Asians.” Competing companies sell similar data.

Acxiom’s data about race or ethnicity is “used for engaging those communities for marketing purposes,” said Ms. Barrett Glasgow, the privacy officer, in an e-mail response to questions.

There may be a legitimate commercial need for some businesses, like ethnic restaurants, to know the race or ethnicity of consumers, says Joel R. Reidenberg, a privacy expert and a professor at the Fordham Law School.

“At the same time, this is ethnic profiling,” he says. “The people on this list, they are being sold based on their ethnic stereotypes. There is a very strong citizen’s right to have a veto over the commodification of their profile.”

He says the sale of such data is troubling because race coding may be incorrect. And even if a data broker has correct information, a person may not want to be marketed to based on race.
“DO you really know your customers?” Acxiom asks in marketing materials for its shopper recognition system, a program that uses ZIP codes to help retailers confirm consumers’ identities — without asking their permission.

 “Simply asking for name and address information poses many challenges: transcription errors, increased checkout time and, worse yet, losing customers who feel that you’re invading their privacy,” Acxiom’s fact sheet explains. In its system, a store clerk need only “capture the shopper’s name from a check or third-party credit card at the point of sale and then ask for the shopper’s ZIP code or telephone number.” With that data Acxiom can identify shoppers within a 10 percent margin of error, it says, enabling stores to reward their best customers with special offers. Other companies offer similar services.

“This is a direct way of circumventing people’s concerns about privacy,” says Mr. Chester of the Center for Digital Democracy.

Ms. Barrett Glasgow of Acxiom says that its program is a “standard practice” among retailers, but that the company encourages its clients to report consumers who wish to opt out.
Acxiom has positioned itself as an industry leader in data privacy, but some of its practices seem to undermine that image. It created the position of chief privacy officer in 1991, well ahead of its rivals. It even offers an online request form, promoted as an easy way for consumers to access information Acxiom collects about them.

But the process turned out to be not so user-friendly for a reporter for The Times.

In early May, the reporter decided to request her record from Acxiom, as any consumer might. Before submitting a Social Security number and other personal information, however, she asked for advice from a cybersecurity expert at The Times. The expert examined Acxiom’s Web site and immediately noticed that the online form did not employ a standard encryption protocol — called https — used by sites like Amazon and American Express. When the expert tested the form, using software that captures data sent over the Web, he could clearly see that the sample Social Security number he had submitted had not been encrypted. At that point, the reporter was advised not to request her file, given the risk that the process might expose her personal information.

Later in May, Ashkan Soltani, an independent security researcher and former technologist in identity protection at the F.T.C., also examined Acxiom’s site and came to the same conclusion. “Parts of the site for corporate clients are encrypted,” he says. “But for consumers, who this information is about and who stand the most to lose from data collection, they don’t provide security.”

Ms. Barrett Glasgow says that the form has always been encrypted with https but that on May 11, its security monitoring system detected a “broken redirect link” that allowed unencrypted access. Since then, she says, Acxiom has fixed the link and determined that no unauthorized person had gained access to information sent using the form.

On May 25, the reporter submitted an online request to Acxiom for her file, along with a personal check, sent by Express Mail, for the $5 processing fee. Three weeks later, no response had arrived.
Regulators at the F.T.C. declined to comment on the practices of individual companies. But Jon Leibowitz, the commission chairman, said consumers should have the right to see and correct personal details about them collected and sold by data aggregators.

After all, he said, “they are the unseen cyberazzi who collect information on all of us.”

Wednesday, June 13, 2012

Big Data: What Facebook Knows

What Facebook Knows

The company's social scientists are hunting for insights about human behavior. What they find could give Facebook new ways to cash in on our data—and remake our view of society.

Tom Simonite, Technology Review, July/August 2012

Laws haven't kept up with the company's ability to mine its users' data.
If Facebook were a country, a conceit that founder Mark Zuckerberg has entertained in public, its 900 million members would make it the third largest in the world. 

It would far outstrip any regime past or present in how intimately it records the lives of its citizens. Private conversations, family photos, and records of road trips, births, marriages, and deaths all stream into the company's servers and lodge there. Facebook has collected the most extensive data set ever assembled on human social behavior. Some of your personal information is probably part of it. 

And yet, even as Facebook has embedded itself into modern life, it hasn't actually done that much with what it knows about us. Now that the company has gone public, the pressure to develop new sources of profit (see "The Facebook Fallacy") is likely to force it to do more with its hoard of information. That stash of data looms like an oversize shadow over what today is a modest online advertising business, worrying privacy-conscious Web users (see "Few Privacy Regulations Inhibit Facebook") and rivals such as Google. Everyone has a feeling that this unprecedented resource will yield something big, but nobody knows quite what.
Even as Facebook has embedded itself into modern life, it hasn't done that much with what it knows about us. Its stash of data looms like an oversize shadow. Everyone has a feeling that this resource will yield something big, but nobody knows quite what.

Heading Facebook's effort to figure out what can be learned from all our data is Cameron Marlow, a tall 35-year-old who until recently sat a few feet away from Zuckerberg. The group Marlow runs has escaped the public attention that dogs Facebook's founders and the more headline-grabbing features of its business. Known internally as the Data Science Team, it is a kind of Bell Labs for the social-networking age. The group has 12 researchers—but is expected to double in size this year. They apply math, programming skills, and social science to mine our data for insights that they hope will advance Facebook's business and social science at large. Whereas other analysts at the company focus on information related to specific online activities, Marlow's team can swim in practically the entire ocean of personal data that Facebook maintains. Of all the people at Facebook, perhaps even including the company's leaders, these researchers have the best chance of discovering what can really be learned when so much personal information is compiled in one place. 

Facebook has all this information because it has found ingenious ways to collect data as people socialize. Users fill out profiles with their age, gender, and e-mail address; some people also give additional details, such as their relationship status and mobile-phone number. A redesign last fall introduced profile pages in the form of time lines that invite people to add historical information such as places they have lived and worked. Messages and photos shared on the site are often tagged with a precise location, and in the last two years Facebook has begun to track activity elsewhere on the Internet, using an addictive invention called the "Like" button. It appears on apps and websites outside Facebook and allows people to indicate with a click that they are interested in a brand, product, or piece of digital content. Since last fall, Facebook has also been able to collect data on users' online lives beyond its borders automatically: in certain apps or websites, when users listen to a song or read a news article, the information is passed along to Facebook, even if no one clicks "Like." Within the feature's first five months, Facebook catalogued more than five billion instances of people listening to songs online. Combine that kind of information with a map of the social connections Facebook's users make on the site, and you have an incredibly rich record of their lives and interactions. 

"This is the first time the world has seen this scale and quality of data about human communication," Marlow says with a characteristically serious gaze before breaking into a smile at the thought of what he can do with the data. For one thing, Marlow is confident that exploring this resource will revolutionize the scientific understanding of why people behave as they do. His team can also help Facebook influence our social behavior for its own benefit and that of its advertisers. This work may even help Facebook invent entirely new ways to make money. 

Contagious Information
Marlow eschews the collegiate programmer style of Zuckerberg and many others at Facebook, wearing a dress shirt with his jeans rather than a hoodie or T-shirt. Meeting me shortly before the company's initial public offering in May, in a conference room adorned with a six-foot caricature of his boss's dog spray-painted on its glass wall, he comes across more like a young professor than a student. He might have become one had he not realized early in his career that Web companies would yield the juiciest data about human interactions.
In 2001, undertaking a PhD at MIT's Media Lab, Marlow created a site called Blogdex that automatically listed the most "contagious" information spreading on weblogs. Although it was just a research project, it soon became so popular that Marlow's servers crashed. Launched just as blogs were exploding into the popular consciousness and becoming so numerous that Web users felt overwhelmed with information, it prefigured later aggregator sites such as Digg and Reddit. But Marlow didn't build it just to help Web users track what was popular online. Blogdex was intended as a scientific instrument to uncover the social networks forming on the Web and study how they spread ideas. Marlow went on to Yahoo's research labs to study online socializing for two years. In 2007 he joined Facebook, which he considers the world's most powerful instrument for studying human society. "For the first time," Marlow says, "we have a microscope that not only lets us examine social behavior at a very fine level that we've never been able to see before but allows us to run experiments that millions of users are exposed to." 

Marlow's team works with managers across Facebook to find patterns that they might make use of. For instance, they study how a new feature spreads among the social network's users. They have helped Facebook identify users you may know but haven't "friended," and recognize those you may want to designate mere "acquaintances" in order to make their updates less prominent. Yet the group is an odd fit inside a company where software engineers are rock stars who live by the mantra "Move fast and break things." Lunch with the data team has the feel of a grad-student gathering at a top school; the typical member of the group joined fresh from a PhD or junior academic position and prefers to talk about advancing social science than about Facebook as a product or company. Several members of the team have training in sociology or social psychology, while others began in computer science and started using it to study human behavior. They are free to use some of their time, and Facebook's data, to probe the basic patterns and motivations of human behavior and to publish the results in academic journals—much as Bell Labs researchers advanced both AT&T's technologies and the study of fundamental physics. 

It may seem strange that an eight-year-old company without a proven business model bothers to support a team with such an academic bent, but Marlow says it makes sense. "The biggest challenges Facebook has to solve are the same challenges that social science has," he says. Those challenges include understanding why some ideas or fashions spread from a few individuals to become universal and others don't, or to what extent a person's future actions are a product of past communication with friends. Publishing results and collaborating with university researchers will lead to findings that help Facebook improve its products, he adds. 

For one example of how Facebook can serve as a proxy for examining society at large, consider a recent study of the notion that any person on the globe is just six degrees of separation from any other. The best-known real-world study, in 1967, involved a few hundred people trying to send postcards to a particular Boston stockholder. Facebook's version, conducted in collaboration with researchers from the University of Milan, involved the entire social network as of May 2011, which amounted to more than 10 percent of the world's population. Analyzing the 69 billion friend connections among those 721 million people showed that the world is smaller than we thought: four intermediary friends are usually enough to introduce anyone to a random stranger. "When considering another person in the world, a friend of your friend knows a friend of their friend, on average," the technical paper pithily concluded. That result may not extend to everyone on the planet, but there's good reason to believe that it and other findings from the Data Science Team are true to life outside Facebook. Last year the Pew Research Center's Internet & American Life Project found that 93 percent of Facebook friends had met in person. One of Marlow's researchers has developed a way to calculate a country's "gross national happiness" from its Facebook activity by logging the occurrence of words and phrases that signal positive or negative emotion. Gross national happiness fluctuates in a way that suggests the measure is accurate: it jumps during holidays and dips when popular public figures die. After a major earthquake in Chile in February 2010, the country's score plummeted and took many months to return to normal. That event seemed to make the country as a whole more sympathetic when Japan suffered its own big earthquake and subsequent tsunami in March 2011; while Chile's gross national happiness dipped, the figure didn't waver in any other countries tracked (Japan wasn't among them). Adam Kramer, who created the index, says he intended it to show that Facebook's data could provide cheap and accurate ways to track social trends—methods that could be useful to economists and other researchers. 

Other work published by the group has more obvious utility for Facebook's basic strategy, which involves encouraging us to make the site central to our lives and then using what it learns to sell ads. An early study looked at what types of updates from friends encourage newcomers to the network to add their own contributions. Right before Valentine's Day this year a blog post from the Data Science Team listed the songs most popular with people who had recently signaled on Facebook that they had entered or left a relationship. It was a hint of the type of correlation that could help Facebook make useful predictions about users' behavior—knowledge that could help it make better guesses about which ads you might be more or less open to at any given time. Perhaps people who have just left a relationship might be interested in an album of ballads, or perhaps no company should associate its brand with the flood of emotion attending the death of a friend. The most valuable online ads today are those displayed alongside certain Web searches, because the searchers are expressing precisely what they want. This is one reason why Google's revenue is 10 times Facebook's. But Facebook might eventually be able to guess what people want or don't want even before they realize it. 

Recently the Data Science Team has begun to use its unique position to experiment with the way Facebook works, tweaking the site—the way scientists might prod an ant's nest—to see how users react. Eytan Bakshy, who joined Facebook last year after collaborating with Marlow as a PhD student at the University of Michigan, wanted to test whether our Facebook friends create an "echo chamber" that amplifies news and opinions we have already heard about. So he messed with how Facebook operated for a quarter of a billion users. Over a seven-week period, the 76 million links that those users shared with each other were logged.

Then, on 219 million randomly chosen occasions, Facebook prevented someone from seeing a link shared by a friend. Hiding links this way created a control group so that Bakshy could assess how often people end up promoting the same links because they have similar information sources and interests. 

He found that our close friends strongly sway which information we share, but overall their impact is dwarfed by the collective influence of numerous more distant contacts—what sociologists call "weak ties." It is our diverse collection of weak ties that most powerfully determines what information we're exposed to. 

That study provides strong evidence against an idea nagging many people: that social networking creates harmful "filter bubbles," to use activist Eli Pariser's term for the effects of tuning the information we receive to match our expectations. But the study also reveals the power Facebook has. "If [Facebook's] News Feed is the thing that everyone sees and it controls how information is disseminated, it's controlling how information is revealed to society, and it's something we need to pay very close attention to," Marlow says. He points out that his team helps Facebook understand what it is doing to society and publishes its findings to fulfill a public duty to transparency. Another recent study, which investigated which types of Facebook activity cause people to feel a greater sense of support from their friends, falls into the same category. 

Facebook is not above using its platform to tweak users' behavior, as it did by nudging them to register as organ donors. Unlike academic social scientists, Facebook's employees have a short path from an idea to an experiment on hundreds of millions of people.
But Marlow speaks as an employee of a company that will prosper largely by catering to advertisers who want to control the flow of information between its users. And indeed, Bakshy is working with managers outside the Data Science Team to extract advertising-related findings from the results of experiments on social influence. "Advertisers and brands are a part of this network as well, so giving them some insight into how people are sharing the content they are producing is a very core part of the business model," says Marlow. 

Facebook told prospective investors before its IPO that people are 50 percent more likely to remember ads on the site if they're visibly endorsed by a friend. Figuring out how influence works could make ads even more memorable or help Facebook find ways to induce more people to share or click on its ads. 

Social Engineering
Marlow says his team wants to divine the rules of online social life to understand what's going on inside Facebook, not to develop ways to manipulate it. "Our goal is not to change the pattern of communication in society," he says. "Our goal is to understand it so we can adapt our platform to give people the experience that they want." But some of his team's work and the attitudes of Facebook's leaders show that the company is not above using its platform to tweak users' behavior. Unlike academic social scientists, Facebook's employees have a short path from an idea to an experiment on hundreds of millions of people.
In April, influenced in part by conversations over dinner with his med-student girlfriend (now his wife), Zuckerberg decided that he should use social influence within Facebook to increase organ donor registrations. Users were given an opportunity to click a box on their Timeline pages to signal that they were registered donors, which triggered a notification to their friends. The new feature started a cascade of social pressure, and organ donor enrollment increased by a factor of 23 across 44 states.
Marlow's team is in the process of publishing results from the last U.S. midterm election that show another striking example of Facebook's potential to direct its users' influence on one another. Since 2008, the company has offered a way for users to signal that they have voted; Facebook promotes that to their friends with a note to say that they should be sure to vote, too. Marlow says that in the 2010 election his group matched voter registration logs with the data to see which of the Facebook users who got nudges actually went to the polls. (He stresses that the researchers worked with cryptographically "anonymized" data and could not match specific users with their voting records.)
This is just the beginning. By learning more about how small changes on Facebook can alter users' behavior outside the site, the company eventually "could allow others to make use of Facebook in the same way," says Marlow. If the American Heart Association wanted to encourage healthy eating, for example, it might be able to refer to a playbook of Facebook social engineering. "We want to be a platform that others can use to initiate change," he says.
Advertisers, too, would be eager to know in greater detail what could make a campaign on Facebook affect people's actions in the outside world, even though they realize there are limits to how firmly human beings can be steered. "It's not clear to me that social science will ever be an engineering science in a way that building bridges is," says Duncan Watts, who works on computational social science at Microsoft's recently opened New York research lab and previously worked alongside Marlow at Yahoo's labs. "Nevertheless, if you have enough data, you can make predictions that are better than simply random guessing, and that's really lucrative."
Doubling Data
Like other social-Web companies, such as Twitter, Facebook has never attained the reputation for technical innovation enjoyed by such Internet pioneers as Google. If Silicon Valley were a high school, the search company would be the quiet math genius who didn't excel socially but invented something indispensable. Facebook would be the annoying kid who started a club with such social momentum that people had to join whether they wanted to or not. In reality, Facebook employs hordes of talented software engineers (many poached from Google and other math-genius companies) to build and maintain its irresistible club. The technology built to support the Data Science Team's efforts is particularly innovative. The scale at which Facebook operates has led it to invent hardware and software that are the envy of other companies trying to adapt to the world of "big data." 

In a kind of passing of the technological baton, Facebook built its data storage system by expanding the power of open-source software called Hadoop, which was inspired by work at Google and built at Yahoo. Hadoop can tame seemingly impossible computational tasks—like working on all the data Facebook's users have entrusted to it—by spreading them across many machines inside a data center. But Hadoop wasn't built with data science in mind, and using it for that purpose requires specialized, unwieldy programming. Facebook's engineers solved that problem with the invention of Hive, open-source software that's now independent of Facebook and used by many other companies. Hive acts as a translation service, making it possible to query vast Hadoop data stores using relatively simple code. To cut down on computational demands, it can request random samples of an entire data set, a feature that's invaluable for companies swamped by data. Much of Facebook's data resides in one Hadoop store more than 100 petabytes (a million gigabytes) in size, says Sameet Agarwal, a director of engineering at Facebook who works on data infrastructure, and the quantity is growing exponentially. "Over the last few years we have more than doubled in size every year," he says. That means his team must constantly build more efficient systems. 

One potential use of Facebook's data storehouse would be to sell insights mined from it. Such information could be the basis for any kind of business. Assuming Facebook can do this without upsetting users and regulators, it could be lucrative.

All this has given Facebook a unique level of expertise, says Jeff Hammerbacher, Marlow's predecessor at Facebook, who initiated the company's effort to develop its own data storage and analysis technology. (He left Facebook in 2008 to found Cloudera, which develops Hadoop-based systems to manage large collections of data.) Most large businesses have paid established software companies such as Oracle a lot of money for data analysis and storage. But now, big companies are trying to understand how Facebook handles its enormous information trove on open-source systems, says Hammerbacher. "I recently spent the day at Fidelity helping them understand how the 'data scientist' role at Facebook was conceived ... and I've had the same discussion at countless other firms," he says. 

As executives in every industry try to exploit the opportunities in "big data," the intense interest in Facebook's data technology suggests that its ad business may be just an offshoot of something much more valuable. The tools and techniques the company has developed to handle large volumes of information could become a product in their own right. 

Mining for Gold
Facebook needs new sources of income to meet investors' expectations. Even after its disappointing IPO, it has a staggeringly high price-to-earnings ratio that can't be justified by the barrage of cheap ads the site now displays. Facebook's new campus in Menlo Park, California, previously inhabited by Sun Microsystems, makes that pressure tangible. The company's 3,500 employees rattle around in enough space for 6,600. I walked past expanses of empty desks in one building; another, next door, was completely uninhabited. A vacant lot waited nearby, presumably until someone invents a use of our data that will justify the expense of developing the space. 

One potential use would be simply to sell insights mined from the information. DJ Patil, data scientist in residence with the venture capital firm Greylock Partners and previously leader of LinkedIn's data science team, believes Facebook could take inspiration from Gil Elbaz, the inventor of Google's AdSense ad business, which provides over a quarter of Google's revenue. He has moved on from advertising and now runs a fast-growing startup, Factual, that charges businesses to access large, carefully curated collections of data ranging from restaurant locations to celebrity body-mass indexes, which the company collects from free public sources and by buying private data sets. Factual cleans up data and makes the result available over the Internet as an on-demand knowledge store to be tapped by software, not humans. 

Customers use it to fill in the gaps in their own data and make smarter apps or services; for example, Facebook itself uses Factual for information about business locations. Patil points out that Facebook could become a data source in its own right, selling access to information compiled from the actions of its users. Such information, he says, could be the basis for almost any kind of business, such as online dating or charts of popular music. Assuming Facebook can take this step without upsetting users and regulators, it could be lucrative. An online store wishing to target its promotions, for example, could pay to use Facebook as a source of knowledge about which brands are most popular in which places, or how the popularity of certain products changes through the year. 

Hammerbacher agrees that Facebook could sell its data science and points to its currently free Insights service for advertisers and website owners, which shows how their content is being shared on Facebook. That could become much more useful to businesses if Facebook added data obtained when its "Like" button tracks activity all over the Web, or demographic data or information about what people read on the site. There's precedent for offering such analytics for a fee: at the end of 2011 Google started charging $150,000 annually for a premium version of a service that analyzes a business's Web traffic. 

Back at Facebook, Marlow isn't the one who makes decisions about what the company charges for, even if his work will shape them. Whatever happens, he says, the primary goal of his team is to support the well-being of the people who provide Facebook with their data, using it to make the service smarter. Along the way, he says, he and his colleagues will advance humanity's understanding of itself. That echoes Zuckerberg's often doubted but seemingly genuine belief that Facebook's job is to improve how the world communicates. Just don't ask yet exactly what that will entail. "It's hard to predict where we'll go, because we're at the very early stages of this science," says Marlow. "The number of potential things that we could ask of Facebook's data is enormous." 

Tom Simonite is Technology Review's senior IT editor.