Paul Taylor, The Financial Times, June 19, 2012
Companies are awash with data, some generated by their customers or systems, some by third parties. These data are growing so fast – by about 2.5 exabytes a day – that 90 per cent of the stored data in the world today has been created in just the past two years, earning it the geeky moniker “big data”.
For the uninitiated, one exabyte is 1bn gigabytes.
Some of this material is traditional structured data, such as store point-of-sale information, bank cash machine transactions or mobile phone records. But much of it is unstructured information gleaned from non-traditional sources such as blogs, Facebook posts, tweets, email messages, smartphone apps, electronic sensors, pictures and YouTube video clips.
This high velocity, high volume and high variety make big data difficult to interpret using traditional database and data analytics, says Cyrus Mewawalla, an independent investment researcher.
But by combing through it with more sophisticated data analytics tools and techniques such as in-memory computing, it can provide companies with a better understanding of their customers and partners. It can help them spot trends in near real time, make more accurate forecasts and adjust their operations quickly to changing demand or new business opportunities.
“Companies benefit from a multidimensional view of their business when they add insight from big data to the traditional types of information they collect and analyse,” says Mr Mewawalla.
For example, he says, a company that operates a retail website can use big data to understand site visitors’ activities, such as paths through the site, pages viewed and comments posted. This knowledge can be combined with purchasing history. From this, the company gains a better understanding of customers and can fine-tune offers to target their interests.
Since companies such as Google, Amazon and Facebook pioneered the collection, processing and analysis of big data, it has become one of the hottest trends in corporate information technology alongside cloud computing, mobility and enterprise social networking.
WinterCorp, a specialist big data consultancy, says: “As enterprises harness big data, they are discovering opportunities better to understand and predict the interests and behaviour of their customers, especially in connection with ecommerce and social networking,”.
Meanwhile high-performance analytics is helping industries from banking to retail, healthcare and insurance to glean insights from big data that once took days or weeks in just hours, minutes or seconds.
In engineering and manufacturing, for example, companies are finding new opportunities to predict maintenance problems, enhance manufacturing quality and manage costs using big data. In healthcare, there are new opportunities to predict and react more rapidly to critical clinical events, resulting in better care for patients and more effective cost management.
Many retailers are already using big data analytics to improve the accuracy of forecasts, anticipate changes in demand and then react accordingly. For example, Brooks Brothers, one of the oldest retailers in the US, introduced business analytics developed by SAS, the US-based business analytics software group, to help improve stock management.
Using analytics to forecast accurately its global stock, store managers were able to make better decisions about stock levels and pricing. As a result, the number of times stores were out of stock when customers came in to buy an item were reduced, and stock decreased by 27 per cent.
In the financial services sector, most investment banks still rely on overnight batch data to make trading decisions. This means their risk management models constantly rely on out-of-date data. By using big data analytics in real time, banks can make better trading and risk decisions, safeguarding them against the threat of collapse and, subsequently, protecting the financial markets.
Using technology from SAS, air traffic controllers at Frankfurt airport in Germany receive early warnings of storms, and managers can access an overview of all key performance indicators in near real time, including average times for luggage delivery, delays and airport security levels.
All of this happens on the go – business data are refreshed every five minutes and both managers and operations experts monitor reports via their PCs and mobile devices.
Even governments, including those of the UK and US, are jumping on the big data bandwagon. A recent study undertaken by SAS and the Centre for Economics and Business Research, the UK-based think-tank, suggested that if the UK government capitalised on big data it could save £2bn in fraud detection, create 2,000 new jobs and generate £3.6bn in savings through better management of processes by, for example, integrating patient data to improve healthcare IT systems.
In March, the US government and six federal agencies launched their own big data initiative backed by a $200m investment. Calling it one of the most important public investments in technology since the rise of supercomputing and the internet, the White House Office of Science and Technology Policy (OSTP) said the investment was aimed at “greatly improving the tools and techniques needed to access, organise and glean discoveries from huge volumes of digital data”.
“The way we look at big data is that it is a confluence of three technology trends: big transaction data, big interaction data and big data processing,” says Sohaib Abbasi, chief executive of Informatica, whose software products help companies to clean, integrate and sift through huge volumes of data.
“Big data is not just about the volume of data,” he says, “it is also about new types and sources of data that can be used to gain new insight and deliver business advantage.”
For years, he explains, corporate IT departments have managed transactional data held within relational databases. “The promise of big data is to do better analysis of transactional data – the more data, the more reliable the data and the higher the quality of the analysis,” he says. But while this has been growing in scale and complexity, there is now an additional source of data that enterprises need to take notice: big interaction data.
This is a new type of data that represents social media interactions (human-generated interactions) and machine interactions (device-generated interactions). In both cases, the data are extremely large and continuously growing, and exist both within and beyond the corporate security firewall. “The challenge for businesses is how to understand and extract useful intelligence from these complex, unstructured data sources.”
T-Mobile USA, Deutsche Telekom’s US-based mobile unit, has used Informatica’s PowerCenter to integrate big data across its disparate federated architecture and predict customer defections based on the analysis of its 33m customer data records, web logs, billing data and social media information. By combining big transaction and big interaction data, T-Mobile has gained a better view of the reasons behind customer defections, which it was able to cut in half in a single quarter.
Similarly, US Xpress, the US trucking company, collects 900 data elements from tens of thousands of trucking systems: sensor data for tyre and fuel usage and engine operation, geospatial data for fleet tracking, and complaints posted on trucker blogs. Using Hadoop, a type of open-source database that is often used for big data projects, and Informatica, US Xpress processes and analyses this data to optimise fleet usage, saving millions of dollars a year.
It is not just big companies that are using big data. As McLaren’s Formula One cars speed round the track they send a stream of data back to the team that are processed and analysed in real time using SAP’s Hana in-memory technology. Hana uses sophisticated data compression to store information in Ram, which is 10,000 times faster than hard disks, enabling analysis of the data in seconds rather than hours.
This real-time analysis of car sensor data is compared with historical data and predictive models, helping the McLaren team to make immediate proactive corrections, avoid costly, dangerous incidents and win races.
“On every lap of every Grand Prix, practice or test session, our cars generate vast quantities of performance data. Our ability to process that data and act on it rapidly is crucial to creating the kind of prescriptive intelligence that enables us to transform the outcome of races. And that need resonates through every other facet of our business,” says Ron Dennis, executive chairman of McLaren.
As McLaren diversifies its business, its electronic systems, including telemetry, modelling and real-time simulations, are being used widely in other areas, such as to help Olympic athletes hone their performance and in rapid transit systems in the US to optimise traffic flow.
By harnessing big data, organisations can improve operational efficiency, reduce data management costs and better manage brands and customer relationships. But some IT leaders and analysts warn that few organisations are equipped to capitalise on the business value of big data, and that by neglecting to adapt effectively to big data, organisations also invite unforeseen cost, complexity and risk.
“Just collecting and storing big data doesn’t drive a cent of value to an organisation’s bottom line,” says Stephen Brobst, chief technology officer at Teradata, the analytics company.
So while many businesses are already capturing big data from sources such as web logs, machine data and text, and inexpensively storing it in open source “Hadoop” systems, complexity and incompatibility issues often make it difficult for them to use standard business intelligence applications and tools to access and analyse the many types of data.
“Business users are clamouring for more access to big data analytics,” says Scott Gnau, president of Teradata Labs, which recently rolled out new software called Aster SQL-H, which is designed to make it easier for business analysts to use raw, multi-structured data in Hadoop files to develop new insights for competitive advantage.
Jim Hagemann Snabe, SAP’s co-chief executive, believes that transforming information into intelligence in real time is increasingly critical for the future of every company. Similarly Informatica’s Mr Abbasi suggests that the wealth of new data now available to companies can bring unprecedented business opportunity.
Whether big data becomes an organisation’s greatest asset or one of its gravest liabilities depends on the strategies and solutions it puts in place to deal with the epic growth in data volumes, complexity, diversity and velocity.
This message seems to be getting through. In a global survey of 600 executives this month by Capgemini and the Economist Intelligence Unit, nine out of 10 respondents identified data as being the fourth factor of production – as fundamental to business as land, labour and capital.
Among the survey’s other findings, respondents said the use of big data has improved businesses’ performance, on average, by 26 per cent and that the impact will grow to 41 per cent over the next three years.
Almost 60 per cent of companies said they planned to make a bigger investment in big data over the next three years, suggesting that the era of big data and big data analytics has already arrived.
Copyright The Financial Times Limited 2012. You may share using our article tools.
Please don't cut articles from FT.com and redistribute by email or post to the web.
Wednesday, June 20, 2012
Tuesday, June 19, 2012
The Era of Big Data Is Here
Tom Silva, Huffington Post, June 18, 2012
As we see the recovery sputtering because of the European debt crisis, we await the next big thing. What will power this economy the way returning servicemen, the housing boom, urbanization and the Keynesianism of presidents from Eisenhower to Nixon powered the U.S. from the 1950s to the early 1970s? Or the way liberalization, low oil prices and the tech boom created 21 million jobs in the 1990s? A clue may lie in my industry. As our economy has moved from manufacturing to service-based over the last century, commercial real estate has traced the same arc, mutating from a sector focused largely around industrial buildings to one that's about the high-rise and suburban offices that dot our commutes home. But recently, a new sector and catchphrase has emerged that indicates a major new spur in our country's growth: Big Data.
Our nation's publicly traded REITs now see cellular towers and data centers as the stars of their real estate portfolios (along with high-end malls and apartments). CALPERS, the largest public pension fund in the U.S., recently announced that it is setting up a $500 million fund to invest in data centers. McKinsey & Company predicts a 40 percent growth annually in the data being generated. Among companies of more than 1,000 employees in 15 out of the economy's 17 sectors, the average amount of data is a surreal 235 terabytes. That's right -- each of these companies has more info than the Library of Congress. And so, why should we care? Because data is valuable. The growth of digital networks and the networked sensors in everything from phones to cars to heavy machinery mean that data has a reach and sweep it has never had before. The key to Big Data is connecting these sensors to computing intelligence which can make sense of all this information (in pure Wall-E style, some theorists call this the Internet of Things). The sexiest manifestation of all of this is natural-language processing, pattern recognition and machine learning, all of which is crystallized in Siri, the backtalking, mind-reading application in iPhones -- a kind of new-millennium sassy personal assistant.
Data Driven
So, why is there so much data out there? Firstly, there's the data companies generate as they go about their business called "exhaust data" (hello, Amazon); the updates and pictures of our vacations that we breathlessly upload onto our social networks; and then all of that multimedia content we stream. With 60 percent of the world owning a mobile phone (and 12 percent of them tapping a smartphone) this is a worldwide phenomenon that has enormous implications for U.S. companies and our economy. In health care, electronic health records have been vaunted as the best way to integrate care, predict the onset of disease, eliminate medical errors, and optimize follow-up care. McKinsey thinks data could be worth $300 billion to health care. Currently, health care providers throw away 90 percent of their data partly because they don't have a place to keep it. No surprise, then, that a venture of the Blue Cross Blue Shield Association (which has health care information on 110 million people) has raised more than $37 million to create an information warehouse with 3.5 billion pieces at a time when insurers are combing patient records for ways to cut costs and improve medical care. And, there are jobs to be had. The Bureau of Labor Statistics estimates that the number of health IT jobs across the country will increase by 20 percent from 2008 to 2018, a pace much faster than the average for all occupations through 2018. The same goes for retailers who McKinsey says could potentially increase their margins by as much as 60 percent. A report by the World Economic Forum in Davos, Switzerland, titled "Big Data, Big Impact" goes as far as to declare data a new class of economic asset, like currency or gold.
As we try to get our arms around all of this, we inevitably run into exponential math. Firstly, according to IDC, a technology research firm, the aggregate amount of data is growing at 50 percent a year, or more than doubling every two years. Then, there's Moore's Law, named after Intel co-founder Gordon Moore, which states that the amount of computing power that can be purchased for a certain amount of money doubles every two years. Also, the volume of business data worldwide, across all companies, doubles every 1.2 years, according to estimates. If all these numbers and hyperbole give you a touch of vertigo, you're not alone.
This is the dromosphere described by the French dystopian theorist Paul Virilio, which depicts modern society mapping the world in tiny fragments in an unending quest for speed and progress driven by technology. Putting aside the concerns about all this data, the hopes for Big Data are not unlike those for the Human Genome Project or for science itself -- a provable, positivist system to unlock age-old mysteries and a nifty way to raise profits. Incidentally, decoding the human genome originally took 10 years to process; now it can be achieved in one week.
What's in it for us?
Erik Brynjolfsson, an economist at Massachusetts Institute of Technology's Sloan School of Management, published a 2011 report with two colleagues that suggests that data-guided management is spreading across corporate America and starting to pay off. Looking at 179 large companies, they found that those adopting "data-driven decision-making" achieved productivity gains that were 5 percent to 6 percent higher than other factors could explain. Retailers now mine huge data sets to analyze sales, pricing, customer profiles, even weather data to tailor their pricing and markdowns and to make supply-chain decisions about how to get the product to the point of sale. But the more exciting frontier may be in solving humanitarian crises. As the World Economic Forum puts it, "By analyzing patterns from mobile phone usage, a team of researchers in San Francisco is able to predict the magnitude of a disease outbreak half way around the world. Similarly, an aid agency sees early warning signs of a drought condition in a remote Sub-Saharan region, allowing the agency to get a head start on mobilizing its resources and save many more lives."
The Forum sees Big Data affecting and intervening in education, agriculture, health and global finance. One of the most extraordinary stories emerged from Haiti after the 2010 earthquake when researchers at the Karolinska Institute and Columbia University obtained data on people fleeing Port-au-Prince by tracking nearly 2 million cell-phone SIM cards in the country. By reading the sensors, they were able to pinpoint the location of more than 600,000 people, and made this information available to government and humanitarian organizations. Later that year, the same team tracked the movements of people during a cholera outbreak allowing aid organizations to mobilize.
Despite concerns with privacy and fraud, you can't argue with results like that. Data centers continue to be built in places like Phoenix, North Carolina, Santa Clara and Northern Virginia. And the U.S. is short by close to 200,000 of people with the deep analytical skills that Big Data requires, according to McKinsey. And that doesn't include the hundreds of thousands of jobs to build, equip and manage these facilities. Big Data could be the next Big Thing
As we see the recovery sputtering because of the European debt crisis, we await the next big thing. What will power this economy the way returning servicemen, the housing boom, urbanization and the Keynesianism of presidents from Eisenhower to Nixon powered the U.S. from the 1950s to the early 1970s? Or the way liberalization, low oil prices and the tech boom created 21 million jobs in the 1990s? A clue may lie in my industry. As our economy has moved from manufacturing to service-based over the last century, commercial real estate has traced the same arc, mutating from a sector focused largely around industrial buildings to one that's about the high-rise and suburban offices that dot our commutes home. But recently, a new sector and catchphrase has emerged that indicates a major new spur in our country's growth: Big Data.
Our nation's publicly traded REITs now see cellular towers and data centers as the stars of their real estate portfolios (along with high-end malls and apartments). CALPERS, the largest public pension fund in the U.S., recently announced that it is setting up a $500 million fund to invest in data centers. McKinsey & Company predicts a 40 percent growth annually in the data being generated. Among companies of more than 1,000 employees in 15 out of the economy's 17 sectors, the average amount of data is a surreal 235 terabytes. That's right -- each of these companies has more info than the Library of Congress. And so, why should we care? Because data is valuable. The growth of digital networks and the networked sensors in everything from phones to cars to heavy machinery mean that data has a reach and sweep it has never had before. The key to Big Data is connecting these sensors to computing intelligence which can make sense of all this information (in pure Wall-E style, some theorists call this the Internet of Things). The sexiest manifestation of all of this is natural-language processing, pattern recognition and machine learning, all of which is crystallized in Siri, the backtalking, mind-reading application in iPhones -- a kind of new-millennium sassy personal assistant.
Data Driven
So, why is there so much data out there? Firstly, there's the data companies generate as they go about their business called "exhaust data" (hello, Amazon); the updates and pictures of our vacations that we breathlessly upload onto our social networks; and then all of that multimedia content we stream. With 60 percent of the world owning a mobile phone (and 12 percent of them tapping a smartphone) this is a worldwide phenomenon that has enormous implications for U.S. companies and our economy. In health care, electronic health records have been vaunted as the best way to integrate care, predict the onset of disease, eliminate medical errors, and optimize follow-up care. McKinsey thinks data could be worth $300 billion to health care. Currently, health care providers throw away 90 percent of their data partly because they don't have a place to keep it. No surprise, then, that a venture of the Blue Cross Blue Shield Association (which has health care information on 110 million people) has raised more than $37 million to create an information warehouse with 3.5 billion pieces at a time when insurers are combing patient records for ways to cut costs and improve medical care. And, there are jobs to be had. The Bureau of Labor Statistics estimates that the number of health IT jobs across the country will increase by 20 percent from 2008 to 2018, a pace much faster than the average for all occupations through 2018. The same goes for retailers who McKinsey says could potentially increase their margins by as much as 60 percent. A report by the World Economic Forum in Davos, Switzerland, titled "Big Data, Big Impact" goes as far as to declare data a new class of economic asset, like currency or gold.
As we try to get our arms around all of this, we inevitably run into exponential math. Firstly, according to IDC, a technology research firm, the aggregate amount of data is growing at 50 percent a year, or more than doubling every two years. Then, there's Moore's Law, named after Intel co-founder Gordon Moore, which states that the amount of computing power that can be purchased for a certain amount of money doubles every two years. Also, the volume of business data worldwide, across all companies, doubles every 1.2 years, according to estimates. If all these numbers and hyperbole give you a touch of vertigo, you're not alone.
This is the dromosphere described by the French dystopian theorist Paul Virilio, which depicts modern society mapping the world in tiny fragments in an unending quest for speed and progress driven by technology. Putting aside the concerns about all this data, the hopes for Big Data are not unlike those for the Human Genome Project or for science itself -- a provable, positivist system to unlock age-old mysteries and a nifty way to raise profits. Incidentally, decoding the human genome originally took 10 years to process; now it can be achieved in one week.
What's in it for us?
Erik Brynjolfsson, an economist at Massachusetts Institute of Technology's Sloan School of Management, published a 2011 report with two colleagues that suggests that data-guided management is spreading across corporate America and starting to pay off. Looking at 179 large companies, they found that those adopting "data-driven decision-making" achieved productivity gains that were 5 percent to 6 percent higher than other factors could explain. Retailers now mine huge data sets to analyze sales, pricing, customer profiles, even weather data to tailor their pricing and markdowns and to make supply-chain decisions about how to get the product to the point of sale. But the more exciting frontier may be in solving humanitarian crises. As the World Economic Forum puts it, "By analyzing patterns from mobile phone usage, a team of researchers in San Francisco is able to predict the magnitude of a disease outbreak half way around the world. Similarly, an aid agency sees early warning signs of a drought condition in a remote Sub-Saharan region, allowing the agency to get a head start on mobilizing its resources and save many more lives."
The Forum sees Big Data affecting and intervening in education, agriculture, health and global finance. One of the most extraordinary stories emerged from Haiti after the 2010 earthquake when researchers at the Karolinska Institute and Columbia University obtained data on people fleeing Port-au-Prince by tracking nearly 2 million cell-phone SIM cards in the country. By reading the sensors, they were able to pinpoint the location of more than 600,000 people, and made this information available to government and humanitarian organizations. Later that year, the same team tracked the movements of people during a cholera outbreak allowing aid organizations to mobilize.
Despite concerns with privacy and fraud, you can't argue with results like that. Data centers continue to be built in places like Phoenix, North Carolina, Santa Clara and Northern Virginia. And the U.S. is short by close to 200,000 of people with the deep analytical skills that Big Data requires, according to McKinsey. And that doesn't include the hundreds of thousands of jobs to build, equip and manage these facilities. Big Data could be the next Big Thing
Monday, June 18, 2012
'Big Data' disguises digital doubts (USA today)
Dan Vergano, USA TODAY, June 16, 2012
Buzzwords don't come any bigger than "Big Data," which promises to reveal the secrets hidden within big blocks of data held by companies, governments and musty old archives.
But maybe Big Data has an Achilles' heel, some experts warn, despite its Big promises.
"The initiative we are launching today promises to transform our ability to use Big Data for scientific discovery, environmental and biomedical research, education and national security," said presidential science adviser John Holdren, announcing a $200 million effort in March by six federal agencies to uncork the power of Big Data.
Holdren compared Big Data's advent to the invention of supercomputers and the Internet. But what is Big Data really? Starting from scientists struggling to analyze massive amounts of genetic data, as they did in the Human Genome Project a decade ago, or astronomical data, such as the survey of more than 930,000 galaxies undertaken by the Sloan Digital Sky Survey, Big Data has blossomed into a constellation of computer science approaches to handling, visualizing and blending together "big" sets of data.
For example, police forces from Honolulu to New York have looked at combinations of crime tips submitted via Facebook, Twitter and text messages to identify "hotspots" for muggings and other felonies. Amazon famously tracks masses of book purchases to suggest new buys to like-minded readers. The Defense Department hopes to weave together information from a new generation of battlefield sensors at speeds 100 times faster than today using Big Data techniques.
Such efforts have blossomed in the Facebook era, where poking through troves of customer data is seen as the key to unlocking sales. In medicine, a two-day "Health Datapalooza" held this month in the nation's capital drew together federal officials, former Senate majority leader Bill Frist, R-Tenn., and Wired Magazine executive editor Thomas Goetz, to talk health data. If your gene map can be compared instantly to the genomes of millions of other folks in coming decades, for example, the hope is that medicine finely tuned to your medical needs will result.
A Sciencejournal study last year introduced the notion of "culturomics," using Big Data — Google's millions of searchable digitized books in its case — to reveal, "linguistic and cultural phenomena that were reflected in the English language between 1800 and 2000."
Data. Data. Data. So much of it is out there, tracked from the moment you look up a dentist on a website, take a trip through a highway tollbooth to sit in the chair, pay your bill at the reception desk and post your toothache experience afterward on Facebook. Can an ad for a toothbrush be far from your in-box?
"The only problem is that a lot of the Big Data isn't really data," says anthropologist Robert Albro of American University in Washington D.C., who studies how culture affects public policy. "It's a mash-up of all kinds of numbers that started out as data, but they don't necessarily mean anything once they have been removed from where they started out." In the social sciences, he says, researchers have learned over the last century that half the battle in any study is carefully explaining your data's origins. "Once you leave that behind, there is a risk you'll be wrong, and a risk that the decisions you make based on being wrong will affect people in negative ways."
In anthropology, one historical example of a problem comes from turn-of-the-century attempts to pigeonhole people in far-off nations into tribes or countries, using data in categories now understood as far too simplistic. That effort contributed to European countries inventing imaginary borders or people in nations such as Rwanda. Public health researchers in the 20th century pigeonholed poor people into "defective" categories, based on bogus data, during the "eugenics" movement aimed at breeding better human beings that led to the involuntary sterilization of perhaps 60,000 people nationwide by the 1960s.
More recently, University of Wisconsin-Madison, communications scholars have warned that Google's search recommendations (the list of suggested searches that pop up when you start typing a word using the popular search engine) actually bend people's perception. Looking at nanotechnology, for example, the study showed that top search suggestions over a few years turned away from business to health concerns. The search recommendations were actually steering more people to look into less-reliable nanotechnology health-issue websites, they found. "Google is shaping the reality we experience in the suggestions it makes, pointing us away from the most accurate information and towards the most popular," study lead author Dietram Scheufele told USA TODAY in 2010.
Still, what's so wrong about using Big Data to find crime hotspots or books you might like? "Nothing. There is obviously immense promise there, as long as the data is kept to uses for which its limits are understood," Albro suggests. However, he worries that since so much of the data out there start out as "market research" — likes or dislikes when it comes to buying things — that removing data from advertising-focused troves and translating it into health care or planning for new roads or "culture" will essentially turn everyone into consumers, rather than citizens, in the minds of planners.
"We can't even agree on what 'culture' is, and now we're going to have 'culturomics.' Isn't that a little ambitious?" Albro asks. "Now we have claims that tweets predicted the 'Arab Spring,' which turns out to be questionable, or can detect 'sentiment' or 'mood,' which are even fuzzier or lazier words. We need to be a little cautious here." (In their defense, the "culturomics" study authors do urge caution on folks using their approach.)
But Albro worries that the hype over Big Data is warming up for Big Disappointment down the road. Much of the criticism of Big Data heard now focuses on privacy issues: who is using your data, or whether it will really help sell stuff. "Those are useful discussions, but we really need to talk a little more deeply about data," Albro says. "It's more than 'Garbage In, Garbage Out,' it's about how we shape the digital world."
Buzzwords don't come any bigger than "Big Data," which promises to reveal the secrets hidden within big blocks of data held by companies, governments and musty old archives.
But maybe Big Data has an Achilles' heel, some experts warn, despite its Big promises.
"The initiative we are launching today promises to transform our ability to use Big Data for scientific discovery, environmental and biomedical research, education and national security," said presidential science adviser John Holdren, announcing a $200 million effort in March by six federal agencies to uncork the power of Big Data.
Holdren compared Big Data's advent to the invention of supercomputers and the Internet. But what is Big Data really? Starting from scientists struggling to analyze massive amounts of genetic data, as they did in the Human Genome Project a decade ago, or astronomical data, such as the survey of more than 930,000 galaxies undertaken by the Sloan Digital Sky Survey, Big Data has blossomed into a constellation of computer science approaches to handling, visualizing and blending together "big" sets of data.
For example, police forces from Honolulu to New York have looked at combinations of crime tips submitted via Facebook, Twitter and text messages to identify "hotspots" for muggings and other felonies. Amazon famously tracks masses of book purchases to suggest new buys to like-minded readers. The Defense Department hopes to weave together information from a new generation of battlefield sensors at speeds 100 times faster than today using Big Data techniques.
Such efforts have blossomed in the Facebook era, where poking through troves of customer data is seen as the key to unlocking sales. In medicine, a two-day "Health Datapalooza" held this month in the nation's capital drew together federal officials, former Senate majority leader Bill Frist, R-Tenn., and Wired Magazine executive editor Thomas Goetz, to talk health data. If your gene map can be compared instantly to the genomes of millions of other folks in coming decades, for example, the hope is that medicine finely tuned to your medical needs will result.
A Sciencejournal study last year introduced the notion of "culturomics," using Big Data — Google's millions of searchable digitized books in its case — to reveal, "linguistic and cultural phenomena that were reflected in the English language between 1800 and 2000."
Data. Data. Data. So much of it is out there, tracked from the moment you look up a dentist on a website, take a trip through a highway tollbooth to sit in the chair, pay your bill at the reception desk and post your toothache experience afterward on Facebook. Can an ad for a toothbrush be far from your in-box?
"The only problem is that a lot of the Big Data isn't really data," says anthropologist Robert Albro of American University in Washington D.C., who studies how culture affects public policy. "It's a mash-up of all kinds of numbers that started out as data, but they don't necessarily mean anything once they have been removed from where they started out." In the social sciences, he says, researchers have learned over the last century that half the battle in any study is carefully explaining your data's origins. "Once you leave that behind, there is a risk you'll be wrong, and a risk that the decisions you make based on being wrong will affect people in negative ways."
In anthropology, one historical example of a problem comes from turn-of-the-century attempts to pigeonhole people in far-off nations into tribes or countries, using data in categories now understood as far too simplistic. That effort contributed to European countries inventing imaginary borders or people in nations such as Rwanda. Public health researchers in the 20th century pigeonholed poor people into "defective" categories, based on bogus data, during the "eugenics" movement aimed at breeding better human beings that led to the involuntary sterilization of perhaps 60,000 people nationwide by the 1960s.
More recently, University of Wisconsin-Madison, communications scholars have warned that Google's search recommendations (the list of suggested searches that pop up when you start typing a word using the popular search engine) actually bend people's perception. Looking at nanotechnology, for example, the study showed that top search suggestions over a few years turned away from business to health concerns. The search recommendations were actually steering more people to look into less-reliable nanotechnology health-issue websites, they found. "Google is shaping the reality we experience in the suggestions it makes, pointing us away from the most accurate information and towards the most popular," study lead author Dietram Scheufele told USA TODAY in 2010.
Still, what's so wrong about using Big Data to find crime hotspots or books you might like? "Nothing. There is obviously immense promise there, as long as the data is kept to uses for which its limits are understood," Albro suggests. However, he worries that since so much of the data out there start out as "market research" — likes or dislikes when it comes to buying things — that removing data from advertising-focused troves and translating it into health care or planning for new roads or "culture" will essentially turn everyone into consumers, rather than citizens, in the minds of planners.
"We can't even agree on what 'culture' is, and now we're going to have 'culturomics.' Isn't that a little ambitious?" Albro asks. "Now we have claims that tweets predicted the 'Arab Spring,' which turns out to be questionable, or can detect 'sentiment' or 'mood,' which are even fuzzier or lazier words. We need to be a little cautious here." (In their defense, the "culturomics" study authors do urge caution on folks using their approach.)
But Albro worries that the hype over Big Data is warming up for Big Disappointment down the road. Much of the criticism of Big Data heard now focuses on privacy issues: who is using your data, or whether it will really help sell stuff. "Those are useful discussions, but we really need to talk a little more deeply about data," Albro says. "It's more than 'Garbage In, Garbage Out,' it's about how we shape the digital world."
Privacy and Big Data
You for Sale: Mapping, and Sharing, the Consumer Genome
Natasha Singer, The New York Times, June 17, 2012
IT knows who you are. It knows where you live. It knows what you do.
It peers deeper into American life than the F.B.I. or the I.R.S., or those prying digital eyes at Facebook and Google. If you are an American adult, the odds are that it knows things like your age, race, sex, weight, height, marital status, education level, politics, buying habits, household health worries, vacation dreams — and on and on.
Right now in Conway, Ark., north of Little Rock, more than 23,000 computer servers are collecting, collating and analyzing consumer data for a company that, unlike Silicon Valley’s marquee names, rarely makes headlines. It’s called the Acxiom Corporation, and it’s the quiet giant of a multibillion-dollar industry known as database marketing.
Few consumers have ever heard of Acxiom. But analysts say it has amassed the world’s largest commercial database on consumers — and that it wants to know much, much more. Its servers process more than 50 trillion data “transactions” a year. Company executives have said its database contains information about 500 million active consumers worldwide, with about 1,500 data points per person. That includes a majority of adults in the United States.
Such large-scale data mining and analytics — based on information available in public records, consumer surveys and the like — are perfectly legal. Acxiom’s customers have included big banks like Wells Fargo and HSBC, investment services like E*Trade, automakers like Toyota and Ford, department stores like Macy’s — just about any major company looking for insight into its customers.
For Acxiom, based in Little Rock, the setup is lucrative. It posted profit of $77.26 million in its latest fiscal year, on sales of $1.13 billion.
But such profits carry a cost for consumers. Federal authorities say current laws may not be equipped to handle the rapid expansion of an industry whose players often collect and sell sensitive financial and health information yet are nearly invisible to the public. In essence, it’s as if the ore of our data-driven lives were being mined, refined and sold to the highest bidder, usually without our knowledge — by companies that most people rarely even know exist.
Julie Brill, a member of the Federal Trade Commission, says she would like data brokers in general to tell the public about the data they collect, how they collect it, whom they share it with and how it is used. “If someone is listed as diabetic or pregnant, what is happening with this information? Where is the information going?” she asks. “We need to figure out what the rules should be as a society.”
Although Acxiom employs a chief privacy officer, Jennifer Barrett Glasgow, she and other executives declined requests to be interviewed for this article, said Ines Rodriguez Gutzmer, director of corporate communications.
In March, however, Ms. Barrett Glasgow endorsed increased industry openness. “It’s not an unreasonable request to have more transparency among data brokers,” she said in an interview with The New York Times. In marketing materials, Acxiom promotes itself as “a global thought leader in addressing consumer privacy issues and earning the public trust.”
But, in interviews, security experts and consumer advocates paint a portrait of a company with practices that privilege corporate clients’ interests over those of consumers and contradict the company’s stance on transparency. Acxiom’s marketing materials, for example, promote a special security system for clients and associates to encrypt the data they send. Yet cybersecurity experts who examined Acxiom’s Web site for The Times found basic security lapses on an online form for consumers seeking access to their own profiles. (Acxiom says it has fixed the broken link that caused the problem.)
In a fast-changing digital economy, Acxiom is developing even more advanced techniques to mine and refine data. It has recruited talent from Microsoft, Google, Amazon.com and Myspace and is using a powerful, multiplatform approach to predicting consumer behavior that could raise its standing among investors and clients.
Of course, digital marketers already customize pitches to users, based on their past activities. Just think of “cookies,” bits of computer code placed on browsers to keep track of online activity. But Acxiom, analysts say, is pursuing far more comprehensive techniques in an effort to influence consumer decisions. It is integrating what it knows about our offline, online and even mobile selves, creating in-depth behavior portraits in pixilated detail. Its executives have called this approach a “360-degree view” on consumers.
“There’s a lot of players in the digital space trying the same thing,” says Mark Zgutowicz, a Piper Jaffray analyst. “But Acxiom’s advantage is they have a database of offline information that they have been collecting for 40 years and can leverage that expertise in the digital world.”
Yet some prominent privacy advocates worry that such techniques could lead to a new era of consumer profiling.
Jeffrey Chester, executive director of the Center for Digital Democracy, a nonprofit group in Washington, says: “It is Big Brother in Arkansas.”
SCOTT HUGHES, an up-and-coming small-business owner and Facebook denizen, is Acxiom’s ideal consumer. Indeed, it created him.
Mr. Hughes is a fictional character who appeared in an Acxiom investor presentation in 2010. A frequent shopper, he was designed to show the power of Acxiom’s multichannel approach.
In the presentation, he logs on to Facebook and sees that his friend Ella has just become a fan of Bryce Computers, an imaginary electronics retailer and Acxiom client. Ella’s update prompts Mr. Hughes to check out Bryce’s fan page and do some digital window-shopping for a fast inkjet printer.
Such browsing seems innocuous — hardly data mining. But it cues an Acxiom system designed to recognize consumers, remember their actions, classify their behaviors and influence them with tailored marketing.
When Mr. Hughes follows a link to Bryce’s retail site, for example, the system recognizes him from his Facebook activity and shows him a printer to match his interest. He registers on the site, but doesn’t buy the printer right away, so the system tracks him online. Lo and behold, the next morning, while he scans baseball news on ESPN.com, an ad for the printer pops up again.
That evening, he returns to the Bryce site where, the presentation says, “he is instantly recognized” as having registered. It then offers a sweeter deal: a $10 rebate and free shipping.
It’s not a random offer. Acxiom has its own classification system, PersonicX, which assigns consumers to one of 70 detailed socioeconomic clusters and markets to them accordingly. In this situation, it pegs Mr. Hughes as a “savvy single” — meaning he’s in a cluster of mobile, upper-middle-class people who do their banking online, attend pro sports events, are sensitive to prices — and respond to free-shipping offers.
Correctly typecast, Mr. Hughes buys the printer.
But the multichannel system of Acxiom and its online partners is just revving up. Later, it sends him coupons for ink and paper, to be redeemed via his cellphone, and a personalized snail-mail postcard suggesting that he donate his old printer to a nearby school.
Analysts say companies design these sophisticated ecosystems to prompt consumers to volunteer enough personal data — like their names, e-mail addresses and mobile numbers — so that marketers can offer them customized appeals any time, anywhere.
Still, there is a fine line between customization and stalking. While many people welcome the convenience of personalized offers, others may see the surveillance engines behind them as intrusive or even manipulative.
“If you look at it in cold terms, it seems like they are really out to trick the customer,” says Dave Frankland, the research director for customer intelligence at Forrester Research. “But they are actually in the business of helping marketers make sure that the right people are getting offers they are interested in and therefore establish a relationship with the company.”
DECADES before the Internet as we know it, a businessman named Charles Ward planted the seeds of Acxiom. It was 1969, and Mr. Ward started a data processing company in Conway called Demographics Inc., in part to help the Democratic Party reach voters. In a time when Madison Avenue was deploying one-size-fits-all national ad campaigns, Demographics and its lone computer used public phone books to compile lists for direct mailing of campaign material.
Today, Acxiom maintains its own database on about 190 million individuals and 126 million households in the United States. Separately, it manages customer databases for or works with 47 of the Fortune 100 companies. It also worked with the government after the September 2001 terrorist attacks, providing information about 11 of the 19 hijackers.
To beef up its digital services, Acxiom recently mounted an aggressive hiring campaign. Last July, it named Scott E. Howe, a former corporate vice president for Microsoft’s advertising business group, as C.E.O. Last month, it hired Phil Mui, formerly group product manager for Google Analytics, as its chief product and engineering officer.
In interviews, Mr. Howe has laid out a vision of Acxiom as a new-millennium “data refinery” rather than a data miner. That description posits Acxiom as a nimble provider of customer analytics services, able to compete with Facebook and Google, rather than as a stealth engine of consumer espionage.
Still, the more that information brokers mine powerful consumer data, the more they become attractive targets for hackers — and draw scrutiny from consumer advocates.
This year, Advertising Age ranked Epsilon, another database marketing firm, as the biggest advertising agency in the United States, with Acxiom second. Most people know Epsilon, if they know it at all, because it experienced a major security breach last year, exposing the e-mail addresses of millions of customers of Citibank, JPMorgan Chase, Target, Walgreens and others. In 2003, Acxiom had its own security breaches.
But privacy advocates say they are more troubled by data brokers’ ranking systems, which classify some people as high-value prospects, to be offered marketing deals and discounts regularly, while dismissing others as low-value — known in industry slang as “waste.”
Exclusion from a vacation offer may not matter much, says Pam Dixon, the executive director of the World Privacy Forum, a nonprofit group in San Diego, but if marketing algorithms judge certain people as not worthy of receiving promotions for higher education or health services, they could have a serious impact.
"Over time, that can really turn into a mountain of pathways not offered, not seen and not known about,” Ms. Dixon says.
Until now, database marketers operated largely out of the public eye. Unlike consumer reporting agencies that sell sensitive financial information about people for credit or employment purposes, database marketers aren’t required by law to show consumers their own reports and allow them to correct errors. That may be about to change. This year, the F.T.C. published a report calling for greater transparency among data brokers and asking Congress to give consumers the right to access information these firms hold about them.
ACXIOM’S Consumer Data Products Catalog offers hundreds of details — called “elements” — that corporate clients can buy about individuals or households, to augment their own marketing databases. Companies can buy data to pinpoint households that are concerned, say, about allergies, diabetes or “senior needs.” Also for sale is information on sizes of home loans and household incomes.
Clients generally buy this data because they want to hold on to their best customers or find new ones — or both.
A bank that wants to sell its best customers additional services, for example, might buy details about those customers’ social media, Web and mobile habits to identify more efficient ways to market to them. Or, says Mr. Frankland at Forrester, a sporting goods chain whose best customers are 25- to 34-year-old men living near mountains or beaches could buy a list of a million other people with the same characteristics. The retailer could hire Acxiom, he says, to manage a campaign aimed at that new group, testing how factors like consumers’ locations or sports preferences affect responses.
But the catalog also offers delicate information that has set off alarm bells among some privacy advocates, who worry about the potential for misuse by third parties that could take aim at vulnerable groups. Such information includes consumers’ interests — derived, the catalog says, “from actual purchases and self-reported surveys” — like “Christian families,” “Dieting/Weight Loss,” “Gaming-Casino,” “Money Seekers” and “Smoking/Tobacco.” Acxiom also sells data about an individual’s race, ethnicity and country of origin. “Our Race model,” the catalog says, “provides information on the major racial category: Caucasians, Hispanics, African-Americans, or Asians.” Competing companies sell similar data.
Acxiom’s data about race or ethnicity is “used for engaging those communities for marketing purposes,” said Ms. Barrett Glasgow, the privacy officer, in an e-mail response to questions.
There may be a legitimate commercial need for some businesses, like ethnic restaurants, to know the race or ethnicity of consumers, says Joel R. Reidenberg, a privacy expert and a professor at the Fordham Law School.
“At the same time, this is ethnic profiling,” he says. “The people on this list, they are being sold based on their ethnic stereotypes. There is a very strong citizen’s right to have a veto over the commodification of their profile.”
He says the sale of such data is troubling because race coding may be incorrect. And even if a data broker has correct information, a person may not want to be marketed to based on race.
“DO you really know your customers?” Acxiom asks in marketing materials for its shopper recognition system, a program that uses ZIP codes to help retailers confirm consumers’ identities — without asking their permission.
“Simply asking for name and address information poses many challenges: transcription errors, increased checkout time and, worse yet, losing customers who feel that you’re invading their privacy,” Acxiom’s fact sheet explains. In its system, a store clerk need only “capture the shopper’s name from a check or third-party credit card at the point of sale and then ask for the shopper’s ZIP code or telephone number.” With that data Acxiom can identify shoppers within a 10 percent margin of error, it says, enabling stores to reward their best customers with special offers. Other companies offer similar services.
“This is a direct way of circumventing people’s concerns about privacy,” says Mr. Chester of the Center for Digital Democracy.
Ms. Barrett Glasgow of Acxiom says that its program is a “standard practice” among retailers, but that the company encourages its clients to report consumers who wish to opt out.
Acxiom has positioned itself as an industry leader in data privacy, but some of its practices seem to undermine that image. It created the position of chief privacy officer in 1991, well ahead of its rivals. It even offers an online request form, promoted as an easy way for consumers to access information Acxiom collects about them.
But the process turned out to be not so user-friendly for a reporter for The Times.
In early May, the reporter decided to request her record from Acxiom, as any consumer might. Before submitting a Social Security number and other personal information, however, she asked for advice from a cybersecurity expert at The Times. The expert examined Acxiom’s Web site and immediately noticed that the online form did not employ a standard encryption protocol — called https — used by sites like Amazon and American Express. When the expert tested the form, using software that captures data sent over the Web, he could clearly see that the sample Social Security number he had submitted had not been encrypted. At that point, the reporter was advised not to request her file, given the risk that the process might expose her personal information.
Later in May, Ashkan Soltani, an independent security researcher and former technologist in identity protection at the F.T.C., also examined Acxiom’s site and came to the same conclusion. “Parts of the site for corporate clients are encrypted,” he says. “But for consumers, who this information is about and who stand the most to lose from data collection, they don’t provide security.”
Ms. Barrett Glasgow says that the form has always been encrypted with https but that on May 11, its security monitoring system detected a “broken redirect link” that allowed unencrypted access. Since then, she says, Acxiom has fixed the link and determined that no unauthorized person had gained access to information sent using the form.
On May 25, the reporter submitted an online request to Acxiom for her file, along with a personal check, sent by Express Mail, for the $5 processing fee. Three weeks later, no response had arrived.
Regulators at the F.T.C. declined to comment on the practices of individual companies. But Jon Leibowitz, the commission chairman, said consumers should have the right to see and correct personal details about them collected and sold by data aggregators.
After all, he said, “they are the unseen cyberazzi who collect information on all of us.”
Natasha Singer, The New York Times, June 17, 2012
IT knows who you are. It knows where you live. It knows what you do.
It peers deeper into American life than the F.B.I. or the I.R.S., or those prying digital eyes at Facebook and Google. If you are an American adult, the odds are that it knows things like your age, race, sex, weight, height, marital status, education level, politics, buying habits, household health worries, vacation dreams — and on and on.
Right now in Conway, Ark., north of Little Rock, more than 23,000 computer servers are collecting, collating and analyzing consumer data for a company that, unlike Silicon Valley’s marquee names, rarely makes headlines. It’s called the Acxiom Corporation, and it’s the quiet giant of a multibillion-dollar industry known as database marketing.
Few consumers have ever heard of Acxiom. But analysts say it has amassed the world’s largest commercial database on consumers — and that it wants to know much, much more. Its servers process more than 50 trillion data “transactions” a year. Company executives have said its database contains information about 500 million active consumers worldwide, with about 1,500 data points per person. That includes a majority of adults in the United States.
Such large-scale data mining and analytics — based on information available in public records, consumer surveys and the like — are perfectly legal. Acxiom’s customers have included big banks like Wells Fargo and HSBC, investment services like E*Trade, automakers like Toyota and Ford, department stores like Macy’s — just about any major company looking for insight into its customers.
For Acxiom, based in Little Rock, the setup is lucrative. It posted profit of $77.26 million in its latest fiscal year, on sales of $1.13 billion.
But such profits carry a cost for consumers. Federal authorities say current laws may not be equipped to handle the rapid expansion of an industry whose players often collect and sell sensitive financial and health information yet are nearly invisible to the public. In essence, it’s as if the ore of our data-driven lives were being mined, refined and sold to the highest bidder, usually without our knowledge — by companies that most people rarely even know exist.
Julie Brill, a member of the Federal Trade Commission, says she would like data brokers in general to tell the public about the data they collect, how they collect it, whom they share it with and how it is used. “If someone is listed as diabetic or pregnant, what is happening with this information? Where is the information going?” she asks. “We need to figure out what the rules should be as a society.”
Although Acxiom employs a chief privacy officer, Jennifer Barrett Glasgow, she and other executives declined requests to be interviewed for this article, said Ines Rodriguez Gutzmer, director of corporate communications.
In March, however, Ms. Barrett Glasgow endorsed increased industry openness. “It’s not an unreasonable request to have more transparency among data brokers,” she said in an interview with The New York Times. In marketing materials, Acxiom promotes itself as “a global thought leader in addressing consumer privacy issues and earning the public trust.”
But, in interviews, security experts and consumer advocates paint a portrait of a company with practices that privilege corporate clients’ interests over those of consumers and contradict the company’s stance on transparency. Acxiom’s marketing materials, for example, promote a special security system for clients and associates to encrypt the data they send. Yet cybersecurity experts who examined Acxiom’s Web site for The Times found basic security lapses on an online form for consumers seeking access to their own profiles. (Acxiom says it has fixed the broken link that caused the problem.)
In a fast-changing digital economy, Acxiom is developing even more advanced techniques to mine and refine data. It has recruited talent from Microsoft, Google, Amazon.com and Myspace and is using a powerful, multiplatform approach to predicting consumer behavior that could raise its standing among investors and clients.
Of course, digital marketers already customize pitches to users, based on their past activities. Just think of “cookies,” bits of computer code placed on browsers to keep track of online activity. But Acxiom, analysts say, is pursuing far more comprehensive techniques in an effort to influence consumer decisions. It is integrating what it knows about our offline, online and even mobile selves, creating in-depth behavior portraits in pixilated detail. Its executives have called this approach a “360-degree view” on consumers.
“There’s a lot of players in the digital space trying the same thing,” says Mark Zgutowicz, a Piper Jaffray analyst. “But Acxiom’s advantage is they have a database of offline information that they have been collecting for 40 years and can leverage that expertise in the digital world.”
Yet some prominent privacy advocates worry that such techniques could lead to a new era of consumer profiling.
Jeffrey Chester, executive director of the Center for Digital Democracy, a nonprofit group in Washington, says: “It is Big Brother in Arkansas.”
SCOTT HUGHES, an up-and-coming small-business owner and Facebook denizen, is Acxiom’s ideal consumer. Indeed, it created him.
Mr. Hughes is a fictional character who appeared in an Acxiom investor presentation in 2010. A frequent shopper, he was designed to show the power of Acxiom’s multichannel approach.
In the presentation, he logs on to Facebook and sees that his friend Ella has just become a fan of Bryce Computers, an imaginary electronics retailer and Acxiom client. Ella’s update prompts Mr. Hughes to check out Bryce’s fan page and do some digital window-shopping for a fast inkjet printer.
Such browsing seems innocuous — hardly data mining. But it cues an Acxiom system designed to recognize consumers, remember their actions, classify their behaviors and influence them with tailored marketing.
When Mr. Hughes follows a link to Bryce’s retail site, for example, the system recognizes him from his Facebook activity and shows him a printer to match his interest. He registers on the site, but doesn’t buy the printer right away, so the system tracks him online. Lo and behold, the next morning, while he scans baseball news on ESPN.com, an ad for the printer pops up again.
That evening, he returns to the Bryce site where, the presentation says, “he is instantly recognized” as having registered. It then offers a sweeter deal: a $10 rebate and free shipping.
It’s not a random offer. Acxiom has its own classification system, PersonicX, which assigns consumers to one of 70 detailed socioeconomic clusters and markets to them accordingly. In this situation, it pegs Mr. Hughes as a “savvy single” — meaning he’s in a cluster of mobile, upper-middle-class people who do their banking online, attend pro sports events, are sensitive to prices — and respond to free-shipping offers.
Correctly typecast, Mr. Hughes buys the printer.
But the multichannel system of Acxiom and its online partners is just revving up. Later, it sends him coupons for ink and paper, to be redeemed via his cellphone, and a personalized snail-mail postcard suggesting that he donate his old printer to a nearby school.
Analysts say companies design these sophisticated ecosystems to prompt consumers to volunteer enough personal data — like their names, e-mail addresses and mobile numbers — so that marketers can offer them customized appeals any time, anywhere.
Still, there is a fine line between customization and stalking. While many people welcome the convenience of personalized offers, others may see the surveillance engines behind them as intrusive or even manipulative.
“If you look at it in cold terms, it seems like they are really out to trick the customer,” says Dave Frankland, the research director for customer intelligence at Forrester Research. “But they are actually in the business of helping marketers make sure that the right people are getting offers they are interested in and therefore establish a relationship with the company.”
DECADES before the Internet as we know it, a businessman named Charles Ward planted the seeds of Acxiom. It was 1969, and Mr. Ward started a data processing company in Conway called Demographics Inc., in part to help the Democratic Party reach voters. In a time when Madison Avenue was deploying one-size-fits-all national ad campaigns, Demographics and its lone computer used public phone books to compile lists for direct mailing of campaign material.
Today, Acxiom maintains its own database on about 190 million individuals and 126 million households in the United States. Separately, it manages customer databases for or works with 47 of the Fortune 100 companies. It also worked with the government after the September 2001 terrorist attacks, providing information about 11 of the 19 hijackers.
To beef up its digital services, Acxiom recently mounted an aggressive hiring campaign. Last July, it named Scott E. Howe, a former corporate vice president for Microsoft’s advertising business group, as C.E.O. Last month, it hired Phil Mui, formerly group product manager for Google Analytics, as its chief product and engineering officer.
In interviews, Mr. Howe has laid out a vision of Acxiom as a new-millennium “data refinery” rather than a data miner. That description posits Acxiom as a nimble provider of customer analytics services, able to compete with Facebook and Google, rather than as a stealth engine of consumer espionage.
Still, the more that information brokers mine powerful consumer data, the more they become attractive targets for hackers — and draw scrutiny from consumer advocates.
This year, Advertising Age ranked Epsilon, another database marketing firm, as the biggest advertising agency in the United States, with Acxiom second. Most people know Epsilon, if they know it at all, because it experienced a major security breach last year, exposing the e-mail addresses of millions of customers of Citibank, JPMorgan Chase, Target, Walgreens and others. In 2003, Acxiom had its own security breaches.
But privacy advocates say they are more troubled by data brokers’ ranking systems, which classify some people as high-value prospects, to be offered marketing deals and discounts regularly, while dismissing others as low-value — known in industry slang as “waste.”
Exclusion from a vacation offer may not matter much, says Pam Dixon, the executive director of the World Privacy Forum, a nonprofit group in San Diego, but if marketing algorithms judge certain people as not worthy of receiving promotions for higher education or health services, they could have a serious impact.
"Over time, that can really turn into a mountain of pathways not offered, not seen and not known about,” Ms. Dixon says.
Until now, database marketers operated largely out of the public eye. Unlike consumer reporting agencies that sell sensitive financial information about people for credit or employment purposes, database marketers aren’t required by law to show consumers their own reports and allow them to correct errors. That may be about to change. This year, the F.T.C. published a report calling for greater transparency among data brokers and asking Congress to give consumers the right to access information these firms hold about them.
ACXIOM’S Consumer Data Products Catalog offers hundreds of details — called “elements” — that corporate clients can buy about individuals or households, to augment their own marketing databases. Companies can buy data to pinpoint households that are concerned, say, about allergies, diabetes or “senior needs.” Also for sale is information on sizes of home loans and household incomes.
Clients generally buy this data because they want to hold on to their best customers or find new ones — or both.
A bank that wants to sell its best customers additional services, for example, might buy details about those customers’ social media, Web and mobile habits to identify more efficient ways to market to them. Or, says Mr. Frankland at Forrester, a sporting goods chain whose best customers are 25- to 34-year-old men living near mountains or beaches could buy a list of a million other people with the same characteristics. The retailer could hire Acxiom, he says, to manage a campaign aimed at that new group, testing how factors like consumers’ locations or sports preferences affect responses.
But the catalog also offers delicate information that has set off alarm bells among some privacy advocates, who worry about the potential for misuse by third parties that could take aim at vulnerable groups. Such information includes consumers’ interests — derived, the catalog says, “from actual purchases and self-reported surveys” — like “Christian families,” “Dieting/Weight Loss,” “Gaming-Casino,” “Money Seekers” and “Smoking/Tobacco.” Acxiom also sells data about an individual’s race, ethnicity and country of origin. “Our Race model,” the catalog says, “provides information on the major racial category: Caucasians, Hispanics, African-Americans, or Asians.” Competing companies sell similar data.
Acxiom’s data about race or ethnicity is “used for engaging those communities for marketing purposes,” said Ms. Barrett Glasgow, the privacy officer, in an e-mail response to questions.
There may be a legitimate commercial need for some businesses, like ethnic restaurants, to know the race or ethnicity of consumers, says Joel R. Reidenberg, a privacy expert and a professor at the Fordham Law School.
“At the same time, this is ethnic profiling,” he says. “The people on this list, they are being sold based on their ethnic stereotypes. There is a very strong citizen’s right to have a veto over the commodification of their profile.”
He says the sale of such data is troubling because race coding may be incorrect. And even if a data broker has correct information, a person may not want to be marketed to based on race.
“DO you really know your customers?” Acxiom asks in marketing materials for its shopper recognition system, a program that uses ZIP codes to help retailers confirm consumers’ identities — without asking their permission.
“Simply asking for name and address information poses many challenges: transcription errors, increased checkout time and, worse yet, losing customers who feel that you’re invading their privacy,” Acxiom’s fact sheet explains. In its system, a store clerk need only “capture the shopper’s name from a check or third-party credit card at the point of sale and then ask for the shopper’s ZIP code or telephone number.” With that data Acxiom can identify shoppers within a 10 percent margin of error, it says, enabling stores to reward their best customers with special offers. Other companies offer similar services.
“This is a direct way of circumventing people’s concerns about privacy,” says Mr. Chester of the Center for Digital Democracy.
Ms. Barrett Glasgow of Acxiom says that its program is a “standard practice” among retailers, but that the company encourages its clients to report consumers who wish to opt out.
Acxiom has positioned itself as an industry leader in data privacy, but some of its practices seem to undermine that image. It created the position of chief privacy officer in 1991, well ahead of its rivals. It even offers an online request form, promoted as an easy way for consumers to access information Acxiom collects about them.
But the process turned out to be not so user-friendly for a reporter for The Times.
In early May, the reporter decided to request her record from Acxiom, as any consumer might. Before submitting a Social Security number and other personal information, however, she asked for advice from a cybersecurity expert at The Times. The expert examined Acxiom’s Web site and immediately noticed that the online form did not employ a standard encryption protocol — called https — used by sites like Amazon and American Express. When the expert tested the form, using software that captures data sent over the Web, he could clearly see that the sample Social Security number he had submitted had not been encrypted. At that point, the reporter was advised not to request her file, given the risk that the process might expose her personal information.
Later in May, Ashkan Soltani, an independent security researcher and former technologist in identity protection at the F.T.C., also examined Acxiom’s site and came to the same conclusion. “Parts of the site for corporate clients are encrypted,” he says. “But for consumers, who this information is about and who stand the most to lose from data collection, they don’t provide security.”
Ms. Barrett Glasgow says that the form has always been encrypted with https but that on May 11, its security monitoring system detected a “broken redirect link” that allowed unencrypted access. Since then, she says, Acxiom has fixed the link and determined that no unauthorized person had gained access to information sent using the form.
On May 25, the reporter submitted an online request to Acxiom for her file, along with a personal check, sent by Express Mail, for the $5 processing fee. Three weeks later, no response had arrived.
Regulators at the F.T.C. declined to comment on the practices of individual companies. But Jon Leibowitz, the commission chairman, said consumers should have the right to see and correct personal details about them collected and sold by data aggregators.
After all, he said, “they are the unseen cyberazzi who collect information on all of us.”
Friday, June 15, 2012
NYT: How Big Data Sees Wikipedia
Quentin Hardy, The New York Times, June 14, 2012
(Screenshots from a video showing the sentiment Wikipedia users expressed when talking about a specific place and time. From left, 1900, 1945 and 2007.)
You can learn a lot about the world from Wikipedia, sometimes without reading the articles.
Kalev Leetaru, a researcher at the University of Illinois, has been looking at the capacious volunteer-written encyclopedia as a Big Data resource, concentrating on the connections between cities around the globe over time. To understand these connections, he focuses on the type of language used to talk about a particular place, to see whether the writers have a generally positive or negative sentiment toward the place at that time.
The result is an interesting historical atlas of the rise of globalization and warfare. His technical sponsor in data mining, Silicon Graphics International, hopes the work is also an advertisement for S.G.I.’s decidedly noncloud style of technical computing for some kinds of number crunching.
Mr. Leetaru scanned Wikipedia’s 37 gigabytes of data in English, securing mention of 80 million locations and 40 million dates, scattered across four million pages of copy. (Wikipedia has many more pages than that, but most of the rest are redirects.)
“I put every coordinate on a map with a date stamp,” he said, adding that he then linked it with every other location mentioned that year. “It gave a map of how the world is connected.”
He then color-coded it for the sentiment used to describe a place. Red meant the writer was describing something bad, green meant something good, or at least neutral.
He then color-coded it for the sentiment used to describe a place. Red meant the writer was describing something bad, green meant something good, or at least neutral.
The result, which was later laid out as a time-lapse movie, was generated in 30 minutes, he said, then followed by a day of tweaking. That is a fast return, made possible, he explained, because the S.G.I. machine, which has 4,096 computing cores and can store 64 terabytes of data in its main memory, does not send data and computing resources to several different locations for cloud-style parallel processing. Such a computer, which costs from $30,000 to more than $1 million, could be cost-effective for certain data-intensive tasks.
Mr. Leetaru’s research is as interesting for what it says about Wikipedia as for what it says about the world.
For one thing, the connections between places build slowly, tracing the course of immigration and empire, up to the present era, where globalization makes so many lines that the map is an unreadable solid green.
In addition, the areas of red are notably few: a little bit at the Napoleonic Wars, a lot at the Civil War, and less during World Wars I and II. That seems to show a United States preoccupation: as bad as the Civil War was, in terms of loss of life it does not even rank in the top 20 conflicts worldwide since 1800, and it had a relatively small effect on other nations.
World War II has far more red in the United States, which had virtually no fighting on its territory, and more in Europe than in Asia, except for the Philippines, where English is more commonly spoken and there were strong ties to the United States.
World War II has far more red in the United States, which had virtually no fighting on its territory, and more in Europe than in Asia, except for the Philippines, where English is more commonly spoken and there were strong ties to the United States.
Mr. Leetaru said his work was meant to be a snapshot of how Wikipedia writers of today view the world, and not a decisive verdict of history. Examining the content of books and other print media written over a longer period might well register different, and changing, sentiments about historical events over time.
Thursday, June 14, 2012
Big data visualized When you line up all the tweets, games, and apps end to end, what do you get?
Alex Wilhelm, The Next Web.com, June 13, 2012
Data is an amazing thing, and as you know, we’re making more of it than ever before. The idea of ‘big data’ has gone from buzzword to reality, as cloud solutions to handling massive quantities of information have become more than the norm, they are now the lifeblood of the Internet.
However, the scale at which information is being created is hard to understand. According to Autonomy, every minute is exceptionally busy, with “98,000 new Tweets, 23,148 apps downloaded, 400,710 ad requests and [a total of] 208,333 minutes of Angry Birds played.” That’s far more Angry Birds than I am comfortable with, as it implies that there are more than 200,000 players on that game every minute of every day.
Still, approve or not, this is our reality. The infographic below is a very fun trip into big data, and one that we think is worth examining. Dig in, and then ask yourself how the modern Web would work without Amazon AWS. It’s almost hard to imagine.
Wednesday, June 13, 2012
Big Data applied: Robots and humans to collaborate in future factories and operating rooms
Kurtzweil Accelerating Intelligence, June 13, 2012
Algorithm lets robots "understand" and adapt to individual workers
Humans and robots may be working side by side in the factory floor or operating room of the future, according to Julie Shah, the MIT Boeing Career Development Assistant Professor of Aeronautics and Astronautics.
Shaw and her colleagues at MIT have devised an algorithm that enables a robot to quickly learn an individual’s preference for a certain task, and adapt accordingly to help complete the task.
She envisions robotic assistants performing tasks that would otherwise hinder a human’s efficiency, particularly in airplane manufacturing.
“If the robot can provide tools and materials so the person doesn’t have to walk over to pick up parts and walk back to the plane, you can significantly reduce the idle time of the person,” says Shah, who leads the Interactive Robotics Group in MIT’s Computer Science and Artificial Intelligence Laboratory (CSAIL).
“It’s really hard to make robots do careful refinishing tasks that people do really well. But providing robotic assistants to do the non-value-added work can actually increase the productivity of the overall factory.”
A robot working in isolation has to simply follow a set of preprogrammed instructions to perform a repetitive task. But working with humans is a different matter.
For example, each mechanic working at the same station at an aircraft assembly plant may prefer to work differently — and Shah says a robotic assistant would have to effortlessly adapt to an individual’s particular style to be of any practical use.
“It’s an interesting machine-learning human-factors problem,” Shah says. “Using this algorithm, we can significantly improve the robot’s understanding of what the person’s next likely actions are.”
Test case: wing assembly
As a test case, Shah’s team looked at spar assembly, a process of building the main structural element of an aircraft’s wing. In the typical manufacturing process, two pieces of the wing are aligned. Once in place, a mechanic applies sealant to predrilled holes, hammers bolts into the holes to secure the two pieces, then wipes away excess sealant.
The entire process can be highly individualized. For example, one mechanic may choose to apply sealant to every hole before hammering in bolts, while another may like to completely finish one hole before moving on to the next. The only constraint is the sealant, which dries within three minutes.
The researchers say robots such as FRIDA, designed by Swiss robotics company ABB, may be programmed to help in the spar-assembly process. FRIDA is a flexible robot with two arms capable of a wide range of motion that Shah says can be manipulated to either fasten bolts or paint sealant into holes, depending on a human’s preferences.
To enable such a robot to anticipate a human’s actions, the group first developed a computational model in the form of a decision tree. Each branch along the tree represents a choice that a mechanic may make — for example, continue to hammer a bolt after applying sealant, or apply sealant to the next hole?
“If the robot places the bolt, how sure is it that the person will then hammer the bolt, or just wait for the robot to place the next bolt?” Shah says. “There are many branches.”
Using the model, the group performed human experiments, training a laboratory robot to observe an individual’s chain of preferences. Once the robot learned a person’s preferred order of tasks, it then quickly adapted, either applying sealant or fastening a bolt according to a person’s particular style of work.
Working side by side
Shah envisions robots and humans undergoing an initial training session off the factory floor. Once the robot learns a person’s work habits, its factory counterpart can be programmed to recognize that same person, and initialize the appropriate task plan. Many workers in existing plants wear RFID tags — a potential way for robots to identify individuals, she adds.
Steve Derby, associate professor and co-director of the Flexible Manufacturing Center at Rensselaer Polytechnic Institute, says the group’s adaptive algorithm moves the field of robotics one step closer to true collaboration between humans and robots.
“The evolution of the robot itself has been way too slow on all fronts, whether on mechanical design, controls or programming interface,” Derby says. “I think this paper is important — it fits in with the whole spectrum of things that need to happen in getting people and robots to work next to each other.”
Uses in the medical operating room
Shah says robotic assistants may also be programmed to help in medical settings. For instance, a robot may be trained to monitor lengthy procedures in an operating room and anticipate a surgeon’s needs, handing over scalpels and gauze, depending on a doctor’s preference. While such a scenario may be years away, robots and humans may eventually work side by side, with the right algorithms.
“We have hardware, sensing, and can do manipulation and vision, but unless the robot really develops an almost seamless understanding of how it can help the person, the person’s just going to get frustrated and say, ‘Never mind, I’ll just go pick up the piece myself,’” Shah says.
This research was supported in part by Boeing Research and Technology and conducted in collaboration with ABB.
Big Data: What Facebook Knows
What Facebook Knows
The company's social scientists are hunting for insights about human behavior. What they find could give Facebook new ways to cash in on our data—and remake our view of society.
Tom Simonite, Technology Review, July/August 2012
Laws haven't kept up with the company's ability to mine its users' data.
If Facebook were a country, a conceit that founder Mark Zuckerberg has entertained in public, its 900 million members would make it the third largest in the world.
It would far outstrip any regime past or present in how intimately it records the lives of its citizens. Private conversations, family photos, and records of road trips, births, marriages, and deaths all stream into the company's servers and lodge there. Facebook has collected the most extensive data set ever assembled on human social behavior. Some of your personal information is probably part of it.
And yet, even as Facebook has embedded itself into modern life, it hasn't actually done that much with what it knows about us. Now that the company has gone public, the pressure to develop new sources of profit (see "The Facebook Fallacy") is likely to force it to do more with its hoard of information. That stash of data looms like an oversize shadow over what today is a modest online advertising business, worrying privacy-conscious Web users (see "Few Privacy Regulations Inhibit Facebook") and rivals such as Google. Everyone has a feeling that this unprecedented resource will yield something big, but nobody knows quite what.
Even as Facebook has embedded itself into modern life, it hasn't done that much with what it knows about us. Its stash of data looms like an oversize shadow. Everyone has a feeling that this resource will yield something big, but nobody knows quite what.
Heading Facebook's effort to figure out what can be learned from all our data is Cameron Marlow, a tall 35-year-old who until recently sat a few feet away from Zuckerberg. The group Marlow runs has escaped the public attention that dogs Facebook's founders and the more headline-grabbing features of its business. Known internally as the Data Science Team, it is a kind of Bell Labs for the social-networking age. The group has 12 researchers—but is expected to double in size this year. They apply math, programming skills, and social science to mine our data for insights that they hope will advance Facebook's business and social science at large. Whereas other analysts at the company focus on information related to specific online activities, Marlow's team can swim in practically the entire ocean of personal data that Facebook maintains. Of all the people at Facebook, perhaps even including the company's leaders, these researchers have the best chance of discovering what can really be learned when so much personal information is compiled in one place.
Facebook has all this information because it has found ingenious ways to collect data as people socialize. Users fill out profiles with their age, gender, and e-mail address; some people also give additional details, such as their relationship status and mobile-phone number. A redesign last fall introduced profile pages in the form of time lines that invite people to add historical information such as places they have lived and worked. Messages and photos shared on the site are often tagged with a precise location, and in the last two years Facebook has begun to track activity elsewhere on the Internet, using an addictive invention called the "Like" button. It appears on apps and websites outside Facebook and allows people to indicate with a click that they are interested in a brand, product, or piece of digital content. Since last fall, Facebook has also been able to collect data on users' online lives beyond its borders automatically: in certain apps or websites, when users listen to a song or read a news article, the information is passed along to Facebook, even if no one clicks "Like." Within the feature's first five months, Facebook catalogued more than five billion instances of people listening to songs online. Combine that kind of information with a map of the social connections Facebook's users make on the site, and you have an incredibly rich record of their lives and interactions.
"This is the first time the world has seen this scale and quality of data about human communication," Marlow says with a characteristically serious gaze before breaking into a smile at the thought of what he can do with the data. For one thing, Marlow is confident that exploring this resource will revolutionize the scientific understanding of why people behave as they do. His team can also help Facebook influence our social behavior for its own benefit and that of its advertisers. This work may even help Facebook invent entirely new ways to make money.
Contagious Information
Marlow eschews the collegiate programmer style of Zuckerberg and many others at Facebook, wearing a dress shirt with his jeans rather than a hoodie or T-shirt. Meeting me shortly before the company's initial public offering in May, in a conference room adorned with a six-foot caricature of his boss's dog spray-painted on its glass wall, he comes across more like a young professor than a student. He might have become one had he not realized early in his career that Web companies would yield the juiciest data about human interactions.
In 2001, undertaking a PhD at MIT's Media Lab, Marlow created a site called Blogdex that automatically listed the most "contagious" information spreading on weblogs. Although it was just a research project, it soon became so popular that Marlow's servers crashed. Launched just as blogs were exploding into the popular consciousness and becoming so numerous that Web users felt overwhelmed with information, it prefigured later aggregator sites such as Digg and Reddit. But Marlow didn't build it just to help Web users track what was popular online. Blogdex was intended as a scientific instrument to uncover the social networks forming on the Web and study how they spread ideas. Marlow went on to Yahoo's research labs to study online socializing for two years. In 2007 he joined Facebook, which he considers the world's most powerful instrument for studying human society. "For the first time," Marlow says, "we have a microscope that not only lets us examine social behavior at a very fine level that we've never been able to see before but allows us to run experiments that millions of users are exposed to."
Marlow's team works with managers across Facebook to find patterns that they might make use of. For instance, they study how a new feature spreads among the social network's users. They have helped Facebook identify users you may know but haven't "friended," and recognize those you may want to designate mere "acquaintances" in order to make their updates less prominent. Yet the group is an odd fit inside a company where software engineers are rock stars who live by the mantra "Move fast and break things." Lunch with the data team has the feel of a grad-student gathering at a top school; the typical member of the group joined fresh from a PhD or junior academic position and prefers to talk about advancing social science than about Facebook as a product or company. Several members of the team have training in sociology or social psychology, while others began in computer science and started using it to study human behavior. They are free to use some of their time, and Facebook's data, to probe the basic patterns and motivations of human behavior and to publish the results in academic journals—much as Bell Labs researchers advanced both AT&T's technologies and the study of fundamental physics.
It may seem strange that an eight-year-old company without a proven business model bothers to support a team with such an academic bent, but Marlow says it makes sense. "The biggest challenges Facebook has to solve are the same challenges that social science has," he says. Those challenges include understanding why some ideas or fashions spread from a few individuals to become universal and others don't, or to what extent a person's future actions are a product of past communication with friends. Publishing results and collaborating with university researchers will lead to findings that help Facebook improve its products, he adds.
For one example of how Facebook can serve as a proxy for examining society at large, consider a recent study of the notion that any person on the globe is just six degrees of separation from any other. The best-known real-world study, in 1967, involved a few hundred people trying to send postcards to a particular Boston stockholder. Facebook's version, conducted in collaboration with researchers from the University of Milan, involved the entire social network as of May 2011, which amounted to more than 10 percent of the world's population. Analyzing the 69 billion friend connections among those 721 million people showed that the world is smaller than we thought: four intermediary friends are usually enough to introduce anyone to a random stranger. "When considering another person in the world, a friend of your friend knows a friend of their friend, on average," the technical paper pithily concluded. That result may not extend to everyone on the planet, but there's good reason to believe that it and other findings from the Data Science Team are true to life outside Facebook. Last year the Pew Research Center's Internet & American Life Project found that 93 percent of Facebook friends had met in person. One of Marlow's researchers has developed a way to calculate a country's "gross national happiness" from its Facebook activity by logging the occurrence of words and phrases that signal positive or negative emotion. Gross national happiness fluctuates in a way that suggests the measure is accurate: it jumps during holidays and dips when popular public figures die. After a major earthquake in Chile in February 2010, the country's score plummeted and took many months to return to normal. That event seemed to make the country as a whole more sympathetic when Japan suffered its own big earthquake and subsequent tsunami in March 2011; while Chile's gross national happiness dipped, the figure didn't waver in any other countries tracked (Japan wasn't among them). Adam Kramer, who created the index, says he intended it to show that Facebook's data could provide cheap and accurate ways to track social trends—methods that could be useful to economists and other researchers.
Other work published by the group has more obvious utility for Facebook's basic strategy, which involves encouraging us to make the site central to our lives and then using what it learns to sell ads. An early study looked at what types of updates from friends encourage newcomers to the network to add their own contributions. Right before Valentine's Day this year a blog post from the Data Science Team listed the songs most popular with people who had recently signaled on Facebook that they had entered or left a relationship. It was a hint of the type of correlation that could help Facebook make useful predictions about users' behavior—knowledge that could help it make better guesses about which ads you might be more or less open to at any given time. Perhaps people who have just left a relationship might be interested in an album of ballads, or perhaps no company should associate its brand with the flood of emotion attending the death of a friend. The most valuable online ads today are those displayed alongside certain Web searches, because the searchers are expressing precisely what they want. This is one reason why Google's revenue is 10 times Facebook's. But Facebook might eventually be able to guess what people want or don't want even before they realize it.
Recently the Data Science Team has begun to use its unique position to experiment with the way Facebook works, tweaking the site—the way scientists might prod an ant's nest—to see how users react. Eytan Bakshy, who joined Facebook last year after collaborating with Marlow as a PhD student at the University of Michigan, wanted to test whether our Facebook friends create an "echo chamber" that amplifies news and opinions we have already heard about. So he messed with how Facebook operated for a quarter of a billion users. Over a seven-week period, the 76 million links that those users shared with each other were logged.
Then, on 219 million randomly chosen occasions, Facebook prevented someone from seeing a link shared by a friend. Hiding links this way created a control group so that Bakshy could assess how often people end up promoting the same links because they have similar information sources and interests.
He found that our close friends strongly sway which information we share, but overall their impact is dwarfed by the collective influence of numerous more distant contacts—what sociologists call "weak ties." It is our diverse collection of weak ties that most powerfully determines what information we're exposed to.
That study provides strong evidence against an idea nagging many people: that social networking creates harmful "filter bubbles," to use activist Eli Pariser's term for the effects of tuning the information we receive to match our expectations. But the study also reveals the power Facebook has. "If [Facebook's] News Feed is the thing that everyone sees and it controls how information is disseminated, it's controlling how information is revealed to society, and it's something we need to pay very close attention to," Marlow says. He points out that his team helps Facebook understand what it is doing to society and publishes its findings to fulfill a public duty to transparency. Another recent study, which investigated which types of Facebook activity cause people to feel a greater sense of support from their friends, falls into the same category.
Facebook is not above using its platform to tweak users' behavior, as it did by nudging them to register as organ donors. Unlike academic social scientists, Facebook's employees have a short path from an idea to an experiment on hundreds of millions of people.
But Marlow speaks as an employee of a company that will prosper largely by catering to advertisers who want to control the flow of information between its users. And indeed, Bakshy is working with managers outside the Data Science Team to extract advertising-related findings from the results of experiments on social influence. "Advertisers and brands are a part of this network as well, so giving them some insight into how people are sharing the content they are producing is a very core part of the business model," says Marlow.
Facebook told prospective investors before its IPO that people are 50 percent more likely to remember ads on the site if they're visibly endorsed by a friend. Figuring out how influence works could make ads even more memorable or help Facebook find ways to induce more people to share or click on its ads.
Social Engineering
Marlow says his team wants to divine the rules of online social life to understand what's going on inside Facebook, not to develop ways to manipulate it. "Our goal is not to change the pattern of communication in society," he says. "Our goal is to understand it so we can adapt our platform to give people the experience that they want." But some of his team's work and the attitudes of Facebook's leaders show that the company is not above using its platform to tweak users' behavior. Unlike academic social scientists, Facebook's employees have a short path from an idea to an experiment on hundreds of millions of people.
In April, influenced in part by conversations over dinner with his med-student girlfriend (now his wife), Zuckerberg decided that he should use social influence within Facebook to increase organ donor registrations. Users were given an opportunity to click a box on their Timeline pages to signal that they were registered donors, which triggered a notification to their friends. The new feature started a cascade of social pressure, and organ donor enrollment increased by a factor of 23 across 44 states.
Marlow's team is in the process of publishing results from the last U.S. midterm election that show another striking example of Facebook's potential to direct its users' influence on one another. Since 2008, the company has offered a way for users to signal that they have voted; Facebook promotes that to their friends with a note to say that they should be sure to vote, too. Marlow says that in the 2010 election his group matched voter registration logs with the data to see which of the Facebook users who got nudges actually went to the polls. (He stresses that the researchers worked with cryptographically "anonymized" data and could not match specific users with their voting records.)
This is just the beginning. By learning more about how small changes on Facebook can alter users' behavior outside the site, the company eventually "could allow others to make use of Facebook in the same way," says Marlow. If the American Heart Association wanted to encourage healthy eating, for example, it might be able to refer to a playbook of Facebook social engineering. "We want to be a platform that others can use to initiate change," he says.
Advertisers, too, would be eager to know in greater detail what could make a campaign on Facebook affect people's actions in the outside world, even though they realize there are limits to how firmly human beings can be steered. "It's not clear to me that social science will ever be an engineering science in a way that building bridges is," says Duncan Watts, who works on computational social science at Microsoft's recently opened New York research lab and previously worked alongside Marlow at Yahoo's labs. "Nevertheless, if you have enough data, you can make predictions that are better than simply random guessing, and that's really lucrative."
Doubling Data
Like other social-Web companies, such as Twitter, Facebook has never attained the reputation for technical innovation enjoyed by such Internet pioneers as Google. If Silicon Valley were a high school, the search company would be the quiet math genius who didn't excel socially but invented something indispensable. Facebook would be the annoying kid who started a club with such social momentum that people had to join whether they wanted to or not. In reality, Facebook employs hordes of talented software engineers (many poached from Google and other math-genius companies) to build and maintain its irresistible club. The technology built to support the Data Science Team's efforts is particularly innovative. The scale at which Facebook operates has led it to invent hardware and software that are the envy of other companies trying to adapt to the world of "big data."
In a kind of passing of the technological baton, Facebook built its data storage system by expanding the power of open-source software called Hadoop, which was inspired by work at Google and built at Yahoo. Hadoop can tame seemingly impossible computational tasks—like working on all the data Facebook's users have entrusted to it—by spreading them across many machines inside a data center. But Hadoop wasn't built with data science in mind, and using it for that purpose requires specialized, unwieldy programming. Facebook's engineers solved that problem with the invention of Hive, open-source software that's now independent of Facebook and used by many other companies. Hive acts as a translation service, making it possible to query vast Hadoop data stores using relatively simple code. To cut down on computational demands, it can request random samples of an entire data set, a feature that's invaluable for companies swamped by data. Much of Facebook's data resides in one Hadoop store more than 100 petabytes (a million gigabytes) in size, says Sameet Agarwal, a director of engineering at Facebook who works on data infrastructure, and the quantity is growing exponentially. "Over the last few years we have more than doubled in size every year," he says. That means his team must constantly build more efficient systems.
One potential use of Facebook's data storehouse would be to sell insights mined from it. Such information could be the basis for any kind of business. Assuming Facebook can do this without upsetting users and regulators, it could be lucrative.
All this has given Facebook a unique level of expertise, says Jeff Hammerbacher, Marlow's predecessor at Facebook, who initiated the company's effort to develop its own data storage and analysis technology. (He left Facebook in 2008 to found Cloudera, which develops Hadoop-based systems to manage large collections of data.) Most large businesses have paid established software companies such as Oracle a lot of money for data analysis and storage. But now, big companies are trying to understand how Facebook handles its enormous information trove on open-source systems, says Hammerbacher. "I recently spent the day at Fidelity helping them understand how the 'data scientist' role at Facebook was conceived ... and I've had the same discussion at countless other firms," he says.
As executives in every industry try to exploit the opportunities in "big data," the intense interest in Facebook's data technology suggests that its ad business may be just an offshoot of something much more valuable. The tools and techniques the company has developed to handle large volumes of information could become a product in their own right.
Mining for Gold
Facebook needs new sources of income to meet investors' expectations. Even after its disappointing IPO, it has a staggeringly high price-to-earnings ratio that can't be justified by the barrage of cheap ads the site now displays. Facebook's new campus in Menlo Park, California, previously inhabited by Sun Microsystems, makes that pressure tangible. The company's 3,500 employees rattle around in enough space for 6,600. I walked past expanses of empty desks in one building; another, next door, was completely uninhabited. A vacant lot waited nearby, presumably until someone invents a use of our data that will justify the expense of developing the space.
One potential use would be simply to sell insights mined from the information. DJ Patil, data scientist in residence with the venture capital firm Greylock Partners and previously leader of LinkedIn's data science team, believes Facebook could take inspiration from Gil Elbaz, the inventor of Google's AdSense ad business, which provides over a quarter of Google's revenue. He has moved on from advertising and now runs a fast-growing startup, Factual, that charges businesses to access large, carefully curated collections of data ranging from restaurant locations to celebrity body-mass indexes, which the company collects from free public sources and by buying private data sets. Factual cleans up data and makes the result available over the Internet as an on-demand knowledge store to be tapped by software, not humans.
Customers use it to fill in the gaps in their own data and make smarter apps or services; for example, Facebook itself uses Factual for information about business locations. Patil points out that Facebook could become a data source in its own right, selling access to information compiled from the actions of its users. Such information, he says, could be the basis for almost any kind of business, such as online dating or charts of popular music. Assuming Facebook can take this step without upsetting users and regulators, it could be lucrative. An online store wishing to target its promotions, for example, could pay to use Facebook as a source of knowledge about which brands are most popular in which places, or how the popularity of certain products changes through the year.
Hammerbacher agrees that Facebook could sell its data science and points to its currently free Insights service for advertisers and website owners, which shows how their content is being shared on Facebook. That could become much more useful to businesses if Facebook added data obtained when its "Like" button tracks activity all over the Web, or demographic data or information about what people read on the site. There's precedent for offering such analytics for a fee: at the end of 2011 Google started charging $150,000 annually for a premium version of a service that analyzes a business's Web traffic.
Back at Facebook, Marlow isn't the one who makes decisions about what the company charges for, even if his work will shape them. Whatever happens, he says, the primary goal of his team is to support the well-being of the people who provide Facebook with their data, using it to make the service smarter. Along the way, he says, he and his colleagues will advance humanity's understanding of itself. That echoes Zuckerberg's often doubted but seemingly genuine belief that Facebook's job is to improve how the world communicates. Just don't ask yet exactly what that will entail. "It's hard to predict where we'll go, because we're at the very early stages of this science," says Marlow. "The number of potential things that we could ask of Facebook's data is enormous."
Tom Simonite is Technology Review's senior IT editor.
Subscribe to:
Posts (Atom)


