Friday, May 6, 2011

FT: A binary goldmine

Financial Times, May 5 2011

Displaying up-to-the-minute information on everything from train times to cinema schedules, apps have in short order become a ubiquitous feature of smartphones. To most users, they are simply useful and entertaining tools.

As well as providing users with information, however, these mobile software applications are also insatiable data-gatherers. Even the most mundane apps often collect a surprising amount from handsets just to do their jobs.

This has put them at the forefront of a fast-evolving science based on the business use of consumer data. If there is commercial advantage to be gained, it seems, almost nothing is too insignificant to be collected and analysed.

Take an app launched recently by Color, one of the most ambitious and best financed of the crop of start-ups that has sprung up in Silicon Valley to cash in on the smartphone boom. Pictures taken by users are mixed into streams with those taken by others who are nearby, or with whom users are often in contact, building ad hoc social networks.

The software taps deeply into handsets, drawing on components such as Global Positioning System chips, gyroscopes and accelerometers to pinpoint where they are, how fast they are moving and which way up they are being held. The lighting conditions in pictures taken with the gadgets, along with the digital “fingerprints” of surrounding noises coming through their microphones, provides other useful crumbs of information.

Thus informed, Color can work out precisely who the user is walking down the street with, says Bill Nguyen, the serial entrepreneur behind the company.

Such innovations are the tip of a data iceberg. Smartphones, social networks and other accoutrements of modern digital life are generating vast new data sets that are revving up the digital economy.

Accompanying all this is a trend that has given the technology lexicon a new term: big data. Rather than sampling only small parts of the digital data deluge, modern companies have a new option: they can study all of it.

The ability to capture and analyse this mass of information is throwing up business ideas and altering the relationship between businesses and their customers.

TECHNOLOGICAL TOOLS
Messy, lacks structure but holds promise: an interim report card on big data

The idea of amassing a large amount of data and scrutinising it for clues about customer behaviour is not new in business. Data mining has long been used by big companies with access to enough computing power, such as credit card issuers and retailers.

But the falling cost of technology and the generation of much more digital information have opened the field to a much wider group of businesses and made it possible to make more informed judgments about customer behaviour.

While “big data” has become the buzzword, a better description would be “messy data”, says Roger Ehrenberg of IA Ventures, an early-stage investor. Harvesting, cleaning up and organising raw data in a way that it can be processed is a large part of the battle, he says.

This has been complicated further by the big growth in unstructured data – information, such as text, that is not organised in a way that a computer can easily process. With the volume of user-generated text and video growing rapidly, this has become one of the main focuses of technological development.

Chief among the new tools are natural language processing, which enables a computer to extract meaning from text, and machine learning, the feedback loops through which computers can test their conclusions on large amounts of data in order progressively to refine their results.

Subjecting large data sets to analysis has also been made easier by two of the forces that have reshaped information technology more widely: the spread of low-cost, standardised computer hardware and the emergence of open-source software.

This has created a cheap computing platform for new technologies such as Hadoop – a piece of software architecture that is designed to handle massive amounts of data. The idea was based on breakthroughs at Google, which needed to find ways to conduct large volumes of intensive web searches simultaneously. It has since been taken up by companies including Facebook and Yahoo.

The rise of cloud computing – which centralises storage and processing power in larger data centres – has also brought big data within the reach of more companies. By tapping into the cloud computing services offered by Amazon, say, a company such as Color can get instant access to all the analytical power it needs without needing to take on the fixed costs of buying its own servers, says D.J. Patil, chief product officer at the IT start-up.

It is also stoking simmering privacy concerns. When Steve Jobs, Apple chief executive, was forced to apologise last week over the handling of data about the location of iPhone and iPad owners, it touched a raw public nerve and resulted in immediate Congressional hearings in Washington.

Color says it plans to use the information it collects to create new services for its customers. By combining it with data from social networks, says DJ Patil, chief product officer, it can tell its users: “Here are people who are near you, and here is how you might know them.” He says the company has no plans, at least for now, to use the information for other purposes, such as sending targeted advertising to customers.

However, with the tide of digital information rising fast – and more sophisticated ways being found to make business use of it – many companies are already being drawn into the new world of sophisticated data collection and analysis.

Some are using it to tailor their own products more precisely to the preferences of their users; others to target advertising of their products more accurately. Some are also selling the data they gather from their customers to the brokers and aggregators who act as middlemen in data markets that have sprung up to recycle such information.

A new consensus is needed to govern the use of this increasingly valuable commodity, says Michele Luzi of management consultancy Bain & Company, which conducted a study for the World Economic Forum on the issue. “Ultimately, you have to have a system of rights,” he says – something that balances the valid, but often conflicting, interests of individuals, governments and businesses.

While lawmakers and regulators on both sides of the Atlantic are becoming more exercised, such an agreement – not to mention the infrastructure and regulations to support it – remains some way off.
Meanwhile, as the analysis of digital information develops, the traditional management virtues of gut instinct and seat-of-the-pants decision-making are being replaced by reliance on intensive number-crunching and the objective testing of multiple potential courses of action.

For business leaders, “the big skill in future will be to ask the right question”, says Tim O’Reilly, a technology commentator and publisher.

Besides smartphones, new sources of data include social networks, blogs and other sources of user-generated content; sensors collecting everything from traffic patterns to a user’s heart rhythm; and click streams generated by people spending an increasing amount of their lives online.

Much of the information is in unstructured form. It has never been collated in a traditional relational database, where it could be queried at will. Without techniques to harvest, verify and analyse it – often in real time – valuable commercial signals are lost in the noise.

It sometimes takes the analysis of massive data sets to detect useful patterns, says Michael Olson. His California start-up, Cloudera, is commercialising the type of technology used by companies such as Facebook and Yahoo to crunch through vast bodies of information. Retailers, for instance, might learn far more from the 10 years’ worth of customer data they can now analyse in one go than from the more limited runs to which they were once restricted, he says.
. . .
Companies born in the digital age are often highly attuned to the possibilities presented by these untapped reservoirs of digital information. Like Color, they place data collection and analysis at their core, and build their business processes on their skills in these areas.

Even businesses whose roots appear to lie in the creative industries now treat data collection and analysis on a vast scale as a core skill. Zynga, which has produced online gaming hits such as FarmVille, believes that by analysing in painstaking detail what users do on its site it can perfect the experience. It mines these data, for instance, to model how users are likely to respond to new features. Zynga’s executives suggest this will ultimately leave it less exposed to the hit-and-miss nature of the games industry (a theory that has yet to be put to the test.)

“Their uniqueness is really in the massive pool of data that is growing every day,” says Theresia Gouw Ranzetta of Accel Partners, a venture capital firm that has backed a Zynga rival that uses similar techniques.

Data mining has long been central to fields such as credit-card marketing. The difference today is that it is becoming available to many more companies at lower cost. In addition, highly valuable new classes of information are emerging.

Social data, in particular, are at the centre of something of a gold rush. Gaining insight into the new rubric of online behaviour ushered in by sites such as Facebook and Twitter has the potential to create business fortunes.

Better-established companies are jumping on the bandwagon. Giant US retailer Walmart is the latest to join the fray, last month buying Kosmix, a Silicon Valley company that filters the deluge of messages on Twitter. Walmart’s understanding of its customers has hitherto been limited to data about purchasing histories and browsing habits, says Ms Ranzetta, who is a member of the Kosmix board. In future, it will be able to tap into information on their personal preferences and interests as well.

Such mining of Twitter and other social sites is being used in a wide range of industries. Roger Ehrenberg, a former hedge fund manager who invests in technology start-ups, says demand is high among financial traders for help with assembling masses of data, or for refining them so that “what’s coming through is a signal, rather than a raw feed”.

Filtering tweets in real time for practical information is one of the most challenging of these tasks, both Mr Ehrenberg and Ms Ranzetta say. Often, it is only when data from such sources are combined with other information that their value emerges.

For instance, combining details of senior management moves revealed in companies’ regulatory filings with changes to profiles on LinkedIn, the business networking site for professionals, may yield valuable insights into what is happening inside companies, says Mr Ehrenberg.

Crunching through vast data sets can reveal patterns in fields far removed from the financial markets that would not otherwise be visible. The result is the rise of techniques such as behavioural clustering (grouping people on the basis of common behavioural characteristics, rather than more traditional demographics) and look-alike marketing (marketing to a particular user based on previous successes in marketing to others with similar profiles).

Comparing people in this way may have many uses. Analysing the detailed financial behaviour of very large groups of customers over a protracted period, for instance, could give banks a clue as to which are most likely to default next, says Mr Olson at Cloudera.

Such uses of predictive analytics – a marriage of statistical modelling and data mining – raise troubling questions. Is it fair, for instance, to judge a person merely on a prediction of their future behaviour?
And what are the long-term consequences of using such analyses to categorise people ever more narrowly, shaping the types of information and advertising they are fed online? Will this lead to a form of digital determinism, in which it becomes hard to escape a life that has been preordained by some giant bank, retailer or government department?

While the use of these techniques is still in its infancy, the digital crumbs of personal information left scattered across the web are already being swept up and used with surprising results.

“If I examine any new data set, the chances are I can find something in that data that has predictive value,” says Frank Rotman, a former head of analytics at Capital One, a US financial company that was a pioneer in the field. He says existing laws about how credit decisions are made, along with current social norms, place limits on how this information is used.

The rules are laxer, however, when it comes to how credit is marketed in the first place. And, ultimately, the opacity of this largely unregulated field makes it hard to tell exactly which signals from the digital morass are being used to inform business or government decisions that have a direct bearing on many lives.

“Where it gets murky and scary is the stuff that’s being sucked out of the social system, where you have no idea how it is being used,” says Chris Larsen, chief executive of Prosper, a web service through which individuals lend to each other directly.

Innocent actions from everyday life, revealed on social networks, may become significant signals when used in a different context.

Simply failing to meet a social commitment, for instance, might turn out to be a leading indicator of a person’s reliability or otherwise as a borrower, says Mr Larsen. “There’s no question people are working on these algorithms, and trying to sell [them].”

As with many other uses of big data, the full potential of this science is still only dimly understood.
However, one thing is clear. Few consumers today are likely to welcome some of the applications that are already possible. As Mr Rotman says: “If you say you decline someone [for credit] because they have blue eyes, how will that go down?”