Tuesday, February 19, 2013

Why job creation is so hard

Why job creation is so hard

By Robert J. Samuelson, Published: February 17, WaPo

President Obama and the Democrats want more jobs. So do Republicans. Heck, everyone does. Yet, job creation is weak. It's true that the economy has generated 5.5 million jobs from its low point. Still, there are 3.2 million fewer jobs now than at the previous high. The official unemployment rate is 7.9 percent, but it would be 14.4 percent if it included part-timers who would like full-time work and discouraged workers who have stopped looking, notes Janet Yellen, vice chair of the Federal Reserve Board. Scarce jobs are the nation's first, second and third most important economic and social problem.

What's especially disheartening and mystifying is that, until now, job creation was considered an inherent strength of the U.S. economy. Despite some years of recession-induced joblessness, unemployment averaged 5.6 percent from 1950 to 2007. The Congressional Budget Office doesn't expect it to fall below 7.5 percent until 2015. That would make six years above 7.5 percent — the longest stretch of high joblessness in 70 years. It has defied massive budget deficits and ultra-low interest rates.

Something's changed in how the economy works. One theory is "deleveraging": Americans paying down their high debt. The economy won't accelerate until this process is complete, the argument goes; the fact that debt-service ratios have dropped to early 1990s levels is considered a good omen. Another approach is to examine the economy by sectors and see which ones are lagging compared with past recoveries. Yellen did this and indicted housing (its deep slump) and state and local governments (spending cuts). Again, there are said to be encouraging signs. Home construction, prices and sales are up; state and local spending is stabilizing.

This analysis helps but misses the main story. To overgeneralize slightly: We have gone from being an expansive, risk-taking society to a skittish, risk-averse one. Before the 2008-09 financial crisis, the bias was toward more spending. The inclination was to surrender to immediate gratification. Want a new car? Sure, why not? More meals out? Great idea! Businesses behaved similarly. Banks made the next loan; companies hired the next worker and approved the next investment project. An ever-expanding economy justified optimism, and optimism supported an ever-expanding economy. Hello, bubble.

The psychology has now reversed. The bias is against extra spending. Eat out? Try leftovers. Remodel the basement? Oh, leave it alone. In the boom years, the personal saving rate (savings as a share of after-tax income) fell from 10.9 percent in 1982 to 1.5 percent in 2005. Now it's edging up; from 2010 to 2012, it averaged 4.4 percent. It could go higher, imposing a further drag on the economy.

Businesses have also retreated. They resist approving the next loan, job hire or investment. Since 1959, business investment in factories, offices and equipment has averaged 11 percent of the economy (gross domestic product) and peaked at nearly 13 percent. It's now a shade over 10 percent, reports economist Nigel Gault of IHS Global Insight.

Note that these attitudes govern sectors accounting for roughly four-fifths of the economy: Consumer spending is about 70 percent of GDP; business investment is the rest. They dwarf housing construction, which is about 2.5 percent of GDP. The caution and risk-aversion aren't so great as to cause a recession, but on the margin they have limited the economy's expansion to rates — lately, 1 percent to 2 percent — too weak to absorb most jobless. Pessimism produces a sluggish economy; a sluggish economy produces pessimism. That's the main explanation of poor job creation.

As I've written before, this psychological shift stemmed from the fact that the financial crisis and Great Recession were largely unpredicted. Americans aren't just deleveraging. They're also building wealth to protect themselves against unknown dangers. Perhaps the stock market's recent assault on record highs signals restored confidence, but remember: The market is simply regaining levels of late 2007. A report from Credit Suisse argues that returns to stocks will average about 3.5 percent annually (after inflation) in the next 20 years, down sharply from 6 percent since 1950. To compensate for lower returns, companies would need to contribute more to pensions. Wages would suffer. Consumption spending would weaken.

We are hostage to a stubborn, restraining psychology. There's no obvious fix for slow job growth, precisely because it requires a change in public mood or some autonomous source of added demand — a burst of exports, investment in new technologies — not easily predicted or controlled. It could happen but is hardly guaranteed. Politics does matter, to a point. Constant budget and tax feuds between the White House and Congress spawn uncertainty and subvert confidence. Obamacare's disincentives to hiring hurt, though how much is unclear. But grandiose solutions, say infrastructure spending, founder on practicality. A meaningful level of projects would take time to start and add excessively to budget deficits. We are waiting and hoping.

Read more from Robert Samuelson's archive.

 

Big hopes for big data spending

http://www.ft.com/cms/s/2/1a780a6c-648b-11e2-934b-00144feab49a.html#ixzz2LLfIMtGV

Big hopes for big data spending

Data are often described as the oil of the 21st century, so it should not be surprising that governments around the world are looking to tap that resource through the use of big-data technologies.

The potential is huge – big data is a market so large and diverse that it has yet to be valued in any meaningful way. But in an age of austerity and increased privacy awareness, can the spending be as big as the data?

Forrester, the technology research company, defines big data as “techniques and technologies that make capturing value from data at extreme scale economical”. In other words, number crunching that a few years ago would have taken weeks can now be done in real time.

In the US, President Barack Obama, whose use of big data on the campaign trail helped him to a second term in office, last year announced $200m in funding to establish the Big Data Research and Development Initiative with input from the country’s science, health and defence agencies.

In 2009, the UN established Global Pulse to track the effects of socioeconomic crises around the world by using huge amounts of data and real-time analytics technology. The idea is that trends will be identified that can better protect populations from such shocks as the global financial crisis. So far, Global Pulse has managed only consultations and proofs of concept, but the organisation says it will begin new projects this year.

In the UK, where cutting the budget deficit is high on the agenda, a recent report by Policy Exchange, a think-tank, estimated that using big data could save the public sector up to £33bn a year in lost time – equivalent to between £250 and £500 per head of the population.

Stephen Kelly, the UK chief operating officer for government, says big data is a key part of his strategy to take £20bn off the £140bn spent each year on information technology, telecoms and property management.

Kelly says governments have lost their position as early adopters of technology. “In the 1980s, government was a pioneer of data innovation – in the US too – and the private sector followed. There was a dearth of technology in consumer markets. In the past couple of years it’s flipped on its head, but the good news for us is that we can leapfrog into 2013.”

Now, the UK government is using the latest analytics technology to zero in on fraud, errors caused by the re-keying of data, and chasing debts.

“We see a £20bn-£40bn leakage from fraud, error and debt. That’s an area where we have started looking at prototypes around analytics and data, and that will play a bigger role,” Mr Kelly says.

Big data is transforming procurement costs, too, he says. A £350m Oracle-based database, for example, can now be replaced by a cloud service at a 10th of the cost.

The change is even more rapid in the health sector, where the NHS has been asked to save £20bn a year by 2015.

“For most hospitals one of the things we need to focus on is how do we get down to almost patient-level costing – how do we understand the true cost of every intervention, every bed day of a patient’s stay?” says Orlando Agrippa, director of business informatics at Colchester Hospital University NHS Foundation Trust.

Using a cloud-based tool called Qlikview, Colchester staff analyse information about diagnostics, operating theatre throughput, outpatient and inpatient stays, waiting times and emergency services.

“Every single day the service managers accountable for wait times use Qlikview as the only source of validating the patients who are on or off the list and need to be treated quickly,” Mr Agrippa says.

The UK is not alone. In a trial in Stockholm, Sweden, authorities were able to cut road travel times in half and reduce emissions by 20 per cent by analysing 250,000 GPS signals per second to direct traffic. A permanent gain on this scale would have a marked effect on productivity and, in theory, economic output.

The UK government’s hopes for big data are not restricted to the public sector. With the manufacturing and finance industries in retreat, many hope data will spawn a new industry to power economic growth. The UK’s Open Data Institute was founded with this in mind.

According to Gartner, the technology research company, big data drove $28bn of IT spending in 2012, and will lead to the creation of 4.4m jobs by 2015.

Further research by Forrester suggests 20 per cent of companies have implemented big data technology, and 37 per cent are planning a big data project.

Nevertheless, economic growth is not assured. Researchers say some big data technologies are between two and five years from mainstream adoption, and privacy remains the biggest obstacle to potential business models.

Holger Kisker, a principal analyst at Forrester and author of its forthcoming Forrsights big data survey, says practical privacy laws are the key to unlocking a range of new products and services that can drive growth.

“We are making a lot of problems for ourselves by going into this with a very high sensitivity of data privacy issues,” he says.

Mr Kisker says policy makers in Europe have given no timeline for proposed privacy and cybersecurity legislation, and that this is holding back investment in new business models around customer data.

“In the US, companies are much more open to taking risks, but in the UK and Europe, companies are holding back and we are moving to a situation of competitive disadvantage over time.”

 

Monday, February 18, 2013

David Brooks: What Data Can’t Do

February 18, 2013

What Data Can't Do

By DAVID BROOKS

Not long ago, I was at a dinner with the chief executive of a large bank. He had just had to decide whether to pull out of Italy, given the weak economy and the prospect of a future euro crisis.

The C.E.O. had his economists project out a series of downside scenarios and calculate what they would mean for his company. But, in the end, he made his decision on the basis of values.

His bank had been in Italy for decades. He didn't want Italians to think of the company as a fair-weather friend. He didn't want people inside the company thinking they would cut and run when times got hard. He decided to stay in Italy and ride out any potential crisis, even with the short-term costs.

He wasn't oblivious to data in making this decision, but ultimately, he was guided by a different way of thinking. And, of course, he was right to be. Commerce depends on trust. Trust is reciprocity coated by emotion. People and companies that behave well in tough times earn affection and self-respect that is extremely valuable, even if it is hard to capture in data.

I tell this story because it hints at the strengths and limitations of data analysis. The big novelty of this historic moment is that our lives are now mediated through data-collecting computers. In this world, data can be used to make sense of mind-bogglingly complex situations. Data can help compensate for our overconfidence in our own intuitions and can help reduce the extent to which our desires distort our perceptions.

But there are many things big data does poorly. Let's note a few in rapid-fire fashion:

Data struggles with the social. Your brain is pretty bad at math (quick, what's the square root of 437), but it's excellent at social cognition. People are really good at mirroring each other's emotional states, at detecting uncooperative behavior and at assigning value to things through emotion.

Computer-driven data analysis, on the other hand, excels at measuring the quantity of social interactions but not the quality. Network scientists can map your interactions with the six co-workers you see during 76 percent of your days, but they can't capture your devotion to the childhood friends you see twice a year, let alone Dante's love for Beatrice, who he met twice.

Therefore, when making decisions about social relationships, it's foolish to swap the amazing machine in your skull for the crude machine on your desk.

Data struggles with context. Human decisions are not discrete events. They are embedded in sequences and contexts. The human brain has evolved to account for this reality. People are really good at telling stories that weave together multiple causes and multiple contexts. Data analysis is pretty bad at narrative and emergent thinking, and it cannot match the explanatory suppleness of even a mediocre novel.

Data creates bigger haystacks. This is a point Nassim Taleb, the author of "Antifragile," has made. As we acquire more data, we have the ability to find many, many more statistically significant correlations. Most of these correlations are spurious and deceive us when we're trying to understand a situation. Falsity grows exponentially the more data we collect. The haystack gets bigger, but the needle we are looking for is still buried deep inside.

One of the features of the era of big data is the number of "significant" findings that don't replicate the expansion, as Nate Silver would say, of noise to signal.

Big data has trouble with big problems. If you are trying to figure out which e-mail produces the most campaign contributions, you can do a randomized control experiment. But let's say you are trying to stimulate an economy in a recession. You don't have an alternate society to use as a control group. For example, we've had huge debates over the best economic stimulus, with mountains of data, and as far as I know not a single major player in this debate has been persuaded by data to switch sides.

Data favors memes over masterpieces. Data analysis can detect when large numbers of people take an instant liking to some cultural product. But many important (and profitable) products are hated initially because they are unfamiliar.

Data obscures values. I recently saw an academic book with the excellent title, " 'Raw Data' Is an Oxymoron." One of the points was that data is never raw; it's always structured according to somebody's predispositions and values. The end result looks disinterested, but, in reality, there are value choices all the way through, from construction to interpretation.

This is not to argue that big data isn't a great tool. It's just that, like any tool, it's good at some things and not at others. As the Yale professor Edward Tufte has said, "The world is much more interesting than any one discipline."

 

Big Data: Obama Seeking to Boost Study of Human Brain

February 17, 2013

Obama Seeking to Boost Study of Human Brain

By JOHN MARKOFF

The Obama administration is planning a decade-long scientific effort to examine the workings of the human brain and build a comprehensive map of its activity, seeking to do for the brain what the Human Genome Project did for genetics.

The project, which the administration has been looking to unveil as early as March, will include federal agencies, private foundations and teams of neuroscientists and nanoscientists in a concerted effort to advance the knowledge of the brain’s billions of neurons and gain greater insights into perception, actions and, ultimately, consciousness.

Scientists with the highest hopes for the project also see it as a way to develop the technology essential to understanding diseases like Alzheimer’s and Parkinson’s, as well as to find new therapies for a variety of mental illnesses.

Moreover, the project holds the potential of paving the way for advances in artificial intelligence.

The project, which could ultimately cost billions of dollars, is expected to be part of the president’s budget proposal next month. And, four scientists and representatives of research institutions said they had participated in planning for what is being called the Brain Activity Map project.

The details are not final, and it is not clear how much federal money would be proposed or approved for the project in a time of fiscal constraint or how far the research would be able to get without significant federal financing.

In his State of the Union address, President Obama cited brain research as an example of how the government should “invest in the best ideas.”

“Every dollar we invested to map the human genome returned $140 to our economy — every dollar,” he said. “Today our scientists are mapping the human brain to unlock the answers to Alzheimer’s. They’re developing drugs to regenerate damaged organs, devising new materials to make batteries 10 times more powerful. Now is not the time to gut these job-creating investments in science and innovation.”

Story C. Landis, the director of the National Institute of Neurological Disorders and Stroke, said that when she heard Mr. Obama’s speech, she thought he was referring to an existing National Institutes of Health project to map the static human brain. “But he wasn’t,” she said. “He was referring to a new project to map the active human brain that the N.I.H. hopes to fund next year.”

Indeed, after the speech, Francis S. Collins, the director of the National Institutes of Health, may have inadvertently confirmed the plan when he wrote in a Twitter message: “Obama mentions the #NIH Brain Activity Map in #SOTU.”

A spokesman for the White House Office of Science and Technology Policy declined to comment about the project.

The initiative, if successful, could provide a lift for the economy. “The Human Genome Project was on the order of about $300 million a year for a decade,” said George M. Church, a Harvard University molecular biologist who helped create that project and said he was helping to plan the Brain Activity Map project. “If you look at the total spending in neuroscience and nanoscience that might be relative to this today, we are already spending more than that. We probably won’t spend less money, but we will probably get a lot more bang for the buck.”

Scientists involved in the planning said they hoped that federal financing for the project would be more than $300 million a year, which if approved by Congress would amount to at least $3 billion over the 10 years.

The Human Genome Project cost $3.8 billion. It was begun in 1990 and its goal, the mapping of the complete human genome, or all the genes in human DNA, was achieved ahead of schedule, in April 2003. A federal government study of the impact of the project indicated that it returned $800 billion by 2010.

The advent of new technology that allows scientists to identify firing neurons in the brain has led to numerous brain research projects around the world. Yet the brain remains one of the greatest scientific mysteries.

Composed of roughly 100 billion neurons that each electrically “spike” in response to outside stimuli, as well as in vast ensembles based on conscious and unconscious activity, the human brain is so complex that scientists have not yet found a way to record the activity of more than a small number of neurons at once, and in most cases that is done invasively with physical probes.

But a group of nanotechnologists and neuroscientists say they believe that technologies are at hand to make it possible to observe and gain a more complete understanding of the brain, and to do it less intrusively.

In June in the journal Neuron, six leading scientists proposed pursuing a number of new approaches for mapping the brain.

One possibility is to build a complete model map of brain activity by creating fleets of molecule-size machines to noninvasively act as sensors to measure and store brain activity at the cellular level. The proposal envisions using synthetic DNA as a storage mechanism for brain activity.

“Not least, we might expect novel understanding and therapies for diseases such as schizophrenia and autism,” wrote the scientists, who include Dr. Church; Ralph J. Greenspan, the associate director of the Kavli Institute for Brain and Mind at the University of California, San Diego; A. Paul Alivisatos, the director of the Lawrence Berkeley National Laboratory; Miyoung Chun, a molecular geneticist who is the vice president for science programs at the Kavli Foundation; Michael L. Roukes, a physicist at the California Institute of Technology; and Rafael Yuste, a neuroscientist at Columbia University.

The Obama initiative is markedly different from a recently announced European project that will invest 1 billion euros in a Swiss-led effort to build a silicon-based “brain.” The project seeks to construct a supercomputer simulation using the best research about the inner workings of the brain.

Critics, however, say the simulation will be built on knowledge that is still theoretical, incomplete or inaccurate.

The Obama proposal seems to have evolved in a manner similar to the Human Genome Project, scientists said. “The genome project arguably began in 1984, where there were a dozen of us who were kind of independently moving in that direction but didn’t really realize there were other people who were as weird as we were,” Dr. Church said.

However, a number of scientists said that mapping and understanding the human brain presented a drastically more significant challenge than mapping the genome.

“It’s different in that the nature of the question is a much more intricate question,” said Dr. Greenspan, who said he is involved in the brain project. “It was very easy to define what the genome project’s goal was. In this case, we have a more difficult and fascinating question of what are brainwide activity patterns and ultimately how do they make things happen?”

The initiative will be organized by the Office of Science and Technology Policy, according to scientists who have participated in planning meetings.

The National Institutes of Health, the Defense Advanced Research Projects Agency and the National Science Foundation will also participate in the project, the scientists said, as will private foundations like the Howard Hughes Medical Institute in Chevy Chase, Md., and the Allen Institute for Brain Science in Seattle.

A meeting held on Jan. 17 at the California Institute of Technology was attended by the three government agencies, as well as neuroscientists, nanoscientists and representatives from Google, Microsoft and Qualcomm. According to a summary of the meeting, it was held to determine whether computing facilities existed to capture and analyze the vast amounts of data that would come from the project. The scientists and technologists concluded that they did.

They also said that a series of national brain “observatories” should be created as part of the project, like astronomical observatories.

 

Friday, February 15, 2013

Big Data and Health: Numbers, Numbers and More Numbers

  • JOURNAL REPORTS
  • Updated February 14, 2013, 9:46 p.m. ET
Numbers, Numbers and More Numbers
Health-care players are finding that crunching the numbers can pay off in both better care and lower costs
Under pressure to do more with less, insurers, pharmacy benefit managers and health-care providers are all pushing data analysis to new heights.
More in Innovations in Health Care
Insurers have been crunching numbers for years to figure out which patients are most likely to generate high costs. Now other groups are gauging probabilities of relapses, and the likelihood of a patient's not taking his or her medicine. Using models that draw on massive troves of medical and other data, some are also focusing on seemingly healthy individuals, trying to prevent problems before they occur.
Several insurers, including UnitedHealth Group Inc. and WellPoint Inc., are seeking to pinpoint who will develop conditions such as diabetes. Pharmacy-benefit managers such as Express Scripts Inc. and CVS Caremark Corp. are working on programs to predict medication compliance. Care providers, meanwhile, are trying to identify who is most likely to be admitted—or readmitted—to a hospital, and are adjusting their care to prevent such return visits.
Proponents say each of these activities can improve care and prevent further expensive health problems, reducing costs for companies and patients.
Jad Ryherd
BUILDING COMPLIANCE Express Scripts Bob Nease aims to get more patients to take their medicine.
"There's a huge increase in demand for health care, but the delivery model is inherently inefficient," says International Data Corp. Health Insights analyst Scott Lundstrom. "Companies are looking to how to optimize the quality of almost everything."
Here's a look at new ways in which health-care companies are using data:
Insurers
Perhaps the best-known health-care analytics push is WellPoint's partnership with International Business Machines Corp. IBM created a supercomputer, dubbed Watson, that responds to voice commands and last year won the TV quiz show "Jeopardy!" by quickly scouring its database for the right answers. WellPoint plans to use Watson's data-crunching technology to help suggest treatment options to doctors, based on medical records, research databases and other sources. "Research doubles every five years," says Lori Beer, executive vice president of enterprise business services at WellPoint. "How does a physician keep up with that?"
Memorial Sloan-Kettering Cancer Center in New York, too, is working with IBM to build a tool for cancer treatment that will draw on patient histories, the center's clinical knowledge and molecular and genomic data.
In a separate project, WellPoint nurses and case managers pore over patient records and claims histories to see whether patients are following doctors' orders. This is particularly important for people with chronic conditions like diabetes, which are managed more effectively with preventive care.
Pharmacy Benefit Managers
Express Scripts this summer plans to roll out a program that will study data about patients in its customer base—which includes employers, government agencies, unions and health plans—to identify and intervene with patients less likely to take their prescriptions correctly. The program will start with medicine for high blood pressure, diabetes, high cholesterol, asthma and osteoporosis, and later include multiple sclerosis.
For plan sponsors that choose to participate, Express Scripts first runs its predictive models. Looking at about 400 variables, including enrollees' prescription history and the economic makeup of their neighborhood, the program tries to identify who is likely to stop taking their medicine and why. Bob Nease, chief scientist at Express Scripts, says the company's testing has shown its predictions are correct about 90% of the time.
Patients identified as likely to be noncompliant then receive tailored interventions to help overcome their specific hurdle. Some patients will receive over-the-phone pharmacist consultations and help signing up for auto-refills and other programs that simplify adherence. For patients not taking their medicine because of behavioral reasons, mainly forgetfulness, timers are attached to the caps of prescription bottles to remind them to take their medication.
Among patients tested in a control group, Dr. Nease says, using the beeping caps increased compliance by about 2%. But when timers were given to those identified as forgetful, there was a 16% rise in compliance.
Care Providers
Many care providers try to use analytics to identify who is at risk of a hospital admission.
Heritage Provider Network, a Northridge, Calif., physicians group that develops health-care delivery networks, is sponsoring a competition to figure out the best algorithm to identify which patients are likely to be sent to a hospital within the next year, based on Heritage patient data. The competition, which has a $3 million prize, will end in April 2013.
"Hospitalizations are very expensive and cost this country a lot of the resources we're using in health care," says Richard Merkin, president and chief executive of Heritage. "Imagine if you could effectively predict who was going to be hospitalized," Dr. Merkin says. "You could reallocate resources to prevent unnecessary hospitalization and put those resources to use for cure rather than care."
Health Management Associates Inc., which operates nearly 70 hospitals across the U.S., also is working with technology partners on predicting readmissions. It is looking for patterns in its own patient and operational data and ranking the likelihoods of admission associated with different factors. The company, based in Naples, Fla., is also trying to predict demand for certain services, such as the emergency room and lab work, to help it better staff its hospitals.
"Companies have to do more with less," says Eric Waller, chief marketing and strategy officer for HMA. "We're being incentivized as an industry to improve clinical care and increase customer satisfaction and improve margins with lower reimbursements. It's driving people to be much more interested in the use of analytics."
As more patients seek care, analytics will also help providers offer more efficient preventive care. "There's a huge amount of waste in the system," says Dr. Nease, the Express Scripts chief scientist. "Advanced analytics allows you to be much more sophisticated in where you intervene and with what."
Ms. Tibken is a reporter for Dow Jones Newswires in New York. She can be reached at reports@wsj.com.
 
 

Monday, February 11, 2013

Crowdsourcing, Big data and Bio-sciences

Quote: demonstrated that a crowdsourcing platform pioneered in the commercial sector can solve a complex biological problem more quickly than conventional approaches—and at a fraction of the cost.

Solving Big-Data Bottleneck

Scientists team with business innovators to tackle research hurdles

By DAVID CAMERON

February 7, 2013

http://hms.harvard.edu/news/solving-big-data-bottleneck-2-7-13?utm_source=SilverpopMailing&utm_medium=email&utm_campaign=02.11.13%20%281%29

In a study that represents a potential cultural shift in how basic science research can be conducted, researchers from Harvard Medical School, Harvard Business School and London Business School have demonstrated that a crowdsourcing platform pioneered in the commercial sector can solve a complex biological problem more quickly than conventional approaches—and at a fraction of the cost.

Partnering with TopCoder, a crowdsourcing platform with a global community of 450,000 algorithm specialists and software developers, researchers identified a program that can analyze vast amounts of data, in this case from the genes and gene mutations that build antibodies and T cell receptors. Since the immune system takes a limited number of genes and recombines them to fight a seemingly infinite number of invaders, predicting these genetic configurations has proven a massive challenge with few good solutions.

The program identified through this crowdsourcing experiment succeeded with an unprecedented level of accuracy and remarkable speed.

“This is a proof-of-concept demonstration that we can bring people together not only from different schools and different disciplines, but from entirely different economic sectors, to solve problems that are bigger than one person, department or institution,” said Eva Guinan, HMS associate professor of radiation oncology at Dana-Farber Cancer Institute and director of the Harvard Catalyst Linkages Program. “Given how complicated the immune system is, this has been a particularly formidable biological problem, and building tools for solving it has been hard and time-consuming. We were stunned by the power of these results and their potential application.”

“This study makes us think about how greater efficiencies in academic research can be obtained,” said Karim Lakhani, associate professor in the Technology and Operations Management Unit at Harvard Business School. “In a traditional setting, a life scientist who needs large volumes of data analyzed will hire a postdoc to create a solution, and it could take well over a year. We’re showing that in certain instances, existing platforms and communities might solve these problems better, cheaper and faster.”

“We’re excited to see that ideas from economics and management fields can be so productively applied to medical research,” said Kevin Boudreau, assistant professor of strategy and entrepreneurship at London Business School. “This progress is heartening, particularly in view of the computational challenges we face in understanding so many diseases. We hope this provides a model of how social scientists and medical researchers can collaborate to solve real-world problems that matter to people.”

Picture (Device Independent Bitmap)Image by alengo/iStockphoto

These findings are reported Feb. 7 in Nature Biotechnology.

For several years Boudreau, Guinan and Lakhani—through Harvard Catalyst—have explored the potential applicability of open and distributed innovation approaches to new areas, such as medical research. This has involved bringing insights from social science and economics to processes of medical research. They teamed up with Ramy Arnaout, HMS assistant professor of pathology at Beth Israel Deaconess Medical Center. Arnaout is also a systems biologist whose laboratory studies immune sequencing and other so-called “big-data” problems in biomedicine. Arnaout had developed computational methods for analyzing immune repertoires, but he could foresee having to invest significant computer and personnel resources to keep those methods able to handle the ever-increasing influx of data.

The researchers offered TopCoder what they thought would be an impossible goal: to develop a predictive algorithm that was an order of magnitude better than either Arnaout’s or the NIH’s standard algorithm (known as BLAST) and that could scale up to the mounting data demands. To do this, they had to first reframe the problem, translating it so that it could be accessible to individuals not trained in computational biology.

In only two weeks, viable solutions came from 122 different individuals. Among these, 16 were more accurate—and up to 1,000 times faster—than BLAST. The research team has released the top five performing code submissions under an open source license.

“This is more than just a quick, inexpensive answer,” said Guinan. “It’s uniting different approaches to a problem by taking from Harvard many disparate reservoirs of knowledge and bringing them together to formulate the question, analyze the data and then put it back to use. This draws on our faculty in a very diverse way. By extending the numbers of people who look at our specific problem, we get solutions rapidly. We have a lot of biases about doing that, and we really shouldn’t. In the end this allows researchers to turn their attention to basic science questions and not get caught up in details that they are less well suited to address.”

“In a way, the immune system is really the dark matter of biology,” said Arnaout. “We have all this sequence data, and there's no good way to figure out what it’s doing. Not only did the best entries achieve truly superior performance, but also this kind of crowdsourcing has the potential to be a general solution for a whole class of problems in biology. No single university or institution has the bandwidth and resources to achieve this kind of result so quickly and efficiently.”

According to Lakhani, it is not only the world of basic biomedical research that can benefit from this project, but any organization that is facing significant data analytics and computational challenges.  “Our research with Harvard Catalyst and the NASA Tournament Lab initiative points to the applicability of deploying crowds as an innovation partner for extraordinarily difficult challenges where there are significant personnel and paradigmatic bottlenecks,” he said. “This paper highlights the use of an alternative organizational form that is cost effective and productive. Many more organizations should also be considering how to effectively use crowds for problem solving.”

Co-authors on the study included Po-Ru Loh (Massachusetts Institute of Technology), Lars Backstrom (TopCoder), Carliss Baldwin (HBS), Eric Lonstein (HBS), Mike Lydon (TopCoder) and Alan MacCormack (HBS).

This work was funded by Harvard Business School’s Division of Research and Faculty Development, the NASA Tournament Lab at Harvard’s Institute for Quantitative Social Science, and Harvard Catalyst.

--------------------------------------------------
Stefaan G. Verhulst
Chief of Research
Markle Foundation
10 Rockefeller Plaza, Floor 16
New York, NY 10020-1903

Tel. 212 713 7630 (LinkedIn)

Check out the WEEKLY DIGEST : Developments, Findings and Views of a Connected World

Sunday, February 10, 2013

IBM's Watson Gets Its First Piece Of Business In Healthcare

IBM's Watson Gets Its First Piece Of Business In Healthcare

http://www.forbes.com/sites/bruceupbin/2013/02/08/ibms-watson-gets-its-first-piece-of-business-in-healthcare/

IBM’s Watson, the Jeopardy!-playing supercomputer that scored one for Team Robot Overlord two years ago, just put out its shingle as a doctor or, more specifically, as a combination lung cancer specialist and expert in the arcane branch of health insurance known as utilization management.  Thanks to a business partnership among IBM, Memorial Sloan-Kettering and WellPoint, health care providers will now be able to tap Watson’s expertise in deciding how to treat patients.

Pricing was not disclosed, but hospitals and health care networks who sign up will be able to buy or rent Watson’s advice from the cloud or their own server. Over the past two years, IBM’s researchers have shrunk Watson from the size of a master bedroom to a pizza-box-sized server that can fit in any data center. And they improved its processing speed by 240%. Now what was once was a fun computer-science experiment in natural language processing is becoming a real business for IBM and Wellpoint, which is the exclusive reseller of the technology for now. Initial customers include WestMed Practice Partners and the Maine Center for Cancer Medicine & Blood Disorders.

Even before the Jeopardy! success, IBM began to hatch bigger plans for Watson and there are few areas more in need of supercharged decision-support than health care. Doctors and nurses are drowning in information with new research, genetic data, treatments and procedures popping up daily. They often don’t know what to do, and are guessing as well as they can. WellPoint’s chief medical officer Samuel Nussbaum said at the press event today that health care pros make accurate treatment decisions in lung cancer cases only 50% of the time (a shocker to me). Watson, since being trained in this  medical specialty, can make accurate decisions 90% of the time. Patients, of course, need 100% accuracy, but making the leap from being right half the time to being right 9 out of ten times will be a huge boon for patient care. The best part is the potential for distributing the intelligence anywhere via the cloud, right at the point of care. This could be the most powerful tool we’ve seen to date for improving care and lowering everyone’s costs via standardization and reduced error. Chris Coburn, the Cleveland Clinic’s executive director for innovations, said at the event that he fully expects Watson to be widely deployed wherever the Clinic does business by 2020.   

Watson has made huge strides in its medical prowess in two short years. In May 2011 IBM had already trained Watson to have the knowledge of a second-year medical student. In March 2012 IBM struck a deal with Memorial Sloan Kettering to ingest and analyze tens of thousands of the renowned cancer center’s patient records and histories, as well as all the publicly available clinical research it can get its hard drives on. Today Watson has analyzed 605,000 pieces of medical evidence, 2 million pages of text, 25,000 training cases and had the assist of 14,700 clinician hours fine-tuning its decision accuracy. Six “instances” of Watson have already been installed in the last 12 months.

Watson doesn’t tell a doctor what to do, it provides several options with degrees of confidence for each, along with the supporting evidence it used to arrive at the optimal treatment. Doctors can enter on an iPad a new bit of information in plain text, such as “my patient has blood in her phlegm,” and Watson within half a minute will come back with an entirely different drug regimen that suits the individual. IBM Watson’s business chief Manoj Saxena says that 90% of nurses in the field who use Watson now follow its guidance.

WellPoint will be using the system internally for its nurses and clinicians who handle utilization management, the process by which health insurers determine which treatments are fair, appropriate and efficient and, in turn, what it will cover. The company will also make the intelligence available as a Web portal to other providers as its Interactive Care Reviewer. It is targeting 1,600 providers by the end of 2013 and will split the revenue with IBM. Terms were undisclosed.

Tuesday, February 5, 2013

The End of the Web, Search, and Computer as We Know It

The End of the Web, Search, and Computer as We Know It

David Gelernter

David Gelernter is a professor of computer science at Yale University and chief scientist at Lifestreams.com. He foresaw the World Wide Web and has been described as “brilliant and visionary.” Gelernter’s books include Mirror Worlds, Machine Beauty, and the forthcoming Other Side of the Mind. A former member of the National Endowment for the Arts governing board, Gelernter is also a painter; his works are currently on show at Yeshiva University Gallery in Manhattan.

http://www.wired.com/opinion/2013/02/the-end-of-the-web-computers-and-search-as-we-know-it/

People ask what the next web will be like, but there won’t be a next web.

The space-based web we currently have will gradually be replaced by a time-based worldstream. It’s already happening, and it all began with the lifestream, a phenomenon that I (with Eric Freeman) predicted in the 1990s and shared in the pages of Wired almost exactly 16 years ago.

This lifestream — a heterogeneous, content-searchable, real-time messaging stream — arrived in the form of blog posts and RSS feeds, Twitter and other chatstreams, and Facebook walls and timelines. Its structure represented a shift beyond the “flatland known as the desktop” (where our interfaces ignored the temporal dimension) towards streams, which flow and can therefore serve as a concrete representation of time.

It’s a bit like moving from a desktop to a magic diary: Picture a diary whose pages turn automatically, tracking your life moment to moment … Until you touch it, and then, the page-turning stops. The diary becomes a sort of reference book: a complete and searchable guide to your life. Put it down, and the pages start turning again.

Today, this diary-like structure is supplanting the spatial one as the dominant paradigm of the cybersphere: All the information on the internet will soon be a time-based structure. In the world of bits, space-based structures are static. Time-based structures are dynamic, always flowing — like time itself.

The web will be history.

Until now, the web has been space-based, like a magazine stand; we use spatial terms such as “second from the top on the far left” to identify a particular magazine. A diary, on the other hand, is time-based: One dimension of space has been borrowed to represent time, so we use temporal terms like “Thursday’s entry” or “everything from last spring” to identify entries.

Time as a metaphor may seem obvious now. Especially because it’s natural for us to see our lives as stories, organized by time.

Yet it took us more than 20 years in computing to get here. The field has finally moved from conserving resources ingeniously to squandering them creatively. In this new environment, we can focus on the best way — instead of the cheapest, most conservative way — for the internet to work.

And today, the most important function of the internet is to deliver the latest information, to tell us what’s happening right now. That’s why so many time-based structures have emerged in the cybersphere: to satisfy the need for the newest data. Whether tweet or timeline, all are time-ordered streams designed to tell you what’s new.

Of course, we can still browse or search into the past: Time moves forwards and backwards in the cybersphere. Any information object can be added at  “now,” and flows steadily backwards — like a twig dropped in a brook — into the past. You can drop files, messages, and conventional websites (those will appear as static, single elements) into the stream, which acts as a content-searchable cloud file system.

But what happens if we merge all those blogs, feeds, chatstreams, and so forth? By adding together every timestream on the net — including the private lifestreams that are just beginning to emerge — into a single flood of data, we get the worldstream: a way to picture the cybersphere as a whole.

No one can see the whole worldstream, because much of the information flowing through it is private. But everyone can see part of it.

Imagine an old-fashioned well with a bucket on a rope, with the bucket plunging deeper and deeper into the well. This well of time is infinitely deep, so the bucket will plunge forever — and the rope is always as long as it needs to be, so there will always be more rope to unwind. (The infinite scrolling we now experience on many timestreamed websites is merely the rope unwinding.) The bucket represents the head or start of the worldstream, the oldest data object. The rope-axle represents now, and the rope (plunging deeper and deeper into the past) is the stream itself.

Instead of today’s static web, information will flow constantly and steadily through the worldstream into the past. So what does it all mean?

Today, the most important function of the internet is to tell us what’s happening right now.

Streams Completely Change the Search Game

Today’s operating systems and browsers — and search models — become obsolete, because people no longer want to be connected to computers or “sites” (they probably never did).

What people really want is to tune in to information. Since many millions of separate lifestreams will exist in the cybersphere soon, our basic software will be the stream-browser: like today’s browsers, but designed to add, subtract, and navigate streams.

Searching content in a time stream is a matter of stream algebra, which is easier than the algebra of space-based structures like today’s web. Add two timestreams and get a third (simply merge the AP news feed and my friend Freeman’s blog streams into time-order); and content search is a matter of stream subtraction (simply subtract all entries that don’t mention “cranberries” to yield all the entries that do). The simple, practical features of stream algebra have one huge benefit: giving us made-to-order information.

Every news source is a lifestream. Stream-browsers will help us tune in to the information we want by implementing a type of custom-coffee blender: We’re offered thousands of different stream “flavors,” we choose the flavors we want, and the blender mixes our streams to order.

Every site’s content is liberated from the confines of space. It becomes part of a universal timestream. Instead of relying on Amazon the site to notify me if there’s a new Cynthia Ozick book or new books on the city of Florence, I can blend together several booksellers’ lifestreams and then apply my search since stream algebra allows any streams to be added (new and used books) and content (Florence, Ozick) to be subtracted.

E-commerce changes drastically. We shouldn’t have to work to find what’s new, yet the way the web is currently architected it’s no different logically than having to visit a thousand separate physical shops. The time-based worldstream lets us sit back instead and watch a single, customized fashion show across sites.

People no longer want to be connected to computers or ‘sites’ (they probably never did).

Worldstreams thus let us blend and tune our information any way we like: My preferred Yale football news, book updates, and shopping recommendations are interspersed with all my email, other messages, posts, documents, calendar notes, and so forth. Think these features already exist in an app somewhere? They don’t. They can’t, not until the millions of different streams each telling their own stories share the same interface for the stream browser to draw on.

Does this sort of precise control limit the serendipitous nature of the web? In a way, yes. But it’s about time: “Bring me what I want” is almost always more useful than “Let me rummage around and see what I can find.” No matter how fast it seems, most search is a waste of time. In a way, we are using time (i.e., the time-based structure) to gain time.

Instead of doing an endless series of separate searches, we tune the knobs on our stream-browser to continuously feed us just the information we need.

This future doesn’t just kill the operating system, browser, and search as we know it — it changes the meaning of “computer” as we know it, too. Whether large or small (e.g., a smartphone), a computer’s main function in the near future will be tuning in to — as a car radio tunes in a broadcast station — the constantly flowing global cyberflow. We won’t care much about the computer devices themselves since we’ll be more focused on the world of information … and our lives as attached to it. 

Finally, the web — soon to become the cybersphere — will no longer resemble a chaotic cobweb. It’s already started to happen. Instead, billions of users will spin their own tales, which will merge seamlessly into an ongoing, endless narrative: the earth telling its own story.

David Brooks: The Philosophy of Data

The Philosophy of Data

By DAVID BROOKS

Published: February 4, 2013

http://www.nytimes.com/2013/02/05/opinion/brooks-the-philosophy-of-data.html?hpw

f you asked me to describe the rising philosophy of the day, I’d say it is data-ism. We now have the ability to gather huge amounts of data. This ability seems to carry with it certain cultural assumptions — that everything that can be measured should be measured; that data is a transparent and reliable lens that allows us to filter out emotionalism and ideology; that data will help us do remarkable things — like foretell the future.

Over the next year, I’m hoping to get a better grip on some of the questions raised by the data revolution: In what situations should we rely on intuitive pattern recognition and in which situations should we ignore intuition and follow the data? What kinds of events are predictable using statistical analysis and what sorts of events are not?

I confess I enter this in a skeptical frame of mind, believing that we tend to get carried away in our desire to reduce everything to the quantifiable. But at the outset let me celebrate two things data does really well.

First, it’s really good at exposing when our intuitive view of reality is wrong. For example, every person who plays basketball and nearly every person who watches it believes that players go through hot streaks, when they are in the groove, and cold streaks, when they are just not feeling it.

But Thomas Gilovich, Amos Tversky and Robert Vallone found that a player who has made six consecutive foul shots has the same chance of making his seventh as if he had missed the previous six foul shots.

When a player has hit six shots in a row, we imagine that he has tapped into some elevated performance groove. In fact, it’s just random statistical noise, like having a coin flip come up tails repeatedly. Each individual shot’s success rate will still devolve back to the player’s career shooting percentage.

Similarly, nearly every person who runs for political office has an intuitive sense that they can powerfully influence their odds of winning the election if they can just raise and spend more money. But this, too, is largely wrong.

The data show that in state and national elections that are well-financed, television ad buys barely matter. After the 2004 election, political scientists tried to measure the effectiveness of campaign commercials. They found that if one candidate ran 1,000 more commercials than his opponent in a county — a huge disproportion — that translated into a paltry 0.19 percent advantage in the vote.

After the 2006 election, Sean Trende constructed a graph comparing the incumbent campaign spending advantages with their eventual margins of victory. There was barely any relationship between more spending and a bigger victory.

In May and June of 2012, the Obama campaign unleashed a giant ad barrage against Mitt Romney, but as political scientist John Sides wrote in The Times’s FiveThirtyEight blog recently, the ads had no lasting effect.

Likewise, many teachers have an intuitive sense that different students have different learning styles: some are verbal and some are visual; some are linear, some are holistic. Teachers imagine they will improve outcomes if they tailor their presentations to each student. But there’s no evidence to support this either.

Second, data can illuminate patterns of behavior we haven’t yet noticed. For example, I’ve always assumed that people who frequently use words like “I,” “me,” and “mine” are probably more egotistical than people who don’t.

But as James Pennebaker of the University of Texas notes in his book, “The Secret Life of Pronouns,” when people are feeling confident, they are focused on the task at hand, not on themselves. High status, confident people use fewer “I” words, not more.

Pennebaker analyzed the Nixon tapes. Nixon used few “I” words early in his presidency, but used many more after the Watergate scandal ravaged his self-confidence. Rudy Giuliani used few “I” words through his mayoralty, but used many more later, during the two weeks when his cancer was diagnosed and his marriage dissolved. Barack Obama, a self-confident person, uses fewer “I” words than any other modern president.

Our brains often don’t notice subtle verbal patterns, but Pennebaker’s computers can. Younger writers use more downbeat and past-tense words than older writers who use more positive and future-tense words.

Liars use more upbeat words like “pal” and “friend” but fewer excluding words like “but,” “except” and “without.” (When you are telling a false story, it’s hard to include the things you did not see or think about.)

We think of John Lennon as the most intellectual of the Beatles, but, in fact, Paul McCartney’s lyrics had more flexible and diverse structures and George Harrison’s were more cognitively complex.

In sum, the data revolution is giving us wonderful ways to understand the present and the past. Will it transform our ability to predict and make decisions about the future? We’ll see.

Monday, February 4, 2013

Big Data for Development: From Information to Knowledge Societies?

Big Data for Development: From Information to Knowledge Societies?

Posted on February 4, 2013 | Leave a comment

http://irevolution.net/2013/02/04/big-data-for-development-2/

Unlike analog information, “digital information inherently leaves a trace that can be analyzed (in real-time or later on).” But the “crux of the ‘Big Data’ paradigm is actually not the increasingly large amount of data itself, but its analysis for intelligent decision-making (in this sense, the term ‘Big Data Analysis’ would actually be more fitting than the term ‘Big Data’ by itself).” Martin Hilbert describes this as the “natural next step in the evolution from the ‘Information Age’ & ‘Information Societies’ to ‘Knowledge Societies’ [...].”

Hilbert has just published this study on the prospects of Big Data for inter-national development. “From a macro-perspective, it is expected that Big Data informed decision-making will have a similar positive effect on efficiency and productivity as ICT have had during the recent decade.” Hilbert references a 2011 study that concluded the following: “firms that adopted Big Data Analysis have output and productivity that is 5–6 % higher than what would be expected given their other investments and information technology usage.” Can these efficiency gains be brought to the unruly world of international development?

Picture (Device Independent Bitmap)

To answer this question, Hilbert introduces the above conceptual framework to “systematically review literature and empirical evidence related to the pre-requisites, opportunities and threats of Big Data Analysis for international development.” Words, Locations, Nature and Behavior are types of data that are becoming increasingly available in large volumes.

“Analyzing comments, searches or online posts [i.e., Words] can produce nearly the same results for statistical inference as household surveys and polls.” For example, “the simple number of Google searches for the word ‘unemployment’ in the U.S. correlates very closely with actual unemployment data from the Bureau of Labor Statistics.” Hilbert argues that the tremendous volume of free textual data makes “the work and time-intensive need for statistical sampling seem almost obsolete.” But while the “large amount of data makes the sampling error irrelevant, this does not automatically make the sample representative.” 

The increasing availability of Location data (via GPS-enabled mobile phones or RFIDs) needs no further explanation. Nature refers to data on natural processes such as temperature and rainfall. Behavior denotes activities that can be captured through digital means, such as user-behavior in multiplayer online games or economic affairs, for example. But “studying digital traces might not automatically give us insights into offline dynamics. Besides these biases in the source, the data-cleaning process of unstructured Big Data frequently introduces additional subjectivity.”

The availability and analysis of Big Data is obviously limited in areas with scant access to tangible hardware infrastructure. This corresponds to the “Infra-structure” variable in Hilbert’s framework. “Generic Services” refers to the production, adoption and adaptation of software products, since these are a “key ingredient for a thriving Big Data environment.” In addition, the exploitation of Big Data also requires “data-savvy managers and analysts and deep analytical talent, as well as capabilities in machine learning and computer science.” This corresponds to “Capacities and Knowledge Skills” in the framework.

The third and final side of the framework represents the types of policies that are necessary to actualize the potential of Big Data for international develop-ment. These policies are divided into those that elicit a Positive Feedback Loops such as financial incentives and those that create regulations such as interoperability, that is, Negative Feedback Loops.

The added value of Big Data Analytics is also dependent on the availability of publicly accessible data, i.e., Open Data. Hilbert estimates that a quarter of US government data could be used for Big Data Analysis if it were made available to the public. There is a clear return on investment in opening up this data. On average, governments with “more than 500 publicly available databases on their open data online portals have 2.5 times the per capita income, and 1.5 times more perceived transparency than their counterparts with less than 500 public databases.” The direction of “causality” here is questionable, however.

Hilbert concludes with a warning. The Big Data paradigm “inevitably creates a new dimension of the digital divide: a divide in the capacity to place the analytic treatment of data at the forefront of informed decision-making. This divide does not only refer to the availability of information, but to intelligent decision-making and therefore to a divide in (data-based) knowledge.” While the advent of Big Data Analysis is certainly not a panacea,”in a world where we desperately need further insights into development dynamics, Big Data Analysis can be an important tool to contribute to our understanding of and improve our contributions to manifold development challenges.”

I am troubled by the study’s assumption that we live in a Newtonian world of decision-making in which for every action there is an automatic equal and opposite reaction. The fact of the matter is that the vast majority of development policies and decisions are not based on empirical evidence. Indeed, rigorous evidence-based policy-making and interventions are still very much the exception rather than the rule in international development. Why? “Account-ability is often the unhappy byproduct rather than desirable outcome of innovative analytics. Greater accountability makes people nervous” (Harvard 2013). Moreover, response is always political. But Big Data Analysis runs the risk de-politicize a problem. As Alex de Waal noted over 15 years ago, “one universal tendency stands out: technical solutions are promoted at the expense of political ones.” I hinted at this concern when I first blogged about the UN Global Pulse back in 2009.

In sum, James Scott (one of my heroes) puts it best in his latest book:

“Applying scientific laws and quantitative measurement to most social problems would, modernists believed, eliminate the sterile debates once the ‘facts’ were known. [...] There are, on this account, facts (usually numerical) that require no interpretation. Reliance on such facts should reduce the destructive play of narratives, sentiment, prejudices, habits, hyperbole and emotion generally in public life. [...] Both the passions and the interests would be replaced by neutral, technical judgment. [...] This aspiration was seen as a new ‘civilizing project.’ The reformist, cerebral Progressives in early twentieth-century American and, oddly enough, Lenin as well believed that objective scientific knowledge would allow the ‘administration of things’ to largely replace politics. Their gospel of efficiency, technical training and engineering solutions implied a world directed by a trained, rational, and professional managerial elite. [...].”

“Beneath this appearance, of course, cost-benefit analysis is deeply political. Its politics are buried deep in the techniques [...] how to measure it, in what scale to use, [...] in how observations are translated into numerical values, and in how these numerical values are used in decision making. While fending off charges of bias or favoritism, such techniques [...] succeed brilliantly in entrenching a political agenda at the level of procedures and conventions of calculation that is doubly opaque and inaccessible. [...] Charged with bias, the official can claim, with some truth, that ‘I am just cranking the handle” of a nonpolitical decision-making machine.”

Picture (Device Independent Bitmap)

See also:

·       Big Data for Development: Challenges and Opportunities [Link]

·       How to Build Resilience Through Big Data [Link]

Saturday, February 2, 2013

The Origins of 'Big Data': An Etymological Detective Story

The Origins of ‘Big Data’: An Etymological Detective Story

By STEVE LOHR

http://bits.blogs.nytimes.com/2013/02/01/the-origins-of-big-data-an-etymological-detective-story/

Words and phrases are fundamental building blocks of language and culture, much as genes and cells are to the biology of life. And words are how we express ideas, so tracing their origin, development and spread is not merely an academic pursuit but a window into a society’s intellectual evolution.

Digital technology is changing both how words and ideas are created and proliferate, and how they are studied. Just last month, for example, the Library of Congress said its archive of public Twitter messages has reached 170 billion tweets and rising, by about 500 million tweets a day.

The Library of Congress archive, resulting from a deal struck with Twitter in 2010, is not yet open to researchers. But the plan is that it soon will be. In a white paper, the Library said that social media promises to be a rich resource that provides “a fuller picture of today’s cultural norms, dialogue, trends and events to inform scholarship, the legislative process, new works of authorship, education and other purposes.”

The new digital forms of communication — Web sites, blog posts, tweets — are often very different from the traditional sources for the study of words, like books, news articles and academic journals.

“It’s almost like oral language instead of edited text,” said Fred R. Shapiro, editor of the “Yale Book of Quotations” and an associate librarian at the Yale Law School. “It’s the way of the future.”

The unruly digital data of the Web is a big ingredient in what is now being called “Big Data.” And as it turns out, the term Big Data seems to be most accurately traced not to references in news or journal archives, but to digital artifacts now posted on technical Web sites, appropriately enough.

To our modest tale of word sleuthing: Last August, I wrote a Sunday column about 2012 being the breakout year for Big Data as an idea, in the marketplace, and as a term.

At the time, I did some reporting on the roots of the term, and I asked Mr. Shapiro of Yale to dig into it. He scoured data bases and came up with several references, including in press releases for product announcements and one intriguing use of the term by a now-famous author (more on that later).

But Mr. Shapiro couldn’t find anything as crisp and definitive as he had done for me years earlier when I asked him to try to find the first reference to the word “software” as a computing term. It was in 1958, in an article in “The American Mathematical Monthly,” written by John Tukey, a Princeton mathematician.

So, without a conclusive answer, I didn’t write about the origins of the term Big Data in that Sunday column. But afterward, I heard from people who had ideas on the subject.

Francis X. Diebold, an economist at the University of Pennsylvania, got in touch and even wrote a paper, with the mildly tongue-in-cheek title, “I Coined the Term ‘Big Data’ ” I had not thought of economics as the breeding ground for the term, but it is not unreasonable. Some of the statistical and algorithmic methods now in the Big Data tool kit trace their heritage to economic modeling and Wall Street.

Mr. Diebold staked a claim based on his paper, “Big Data Dynamic Factor Models for Macroeconomic Measurement and Forecasting,” presented in 2000 and published in 2003. The economic-modeling paper was first academic reference found to Big Data, according to research by Marco Pospiech, a Ph. D. candidate at the Technical University of Freiberg in Germany.

By then, I had heard from Douglas Laney, an veteran data analyst at Gartner. His said the father of the term Big Data might well be John Mashey, who was the chief scientist at Silicon Graphics in the 1990s.

I replied to Mr. Diebold that I thought from what I had seen he probably had plenty of competition. And I passed along the e-mail correspondence I had received. Mr. Diebold said thanks much, and added that he had a University of Pennsylvania research librarian looking into it as well.

The term Big Data is so generic that the hunt for its origin was not just an effort to find an early reference to those two words being used together. Instead, the goal was the early use of the term that suggests its present connotation — that is, not just a lot of data, but different types of data handled in new ways.

The credit, it seemed to me, should go to someone who was aware of the computing context. That is why, in my view, a very intriguing reference, discovered by the Yale researcher Mr. Shapiro, does not qualify.

In 1989, Erik Larson, later the author of bestsellers including “The Devil in the White City” and “In The Garden of Beasts,” wrote a piece for Harper’s Magazine, which was reprinted in The Washington Post. The article begins with the author wondering how all that junk mail arrives in his mailbox and moves on to the direct-marketing industry. The article includes these two sentences: “The keepers of big data say they do it for the consumer’s benefit. But data have a way of being used for purposes other than originally intended.”

Prescient indeed. But not, I don’t think, a use of the term that suggests an inkling of the technology we call Big Data today.

Since I first looked at how he used the term, I liked Mr. Mashey as the originator of Big Data. In the 1990s, Silicon Graphics was the giant of computer graphics, used for special-effects in Hollywood and for video surveillance by spy agencies. It was a hot company in the Valley that dealt with new kinds of data, and lots of it.

There are no academic papers to support the attribution to Mr. Mashey. Instead, he gave hundreds of talks to small groups in the middle and late 1990s to explain the concept and, of course, pitch Silicon Graphics products. The case for Mr. Mashey is on the Web sites of technical and professional organizations, like Usenix. There, some of his presentation slides from those talks are posted, including “Big Data and the Next Wave of Infrastress” in 1998.

For me, looking for the origins of Big Data has been a matter of personal curiosity, something to get back to someday and write up on a weekend.

When I called Mr. Mashey recently, he said that Big Data is such a simple term, it’s not much a claim to fame. His role, if any, he said, was to popularize the term within a portion of the high-tech community in the 1990s. “I was using one label for a range of issues, and I wanted the simplest, shortest phrase to convey that the boundaries of computing keep advancing,” said Mr. Mashey, a consultant to tech companies and a trustee of the Computer History Museum in Mountain View, Calif.

At the University of Pennsylvania, Mr. Diebold kept looking into the subject as well. His follow-up inquiries, he said, proved to be “a journey of increasing humility.” He has written to two papers since the first one.

His most recent paper concludes: “The term Big Data, which spans computer science and statistics/econometrics, probably originated in the lunch-table conversations at Silicon Graphics in the mid-1990s, in which John Mashey figured prominently.”

Tracing the origins of Big Data points to the evolution in the field of etymology, according to Mr. Shapiro. The Yale researcher began his word-hunting nearly 35 years ago, as a student at the Harvard Law School, poring through the library stacks. He was an early user of databases of legal documents, news articles and other documents, in computerized archives.

The Web, Mr. Shapiro said, opens up new linguistic terrain. “What you’re seeing is a marriage of structured databases and novel, less structured materials,” he said. “It can be a powerful tool to see far more.”