Tuesday, December 27, 2011

For Start-Ups That Aim at Giants, Sorting the Data Cloud Is the Next Big Thing

Malia Wollan      The New York Times       December 25, 2011

SAN FRANCISCO — The idea of big data goes something like this: In a world of ever-increasing digital connectivity, ever larger mountains of data are produced by our cellphones, computers, digital cameras, RFID readers, smart meters and GPS devices. The huge quantity of data becomes unwieldy and difficult for companies and governments to manage and understand.

 “My smartphone produces a huge amount of data, my car produces ridiculous amounts of really valuable data, my house is throwing off data, everything is making data,” said Erik Swan, 47, co-founder of Splunk, a San Francisco-based start-up whose software indexes vast quantities of machine-generated data into searchable links. Companies search those links, as one searches Google, to analyze customer behavior in real time.

Splunk is among a crop of enterprise software start-up companies that analyze big data and are establishing themselves in territory long controlled by giant business-technology vendors like Oracle and I.B.M.

Founded in 2004, before the term “big data” had worked its way into the vocabulary of Silicon Valley, Splunk now has some 3,200 customers in more than 75 countries, including more than half the Fortune 100 companies.
Customers include the online gaming company Zynga, the maker of FarmVille and Mafia Wars, which uses the software monitor game function to determine where players get stuck or quit playing, allowing Zynga to tweak games in real time to retain players.

Macy’s uses Splunk’s software to observe its Web traffic in order to avoid costly down times, particularly during peak holiday shopping. Edmunds, an automotive research Web site, started using Splunk to troubleshoot its information technology infrastructure and now uses the software to analyze all its customers’ online actions. Hundreds of government agencies use Splunk to monitor suspicious activity on secure sites, and a Japanese tsunami relief organization used it to track aid and monitor road and weather conditions.

The amount of data being generated globally increases by 40 percent a year, according to the McKinsey Global Institute, the consulting firm’s research arm. And while Splunk has a lead in selling software to analyze machine data, big data is big enough to create new opportunities for a multitude of start-ups, many of them using the open-source software Hadoop.

Venture capital is absolutely foaming at the mouth over big data,” said Peter Goldmacher, an analyst and managing director at Cowen & Company. “The volume of data being created now is not 10 times bigger, it is like a thousand times bigger.”

While skyrocketing valuations for social networking sites like Twitter, LinkedIn and Facebook have kept Silicon Valley investors betting heavily on the next social start-up, investors are increasingly looking at companies that build software for other companies. Worldwide revenue from enterprise software reached $244 billion in 2010, according to the research firm Gartner. Splunk is seen by some investors as proof that a wily start-up can chip away at some of that market.

“For a while there, people felt like everything that needed to be solved had been solved and that big companies would inevitably find all of the white space in enterprise,” said David Hornik, an investor at August Capital, which invested $3 million in Splunk in 2004.  “Splunk is really the poster child for thinking differently about an enterprise challenge and creating a platform that ends up really being disruptive and valuable.” The start-up got a total of $40 million in venture capital at that time from August Capital, Ignition Partners, JK&B Capital and Sevin Rosen Funds.

From the start, Splunk’s founders — Mr. Swan and Rob Das, 52, who is the company’s chief architect — set out to shake up what they saw as the stodgy, top-down world of enterprise software. “Big software is sold on the golf course, not sold to the people who actually use it,” said Mr. Das. Instead of aiming at the golf-playing chief information officer, the company took a quirky name that sounded like “spelunking” and zeroed in on the culture and tastes of everyday I.T. employees, the ones who actually had to use, and program around, enterprise software.

In 2005, when Splunk unveiled the first version of its software at the LinuxWorld conference in San Francisco, its booth was in an obscure corner, hidden by “rows and rows of vendors plastered with stock art of guys in suits and ties,” remembered Mr. Das. Nothing about enterprise software seemed hip or even vaguely playful, said Mr. Das, who spent more than a decade working in I.T. at companies like Lotus and Sun Microsystems. “We wanted to make enterprise software cool again.” So they decorated Splunk’s booth in all black and gave away T-shirts that said, “Take the SH out of IT.”
“People were stacked up 10 deep,” said Mr. Swan. Everyone, it seemed, wanted a T-shirt.

“Our customers, especially at the start, were I.T. people,” said Mr. Swan, who had worked at Apple and Disney Online, before becoming a co-founder of Splunk. “We’re talking about the guys in the basement, the guys in kilts and Mohawks. Those are our people.”
The company says it has been profitable for two years, and though executives will not comment on its exact plans to go public, Mr. Swan says, “We will be the first one to get shot out of this big data thing like LinkedIn got shot out of the social media space first.”

In another sign of an impending initial public offering, in 2008, the company hired Godfrey Sullivan, formerly of enterprise software companies like Hyperion, as its chief executive.

“There is a lot of money chasing this new world of unstructured data,” said Mr. Sullivan. “I would call Splunk the first mover in big data because we have been at this for years now.”

Thursday, December 22, 2011

IBM Takes Its Big Data Analytics To Academia

IBM Teams with 500 Universities in India in First-of-a-Kind Faculty Development Program;
Students in China, India, Northern Ireland and Scotland Studying How Analytics Applies to Industries


IBM News Room      December 21, 2011

ARMONK, NY - 21 Dec 2011: To address a growing market demand for analytics savvy graduates, IBM (NYSE:IBM) is working with universities around the world to bring advanced analytics training directly into the classroom. The company is expanding its academic initiatives for business analytics with new programs in China, India, Ireland and Scotland, helping students keep pace with today's competitive job market by gaining skills in this fast-growing field of technology. 

Everyday people create the equivalent of 2.5 quintillion bytes of data from sensors, mobile devices, online transactions, and social networks; so much that 90 percent of the world's data has been generated in the past two years. This amounts to more data than organizations can effectively use without applying analytics.  The new programs are providing students and faculty members, regardless of their course of study, with access to the latest software capabilities and thinking on how advanced analytics can be applied to tackle complex business and societal challenges. 

According to the 2010 IBM Institute for Business Value and MIT Sloan Management Reviewstudy of nearly 3,000 executives worldwide, the biggest challenge is the lack of understanding in how to use analytics to gain insights that can improve business outcomes. In response to market demand, universities are incorporating analytics curricula and courseware into a variety of degree programs to educate college students in this growing field.   

In India, IBM is working with faculty members from 500 universities to help more than 30,000 students develop skills in predictive analytics. As part of the program, IBM will conduct a series of training programs with business school faculty concentrating on predictive and business analytics, in 15 major cities throughout the country of India. The faculty members will complete a certification process in analytics at the end of the program. 
Once certified they will begin to teach students about how analytics can be applied to their topic of study.  The learning will involve access to predictive analytics technology and will focus on how to act on the results the analytics technology uncovers. 

“I have been using IBM predictive analytics technology in a number of programs at Indian Institute of Management Calcutta,” said Sahadeb Sarkar, Professor, Operations Management Group, Indian Institute of Management Calcutta (IIM). “I hope this initiative will help teachers in universities to learn and include analytics in existing courses and design new curriculum that will helps students gain a top-notch education to meet the demands of today’s businesses and government organizations.” 

University of theWest of Scotland (UWS) is introducing several new courses to its School of Computing curriculum including data mining, business intelligence and knowledge management. Plans to expand the analytics course offerings to non-IT and non-finance students are underway.
“Beyond teaching business and IT skills, we are preparing students for future job opportunities with new analytics courses,” said Professor Malcolm Crowe, University of the West of Scotland. “UWS is adding new courses in direct response to the recommendations of regional employers. They have specifically advised the School of Computing that important computing skills such as business analytics are in demand and will help graduates secure jobs.” 

Xi'an Jiao Tong University in China, together in cooperation with IBM’s China Development Lab in Xi'an, has developed business analytics oriented curriculum, project training materials, and planned a series of technical salon and master speech focus on analytics. These activities cover Cognos, SPSS and many best practices and tips integrated and tailored by the China Development Lab, and this analytics curriculum is planned to be replicated to six other Chinese universities in the future. This promotion of business analytics techniques and tools will enable a new generation of students, helping the Xi'an Lab with a pipeline of students with necessary skills, and will help to build up the business analytics ecosystem in China. 

At the University of Ulster, Northern Ireland’s largest university, students are using analytics software in a variety of application areas allowing them to collect hidden data and applying knowledge that seemed impossible to find before that can now be uncovered. 

These universities join schools around the world including Northwestern UniversityYale School of ManagementFordham UniversityDePaul UniversityUniversity of Southern California and University of Ottawa Telfer School of Management, that are working with IBM to develop and implement undergraduate and graduate curriculum and training on business analytics. 
Some of the early analytics projects underway at the university level were inspired by IBM’s Watson technology – the most advanced analytics technology currently available. Through the development of Watson, IBM sparked the interest of many students in the areas of math and computer science. IBM has teamed with universities to work on the sophisticated technology associated with Watson’s deep-Question and Answer capabilities, giving more than 10,000 students exposure to analytics technology. 

“Through IBM’s Academic Initiative, universities are adding analytics to their course offerings, establishing new degree programs and now we are seeing an acceleration in global demand for training in analytics,” said Jim Corgel, general manager of IBM’s Academic Initiative. “By combining IBM’s leadership in analytics with its global reach, we will begin to bridge the gap between to better equip students for new job opportunities.” 

Through its Academic Initiative, IBM is making its software, courseware and curricula available to nearly 6,000 universities and more than 30,000 faculty to advance technology skills. More information about IBM’s University Programs and Academic Initiative is available at www.ibm.com/press/university

For more information on IBM business analytics, please visit www.ibm.com/bao.

For images and picture stories visit:http://www.flickr.com/photos/curiosityshop/sets/72157628487625803/detail/

UK Gov Strategy on Big and Open Data

Gov unveils plans to make tax-funded research freely accessible

OUT-LAW.COM      The A Register         December 22, 2011

All publicly-funded research data should be made freely accessible to benefit business and society, the government has said.

The Department for Business, Innovation and Skills (BIS) said making the information freely available would help stimulate economic growth, innovation and entrepreneurism, and improve public sector transparency.

"The government, in line with our overarching commitment to transparency and open data, is committed to ensuring that publicly-funded research should be accessible free of charge," a BIS report into Innovation and Research Strategy for Growth (104-page/1.15MB PDF) said.

"Free and open access to taxpayer-funded research offers significant social and economic benefits by spreading knowledge, raising the prestige of UK research and encouraging technology transfer," it said. "At the moment, such research is often difficult to find and expensive to access. This can defeat the original purpose of taxpayer-funded academic research and limits understanding and innovation."

"We have already committed, in our response to Ian Hargreaves' review of intellectual property, to facilitate data mining of published research. This could have substantial benefits, for example in tackling diseases. But we need to go much further if, as a nation, we are to gain the full potential benefits of publicly-funded research," the report said.

"Government will work with partners, including the publishing industry, to achieve free access to publicly-funded research as soon as possible and will set an example itself," it said.

Among the plans the government hopes will advance its strategy are proposals to establish an Open Data Institute to "ensure that Open Data research is transformed into commercial advantage for UK companies, work with academic centres to increase the number of trained personnel with extensive Open Data skills and provide expert advice for government," BIS said.

The government will also force Research Councils to "ensure the researchers they fund" comply with an existing requirement to "deposit published articles or conference proceedings in an open access repository at or around the time of publication". Currently this is "unevenly enforced," BIS said.

The Research Councils have also committed to investing £2m to develop a 'Gateway to Research' by 2013, BIS said.

"In the first instance this will allow ready access to Research Council funded research information and related data but it will be designed so that it can also include research funded by others in due course," the report said. "The Research Councils will work with their partners and users to ensure information is presented in a readily reusable form, using common formats and open standards".

The government has previously outlined plans to make NHS patient data available to clinical researchers to improve the development of medical treatments and proposes to make a range of further public data available, including information from the healthcare and transport sectors within the next few years. The Open Data Institute will use this data to "to help industry exploit the opportunities created through the release of this data," BIS said.
"Our goal is a transformation in the accessibility of research and data. As these new initiatives take effect, we will be mindful of the need to protect the national interest – for example, on national security, personal privacy and commercial sensitivity – as well as the reputation of our research base," it said.

BIS said measures such as helping businesses innovate by cutting 'red tape', promoting "curiosity-driven" research and helping businesses benefit from international collaborations were "central elements" of the government's open data and research strategy.

"Large volumes of data remain unused and its value is untapped. The Office of Fair Trading has noted that key barriers to exploiting the value in public sector data included difficulty of access, charging regimes and a simple failure to exploit it. We believe there is an opportunity for the UK to establish a first mover advantage in open data. We will help to facilitate access to public sector data so that maximum value can be derived," BIS said.

http://www.out-law.com/Copyright © 2011, OUT-LAW.com

Monday, December 19, 2011

NYTimes: The Internet Gets Physical

Steve Lohr  The New York Times   December 17, 2011

 THE Internet likes you, really likes you. It offers you so much, just a mouse click or finger tap away. Go Christmas shopping, find restaurants, locate partying friends, tell the world what you’re up to. Some of the finest minds in computer science, working at start-ups and big companies, are obsessed with tracking your online habits to offer targeted ads and coupons, just for you.

 But now — nothing personal, mind you — the Internet is growing up and lifting its gaze to the wider world. To be sure, the economy of Internet self-gratification is thriving. Web start-ups for the consumer market still sprout at a torrid pace. And young corporate stars seeking to cash in for billions by selling shares to the public are consumer services — the online game company Zynga last week, and the social network giant Facebook, whose stock offering is scheduled for next year.

As this is happening, though, the protean Internet technologies of computing and communications are rapidly spreading beyond the lucrative consumer bailiwick. Low-cost sensors, clever software and advancing computer firepower are opening the door to new uses in energy conservation, transportation, health care and food distribution. The consumer Internet can be seen as the warm-up act for these technologies.

The concept has been around for years, sometimes called the Internet of Things or the Industrial Internet. Yet it takes time for the economics and engineering to catch up with the predictions. And that moment is upon us.

“We’re going to put the digital ‘smarts’ into everything,” said Edward D. Lazowska, a computer scientist at the University of Washington. These abundant smart devices, Dr. Lazowska added, will “interact intelligently with people and with the physical world.”

The role of sensors — once costly and clunky, now inexpensive and tiny — was described this month in an essay in The New York Times by Larry Smarr, founding director of the California Institute for Telecommunications and Information Technology; he said the ultimate goal was “the sensor-aware planetary computer.”

That may sound like blue-sky futurism, but evidence shows that the vision is beginning to be realized on the ground, in recent investments, products and services, coming from large industrial and technology corporations and some ambitious start-ups.

One of the hot new ventures in Silicon Valley is Nest Labs, founded by Tony Fadell, a former Apple executive, which has hired more than 100 engineers from Apple, Google, Microsoft and other high-tech companies.

Its product, introduced in late October, is a digital thermostat, combining sensors, machine learning and Web technology. It senses not just air temperature, but the movements of people in a house, their comings and goings, and adjusts room temperatures accordingly to save energy.

At the Nest offices in Palo Alto, Calif., there is a lot of talk of helping the planet, as well as the thrill of creating cool technology. Yoky Matsuoka, a former Google computer scientist and winner of a MacArthur “genius” grant, said, “This is the next wave for me.”

Matt Rogers, 28, a Nest co-founder, led a team of engineers at Apple that wrote software for iPods. He loved his job and working for Apple, he said. But he added: “In essence, we were building toys. I wanted to build a product that could really make a huge impact on a big problem.”

Across many industries, products and practices are being transformed by communicating sensors and computing intelligence. The smart industrial gear includes jet engines, bridges and oil rigs that alert their human minders when they need repairs, before equipment failures occur. Computers track sensor data on operating performance of a jet engine, or slight structural changes in an oil rig, looking for telltale patterns that signal coming trouble.

SENSORS on fruit and vegetable cartons can track location and sniff the produce, warning in advance of spoilage, so shipments can be rerouted or rescheduled. Computers pull GPS data from railway locomotives, taking into account the weight and length of trains, the terrain and turns, to reduce unnecessary braking and curb fuel consumption by up to 10 percent.

Researchers at General Electric, the nation’s largest industrial company, are working on such applications and others. One is a smart hospital room, equipped with three small cameras, mounted inconspicuously on the ceiling. With software for analysis, the room can monitor movements by doctors and nurses in and out of the room, alerting them if they have forgotten to wash their hands before and after touching patients — lapses that contribute significantly to hospital-acquired infections. Computer vision software can analyze facial expressions for signs of severe pain, the onset of delirium or other hints of distress, and send an electronic alert to a nearby nurse.

Last month, G.E. announced that it was opening a new global software center in Northern California and would hire 400 engineers there to write code to accelerate the commercial development of intelligent machines. “Our role is to build the software that enables us to do this industrial Internet,” said William Ruh, who will head the new center.

In 2008, I.B.M. declared that it was going to make a big push into the industrial Internet, using computing intelligence to create more efficient systems for utility grids, traffic management, food distribution, water conservation and health care. Smarter Planet was the label the company tacked on to the initiative, and industry analysts wondered if it was more than a sales campaign.

In a recent interview, Samuel J. Palmisano, chief executive of I.B.M., emphasized that the program’s origins were in the company’s research labs rather than its marketing department. “The timing was right because we had the technology,” he said.

 Today, I.B.M. says it is working on more than 2,000 projects worldwide that fit in the Smarter Planet category.

In Dubuque, Iowa, for example, I.B.M. has embarked on a long-term program with the local government to use sensors, software and Internet computing to improve the city’s use of water, electricity and transportation. In a pilot project this year, digital water meters were installed in 151 homes, and software monitored water use and patterns, informing residents about ways to consume less and alerting them to likely leaks. The savings in the pilot, nearly 7 percent, would translate into curbing water use by 65 million gallons a year in Dubuque, a Midwestern city of 60,000.

In Rio de Janeiro, I.B.M. is employing ground and airborne sensors, along with artificial intelligence software, for neighborhood-level disaster preparedness. The system, which is being developed by I.B.M. researchers, aims to predict heavy rains and mudslides up to 48 hours in advance and conduct evacuations before they occur — and avoid tragedies like the one last year, when a mudslide left more than 70 people dead and thousands homeless.

The next wave of computing does not step away from the consumer Internet so much as build on it for different uses (posing some of the same sorts of privacy and civil liberties concerns). Software techniques like pattern recognition and machine learning used in Internet searches, online advertising and smartphone apps are also ingredients in making smart devices to manage energy consumption, health care and traffic.

Take Google’s robot car program, for example. The automated cars, each with a human along for the ride, have deftly navigated thousands of miles on California highways and city streets. The project — a research effort so far — uses a bundle of artificial intelligence technologies, as does Google’s search-and-ad business.

GLOBAL PULSE is a new initiative by the United Nations to leverage data from the consumer Internet for global development. So-called sentiment analysis of messages in social networks and phone text messages — using natural-language deciphering software — can help predict job losses or lower spending in a region, or disease outbreaks.

In parts of Africa and Asia, where cellphones serve as automated bank tellers, with text messages initiating money transfers, they can also serve as an early warning system. When savings transfers drop to 50 cents or zero from $10 a month, “something is happening that is evident in the digital smoke signals,” said Robert Kirkpatrick, the director of Global Pulse. School feeding programs or government assistance might be stepped up to prevent a region from slipping back into poverty.

Global Pulse, begun in late 2009, is conducting research and trying to forge partnerships with private companies. To really succeed, the program needs the cooperation of Internet companies and cellphone carriers to give it access to social network and text-message communications, which would be stripped of any personally identifying information.

Mr. Kirkpatrick terms such contributions “data philanthropy.” His argument is that cooperating helps companies by nurturing economic health in the markets where they do business.

Global Pulse, Mr. Kirkpatrick said, is exploring new frontiers in knowledge with its real-time tracking of what is happening to people, not to sell them something but to target development efforts. “This is computational behavioral economics,” he said. “We’re part of a whole new science here.”
Steve Lohr is a technology reporter for The New York Times.

Thursday, December 15, 2011

Second Half of the Information Age to Focus on Exploitation of Technology and the Information It Processes

Special Report Examines New Opportunities in IT to Seize a Competitive Advantage

STAMFORD, Conn.–(BUSINESS WIRE)–The Information Age is an 80 year wave of economic and societal change that is in its second half, where business value comes from exploitation of technology rather than from installation, according to Gartner, Inc.

“In the first half of the Information Age, the primary focus was the technology itself; this is where great fortunes were made by companies like IBM and Microsoft,” said Mark Raskino, vice president and Gartner Fellow. “In this period, the majority of companies that gained competitive advantage did so by differential access to the technology from these providers — for example, by having more capital to invest in it or better skills at installing it in their businesses.

“In the second half of the age, as technology becomes ubiquitous, consumerized, cheaper and more equally available to all, the focus for differentiation moves to exploitation of the technology and to the information it processes,” Mr. Raskino said. “It is already noticeable that the great fortunes of the second half of the age are being made by companies like Google and Facebook, which are not traditional makers of technology. In this period, the majority of companies that enjoy competitive advantage will gain it from a differential ability to see and exploit the opportunities of new kinds of information.”

In the report, “Strategic Information Management for Competitive Advantage,” (http://www.gartner.com/resId=1851616) Gartner identified four particular types of generally useful information likely to dominate competition this decade, in the way that process information and customer information did in the past 15 years:

·       Location information, which is now maturing in availability, will offer opportunities to better opt
imize the utilization of almost any movable physical asset (human or inanimate) in almost any business.
·       Sustainability information will be vital in advancing business models in industries that are adapting to the realities of a finite Earth meeting the demands of massive, consumerizing emerging markets.
·       DNA information and the rapidly falling cost of obtaining it will obviously be critical to innovation and productivity leaps in agriculture, medical care and pharmaceuticals, but it will also impact insurance and other sectors.
·       Social graph information will help companies “X-ray” and understand organization, team design, culture and other factors impacting knowledge worker productivity, yielding valuable insights to advance the intellectual service economy the way time and motion study did for manufacturing in the 20th century.

Beyond these, context, gesture, the live state of everyday objects (Internet of Things), inherent identity (untagged, image-recognition-based), human emotional state and even brain response to stimuli are all new types of information that are at the radar’s edge or are starting to be brought into play within businesses.
In the Gartner Special Report, “Seizing Competitive Advantage: New Opportunities in IT” (http://www.gartner.com/technology/research/competitive-advantage/), Gartner analysts examine the key issues around seizing a new competitive advantage by revitalizing leadership, management and culture, and leveraging “big culture technologies.”

Gartner defines “competitive advantage” as a difference between a company and its competitors that matters to customers. It is one of the two key components of corporate profitability:
1.      Industry structure: This determines the range of profitability of the average competitor and can be very difficult to change.
2.      Competitive advantage: This enables a company to outperform the average competitor.
“The Gartner view of competitive advantage is about leadership — that is, how does an organization gain a leadership position from the tools, capabilities and competencies at its disposal? The focus is not on doing something merely to be competitive, but rather on taking action to be the leader,” said Jorge Lopez, vice president and distinguished analyst at Gartner.

“Competitive advantage cannot be sustained other than by ceaselessly pursuing new ways to compete, and changing one’s culture to match the new needs,” Mr. Lopez said.

Mr. Lopez will provide additional analysis during the Gartner webinar, “Seizing Competitive Advantage: New Opportunities in IT” on January 5 at 9 a.m. EST and noon EST.

Monday, December 12, 2011

Chuchil Club Event: The Big Data Effect

The Big Data Effect      
Speakers:
Keith Collins, Senior Vice President and Chief Technology Officer, SAS
Gil Elbaz, Founder and CEO, Factual
Ping Li, Partner, Accel Partners
Luke Lonergan, Chief Technology Officer, Vice President and Co-Founder, Greenplum, an EMC Company
Anand Rajaraman, Senior Vice President, Walmart Global E-Commerce & co-founder, @WalmartLabs
Moderator:
Michael Chui, Senior Fellow, McKinsey Global Institute
Watch event at http://www.brighttalk.com/webcast/6499/39163.
    
Gartner predicts that data will grow by 800% in five years, with 80% of it unstructured. The World Economic Forum recently declared big data as an asset class. We’re only getting started to discover the implications of making better sense of large amounts of unstructured data to uncover business opportunities, strategies, and more.

Is big data really an emerging market with lots of innovation, startups, job creation on the horizon? Why did it suddenly become possible? What are the obstacles, and the most promising areas of opportunity? How do you make it real in your organization?

Join this group of thought leaders from Accel Partners, Factual, Greenplum, SAS, @Walmartlabs, and McKinsey Global Institute for a conversation that gets beyond the hype about big data.


Big Data and Europe: Turning government data into gold

European Commission Press Release December 12, 2011

Internet Governance


From the Release: “The Commission has launched an Open Data Strategy for Europe, which is expected to deliver a EUR 40 billion boost to the EU's economy each year. Europe's public administrations are sitting on a goldmine of unrealised economic potential: the large volumes of information collected by numerous public authorities and services.

Member States such as the United Kingdom and France are already demonstrating this value. The strategy to lift performance EU-wide is three-fold: firstly the Commission will lead by example, opening its vaults of information to the public for free through a new data portal. Secondly, a level playing field for open data across the EU will be established. Finally, these new measures are backed by the EUR 100 million which will be granted in 2011-2013 to fund research into improved data-handling technologies.”


See also: 
http://ec.europa.eu/information_society/policy/psi/index_en.htm

http://www.zdnet.co.uk/news/regulation/2011/12/12/reuse-of-public-data-to-get-easier-under-new-eu-rules-40094628/

Tuesday, December 6, 2011

New Google Blog: Policy by the Numbers (Data for sound policymaking from Google and friends)

Why data matters for public policy

Build Economy Create Jobs Explore Future Improve Lives

Google Blog   November 16, 2011

“The purpose of computing is insight, not numbers.” - Richard Hamming

As a computer scientist and engineer, I’ve always been fascinated by the process that determines how policies and institutions are created. Unlike computing systems, policymaking is anything but binary. An unpredictable combination of special interests, money, hot topics, loyalties and many other factors shape legislation that passes into law.

Now, more than ever, we need to use data to build sound policy frameworks that facilitate innovative breakthroughs. In order to inspire confidence in the future (and the markets), governments have to lead by using today’s facts to place big bets on—not against—a better tomorrow.

To get conversations rolling, Google’s public policy team will be sharing data insights here on this blog. We’ll also be inviting researchers, policymakers and thought leaders to contribute their interpretations of various data sets and what they mean for public policy. This forum will be open to ideas, and we welcome everyone to leave comments discussing their opinions.

Measurement and analysis provide the checks and balances we need to build a better future in the information age. When we don’t examine the numbers, policy is all too often created at the expense of the next generation. The Internet generates 2.6 jobs for every one lost, and today the world’s data is doubling every two years. We need to make sure that we sustain the laws that got us the open Internet we have today, and that sound policies are in place to keep this unparalleled engine of growth going.

Public discussions that are grounded in numbers reveal whether laws are effective and relevant or failing to protect citizens’ interests. We are all entitled to our own opinions, but we are not entitled to our own facts; the facts speak for themselves and it is folly to ignore them. With this blog, we hope to spark policy debates, foster discussions among policymakers and constituents and help citizens exercise their right to hold governments accountable.

posted by Vint Cerf, Internet architect and policy enthusiast

Economist's Technology Quarterly

Technology and society: The old idea of human computers, who work together to perform tricky tasks, is making a comeback

The Economist  December, 3, 2011 from the print edition

IT WAS late summer 1937, and the recovery from the Depression had stalled. American government officials had stimulus money to spend but, with winter looming, there were few construction projects to fund. So the officials created office posts instead. One project was assigned to a floor of a dusty old New York industrial building, not far from Times Square. It would eventually house 300 computers—humans, not machines.

The computers crunched through the calculations necessary to create mathematical tables, then an indispensable reference tool for many scientists. The calculations were complex and the computers, drawn largely from the ranks of New York’s poor, possessed only basic numeracy. So the mathematicians in charge of the project worked out how to break each calculation down into simple operations, the outcomes of which could be combined to give a final result.

It was a technique that had been employed for decades across America and Europe. The field of human computing even had its own journal and trade-union representation. Computing offices calculated ballistics trajectories, processed census statistics and charted the course of comets. They would continue to do so until the 1960s, when electronic computers became cheap enough to consign the profession to history.

Until recently, that is. Over the past few years, human computing has been reborn. The new generation of human computers carry out different tasks, but they mirror their predecessors in many other ways. They are being drafted in to perform tasks that computers cannot. They are employed in large numbers and are organised into streamlined workflows. And, as was the case in the age before electronic computers, their output is combined to generate results that could not easily be produced in any other way.

In one proof-of-principle experiment, published earlier this year, human computers were used to create encyclopedia entries. Like performing mathematical calculations, this is a skilled job, but one that can be broken down into simpler parts, such as initial research, writing and editing. Aniket Kittur and colleagues at Carnegie Mellon University in Pittsburgh, Pennsylvania created software, known as CrowdForge, that manages the process. It hands out tasks to online workers, which it contacts via Mechanical Turk, an outsourcing website run by Amazon. The workers send their work back to CrowdForge, which combines their output to produce surprisingly readable results.

Several American start-ups are operating similar workflows. CastingWords breaks audio files down into five-minute segments and farms each out to a transcriber. Each transcription is automatically bounced back to other workers for checking and, once deemed good enough, an (electronic) computer combines the segments and returns the finished product to the customer. At CloudCrowd a similar system is used to co-ordinate teams of human translators. Others are combining human and artificial intelligences. An app called oMoby, produced by IQ Engines, can identify objects in images snapped by iPhone users. First it applies object-recognition software, which may not be able to cope if the lighting is poor or the image was captured from an unusual angle. When that happens, the image is sent to a human analyst. Either way, the user gets an answer in half a minute or so.

Much more is to come. In old-fashioned computing offices, workflows were co-ordinated by senior staff, often mathematicians, who had worked out how to deconstruct the complex calculations the computers were tackling. Now silicon foremen such as CrowdForge oversee human computers. These algorithms, which co-ordinate workers by plugging into Mechanical Turk and other online piecework platforms, are relatively new and are likely to get considerably more sophisticated. Researchers are, for example, creating software to make it easier to assign tasks to workers—or, to put it another way, to program humans.


Eric Horvitz, a researcher at Microsoft’s research labs in Redmond, Washington, has considered how such software could be put to use. He imagines a future in which algorithms co-ordinate an army of human workers, physical sensors and conventional computers. In the event of a child going missing, for example, an algorithm might assign some volunteers to search duties and ask others to examine CCTV footage for sightings. The system would also trawl local news reports for similar cases. These elements would be combined to create a cyborg detective.


This sounds terribly futuristic, and rather different to the pen-and-paper human computation of the 19th century. But David Alan Grier, a historian of computing at George Washington University in Washington, DC, thinks that the architects of the new systems could learn a lot by studying the old ones. He points out that Charles Babbage, the designer of an early mechanical computer, gave much thought to reducing the errors that human computers made. Babbage realised that duplicating tasks and comparing the results was not enough, because different workers tended to make the same mistakes. A better solution was to find different ways to perform the same calculation. If two methods produce the same answer, the result is much less likely to be flawed, Babbage reasoned.


There are many more such useful tips in the historical record, says Dr Grier. Human-computing pioneers also wrote a lot about how best to break a complex calculation into sub-tasks that are completely independent of each other, for example. “There are all sorts of hints in the old literature about what’s useful,” he says. He is often invited to human-computing conferences at which he likes to chide researchers for overlooking such lessons from this forgotten but intriguing early chapter of computer history.


More than just digital quilting
Technology and society: The “maker” movement could change how science is taught and boost innovation. It may even herald a new industrial revolution


Dec 3rd 2011 | from the print edition


THE scene in the park surrounding New York’s Hall of Science, on a sunny weekend in mid-September, resembles a futuristic craft fair. Booths displaying handmade clothes sit next to a pavilion full of electronics and another populated by toy robots. In one corner visitors can learn how to pick locks, in another how to use a soldering iron. All this and much more was on offer at an event called Maker Faire, which attracted more than 35,000 visitors. This show and an even bigger one in Silicon Valley, held every May, are the most visible manifestations of what has come to be called the “maker” movement. It started on America’s West Coast but is spreading around the globe: a Maker Faire was held in Cairo in October.


The maker movement is both a response to and an outgrowth of digital culture, made possible by the convergence of several trends. New tools and electronic components let people integrate the physical and digital worlds simply and cheaply. Online services and design software make it easy to develop and share digital blueprints.


And many people who spend all day manipulating bits on computer screens are rediscovering the pleasure of making physical objects and interacting with other enthusiasts in person, rather than online. Currently the preserve of hobbyists, the maker movement’s impact may be felt much farther afield.


Start with hardware. The heart of New York’s Maker Faire was a pavilion labelled with an obscure Italian name: “Arduino” (meaning “strong friend”). Inside, visitors were greeted by a dozen stands displaying credit-card-sized circuit boards. These are Arduino micro-controllers, simple computers that make it easy to build all kinds of strange things: plants that send Twitter messages when they need watering, a harp made of lasers, an etch-a-sketch clock, a microphone that serves as a breathalyser, or a vest that displays your speed when riding a bike.


Such projects are taking off because Arduino is affordable (basic boards cost $20), can easily be extended using add-ons called “shields” to add new functions and has a simple programming system that almost anyone can use. “Not knowing what you are doing is an advantage,” says Massimo Banzi, an Italian engineer and designer who started the Arduino project a decade ago to enable students to build all kinds of contraptions. Arduino has since become popular—selling around 200,000 units in 2011—because Mr Banzi made the board’s design “open source” (which means that anyone can download its blueprints and build their own versions), and because he has spent much time and effort getting engineers all over the world involved with the project.


This openness has prompted a sizeable ecosystem of add-ons. They include a touch-screen, an illuminated display and support for Wi-Fi networking. Other firms have built specialised variants of Arduino. SparkFun, for instance, has developed Lilypad, a flexible micro-controller that can be sewn into clothing (think blinking T-shirts), along with many other add-ons.


Applying the open-source approach to hardware has also driven the development of the maker movement’s other favourite piece of kit, which could be found everywhere at the Maker Faire in New York: 3D printers. These machines are another way to connect the digital and the physical realms: they take a digital model of an object and print it out by building it up, one layer at a time, using plastic extruded from a nozzle. The technique is not new, but in recent years 3D printers have become cheap enough for consumers. MakerBot Industries, a start-up based in New York, now sells its machines for $1,300. The output quality is rapidly improving thanks to regular upgrades, many of them suggested by users.


None of this action in hardware would have happened without a second set of powerful drivers: software, standards and online communities. Arduino, for instance, relies on open-source programs that turn simple code into a form that can be understood by the board’s brain. Similarly, MakerBot’s 3D printers depend on a standard way to describe physical objects, called STL, and affordable software to design them. Some basic modelling programs, such as Google SketchUp and Blender, can be downloaded free.


As for online communities, Arduino has an active forum on its website, while MakerBot runs a website called Thingiverse, which lets people share 3D designs. YouTube and other video-sharing sites offer how-to clips for almost everything. On Instructables, users post and discuss recipes to make and do all kinds of things. And then there is Etsy, an online marketplace for handmade goods, from hand-knitted scarves to 3D-printed jewellery.


The ease with which designs for physical things can be shared digitally goes a long way towards explaining why the maker movement has already developed a strong culture—its third driver. “If you are not sharing your designs, you are doing it wrong,” says Bre Pettis, the chief executive of MakerBot. Physical space and tools are being shared, too, in the form of common workshops. Some 400 such “hacker spaces” already operate worldwide, according to Hackerspaces.org. Many are organised like artists’ collectives. At Noisebridge, a hacker space in San Francisco, even non-members can come and tinker—as long as they comply with the group’s main rule: to be “excellent” to each other. “The internet is no substitute for a real community,” says Mitch Altman, a co-founder of Noisebridge.


This sort of thing makes the maker movement sound a lot like the digital equivalent of quilting bees. But it has already had a wider impact, mainly in schools in America. Many have discovered 3D printers and Arduino boards—and are using them to make their science and technology classes more hands-on again, and teach students to be producers as well as users of digital products.


All this will boost innovation, predicts Dale Dougherty, the founder of Make magazine, a central organ of the maker movement. Its tools and culture promote experimentation, collaboration and rapid improvement. Makers can play in niches that big firms ignore—though they are watching the maker movement and will borrow ideas from it, Mr Dougherty believes. The Maker Faire in New York was sponsored by technology companies including HP and Cognizant. Autodesk, which makes computer-aided design software, bought Instructables in August.


Firms may also copy some of the unusual business models that makers, often accidental entrepreneurs, have come up with. Arduino lets other firms copy its designs, for example, but charges them to use its logo. Quirky, an industrial design firm based in New York City, uses crowdsourcing to decide which products to make. MakieLab of London is developing a platform to allow toy shops or individuals to develop customised toys and have them printed. Venture capitalists are nosing around the field. In recent months Quirky raised $16m, MakerBot raised $10m and Shapeways, a firm that offers a 3D-printing service, received $5m.


The parallel with the hobbyist computer movement of the 1970s is striking. In both cases enthusiastic tinkerers, many on America’s West Coast, began playing with new technologies that had huge potential to disrupt business and society. Back then the machines manipulated bits; now the action is in atoms. This has prompted predictions of a new industrial revolution, in which more manufacturing is done by small firms or even by individuals. “The tools of factory production, from electronics assembly to 3D printing, are now available to individuals, in batches as small as a single unit,” writes Chris Anderson, the editor of Wired magazine.


It is easy to laugh at the idea that hobbyists with 3D printers will change the world. But the original industrial revolution grew out of piecework done at home, and look what became of the clunky computers of the 1970s. The maker movement is worth watching.

Joichi Ito in NYTimes (Special Section on Future of Computing) on Open Innovation

In an Open-Source Society, Innovating by the Seat of Our Pants

Joici Ito  The New York Times   December 5, 2011

 The Internet isn’t really a technology. It’s a belief system, a philosophy about the effectiveness of decentralized, bottom-up innovation. And it’s a philosophy that has begun to change how we think about creativity itself.

Almost 20 years ago, I installed on my computer a tiny piece of software called MacPPP, which connected the programs running on it to the Internet. The program immediately transformed my computer from a fancy telex machine to a device running a very early version of the graphical Web.

I was working in entertainment at the time, and I remember thinking that this connection was going to change everything. I left to join the first commercial Internet service provider in Japan, PSINet Japan, as its first chief executive. Our first serious challenge, oddly enough, was a battle over an obscure information-sharing computer protocol called X.25. Most of us laboring to build the new Internet preferred the less regulated and simpler Internet Protocol.

Until then, large intergovernmental agencies had always gathered experts to work on the technical standards that would become the DNA of the telecommunications industry, the standards to which all companies would have to build their networks and products. These researchers had produced X.25, a complex and extremely well-considered standard that seemed to anticipate every possible problem and application.

The Internet, on the other hand, was designed and deployed by small groups of researchers following the credo of one of its chief architects, David Clark: “rough consensus and running code.” Its early standards — uncomplicated, consensual — were stewarded by small organizations that resisted permission or authority. And they won: The Internet Protocol on which every connected device relies was a triumph of distributed innovation over centralized expertise.

The ethos of the Internet is that everyone should have the freedom to connect, to innovate, to program, without asking permission. No one can know the whole of the network, and by design it cannot be centrally controlled. This network was intended to be decentralized, its assets widely distributed. Today most innovation springs from small groups at its “edges.”

This technical strategy has led to the creation of a gigantic network of far-flung innovators who develop standards with one another and share the products of their work in the form of free and open-source software. The architecture of the Internet and its abundance of free software and components has driven down the cost of manufacturing, distribution and collaboration — of innovation. It used to cost millions of dollars to start a software company. Today, for little or no money, entrepreneurs are able to develop and release a “minimum viable product” and test it with real users on the Internet before they have to raise any money from investors. In their earliest iterations, Facebook, Yahoo and Google were running in dorm rooms and labs before the founders had left college or had raised outside money.

In fact, it is now usually cheaper to just try something than to sit around and try to figure out whether to try something. The product map is now often more complex and more expensive to create than trying to figure it out as you go. The compass has replaced the map, and “rough consensus and running code” has become the fundamental philosophy for the so-called lean start-up movement.

Innovators are able to prototype a new product with 3-D printers and cheap laser cutters for nearly nothing. Even complex products can be manufactured with help from supply chain companies that are making their systems available online to anybody. Today we are seeing the emergence of a community of hardware hackers and designers very reminiscent of the developers who wrote the original open standards of the Internet. An explosion of grass-roots innovation in hardware is coming — freely designed and freely shared — as it did in software.

What has been a wildly successful model for consumer Internet start-ups in Silicon Valley turns out to be an extremely good model for learning in a wide variety of fields and disciplines. The students at M.I.T.’s Media Lab experiment, create and iterate; they produce demos and prototypes, and share and collaborate with the rest of the world through the Internet and a distributed network of connections and relationships.

 I don’t think education is about centralized instruction anymore; rather, it is the process establishing oneself as a node in a broad network of distributed creativity.

Neoteny, one of my favorite words, means the retention of childlike attributes in adulthood: idealism, experimentation and wonder. In this new world, not only must we behave more like children, we also must teach the next generation to retain those attributes that will allow them to be world-changing, innovative adults who will help us reinvent the future.

Joichi Ito is the director of the M.I.T. Media Lab.

Larry Smarr in NYTimes: An Evolution Toward a Programmable Universe

Essay      Larry Smarr  The New York Times December 5, 2011

Over the next 10 years, the physical world will become ever more overlaid with devices for sending and receiving information.

Already billions of processors are embedded in our smartphones, cars, appliances and buildings and the environment. These sensors can send out streams of data about their surroundings, and more and more it is anonymously transmitted to remote data centers — the “clouds” of Google, Amazon, Microsoft, Yahoo and Apple.

From these vast clouds, the companies can power apps that are “spatially aware.” For instance, Google Maps now draws on data in the cloud to sample the location and movement of cellphones in cars, producing a real-time picture of traffic congestion.

Smart electric grids are measuring our homes’ use of power; active people are tracking their heart rates; and hundreds of millions of us are uploading geo-tagged data to Flickr, Yelp, Facebook and Google Plus. As we look 10 years ahead, the fastest supercomputer (the “exascale” machine) will be composed of one billion processors, and the clouds will most likely grow to this scale as well, creating a distributed planetary computer of enormous power.

Such computational power, co-located with the gigantic storage that holds the data from all the incoming data streams, will enable faster-than-real-time simulations of many aspects of our physical world. As Mike Liebhold and his colleagues at the Institute for the Future have discussed, computing will have evolved from merely sensing local information to analyzing it to being able to control it. In this evolution, the world gradually becomes programmable.

At the California Institute for Telecommunications and Information Technology, we are using this vision to better understand the coming digital transformation of health, energy, environment and culture. We are experimenting with sensors to monitor electricity use in homes, buildings and data centers; the data can then be analyzed and used to control lighting, heating, cooling, appliances and computers to make them more energy-efficient.

It is logical that the analysis of traffic data, coupled with in-car radar and autopilot electronics, will enable software control of large numbers of robot-driven electric cars. Since buildings and transportation are major sources of greenhouse gas emissions, the sensor-aware planetary computer can be a crucial factor in reducing our carbon footprints.

The same principle applies to our bodies. I wear sensors to measure my steps, caloric burn and sleep patterns, while heart patients can wear sensors that wirelessly notify their doctors of life-threatening conditions. People will soon be able to have their genetic code and medical imaging stored in the cloud, along with charts of vital signs and detailed nutritional analysis of everything they consume.

Using this data, the planetary computer will be able to build a computational model of your body and compare your sensor stream with millions of others. Besides providing early detection of internal changes that could lead to disease, cloud-powered voice-recognition wellness coaches could provide continual personalized support on lifestyle choices, potentially staving off disease and making health care affordable for everyone.

Finally, in culture, the fine-grain streaming provided by Twitter, Facebook and Google Plus enables us to map out phenomena using “human sensors.”
For instance, a vast power failure occurred in Southern California in September; within minutes we could tell from the locations of Twitter messages saying “my power just went out” that it was widespread, long before the official announcement. Similarly, Twitter feeds from large geographic areas have been analyzed to create dynamic “social mood” or “political anger” maps, like the Google traffic maps constructed from GPS feeds.

Conceivably, the coupling of the sensor and human streams with planetary computing power will make it possible to create “social forecasts.” For good or evil, it seems inevitable that individuals, corporations, political leaders and intelligence agencies will come to use planetary computer models of social behavior to inject content into the global attention stream at just the right moment, hoping to steer the social dynamics to a desired outcome.

With the continuing exponential increase in the power of the planetary computer, one has to wonder whether we stand at the beginning of what Isaac Asimov’s “Foundation” series, more than 60 years ago, called “psychohistory.” His visionary genius Hari Seldon believed that statistical forecasting of human society’s actions would be possible with data from enough people throughout the galaxy.

In the next several decades, we will have a glimpse of whether something similar can emerge on planet Earth.

Larry Smarr is the founding director of Calit2.

Thursday, December 1, 2011

Big Data: DNA Sequencing Caught in Deluge of Data

Andrew Pollack  The New York Times  November 30, 2011

BGI, based in China, is the world’s largest genomics research institute, with 167 DNA sequencers producing the equivalent of 2,000 human genomes a day.
BGI churns out so much data that it often cannot transmit its results to clients or collaborators over the Internet or other communications lines because that would take weeks. Instead, it sends computer disks containing the data, via FedEx.

“It sounds like an analog solution in a digital age,” conceded Sifei He, the head of cloud computing for BGI, formerly known as the Beijing Genomics Institute. But for now, he said, there is no better way.

The field of genomics is caught in a data deluge. DNA sequencing is becoming faster and cheaper at a pace far outstripping Moore’s law, which describes the rate at which computing gets faster and cheaper.

The result is that the ability to determine DNA sequences is starting to outrun the ability of researchers to store, transmit and especially to analyze the data.
“Data handling is now the bottleneck,” said David Haussler, director of the center for biomolecular science and engineering at the University of California, Santa Cruz. “It costs more to analyze a genome than to sequence a genome.”

That could delay the day when DNA sequencing is routinely used in medicine. In only a year or two, the cost of determining a person’s complete DNA blueprint is expected to fall below $1,000. But that long-awaited threshold excludes the cost of making sense of that data, which is becoming a bigger part of the total cost as sequencing costs themselves decline.

“The real cost in the sequencing is more than just running the sequencing machine,” said Mark Gerstein, professor of biomedical informatics at Yale. “And now that is becoming more apparent.”

But the data challenges are also creating opportunities. There is demand for people trained in bioinformatics, the convergence of biology and computing. Numerous bioinformatics companies, like SoftGenetics, DNAStar, DNAnexus and NextBio, have sprung up to offer software and services to help analyze the data. EMC, a maker of data storage equipment, has found life sciences a fertile market for products that handle large amounts of information. BGI is starting a journal, GigaScience, to publish data-heavy life science papers.

“We believe the field of bioinformatics for genetic analysis will be one of the biggest areas of disruptive innovation in life science tools over the next few years,” Isaac Ro, an analyst at Goldman Sachs, wrote in a recent report.

Sequencing involves determining the order of the bases, the chemical units represented by the letters A, C, G and T, in a stretch of DNA. The cost has plummeted, particularly in the last four years, as new techniques have been introduced.

The cost of sequencing a human genome — all three billion bases of DNA in a set of human chromosomes — plunged to $10,500 last July from $8.9 million in July 2007, according to the National Human Genome Research Institute.

That is a decline by a factor of more than 800 over four years. By contrast, computing costs would have dropped by perhaps a factor of four in that time span.
The lower cost, along with increasing speed, has led to a huge increase in how much sequencing data is being produced. World capacity is now 13 quadrillion DNA bases a year, an amount that would fill a stack of DVDs two miles high, according to Michael Schatz, assistant professor of quantitative biology at the Cold Spring Harbor Laboratory on Long Island.

There will probably be 30,000 human genomes sequenced by the end of this year, up from a handful a few years ago, according to the journal Nature. And that number will rise to millions in a few years.

In a few cases, human genomes are being sequenced to help diagnose mysterious rare diseases and treat patients. But most are being sequenced as part of studies. The federally financed Cancer Genome Atlas, for instance, is sequencing the genomes of thousands of tumors and of healthy tissue from the same people, looking for genetic causes of cancer.

One near victim of the data explosion has been a federal online archive of raw sequencing data. The amount stored has more than tripled just since the beginning of the year, reaching 300 trillion DNA bases and taking up nearly 700 trillion bytes of computer memory.

Straining under the load and facing budget constraints, federal officials talked earlier this year about shutting the archive, to the dismay of researchers. It will remain open, but certain big sequencing projects will now have to pay to store their data there.

If the problem is tough for human genomes, it is far worse for the field known as metagenomics. This involves sequencing the DNA found in a particular environment, like a sample of soil or the human gut. The idea is to take a census of what microbial species are present.

E. Virginia Armbrust, who studies ocean-dwelling microscopic organisms at the University of Washington, said her lab generated 60 billion bases — as much as 20 human genomes — from just two surface water samples. It took weeks to do the sequencing, but nearly two years to then analyze the data, she said.

“There is more data that is infiltrating lots of different fields that weren’t particularly ready for that,” Professor Armbrust said. “It’s all a little overwhelming.”

The Human Microbiome Project, which is sequencing the microbial populations in the human digestive tract, has generated about a million times as much sequence data as a single human genome, said C. Titus Brown, a bioinformatics specialist at Michigan State University.

“It’s not at all clear what you do with that data,” he said. “Doing a comprehensive analysis of it is essentially impossible at the moment.”

Other scientific fields, like particle physics and astronomy, handle huge amounts of data. In those fields, however, much of the data is generated by a few huge accelerators or observatories, said Eugene Kolker, chief data officer at Seattle Children’s Hospital.

“In the life sciences, anyone can produce so much data, and it’s happening in thousands of different labs throughout the world,” he said.

Moreover, DNA is just part of the story. To truly understand biology, researchers are gathering data on the RNA, proteins and chemicals in cells. That data can be even more voluminous than data on genes. And those different types of data have to be integrated.

“We have these giant piles of data and no way to connect them” said H. Steven Wiley, a biologist at the Pacific Northwest National Laboratory. He added, “I’m sitting in front of a pile of data that we’ve been trying to analyze for the last year and a half.”

Still, many say the situation will be manageable. Jay Flatley, chief executive of Illumina, the leading supplier of sequencing machines, said he did not think information handling was a bottleneck or that it was causing people to hold off on buying new sequencers.

Researchers are increasingly turning to cloud computing so they do not have to buy so many of their own computers and disk drives.

Google might help as well.

“Google has enough capacity to do all of genomics in a day,” said Dr. Schatz of Cold Spring Harbor, who is trying to apply Google’s techniques to genomics data. Prodded by Senator Charles E. Schumer, Democrat of New York, Google is exploring cooperation with Cold Spring Harbor.

Google’s venture capital arm recently invested in DNAnexus, a bioinformatics company. DNAnexus and Google plan to host their own copy of the federal sequence archive that had once looked as if it might be closed.

The amount of data stored for a human genome will drop sharply. Sequencers produce huge amounts of raw data that then has to be analyzed and processed by software to produce the result.

With the field still young, many researchers store all the raw data, so it can be re-analyzed if better software is developed in the future.

In uncertain times, “scientists cling to their data,” said David J. Dooling, assistant director of the genome institute at Washington University in St. Louis.

But there is now so much raw data that it is becoming not feasible to re-analyze it. So researchers will increasingly store just the final results. In the case of human genomes, they might store even less — only the difference between a particular genome and some reference genome.

Professor Brown of Michigan State said: “We are going to have to come up with really clever ways to throw away data so we can see new stuff.”