Gareth Cook The Boston Globe November 11, 211
At the end of the 19th century, a team of British archeologists happened upon what is now one of the world's most treasured trash dumps.
The site, situated west of the main course of the Nile, about five days journey south of Memphis, lay near the city of Oxyrhynchus. Garbage mounds are always a sweet target for those interested in the past, but what made the Oxyrhynchus dump special was its exceptional dryness. The water table lay deep; it never rained. And this meant that the 2,000-year-old papyrus in the mounds, and the text inscribed on it, were remarkably well preserved.
Eventually some half a million pieces of papyrus were drawn from the desert and shipped back to Oxford University, where generations of scholars have been painstakingly transcribing and translating them. The manuscripts are rich, fascinating, and varied. The texts include lost comedies by the great Athenian playwright Menander, and the controversial Gospel of Thomas, along with glimpses of daily life — personal notes, receipts for the purchase of donkeys and dates — and the occasional scrap of sex magic.
The pace, however, has been glacial. After a hundred-plus years, scholars have been able to work through only about 15 percent of the collection. The finish line appeared to lie centuries in the future.
But a few months ago, the papyrologists tried something bold. They put up a website, called Ancient Lives, with a game that allowed members of the public to help transcribe the ancient Greek at home by identifying images from the papyrus. Help began pouring in. In the short time the site has been running, people have contributed 4 million transcriptions. They have helped identify Thucydides, Aristophanes, Plutarch's "On the Cleverness of Animals," and more.
Ancient Lives is part of a new approach to the conduct of modern scholarship, called crowd science or citizen science. The idea is to unlock thorny research projects by tapping the time and enthusiasm of the general public. In just the last few years, crowd science projects have generated notable contributions to fields as disparate as ecology, AIDS research, and astronomy. The approach has already accelerated research in a handful of specialized fields. And it may also accomplish something else: breaking down some of the old divisions between the highly educated mandarins of the academy and the curious amateurs out in the world.
"It may seem intimidating when we say you are going to help transcribe ancient Greek papyri, but it's all about pattern recognition, and the brain excels at pattern recognition," says James Brusuelas, an Oxford classicist who is part of the Ancient Lives team. "The reaction has been fantastic."
One reason for the sudden turn to crowd science is that it offers an imaginative answer to a central problem of 21st-century science: too much information. Oxford's scholars had an overwhelming load of work given them, in the form of a desert trove. More often, though, scientists are themselves creating floods of data that they simply don't have the hours to interpret. Every night, robotic telescopes relentlessly track the sky, pouring terabytes of images into hard drive farms. From biological labs come rivers of genetic code. And in many other fields — from high energy physics to environmental science — researchers are puzzling over how to handle the sudden embarrassment of riches.
For now, the new citizen science has touched only the tiniest fraction of the research conducted around the world. But its early successes, which have shocked even the architects of the approach, suggest that over time pro-am collaborations hold the potential to alter the landscape of science in important ways, harnessing countless able brains to do work that was once the province of a few overwhelmed experts. And as it does, it also offers an uncomfortable insight: There are ways that the structure of modern science may actually be limiting what we can learn.
The idea of recruiting amateur scientists has roots that go back at least a century. In 1900, in the early days of the American conservation movement, ornithologist Frank Chapman organized a Christmas bird census. Teams of avid birders collected observations from Toronto to Baldwin, La.: the American black duck, the red-breasted nuthatch, the common grackle, and 86 other species. It was an unprecedented one-day data dump. The Christmas Bird Count has become an Audubon tradition, with about 60,000 people going out every year, and the data it has generated through the years have proved invaluable to researchers.
Today there are firefly counts, herring counts, and ladybug counts. One can help track spiders or bats or coral reefs. A new iPhone app called Noah (for Networked Organisms and Habitats) allows users to snap pictures of species they come across and share the information with researchers and others. A similar British effort, called iSpot, led to the discovery of two species that had not been recorded before in England, according to a report by the BBC. Some projects use networks of observers to monitor the timing of natural events, such as the arrival of hummingbirds, or the budding of flowers, which provide information on the planet's changing climate. None of these projects would be possible without countless amateurs willing to serve as devoted foot soldiers across the planet.
The advent of the Internet has also opened up a new possibility: that the interested public could offer scholars more than help gathering data. In the best-known early example, they offered up their computers: 1999 saw the launch of SETI@home, an example of "distributed computing" in which volunteers downloaded software so their idling computers could help crunch radio-telescope data for signs of alien life.
More recently, though, has come a truly fascinating turn: the move from people volunteering their computers' down time, to people volunteering their brains' down time — from distributed computing to distributed thinking. Oxford University astronomer Chris Lintott says that his own involvement dates back to a 2007 conversation he had over a pint at the Royal Oak, a traditional watering hole for Oxford astronomers where tables are crammed into small rooms with old fireplaces and ancient wood beams. A student, Kevin Schawinski, had recently finished the exhausting task of categorizing 50,000 galaxy images for a project. As they spoke, though, it became clear that that wasn't nearly enough: What the project really required to succeed was to categorize a million galaxies.
"One look at Kevin's face," says Lintott, "suggested we should find an alternative method."
This led them to create Galaxy Zoo in 2007. The site provided a simple tutorial that trained people to classify galaxies by their appearance, and then served up images that astronomers had not yet categorized. Galaxy Zoo was so popular that soon after it launched, the servers literally caught fire from all the activity. A schoolteacher sitting in an apartment in the southeast of the Netherlands discovered a strange green cloud that had never been observed before. The astronomical data from the project have been used in a growing list of scientific publications.
The approach was so successful that Lintott and the other organizers decided to expand it to other areas, including solar explosions and climate change, under the name Zooniverse. (The Ancient Lives project uses the Zooniverse website.) Meanwhile, many other scientists and organizations are jumping in: One popular website, scienceforcitizens.net, lists more than 400 projects, and the site's founder says she expects to hit 1,000 within a year.
What marks this as an important milestone in the history of science is the new way it harnesses the power of the mind. There are many tasks that are beyond the grasp of even today's computers, particularly those which involve interpreting complex images. Like identifying cancer cells. Or categorizing galaxies. Or picking out letters of ancient Greek, written in a faded ink with a fast, messy hand, without breaks between words. The Internet, it turns out, is a brilliant way to feed those problems into an array of the planet's true supercomputers — human brains.A recent discovery highlights the sophis
tication of the work volunteers can do. Biologists are keenly interested in the three-dimensional shapes assumed by protein molecules inside the human body. Proteins are intimately involved in many aspects of life, but they fold into shapes that can be very difficult to predict, even given their precise chemical makeup. Protein-folding is a roadblock that holds up research into many diseases.
So a team of scientists at the University of Washington created a game called FoldIt, which gives players an image of a protein molecule and video game-like tools for folding the molecule. As the energy required to maintain the molecule in a particular shape drops — meaning it's closer to nature's solution — a player's score increases. FoldIt is a potentially addictive game that requires excellent spatial reasoning. Some players excelled at it — indeed, some became whizzes, and the researchers put their skills to work on unsolved problems. In September, the scientists announced that a team of its players had deciphered the folding of a protein important in AIDS research.
In a paper describing the result for Nature Structural and Molecular Biology, the scientists argued to their colleagues that a line had been crossed: "Although much attention has recently been given to the potential of crowdsourcing and game playing, this is the first instance we are aware of in which online gamers solved a longstanding scientific problem."
FoldIt is the most impressive demonstration yet that the public can make genuine contributions to scientific projects. But its success also stands as a potent critique of the way that the scientific enterprise is currently organized.
Science is, for the most part, a closed society organized into little fiefdoms of highly trained specialists, which means only a few minds engage with any given problem. Before FoldIt, for example, a problem in protein folding was the exclusive province of a relatively small number of experts — even though, it is now clear, there are real contributions to be made by 13-year-old video gamers.
The system is shaped in part by the force of tradition, but the larger challenge is that most scientific data is proprietary. A scientist works long and hard to generate original data, and then expects to reap the reward in the form of publishing the first research paper to describe some new phenomenon. She is not going to want share this data with others, particularly strangers, any more than say, an investigative reporter would want to share his notes before a story has been written. Harnessing 1,000 people requires sending your data out into the world — something that science is loath to do. The scientist's interest in keeping things private and getting credit, in other words, is directly opposed to society's interest in tackling some problems with a hive of the best minds.
There are exceptions, such as large astronomical and biological data sets that are available for anyone to work with. But the last 10 years have seen a boom in technology that allows large numbers of people to do amazing, cooperative things with information, and the scientific establishment has taken only baby steps toward figuring out ways to share it productively, according to Michael Nielsen, a former theoretical physicist and author of "Reinventing Discovery: The New Era of Networked Science."
To encourage this shift, the federal government, which funds the lion's share of the country's research, has been pressuring scientists to work more cooperatively, and share more of what they find faster. And there is a nascent effort within academia to identify ways that scientists might be recognized for their contributions to the community as a whole, beyond the publication of their individual discoveries.
"It is essential that scientists be rewarded when they share," says Nielsen.
It's a difficult problem, and Nielsen says he expects the real rewards of networked science to be tallied over decades, not years. Even if science becomes more open, there are also practical limitations: It takes a certain brilliance, and a lot of work, to recognize problems that can be shared with a crowd, and set up the systems needed for strangers to work together productively. It is not always clear when this tactic will move a project forward, or slow it down.
With time, though, one might expect a new type of scientist to emerge: one who is especially adept at recognizing problems, and designing projects, that tap the brilliance of a dispersed and motley team, whoever they may be.
Science is driven forward by discovery, and we appear to stand at the beginning of a democratization of discovery. An ordinary person can be the one who realizes that a long arm of a protein probably tucks itself just so; a woman who never went to college can provide the crucial transcription that reveals a spidery script to be a love poem from 2,000 years in the past. Nobody can say where the movement will go, but among the new pioneers of crowd science, there is a palpable sense that they have just happened upon a powerful, poorly understood new resource.
"We have used," says Lintott, "just a tiny fraction of the human attention span that goes into an episode of Jerry Springer."
Gareth Cook is a Globe columnist, a Pulitzer Prize-winning journalist, and a former editor of Ideas. He can be reached at cook@globe.com. Follow him on Twitter @garethideas.
Monday, November 21, 2011
Friday, November 18, 2011
Big Data Video/Stats
The Economist online November 18, 2011, 16:09
Drowning in numbers
See http://www.economist.com/blogs/dailychart/2011/11/big-data-0
Digital data will flood the planet—and help us understand it better
More from The World in 2012
Drowning in numbers
See http://www.economist.com/blogs/dailychart/2011/11/big-data-0
Digital data will flood the planet—and help us understand it better
More from The World in 2012
Thursday, November 17, 2011
Can Big Data Fix Healthcare?
Colin Hill Forbes November 17, 2011
What will healthcare look like in the year 2020? One thing is certain: we can’t afford its current trajectory. Left unchecked, our $2.6 trillion in annual spending will grow to $4.6 trillion by 2020, one-fifth of GDP. With almost 80 million Baby Boomers approaching retirement, economists forecast these trends will likely bankrupt Medicare and Medicaid in the near future. And while healthcare reform ignites a number of important changes, alone it does not resolve our issues. It’s critical we fix our system now.
Growth in Literature
Over the past 50 years medicine has grown dramatically. We now have more preventive, diagnostic, and treatment alternatives than ever before, with more being developed all the time.
This proliferation has been accompanied by an explosion in literature. Fully 35 percent of the 20 million articles indexed in MEDLINE were published in the last 10 years, with the annual pace approaching one million articles.
Despite this vast body of literature, the healthcare system remains starved for evidence of what works. While we have many treatment alternatives, in many cases we do not have much of an idea where best to apply them. Treatments that work well for some work poorly – or worse – for others.
We must learn to distinguish what interventions work, and for whom.
Limited Knowledge
Unfortunately, we have critical limitations in our ability to evaluate treatment effectiveness. First, and despite its volume, medical literature is of variable quality and has limited generalizability. Only a small fraction of studies compare the effectiveness of different treatments, and few evaluate effectiveness in real-world settings.
Evidence of what does work is increasingly overturned. A recent study found that 13 percent of articles concerning a clinical practice published in the New England Journal of Medicine in 2009 were reversals of previous findings.
And clinical practice guidelines, whose goal is to synthesize research into evidence for use by clinicians, don’t fare much better. Their quality is highly variable and, in some instances, quite poor – even the best have a limited shelf life. A 2001 review of guidelines estimated that half had become outdated in less than six years. It is unlikely that the situation has improved since then.
Clearly, and despite all the effort and expense, healthcare remains one of our nation’s most well-endowed, yet-poorly-informed, industries, with an approach to creating evidence clearly inadequate to the needs of practitioners and patients.
We must start producing better evidence faster and on a large scale. Before we can reduce costs and deliver meaningful improvements in outcomes, we must have meaningful evidence. Without it, we can never know what works, and for whom.
New Sources of Information
New types and sources of health care data have become available – or soon will – and in overwhelming quantity. The federal government is investing $20 billion in Electronic Health Records; industry is developing new electronic transaction standards; and innovators like PatientsLikeMe, 23andMe, Fitbit and Zeo are helping people generate and share their own data. The era of Big Data in healthcare has arrived.
Can Big Data Fix Healthcare? A recent McKinsey report called Big Data, “the next frontier for innovation, competition and productivity.” The Aspen Institute reported on the “promise and perils” of it. The Economist issued a special report about it. O’Reilly Media hosted two conferences on it this year alone. In all of these, Big Data’s opportunity to transform health care was featured prominently.
Are these expectations justified? Can Big Data fix healthcare? What analytic technologies will be required to actually deliver on Big Data’s promise and discover what works? What other promises need to be met for this to become a reality, and what will reality look like when evidence becomes available at the push of a button?
My name is Colin Hill. I am CEO and co-founder of GNS Healthcare, a healthcare analytics company focused on using observational data directly, and at-scale, to create evidence of what works for whom in healthcare. We’ll take on these questions and more in subsequent entries. Welcome to the conversation!
What will healthcare look like in the year 2020? One thing is certain: we can’t afford its current trajectory. Left unchecked, our $2.6 trillion in annual spending will grow to $4.6 trillion by 2020, one-fifth of GDP. With almost 80 million Baby Boomers approaching retirement, economists forecast these trends will likely bankrupt Medicare and Medicaid in the near future. And while healthcare reform ignites a number of important changes, alone it does not resolve our issues. It’s critical we fix our system now.
Growth in Literature
Over the past 50 years medicine has grown dramatically. We now have more preventive, diagnostic, and treatment alternatives than ever before, with more being developed all the time.
This proliferation has been accompanied by an explosion in literature. Fully 35 percent of the 20 million articles indexed in MEDLINE were published in the last 10 years, with the annual pace approaching one million articles.
Despite this vast body of literature, the healthcare system remains starved for evidence of what works. While we have many treatment alternatives, in many cases we do not have much of an idea where best to apply them. Treatments that work well for some work poorly – or worse – for others.
We must learn to distinguish what interventions work, and for whom.
Limited Knowledge
Unfortunately, we have critical limitations in our ability to evaluate treatment effectiveness. First, and despite its volume, medical literature is of variable quality and has limited generalizability. Only a small fraction of studies compare the effectiveness of different treatments, and few evaluate effectiveness in real-world settings.
Evidence of what does work is increasingly overturned. A recent study found that 13 percent of articles concerning a clinical practice published in the New England Journal of Medicine in 2009 were reversals of previous findings.
And clinical practice guidelines, whose goal is to synthesize research into evidence for use by clinicians, don’t fare much better. Their quality is highly variable and, in some instances, quite poor – even the best have a limited shelf life. A 2001 review of guidelines estimated that half had become outdated in less than six years. It is unlikely that the situation has improved since then.
Clearly, and despite all the effort and expense, healthcare remains one of our nation’s most well-endowed, yet-poorly-informed, industries, with an approach to creating evidence clearly inadequate to the needs of practitioners and patients.
We must start producing better evidence faster and on a large scale. Before we can reduce costs and deliver meaningful improvements in outcomes, we must have meaningful evidence. Without it, we can never know what works, and for whom.
New Sources of Information
New types and sources of health care data have become available – or soon will – and in overwhelming quantity. The federal government is investing $20 billion in Electronic Health Records; industry is developing new electronic transaction standards; and innovators like PatientsLikeMe, 23andMe, Fitbit and Zeo are helping people generate and share their own data. The era of Big Data in healthcare has arrived.
Can Big Data Fix Healthcare? A recent McKinsey report called Big Data, “the next frontier for innovation, competition and productivity.” The Aspen Institute reported on the “promise and perils” of it. The Economist issued a special report about it. O’Reilly Media hosted two conferences on it this year alone. In all of these, Big Data’s opportunity to transform health care was featured prominently.
Are these expectations justified? Can Big Data fix healthcare? What analytic technologies will be required to actually deliver on Big Data’s promise and discover what works? What other promises need to be met for this to become a reality, and what will reality look like when evidence becomes available at the push of a button?
My name is Colin Hill. I am CEO and co-founder of GNS Healthcare, a healthcare analytics company focused on using observational data directly, and at-scale, to create evidence of what works for whom in healthcare. We’ll take on these questions and more in subsequent entries. Welcome to the conversation!
Tuesday, November 15, 2011
Microsoft's Mundie: Big data could cure U.S. healthcare
Barb Darrow GigOM November, 15, 2011, 10:01am PT
Craig Mundie thinks that if the U.S. really wants to solve its massive healthcare problem, it should subject it to big data practices.
Specifically, Microsoft’s chief research and strategy officer said government and/or industry should bring the Internet model to bear on the problem of out-of-control healthcare costs, and that means sharing, not segregating, massive amounts of data: something that flies in the face of current HIPAA requirements. HIPAA is the Healthcare Insurance Portability and Accountability Act passed in 1996, which mandates strict privacy for patient records.
The big data concept, which calls for analyzing huge amounts of collected data — often in different formats — has become a rallying cry for vendors and customers that want to wring more value out of the information they already have but didn’t necessarily know how to capitalize on. Until now.
Speaking at the Techonomy 2011 conference Monday night, Mundie used the extensible business reporting language (XBRL) which underpins all financial reporting as an example of how the government can grease the skids for change. ”XBRL was ginned up by the SEC …. and went from a standing start to full adoption by U.S. companies in less than four years,” said Mundie.
Between Medicare and Medicaid, the U.S. government pays for more than half (54 percent) of aggregate domestic healthcare costs. “If the government said, ‘I don’t care how the private sector does it; we’ll do it this way,’ the whole industry would flip,” he said.
Free data flow is necessary so costs can be compared between the providers that are so reluctant to share their information. The fundamental issue with healthcare is “perverse payment system”: Most people don’t care about what their care really costs because a third party foots the bill. “There is no braking component” to contain runaway costs, he said.
If, however, Americans paid based on outcomes, there needs to be a total view of data to allow comparisons between providers. But until the data is in a form that can be shared — with privacy concerns met — he insisted there’s no basis for comparison and no way to do quality control.
“We’re not serious yet about aggregating this data. If the data is not available, you can’t make wholesale change.”
He didn’t venture a guess as to how much could be saved overall, but said in a smaller test case, Microsoft worked with a system of several hospitals that made their information available to a heath information exchange. “It slurped up all their Medicare data, and as soon as it had that data, they could query it to see just what was costing so much.”
Once the data is in one place, it’s not expensive to work with it, he noted, and it turned out that queries were able to identify a relatively small number of uninsured people who visited various emergency rooms around the system and ran up millions of dollars in charges. “Without this data, there was no way to find these people,” he said, adding that just eliminating redundancy and fraud alone would result in huge savings in a national context.
And, he insisted privacy issues can be addressed. Using metadata, “you can describe the provenance of the data to support sharing and also take privacy constraints into account,” he said. “Technology can be privacy-enhancing,” he noted.
But it’s highly unlikely insurance companies will drive this sort of change given the current privacy regulations and that their business is so profitable the way it is now, hence the need for a strong-spined government mandate.
Asked if these big data queries on patient care represent a “Google-like” capability for healthcare, Mundie said it’s “really a Bing-like system for healthcare, but absolutely.”
Craig Mundie thinks that if the U.S. really wants to solve its massive healthcare problem, it should subject it to big data practices.
Specifically, Microsoft’s chief research and strategy officer said government and/or industry should bring the Internet model to bear on the problem of out-of-control healthcare costs, and that means sharing, not segregating, massive amounts of data: something that flies in the face of current HIPAA requirements. HIPAA is the Healthcare Insurance Portability and Accountability Act passed in 1996, which mandates strict privacy for patient records.
The big data concept, which calls for analyzing huge amounts of collected data — often in different formats — has become a rallying cry for vendors and customers that want to wring more value out of the information they already have but didn’t necessarily know how to capitalize on. Until now.
Speaking at the Techonomy 2011 conference Monday night, Mundie used the extensible business reporting language (XBRL) which underpins all financial reporting as an example of how the government can grease the skids for change. ”XBRL was ginned up by the SEC …. and went from a standing start to full adoption by U.S. companies in less than four years,” said Mundie.
Between Medicare and Medicaid, the U.S. government pays for more than half (54 percent) of aggregate domestic healthcare costs. “If the government said, ‘I don’t care how the private sector does it; we’ll do it this way,’ the whole industry would flip,” he said.
Free data flow is necessary so costs can be compared between the providers that are so reluctant to share their information. The fundamental issue with healthcare is “perverse payment system”: Most people don’t care about what their care really costs because a third party foots the bill. “There is no braking component” to contain runaway costs, he said.
If, however, Americans paid based on outcomes, there needs to be a total view of data to allow comparisons between providers. But until the data is in a form that can be shared — with privacy concerns met — he insisted there’s no basis for comparison and no way to do quality control.
“We’re not serious yet about aggregating this data. If the data is not available, you can’t make wholesale change.”
He didn’t venture a guess as to how much could be saved overall, but said in a smaller test case, Microsoft worked with a system of several hospitals that made their information available to a heath information exchange. “It slurped up all their Medicare data, and as soon as it had that data, they could query it to see just what was costing so much.”
Once the data is in one place, it’s not expensive to work with it, he noted, and it turned out that queries were able to identify a relatively small number of uninsured people who visited various emergency rooms around the system and ran up millions of dollars in charges. “Without this data, there was no way to find these people,” he said, adding that just eliminating redundancy and fraud alone would result in huge savings in a national context.
And, he insisted privacy issues can be addressed. Using metadata, “you can describe the provenance of the data to support sharing and also take privacy constraints into account,” he said. “Technology can be privacy-enhancing,” he noted.
But it’s highly unlikely insurance companies will drive this sort of change given the current privacy regulations and that their business is so profitable the way it is now, hence the need for a strong-spined government mandate.
Asked if these big data queries on patient care represent a “Google-like” capability for healthcare, Mundie said it’s “really a Bing-like system for healthcare, but absolutely.”
Monday, November 14, 2011
BigQuery Service: Big data analytics at Google speed
Posted by Ju-kay Kwek, Product Manager November 14, 2011
(Cross-posted on the Google App Engine Blog and the Google Code Blog.)
Rapidly crunching terabytes of big data can lead to better business decisions, but this has traditionally required tremendous IT investments. Imagine a large online retailer that wants to provide better product recommendations by analyzing website usage and purchase patterns from millions of website visits. Or consider a car manufacturer that wants to maximize its advertising impact by learning how its last global campaign performed across billions of multimedia impressions. Fortune 500 companies struggle to unlock the potential of data, so it’s no surprise that it’s been even harder for smaller businesses.
We developed Google BigQuery Service for large-scale internal data analytics. At Google I/O last year, we opened a preview of the service to a limited number of enterprises and developers. Today we're releasing some big improvements, and putting one of Google's most powerful data analysis systems into the hands of more companies of all sizes.
· We’ve added a graphical user interface for analysts and developers to rapidly explore massive data through a web application.
· We’ve made big improvements for customers accessing the service programmatically through the API. The new REST API lets you run multiple jobs in the background and manage tables and permissions with more granularity.
· Whether you use the BigQuery web application or API, you can now write even more powerful queries with JOIN statements. This lets you run queries across multiple data tables, linked by data that tables have in common.
· It’s also now easy to manage, secure, and share access to your data tables in BigQuery, and export query results to the desktop or to Google Cloud Storage.
Michael J. Franklin, Professor of Computer Science at UC Berkeley, remarked that BigQuery (internally known as Dremel) leverages “thousands of machines to process data at a scale that is simply jaw-dropping given the current state of the art.” We’re looking forward to helping businesses innovate faster by harnessing their own large data sets. BigQuery is available free of charge for now, and we’ll let customers know at least 30 days before the free period ends. We’re bringing on a new batch of pilot customers, so let us know if your business wants to test-drive BigQuery Service.
(Cross-posted on the Google App Engine Blog and the Google Code Blog.)
Rapidly crunching terabytes of big data can lead to better business decisions, but this has traditionally required tremendous IT investments. Imagine a large online retailer that wants to provide better product recommendations by analyzing website usage and purchase patterns from millions of website visits. Or consider a car manufacturer that wants to maximize its advertising impact by learning how its last global campaign performed across billions of multimedia impressions. Fortune 500 companies struggle to unlock the potential of data, so it’s no surprise that it’s been even harder for smaller businesses.
We developed Google BigQuery Service for large-scale internal data analytics. At Google I/O last year, we opened a preview of the service to a limited number of enterprises and developers. Today we're releasing some big improvements, and putting one of Google's most powerful data analysis systems into the hands of more companies of all sizes.
· We’ve added a graphical user interface for analysts and developers to rapidly explore massive data through a web application.
· We’ve made big improvements for customers accessing the service programmatically through the API. The new REST API lets you run multiple jobs in the background and manage tables and permissions with more granularity.
· Whether you use the BigQuery web application or API, you can now write even more powerful queries with JOIN statements. This lets you run queries across multiple data tables, linked by data that tables have in common.
· It’s also now easy to manage, secure, and share access to your data tables in BigQuery, and export query results to the desktop or to Google Cloud Storage.
Michael J. Franklin, Professor of Computer Science at UC Berkeley, remarked that BigQuery (internally known as Dremel) leverages “thousands of machines to process data at a scale that is simply jaw-dropping given the current state of the art.” We’re looking forward to helping businesses innovate faster by harnessing their own large data sets. BigQuery is available free of charge for now, and we’ll let customers know at least 30 days before the free period ends. We’re bringing on a new batch of pilot customers, so let us know if your business wants to test-drive BigQuery Service.
Disruptions: The 3-D Printing Free-for-All
Nick Bilton The New York Times November 13, 2011
Downloading — quite often stealing, in the eyes of the law — music, movies, books and photos is easier than bobbing for apples in a bucket without water. It has kept legions of lawyers employed fighting copyright violations without a whole lot to show for their efforts in the past decade.
You think that was bad? Just wait until we can copy physical things.
It won’t be long before people have a 3-D printer sitting at home alongside its old inkjet counterpart. These 3-D printers, some already costing less than a computer did in 1999, can print objects by spraying layers of plastic, metal or ceramics into shapes. People can download plans for an object, hit print, and a few minutes later have it in their hands.
Call it the Industrial Revolution 2.0. Not only will it change the nature of manufacturing, but it will further challenge our concept of ownership and copyright. Suppose you covet a lovely new mug at a friend’s house. So you snap a few pictures of it. Software renders those photos into designs that you use to print copies of the mug on your home 3-D printer.
Did you break the law by doing this? You might think so, but surprisingly, you didn’t.
What about a lamp, a vase, an iPhone protective cover, board game pieces, wall hooks, even large pieces of furniture? In each of these cases, if you copy them, it’s highly unlikely that you’re breaking any copyright laws.
“Copyright doesn’t necessarily protect useful things,” said Michael Weinberg, a senior staff attorney with Public Knowledge, a Washington digital advocacy group. “If an object is purely aesthetic it will be protected by copyright, but if the object does something, it is not the kind of thing that can be protected.”
When I posed my mug scenario to Mr. Weinberg, he responded: “If you took that mug and went to a pottery class and remade it, would you be asking me the same questions about breaking a copyright law? No.” Just because new tools arrive, like 3-D printers and digital files that make it easier to recreate an object, he said, it doesn’t mean people break the law when using them.
But it could turn design and manufacturing into the Wild West. That’s already happening on Thingiverse, a free online site that offers schematics of more than 15,000 objects. Thomas Lombardi, a 3-D printer owner and regular contributor to Thingiverse, uploaded a free design for a “Lucky Charms Cereal Sifter.” This brilliant piece of American engineering is a cup with several holes in the bottom. When you pour Lucky Charms cereal into the sifter and shake it from side to side, the cereal falls through the holes and the marshmallow charms — clearly the most sought-after part of the product — stay in the sifter, leaving you with nothing but marshmallowy goodness to pour into a bowl.
After Mr. Lombardi posted his invention on Thingiverse, someone else downloaded the design and began selling a finished Lucky Charms Cereal Sifter on a competing Web site for $30.
Because the sifter is a useful object (although some might argue otherwise) and not simply decorative, there was nothing Mr. Lombardi could have done to stop them.
A recent research paper published by the Institute for the Future in Palo Alto, Calif., titled “The Future of Open Fabrication,” says 3-D printing will be “manufacturing’s Big Bang.” as jobs in manufacturing, many overseas, and jobs shipping products around the globe are replaced by companies setting up 3-D fabrication labs in stores to print objects rather than ship them.
The disregard for copyright smoothes the way for this shift. Downloading music online prospered because it was quicker and easier to press a button than go to a store to buy a CD. Given the choice to download a mug, or deal with Ikea on a Saturday afternoon, which one do you think you would choose?
Downloading — quite often stealing, in the eyes of the law — music, movies, books and photos is easier than bobbing for apples in a bucket without water. It has kept legions of lawyers employed fighting copyright violations without a whole lot to show for their efforts in the past decade.
You think that was bad? Just wait until we can copy physical things.
It won’t be long before people have a 3-D printer sitting at home alongside its old inkjet counterpart. These 3-D printers, some already costing less than a computer did in 1999, can print objects by spraying layers of plastic, metal or ceramics into shapes. People can download plans for an object, hit print, and a few minutes later have it in their hands.
Call it the Industrial Revolution 2.0. Not only will it change the nature of manufacturing, but it will further challenge our concept of ownership and copyright. Suppose you covet a lovely new mug at a friend’s house. So you snap a few pictures of it. Software renders those photos into designs that you use to print copies of the mug on your home 3-D printer.
Did you break the law by doing this? You might think so, but surprisingly, you didn’t.
What about a lamp, a vase, an iPhone protective cover, board game pieces, wall hooks, even large pieces of furniture? In each of these cases, if you copy them, it’s highly unlikely that you’re breaking any copyright laws.
“Copyright doesn’t necessarily protect useful things,” said Michael Weinberg, a senior staff attorney with Public Knowledge, a Washington digital advocacy group. “If an object is purely aesthetic it will be protected by copyright, but if the object does something, it is not the kind of thing that can be protected.”
When I posed my mug scenario to Mr. Weinberg, he responded: “If you took that mug and went to a pottery class and remade it, would you be asking me the same questions about breaking a copyright law? No.” Just because new tools arrive, like 3-D printers and digital files that make it easier to recreate an object, he said, it doesn’t mean people break the law when using them.
But it could turn design and manufacturing into the Wild West. That’s already happening on Thingiverse, a free online site that offers schematics of more than 15,000 objects. Thomas Lombardi, a 3-D printer owner and regular contributor to Thingiverse, uploaded a free design for a “Lucky Charms Cereal Sifter.” This brilliant piece of American engineering is a cup with several holes in the bottom. When you pour Lucky Charms cereal into the sifter and shake it from side to side, the cereal falls through the holes and the marshmallow charms — clearly the most sought-after part of the product — stay in the sifter, leaving you with nothing but marshmallowy goodness to pour into a bowl.
After Mr. Lombardi posted his invention on Thingiverse, someone else downloaded the design and began selling a finished Lucky Charms Cereal Sifter on a competing Web site for $30.
Because the sifter is a useful object (although some might argue otherwise) and not simply decorative, there was nothing Mr. Lombardi could have done to stop them.
A recent research paper published by the Institute for the Future in Palo Alto, Calif., titled “The Future of Open Fabrication,” says 3-D printing will be “manufacturing’s Big Bang.” as jobs in manufacturing, many overseas, and jobs shipping products around the globe are replaced by companies setting up 3-D fabrication labs in stores to print objects rather than ship them.
The disregard for copyright smoothes the way for this shift. Downloading music online prospered because it was quicker and easier to press a button than go to a store to buy a CD. Given the choice to download a mug, or deal with Ikea on a Saturday afternoon, which one do you think you would choose?
Google's Lab of Wildest Dreams
Claire Cain Miller & Nick Bilton The New York Times November 13, 2011
MOUNTAIN VIEW, Calif. — In a top-secret lab in an undisclosed Bay Area location where robots run free, the future is being imagined.
It’s a place where your refrigerator could be connected to the Internet, so it could order groceries when they ran low. Your dinner plate could post to a social network what you’re eating. Your robot could go to the office while you stay home in your pajamas. And you could, perhaps, take an elevator to outer space.
These are just a few of the dreams being chased at Google X, the clandestine lab where Google is tackling a list of 100 shoot-for-the-stars ideas. In interviews, a dozen people discussed the list; some work at the lab or elsewhere at Google, and some have been briefed on the project. But none would speak for attribution because Google is so secretive about the effort that many employees do not even know the lab exists.
Although most of the ideas on the list are in the conceptual stage, nowhere near reality, two people briefed on the project said one product would be released by the end of the year, although they would not say what it was.
“They’re pretty far out in front right now,” said Rodney Brooks, a professor emeritus at M.I.T.’s computer science and artificial intelligence lab and founder of Heartland Robotics. “But Google’s not an ordinary company, so almost nothing applies.”
At most Silicon Valley companies, innovation means developing online apps or ads, but Google sees itself as different. Even as Google has grown into a major corporation and tech start-ups are biting at its heels, the lab reflects its ambition to be a place where ground-breaking research and development are happening, in the tradition of Xerox PARC, which developed the modern personal computer in the 1970s.
A Google spokeswoman, Jill Hazelbaker, declined to comment on the lab, but said that investing in speculative projects was an important part of Google’s DNA. “While the possibilities are incredibly exciting, please do keep in mind that the sums involved are very small by comparison to the investments we make in our core businesses,” she said.
At Google, which uses artificial intelligence techniques and machine learning in its search algorithm, some of the outlandish projects may not be as much of a stretch as they first appear, even though they defy the bounds of the company’s main Web search business.
For example, space elevators, a longtime fantasy of Google’s founders and other Silicon Valley entrepreneurs, could collect information or haul things into space. (In theory, they involve rocketless space travel along a cable anchored to Earth.) “Google is collecting the world’s data, so now it could be collecting the solar system’s data,” Mr. Brooks said.
Sergey Brin, Google’s co-founder, is deeply involved in the lab, said several people with knowledge of it, and came up with the list of ideas along with Larry Page, Google’s other founder, who worked on Google X before becoming chief executive in April; Eric E. Schmidt, its chairman; and other top executives. “Where I spend my time is farther afield projects, which we hope will graduate to important key businesses in the future,” Mr. Brin said recently, though he did not mention Google X.
Google may turn one of the ideas — the driverless cars that it unleashed on California’s roads last year — into a new business. Unimpressed by the innovative spirit of Detroit automakers, Google now is considering manufacturing them in the United States, said a person briefed on the effort.
Google could sell navigation or information technology for the cars, and theoretically could show location-based ads to passengers as they zoom by local businesses while playing Angry Birds in the driver’s seat.
Robots figure prominently in many of the ideas. They have long captured the imagination of Google engineers, including Mr. Brin, who has already attended a conference through robot instead of in the flesh.
Fleets of robots could assist Google with collecting information, replacing the humans that photograph streets for Google Maps, say people with knowledge of Google X. Robots born in the lab could be destined for homes and offices, where they could assist with mundane tasks or allow people to work remotely, they say.
Other ideas involve what Google referred to as the “Web of things” at its software developers conference in May — a way of connecting objects to the Internet. Every time anyone uses the Web, it benefits Google, the company argued, so it could be good for Google if home accessories and wearable objects, not just computers, were connected.
Among the items that could be connected: a garden planter (so it could be watered from afar); a coffee pot (so it could be set to brew remotely); or a light bulb (so it could be turned off remotely). Google said in May that by the end of this year another team planned to introduce a Web-connected light bulb that could communicate wirelessly with Android devices.
One Google engineer familiar with Google X said it was run as mysteriously as the C.I.A. — with two offices, a nondescript one for logistics, on the company’s Mountain View campus, and one for robots, in a secret location.
While software engineers toil away elsewhere at Google, the lab is filled with roboticists and electrical engineers. They have been hired from Microsoft, Nokia Labs, Stanford, M.I.T., Carnegie Mellon and New York University.
A leader at Google X is Sebastian Thrun, one of the world’s top robotics and artificial intelligence experts, who teaches computer science at Stanford and invented the world’s first driverless car. Also at the lab is Andrew Ng, another Stanford professor, who specializes in applying neuroscience to artificial intelligence to teach robots and machines to operate like people.
Johnny Chung Lee, a specialist in human-computer interaction, came to Google X from Microsoft this year after helping develop Microsoft’s Kinect, the video game player that responds to human movement and voice. At Google X, where he is working on the Web of things, according to people familiar with his role, he has the mysterious title of rapid evaluator.
Because Google X is a breeding ground for big bets that could turn into colossal failures or Google’s next big business — and it could take years to figure out which — just the idea of these experiments terrifies some shareholders and analysts.
“These moon-shot projects are a very Google-y thing for them to do,” said Colin W. Gillis, an analyst at BGC Partners. “People don’t love it but they tolerate it because their core search business is firing away.”
Mr. Page has tried to appease analysts by saying that crazy projects are a tiny proportion of Google’s work.
“There are a few small, speculative projects happening at any one time, but we are very careful stewards of shareholders’ money,” he told analysts in July. “We are not betting the farm on these.”
MOUNTAIN VIEW, Calif. — In a top-secret lab in an undisclosed Bay Area location where robots run free, the future is being imagined.
It’s a place where your refrigerator could be connected to the Internet, so it could order groceries when they ran low. Your dinner plate could post to a social network what you’re eating. Your robot could go to the office while you stay home in your pajamas. And you could, perhaps, take an elevator to outer space.
These are just a few of the dreams being chased at Google X, the clandestine lab where Google is tackling a list of 100 shoot-for-the-stars ideas. In interviews, a dozen people discussed the list; some work at the lab or elsewhere at Google, and some have been briefed on the project. But none would speak for attribution because Google is so secretive about the effort that many employees do not even know the lab exists.
Although most of the ideas on the list are in the conceptual stage, nowhere near reality, two people briefed on the project said one product would be released by the end of the year, although they would not say what it was.
“They’re pretty far out in front right now,” said Rodney Brooks, a professor emeritus at M.I.T.’s computer science and artificial intelligence lab and founder of Heartland Robotics. “But Google’s not an ordinary company, so almost nothing applies.”
At most Silicon Valley companies, innovation means developing online apps or ads, but Google sees itself as different. Even as Google has grown into a major corporation and tech start-ups are biting at its heels, the lab reflects its ambition to be a place where ground-breaking research and development are happening, in the tradition of Xerox PARC, which developed the modern personal computer in the 1970s.
A Google spokeswoman, Jill Hazelbaker, declined to comment on the lab, but said that investing in speculative projects was an important part of Google’s DNA. “While the possibilities are incredibly exciting, please do keep in mind that the sums involved are very small by comparison to the investments we make in our core businesses,” she said.
At Google, which uses artificial intelligence techniques and machine learning in its search algorithm, some of the outlandish projects may not be as much of a stretch as they first appear, even though they defy the bounds of the company’s main Web search business.
For example, space elevators, a longtime fantasy of Google’s founders and other Silicon Valley entrepreneurs, could collect information or haul things into space. (In theory, they involve rocketless space travel along a cable anchored to Earth.) “Google is collecting the world’s data, so now it could be collecting the solar system’s data,” Mr. Brooks said.
Sergey Brin, Google’s co-founder, is deeply involved in the lab, said several people with knowledge of it, and came up with the list of ideas along with Larry Page, Google’s other founder, who worked on Google X before becoming chief executive in April; Eric E. Schmidt, its chairman; and other top executives. “Where I spend my time is farther afield projects, which we hope will graduate to important key businesses in the future,” Mr. Brin said recently, though he did not mention Google X.
Google may turn one of the ideas — the driverless cars that it unleashed on California’s roads last year — into a new business. Unimpressed by the innovative spirit of Detroit automakers, Google now is considering manufacturing them in the United States, said a person briefed on the effort.
Google could sell navigation or information technology for the cars, and theoretically could show location-based ads to passengers as they zoom by local businesses while playing Angry Birds in the driver’s seat.
Robots figure prominently in many of the ideas. They have long captured the imagination of Google engineers, including Mr. Brin, who has already attended a conference through robot instead of in the flesh.
Fleets of robots could assist Google with collecting information, replacing the humans that photograph streets for Google Maps, say people with knowledge of Google X. Robots born in the lab could be destined for homes and offices, where they could assist with mundane tasks or allow people to work remotely, they say.
Other ideas involve what Google referred to as the “Web of things” at its software developers conference in May — a way of connecting objects to the Internet. Every time anyone uses the Web, it benefits Google, the company argued, so it could be good for Google if home accessories and wearable objects, not just computers, were connected.
Among the items that could be connected: a garden planter (so it could be watered from afar); a coffee pot (so it could be set to brew remotely); or a light bulb (so it could be turned off remotely). Google said in May that by the end of this year another team planned to introduce a Web-connected light bulb that could communicate wirelessly with Android devices.
One Google engineer familiar with Google X said it was run as mysteriously as the C.I.A. — with two offices, a nondescript one for logistics, on the company’s Mountain View campus, and one for robots, in a secret location.
While software engineers toil away elsewhere at Google, the lab is filled with roboticists and electrical engineers. They have been hired from Microsoft, Nokia Labs, Stanford, M.I.T., Carnegie Mellon and New York University.
A leader at Google X is Sebastian Thrun, one of the world’s top robotics and artificial intelligence experts, who teaches computer science at Stanford and invented the world’s first driverless car. Also at the lab is Andrew Ng, another Stanford professor, who specializes in applying neuroscience to artificial intelligence to teach robots and machines to operate like people.
Johnny Chung Lee, a specialist in human-computer interaction, came to Google X from Microsoft this year after helping develop Microsoft’s Kinect, the video game player that responds to human movement and voice. At Google X, where he is working on the Web of things, according to people familiar with his role, he has the mysterious title of rapid evaluator.
Because Google X is a breeding ground for big bets that could turn into colossal failures or Google’s next big business — and it could take years to figure out which — just the idea of these experiments terrifies some shareholders and analysts.
“These moon-shot projects are a very Google-y thing for them to do,” said Colin W. Gillis, an analyst at BGC Partners. “People don’t love it but they tolerate it because their core search business is firing away.”
Mr. Page has tried to appease analysts by saying that crazy projects are a tiny proportion of Google’s work.
“There are a few small, speculative projects happening at any one time, but we are very careful stewards of shareholders’ money,” he told analysts in July. “We are not betting the farm on these.”
Internet Architects Warn of Risks in Ultrafast Networks
Quentin Hardy The New YorkTimes November 13, 2011
SANTA CLARA, Calif. — If nothing else, Arista Networks proves that two people can make more than $1 billion each building the Internet and still be worried about its reliability.
David Cheriton, a computer science professor at Stanford known for his skills in software design, and Andreas Bechtolsheim, one of the founders of Sun Microsystems, have committed $100 million of their money, and spent half that, to shake up the business of connecting computers in the Internet’s big computing centers.
As the Arista founders say, the promise of having access to mammoth amounts of data instantly, anywhere, is matched by the threat of catastrophe. People are creating more data and moving it ever faster on computer networks. The fast networks allow people to pour much more of civilization online, including not just Facebook posts and every book ever written, but all music, live video calls, and most of the information technology behind modern business, into a worldwide “cloud” of data centers. The networks are designed so it will always be available, via phone, tablet, personal computer or an increasing array of connected devices.
Statistics dictate that the vastly greater number of transactions among computers in a world 100 times faster than today will lead to a greater number of unpredictable accidents, with less time in between them. Already, Amazon’s cloud for businesses failed for several hours in April, when normal computer routines faltered and the system overloaded. Google’s cloud of e-mail and document collaboration software has been interrupted several times.
“We think of the Internet as always there. Just because we’ve become dependent on it, that doesn’t mean it’s true,” Mr. Cheriton says. Mr. Bechtolsheim says that because of the Internet’s complexity, the global network is impossible to design without bugs. Very dangerous bugs, as they describe them, capable of halting commerce, destroying financial information or enabling hostile attacks by foreign powers.
Both were among the first investors in Google, which made them billionaires, and, before that, they created and sold a company to the networking giant Cisco Systems for $220 million. Wealth and reputations as technology seers give their arguments about the risks of faster networks rare credibility.
More transactions also mean more system attacks. Even though he says there is no turning back on the online society, Mr. Cheriton worries most about security hazards. “I’ve made the claim that the Chinese military can take it down in 30 seconds, no one can prove me wrong,” he said. By building a new way to run networks in the cloud era, he says, “we have a path to having software that is more sophisticated, can be self-defending, and is able to detect more problems, quicker.”
The common connection among computer servers, one gigabit per second, is giving way to 10-gigabit connections, because of improvements in semiconductor design and software. Speeds of 40 gigabits, even 100 gigabits, are now used for specialty purposes like consolidating huge data streams among hundreds of thousands of computers across the globe, and that technology is headed into the mainstream. An engineering standard for a terabit per second, 1,000 gigabits, is expected in about seven years.
Arista, which is based here, was built with the 10-gigabit world in mind. It now has 250 employees, 167 of them engineers, building a fast data-routing switch that could isolate problems and fix them without ever shutting down the network. It is intended to run on inexpensive mass-produced chips. In terms of software and hardware, it was a big break from the way things had been done in networking for the last quarter-century.
“Companies like Cisco had to build their own specialty chips to work at high speed for the time,” Mr. Bechtolsheim said. Because of improvements in the quality and capability of the kind of chips used in computers, phones and cable television boxes, “we could build a network that is a lot more software-enabled, something that is a lot easier to defend and modify,” he said.
For Mr. Cheriton, who cuts his own hair despite his great wealth, Arista was an opportunity to work on a new style of software he said he had been thinking about since 1989.
No matter how complex, software is essentially a linear system of commands: Do this, and then do that. Sometimes it is divided into “objects” or modules, but these tend to operate sequentially.
From 2004 to 2008, when Arista shipped its first product, Mr. Cheriton developed a five million-line system that breaks operations into a series of tasks, which when completed, other parts of the program can check on and pick up if everything seems fine. If it does not, the problem is rapidly isolated and addressed. Mr. Bechtolsheim worked with him to make the system operate with chips that were already on the market.
The first products were sold to financial traders looking to shave 100 nanoseconds off their high-frequency trades. Arista has more than 1,000 customers now, including telecommunications companies and university research laboratories.
“They have created something that is architecturally unique in networking, with a lot of value for the industry,” says Nicholas Lippis, who tests and evaluates switching equipment. “They built something fast that has a unique value for the industry.”
Kenneth Duda, another founder, said, “What drives us here is finding a new way to do software.” Mr. Duda also worked with Mr. Cheriton and Mr. Bechtolsheim at Granite Systems, the company they sold to Cisco. “The great enemy is complexity, measured in lines of code, or interactions,” he said. In the world of cloud computing, “there is no person alive who can understand 10 percent of the technology involved in my writing and printing out an online shopping list.”
Not surprisingly, Cisco, which dominates the $5 billion network switching business, disagrees.
“You don’t have to reinvent the Internet,” says Ram Velaga, vice president for product management in Cisco’s core technology group. “These protocols were designed to work even if Washington is taken out. That is in the architecture.”
Still, Cisco’s newest data center switches have rewritten software in a way more like Arista’s. A few products are using so-called merchant silicon, instead of its typical custom chips. “Andy made a bet that Cisco would never use merchant silicon,” Mr. Velaga says.
Mr. Cheriton and Mr. Bechtolsheim have known each other since 1981, when Mr. Cheriton arrived from his native Canada to teach at Stanford. Mr. Bechtolsheim, a native of Germany, was studying electrical engineering and building what became Sun’s first product, a computer workstation.
The two became friends and intellectual compatriots, and in 1994 began Granite Networks, which made one of the first gigabit switches. Cisco bought the company two years later.
With no outside investors in Arista, they could take as long as they wanted on the product, Mr. Bechtolsheim said.
“Venture capitalists have no patience for a product to develop.” he said. “Pretty soon they want to bring in their best buddy as the C.E.O. Besides, this looked like a good investment.”
Mr. Cheriton said, “Not being venture funded was definitely a competitive advantage.” Besides, he said, “Andy never told me it would be $100 million.”
SANTA CLARA, Calif. — If nothing else, Arista Networks proves that two people can make more than $1 billion each building the Internet and still be worried about its reliability.
David Cheriton, a computer science professor at Stanford known for his skills in software design, and Andreas Bechtolsheim, one of the founders of Sun Microsystems, have committed $100 million of their money, and spent half that, to shake up the business of connecting computers in the Internet’s big computing centers.
As the Arista founders say, the promise of having access to mammoth amounts of data instantly, anywhere, is matched by the threat of catastrophe. People are creating more data and moving it ever faster on computer networks. The fast networks allow people to pour much more of civilization online, including not just Facebook posts and every book ever written, but all music, live video calls, and most of the information technology behind modern business, into a worldwide “cloud” of data centers. The networks are designed so it will always be available, via phone, tablet, personal computer or an increasing array of connected devices.
Statistics dictate that the vastly greater number of transactions among computers in a world 100 times faster than today will lead to a greater number of unpredictable accidents, with less time in between them. Already, Amazon’s cloud for businesses failed for several hours in April, when normal computer routines faltered and the system overloaded. Google’s cloud of e-mail and document collaboration software has been interrupted several times.
“We think of the Internet as always there. Just because we’ve become dependent on it, that doesn’t mean it’s true,” Mr. Cheriton says. Mr. Bechtolsheim says that because of the Internet’s complexity, the global network is impossible to design without bugs. Very dangerous bugs, as they describe them, capable of halting commerce, destroying financial information or enabling hostile attacks by foreign powers.
Both were among the first investors in Google, which made them billionaires, and, before that, they created and sold a company to the networking giant Cisco Systems for $220 million. Wealth and reputations as technology seers give their arguments about the risks of faster networks rare credibility.
More transactions also mean more system attacks. Even though he says there is no turning back on the online society, Mr. Cheriton worries most about security hazards. “I’ve made the claim that the Chinese military can take it down in 30 seconds, no one can prove me wrong,” he said. By building a new way to run networks in the cloud era, he says, “we have a path to having software that is more sophisticated, can be self-defending, and is able to detect more problems, quicker.”
The common connection among computer servers, one gigabit per second, is giving way to 10-gigabit connections, because of improvements in semiconductor design and software. Speeds of 40 gigabits, even 100 gigabits, are now used for specialty purposes like consolidating huge data streams among hundreds of thousands of computers across the globe, and that technology is headed into the mainstream. An engineering standard for a terabit per second, 1,000 gigabits, is expected in about seven years.
Arista, which is based here, was built with the 10-gigabit world in mind. It now has 250 employees, 167 of them engineers, building a fast data-routing switch that could isolate problems and fix them without ever shutting down the network. It is intended to run on inexpensive mass-produced chips. In terms of software and hardware, it was a big break from the way things had been done in networking for the last quarter-century.
“Companies like Cisco had to build their own specialty chips to work at high speed for the time,” Mr. Bechtolsheim said. Because of improvements in the quality and capability of the kind of chips used in computers, phones and cable television boxes, “we could build a network that is a lot more software-enabled, something that is a lot easier to defend and modify,” he said.
For Mr. Cheriton, who cuts his own hair despite his great wealth, Arista was an opportunity to work on a new style of software he said he had been thinking about since 1989.
No matter how complex, software is essentially a linear system of commands: Do this, and then do that. Sometimes it is divided into “objects” or modules, but these tend to operate sequentially.
From 2004 to 2008, when Arista shipped its first product, Mr. Cheriton developed a five million-line system that breaks operations into a series of tasks, which when completed, other parts of the program can check on and pick up if everything seems fine. If it does not, the problem is rapidly isolated and addressed. Mr. Bechtolsheim worked with him to make the system operate with chips that were already on the market.
The first products were sold to financial traders looking to shave 100 nanoseconds off their high-frequency trades. Arista has more than 1,000 customers now, including telecommunications companies and university research laboratories.
“They have created something that is architecturally unique in networking, with a lot of value for the industry,” says Nicholas Lippis, who tests and evaluates switching equipment. “They built something fast that has a unique value for the industry.”
Kenneth Duda, another founder, said, “What drives us here is finding a new way to do software.” Mr. Duda also worked with Mr. Cheriton and Mr. Bechtolsheim at Granite Systems, the company they sold to Cisco. “The great enemy is complexity, measured in lines of code, or interactions,” he said. In the world of cloud computing, “there is no person alive who can understand 10 percent of the technology involved in my writing and printing out an online shopping list.”
Not surprisingly, Cisco, which dominates the $5 billion network switching business, disagrees.
“You don’t have to reinvent the Internet,” says Ram Velaga, vice president for product management in Cisco’s core technology group. “These protocols were designed to work even if Washington is taken out. That is in the architecture.”
Still, Cisco’s newest data center switches have rewritten software in a way more like Arista’s. A few products are using so-called merchant silicon, instead of its typical custom chips. “Andy made a bet that Cisco would never use merchant silicon,” Mr. Velaga says.
Mr. Cheriton and Mr. Bechtolsheim have known each other since 1981, when Mr. Cheriton arrived from his native Canada to teach at Stanford. Mr. Bechtolsheim, a native of Germany, was studying electrical engineering and building what became Sun’s first product, a computer workstation.
The two became friends and intellectual compatriots, and in 1994 began Granite Networks, which made one of the first gigabit switches. Cisco bought the company two years later.
With no outside investors in Arista, they could take as long as they wanted on the product, Mr. Bechtolsheim said.
“Venture capitalists have no patience for a product to develop.” he said. “Pretty soon they want to bring in their best buddy as the C.E.O. Besides, this looked like a good investment.”
Mr. Cheriton said, “Not being venture funded was definitely a competitive advantage.” Besides, he said, “Andy never told me it would be $100 million.”
Wednesday, November 9, 2011
Techonomy: Digitized Decision Making and the Hidden Second Economy
Stephen Hoover Forbes 11/09/2011 @ 4:26PM
There’s something big happening right now. I’m not referring to any of the popular technology memesper se—big data, social, cloud, mobile, augmented reality, context, post-PC devices, consumerization, 3-D printing, etc. I’m referring to something behind, and beyond, all of these technologies: the digitization of decision making. This increasing trend is creating a “second economy” underneath and alongside the physical economy we know so well, and on a revolutionary scale.
Economist W. Brian Arthur, a longtime visiting researcher at PARC and external professor at the Santa Fe Institute, quantified this phenomenon recently in a PARC Forum video and McKinsey Quarterly article:
Since joining PARC, a Xerox company approaching its 10-year anniversary as a business for open innovation with multiple clients, I have been focused on the following question: just what will happen to invention and innovation in this second economy?More specifically, what will be the role of R&D and innovation organizations in a new global innovation landscape?
On the one hand, aspects of the second economy can empower innovation, given the mass democratization of tools. Some would argue, for example, that “citizen science” is transforming the field of bioscience. Chris Anderson touched on a similar theme in his commentsabout the industrialization of the Maker movement, noting, “What we have discovered…are these extraordinary new ways to work—new ways to communicate, new ways to congregate, new ways to find talent, new ways to build on ideas. These post-institutional organizational models transcend geography, transcend your credentials…and everything else.”
On the other hand, innovation is more than tools and approaches. Based on our experiences and observations, the kind of innovation that transforms businesses requires the intersection of three things: 1. Expertise in technology and its trends (this goes beyond trend-spotting!); 2. A deep understanding of human behavior and context; and 3. Not just new business models, but new models for business—that is, how business is done and the new mindsets under which we’ll need to manage risk.
Just one of these is no longer enough, especially as we move beyond the services economy into what I call the “post-services age.” Our innovation models need to address what’s required today to sustain and grow our businesses, while anticipating the uncertainty ahead.
Stephen Hoover is CEO of PARC, a Xerox company that works with Fortune Global 500 and medium-sized companies, startups, and government agencies and partners to invent, co-develop, and deliver new business opportunities. Hoover oversees PARC’s work in diverse areas, from networking and novel electronics, to ethnography services, cleantech, and intelligent mobile computing.
For more information about the Techonomy 2011 conference (Nov. 13-15), visit www.techonomy.com. You can also follow Techonomy on Twitter andFacebook.
There’s something big happening right now. I’m not referring to any of the popular technology memesper se—big data, social, cloud, mobile, augmented reality, context, post-PC devices, consumerization, 3-D printing, etc. I’m referring to something behind, and beyond, all of these technologies: the digitization of decision making. This increasing trend is creating a “second economy” underneath and alongside the physical economy we know so well, and on a revolutionary scale.
Economist W. Brian Arthur, a longtime visiting researcher at PARC and external professor at the Santa Fe Institute, quantified this phenomenon recently in a PARC Forum video and McKinsey Quarterly article:
- The second economy…is vast, silent, connected, unseen, and autonomous (meaning that human beings may design it but are not directly involved in running it). It is remotely executing and global, always on, and endlessly configurable. It is concurrent…everything happens in parallel. It is self-configuring, meaning it constantly reconfigures itself on the fly, and increasingly, it is also self-organizing, self-architecting, and self-healing.
Since joining PARC, a Xerox company approaching its 10-year anniversary as a business for open innovation with multiple clients, I have been focused on the following question: just what will happen to invention and innovation in this second economy?More specifically, what will be the role of R&D and innovation organizations in a new global innovation landscape?
On the one hand, aspects of the second economy can empower innovation, given the mass democratization of tools. Some would argue, for example, that “citizen science” is transforming the field of bioscience. Chris Anderson touched on a similar theme in his commentsabout the industrialization of the Maker movement, noting, “What we have discovered…are these extraordinary new ways to work—new ways to communicate, new ways to congregate, new ways to find talent, new ways to build on ideas. These post-institutional organizational models transcend geography, transcend your credentials…and everything else.”
On the other hand, innovation is more than tools and approaches. Based on our experiences and observations, the kind of innovation that transforms businesses requires the intersection of three things: 1. Expertise in technology and its trends (this goes beyond trend-spotting!); 2. A deep understanding of human behavior and context; and 3. Not just new business models, but new models for business—that is, how business is done and the new mindsets under which we’ll need to manage risk.
Just one of these is no longer enough, especially as we move beyond the services economy into what I call the “post-services age.” Our innovation models need to address what’s required today to sustain and grow our businesses, while anticipating the uncertainty ahead.
Stephen Hoover is CEO of PARC, a Xerox company that works with Fortune Global 500 and medium-sized companies, startups, and government agencies and partners to invent, co-develop, and deliver new business opportunities. Hoover oversees PARC’s work in diverse areas, from networking and novel electronics, to ethnography services, cleantech, and intelligent mobile computing.
For more information about the Techonomy 2011 conference (Nov. 13-15), visit www.techonomy.com. You can also follow Techonomy on Twitter andFacebook.
Tuesday, November 8, 2011
McKinsey: Measuring the impact of the Internet
James Manyika and Charl Roxburgh, McKinsey Global Inst October 2011
In a paper prepared for the Foreign Commonwealth Office International Cyber Conference, MGI examines what more can be done to fully capture the benefits of the Internet.
The Internet is changing the way we work, socialize, create and share information, and organize the flow of people, ideas, and things around the globe. Yet the magnitude of this transformation is still underappreciated. The Internet accounted for 21 percent of the GDP growth in mature economies over the past 5 years. In that time, we went from a few thousand students accessing Facebook to more than 800 million users around the world, including many leading firms, who regularly update their pages and share content. While large enterprises and national economies have reaped major benefits from this technological revolution, individual consumers and small, upstart entrepreneurs have been some of the greatest beneficiaries from the Internet’s empowering influence. If Internet were a sector, it would have a greater weight in GDP than agriculture or utilities.
And yet we are still in the early stages of the transformations the Internet will unleash and the opportunities it will foster. Many more technological innovations and enabling capabilities such as payments platforms are likely to emerge, while the ability to connect many more people and things and engage them more deeply will continue to expand exponentially.
As a result, governments, policy makers, and businesses must recognize and embrace the enormous opportunities the Internet can create, even as they work to address the risks to security and privacy the Internet brings. As the Internet’s evolution over the past two decades has demonstrated, such work must include helping to nurture the development of a healthy Internet ecosystem, one that boosts infrastructure and access, builds a competitive environment that benefits users and lets innovators and entrepreneurs thrive, and nurtures human capital. Together these elements can maximize the continued impact of the Internet on economic growth and prosperity.
(http://www.mckinsey.com/mgi/publications/great_transformer/pdfs/McKinsey_the_great_transformer.pdf)
Read the full report (PDF - 232 KB)
Ignition, Accel Launches $100 Million Big Data Fund
Arik Hesseldahl All Things D, November 8, 2011
Elephants, it seems, are attracting money. As Hadoop World gets under way in New York today, Cloudera, the start-up company that is putting on the event, has landed an a big new investor.
A day after teaming up with thestorage concern NetApp, Cloudera announced today that it has landed a $40 million series D round of venture capital funding from Ignition Partners, in a round led by its partner, Frank Artale. Previous investors include Accel Partners, Greylock Partners, Meritech Capital Partners and In-Q-Tel. Cloudera says it will use the funds to expand its marketing and sales operations. By my count, the round brings Cloudera’s total capital raised so far to $76 million.
Cloudera has been on a roll — it’s the Hadoop outfit that many companies are turning to when they decide to tackle their big-data problems. Among its customers are eBay, AOL, Facebook and Groupon. While Hadoop itself is free for anyone to download and install from the Apache Software Foundation, Cloudera provides support and training, and an enterprise-ready version of Hadoop that has been tweaked for easier deployment in big companies.
And that’s not all the new money sloshing around the world of Hadoop, the open-source project with the cute cartoon elephant as its mascot. (Hence the money-origami elephant pictured above.)
Accel Partners, which led Cloudera’s last round, is launching a $100 million “Big Data Fund,” with Cloudera as a partner. The point, Accel partner Ping Li told me, is to fund companies working in what he calls the “big data stack,” whether that’s in infrastructure like storage, or security or management, or building applications that run on Hadoop. And the opportunities for that are multiplying, he told me.
“We’re seeing an undercurrent of picks-and-shovels kind of innovation around solving big data problems,” Li told me in an email. The volume of data is exploding at such a rate that it’s breaking traditional data-management technology like relational databases. It’s a problem that touches practically every industry. The fund will be overseen by several Accel partners based in the U.S., Europe, China and India.
Elephants, it seems, are attracting money. As Hadoop World gets under way in New York today, Cloudera, the start-up company that is putting on the event, has landed an a big new investor.
A day after teaming up with thestorage concern NetApp, Cloudera announced today that it has landed a $40 million series D round of venture capital funding from Ignition Partners, in a round led by its partner, Frank Artale. Previous investors include Accel Partners, Greylock Partners, Meritech Capital Partners and In-Q-Tel. Cloudera says it will use the funds to expand its marketing and sales operations. By my count, the round brings Cloudera’s total capital raised so far to $76 million.
Cloudera has been on a roll — it’s the Hadoop outfit that many companies are turning to when they decide to tackle their big-data problems. Among its customers are eBay, AOL, Facebook and Groupon. While Hadoop itself is free for anyone to download and install from the Apache Software Foundation, Cloudera provides support and training, and an enterprise-ready version of Hadoop that has been tweaked for easier deployment in big companies.
And that’s not all the new money sloshing around the world of Hadoop, the open-source project with the cute cartoon elephant as its mascot. (Hence the money-origami elephant pictured above.)
Accel Partners, which led Cloudera’s last round, is launching a $100 million “Big Data Fund,” with Cloudera as a partner. The point, Accel partner Ping Li told me, is to fund companies working in what he calls the “big data stack,” whether that’s in infrastructure like storage, or security or management, or building applications that run on Hadoop. And the opportunities for that are multiplying, he told me.
“We’re seeing an undercurrent of picks-and-shovels kind of innovation around solving big data problems,” Li told me in an email. The volume of data is exploding at such a rate that it’s breaking traditional data-management technology like relational databases. It’s a problem that touches practically every industry. The fund will be overseen by several Accel partners based in the U.S., Europe, China and India.
Friday, November 4, 2011
Big Data: Drugmakers Mine Data for Trial Patients
A Pfizer-led group plans to buy access to hospital records in New York State to help identify and enroll participants in drug studies
Carol Eisenberg Businessweek November 3, 2011
The bottom line: A group of 13 New York hospitals will sell access to patient data to drugmakers for $50,000 to $200,000 per search.
Pharmaceutical companies can easily spend years—and more than $1 billion—bringing a new drug to market, in part because they can’t find enough patients to do the required testing of the compound. Such delays can cost up to $1 million a day, fritter away valuable months of patent protection, and allow rival developers to catch up. One remedy: pay hospitals to sift through the health records of their patients.
Five big drugmakers, led by Pfizer, are planning to use electronic health data gathered from patients of 13 hospital systems across New York State to help them identify and enroll participants in drug studies. The effort, which begins testing this month, is projected to make $75 million a year for the hospitals and save the pharmaceutical companies time and money in developing new products. “This is going to be a game changer, making medicine more of a science and less of an art,” says John Murphy, senior director of clinical analytics for Quintiles Transnational (QTRN), which helps drugmakers conduct trials.
The program, called the Partnership to Advance Clinical Electronic Research, or PACeR, is among scores of such initiatives created nationally as medical providers, software vendors, and health data businesses seek ways to profit from the flood of clinical data now being gathered electronically. Besides Pfizer, companies helping fund PACeR include Merck (MRK), Roche, Johnson & Johnson (JNJ), and Bayer HealthCare. Quintiles and Oracle (ORCL) are developing the system.
Federal law bars medical providers, hospitals, and insurers from disclosing identifying information such as names, addresses, and Social Security numbers. So drug companies that want to test a new product or compound would pay PACeR to query the records systems of participating hospitals to compile a list of patients who match a trial’s requirements. Each query would cost between $50,000 and $200,000.
Once they determine how many patients might qualify and where they’re located, and get approval from a local ethics board, the hospitals would contact the patients’ doctors. A drugmaker would have access to personal information only if the patient consents, says David A. Krusch, head of PACeR’s leadership team and director of medical informatics at the University of Rochester Medical Center. “There is no central database,” Krusch says. “We’re not dumping big buckets of de-identified data anyplace.
Pfizer will not have a network connection, say, to the University of Rochester, or to any other participant in the network.”
Still, some civil liberties groups worry about data breaches and the potential for third parties to reconnect names to the data. “In a world where so much data is being retained, exchanged, and sold, being able to protect the privacy of individuals is a lot more difficult,” says Lillie Coney, associate director of the Electronic Privacy Information Center, a Washington-based advocacy group. Even data scrubbed of someone’s name and Social Security number can be “re-identified,” Coney warned, citing a 2006 case where AOL href="http://investing.businessweek.com/research/stocks/snapshot/snapshot.asp?ticker=AOL">AOL) released information about people’s search queries, which others were able to combine with publicly available data to identify an elderly Georgia woman who had used her AOL account to research medical conditions.
Consumer and ethics groups who advised the New York project say they see patient benefits as long as confidentiality is protected. “One of the problems with clinical trials is it’s very hard to get the information out to the average physician, or to the average cancer patient being treated in the community,” says research scholar Karen Maschke of the Hastings Center, a bioethics group that consulted on PACeR. Maschke says that while she supports the PACeR concept, “I’m also a skeptic. There probably should be conversations at the national level about best practices for these endeavors to guard confidentiality.”
Such discussions are expected to increase around the country as the use of electronic health records increases dramatically over the next few years, spurred by $27.4 billion set aside in the 2009 U.S. stimulus to pay doctors and hospitals to adopt and use them.
Drugmakers are big potential customers of aggregated health data since drug development times have more than doubled in the past 20 years without a marked increase in the percentage of compounds successfully brought to market. “If pharmaceutical companies can make this happen faster and more cheaply, they’re big winners,” says C. William Schroth, a former consultant for the New York State Health Dept., who first broached the concept behind PACeR to a group representing New York hospitals and also to drug companies.
Doctors and hospitals liked the idea because it makes them more attractive research partners. It wasn’t a hard sale for Big Pharma, either: Every day of delay during a Phase 3 trial costs drugmakers more than $1 million, with an average delay of 90 days, says David Leventhal, director of clinical innovation at Pfizer, citing a Deloitte Consulting analysis. “Even if that number is only 25 percent right, it’s still a compelling message.”
Carol Eisenberg Businessweek November 3, 2011
The bottom line: A group of 13 New York hospitals will sell access to patient data to drugmakers for $50,000 to $200,000 per search.
Pharmaceutical companies can easily spend years—and more than $1 billion—bringing a new drug to market, in part because they can’t find enough patients to do the required testing of the compound. Such delays can cost up to $1 million a day, fritter away valuable months of patent protection, and allow rival developers to catch up. One remedy: pay hospitals to sift through the health records of their patients.
Five big drugmakers, led by Pfizer, are planning to use electronic health data gathered from patients of 13 hospital systems across New York State to help them identify and enroll participants in drug studies. The effort, which begins testing this month, is projected to make $75 million a year for the hospitals and save the pharmaceutical companies time and money in developing new products. “This is going to be a game changer, making medicine more of a science and less of an art,” says John Murphy, senior director of clinical analytics for Quintiles Transnational (QTRN), which helps drugmakers conduct trials.
The program, called the Partnership to Advance Clinical Electronic Research, or PACeR, is among scores of such initiatives created nationally as medical providers, software vendors, and health data businesses seek ways to profit from the flood of clinical data now being gathered electronically. Besides Pfizer, companies helping fund PACeR include Merck (MRK), Roche, Johnson & Johnson (JNJ), and Bayer HealthCare. Quintiles and Oracle (ORCL) are developing the system.
Federal law bars medical providers, hospitals, and insurers from disclosing identifying information such as names, addresses, and Social Security numbers. So drug companies that want to test a new product or compound would pay PACeR to query the records systems of participating hospitals to compile a list of patients who match a trial’s requirements. Each query would cost between $50,000 and $200,000.
Once they determine how many patients might qualify and where they’re located, and get approval from a local ethics board, the hospitals would contact the patients’ doctors. A drugmaker would have access to personal information only if the patient consents, says David A. Krusch, head of PACeR’s leadership team and director of medical informatics at the University of Rochester Medical Center. “There is no central database,” Krusch says. “We’re not dumping big buckets of de-identified data anyplace.
Pfizer will not have a network connection, say, to the University of Rochester, or to any other participant in the network.”
Still, some civil liberties groups worry about data breaches and the potential for third parties to reconnect names to the data. “In a world where so much data is being retained, exchanged, and sold, being able to protect the privacy of individuals is a lot more difficult,” says Lillie Coney, associate director of the Electronic Privacy Information Center, a Washington-based advocacy group. Even data scrubbed of someone’s name and Social Security number can be “re-identified,” Coney warned, citing a 2006 case where AOL href="http://investing.businessweek.com/research/stocks/snapshot/snapshot.asp?ticker=AOL">AOL) released information about people’s search queries, which others were able to combine with publicly available data to identify an elderly Georgia woman who had used her AOL account to research medical conditions.
Consumer and ethics groups who advised the New York project say they see patient benefits as long as confidentiality is protected. “One of the problems with clinical trials is it’s very hard to get the information out to the average physician, or to the average cancer patient being treated in the community,” says research scholar Karen Maschke of the Hastings Center, a bioethics group that consulted on PACeR. Maschke says that while she supports the PACeR concept, “I’m also a skeptic. There probably should be conversations at the national level about best practices for these endeavors to guard confidentiality.”
Such discussions are expected to increase around the country as the use of electronic health records increases dramatically over the next few years, spurred by $27.4 billion set aside in the 2009 U.S. stimulus to pay doctors and hospitals to adopt and use them.
Drugmakers are big potential customers of aggregated health data since drug development times have more than doubled in the past 20 years without a marked increase in the percentage of compounds successfully brought to market. “If pharmaceutical companies can make this happen faster and more cheaply, they’re big winners,” says C. William Schroth, a former consultant for the New York State Health Dept., who first broached the concept behind PACeR to a group representing New York hospitals and also to drug companies.
Doctors and hospitals liked the idea because it makes them more attractive research partners. It wasn’t a hard sale for Big Pharma, either: Every day of delay during a Phase 3 trial costs drugmakers more than $1 million, with an average delay of 90 days, says David Leventhal, director of clinical innovation at Pfizer, citing a Deloitte Consulting analysis. “Even if that number is only 25 percent right, it’s still a compelling message.”
Thursday, November 3, 2011
Startup to solve big data problem collaboratively
Max Levchin Becomes Chairman Of Kaggle, A Startup That Helps NASA Solve Impossible Problems
Zachary Lichaa Business Insider November 3, 2011
San Francisco-based Kaggle has raised an $11 million Series A round from Index Ventures, Khosla Ventures, SV Angel, and others.
Slide's Max Levchin has also been named its chairman. Kaggle is a platform for solving some of the world's toughest data problems. NASA, Delloite, and The University of Michigan have all turned to Kaggle's pool of 17,000 PhD-level scientists to solve complex problems and create winning models.
Algorithms, which are at the heart of solving these complex data problems, are posted to Kaggle's site continuously. As they're posted, solutions become clearer and eventually the purveyor of that problem has an improved theory.
Kaggle sets up real-time leaderboards for each problem that's being solved and turns them into competitions to foster community engagement. Whoever solves the problem wins a prize.
When Kaggle is working with sensitive material, competitions are held privately using a selected group of scientists that are pre-screened.
Past projects include mapping HIV trends and Dark Matter in outer space. The latter was commissioned by NASA, and its European counterparts, the European Space Agency and The Royal Astronomical Society.
"The first time we could prove Kaggle worked was with HIV research. In a week and a half, we were able to outdo four years of research," says Kaggle CEO Anthony Goldbloom.
Some of the prizes are pretty unbelievable. Dr. Richard Murkin, the billionaire CEO of Heritage Provider Network, has personally put up $3 million and posted a problem based on data from his company.
The company's revenue currently comes from a platform fee, charged to those who sponsor each study. "Once we scale up, the idea is to take a percentage of each prize. We hope to do tens of thousands of these studies each year," says Kaggle's Chief Data Scientist Jeremy Howard.
Zachary Lichaa Business Insider November 3, 2011
San Francisco-based Kaggle has raised an $11 million Series A round from Index Ventures, Khosla Ventures, SV Angel, and others.
Slide's Max Levchin has also been named its chairman. Kaggle is a platform for solving some of the world's toughest data problems. NASA, Delloite, and The University of Michigan have all turned to Kaggle's pool of 17,000 PhD-level scientists to solve complex problems and create winning models.
Algorithms, which are at the heart of solving these complex data problems, are posted to Kaggle's site continuously. As they're posted, solutions become clearer and eventually the purveyor of that problem has an improved theory.
Kaggle sets up real-time leaderboards for each problem that's being solved and turns them into competitions to foster community engagement. Whoever solves the problem wins a prize.
When Kaggle is working with sensitive material, competitions are held privately using a selected group of scientists that are pre-screened.
Past projects include mapping HIV trends and Dark Matter in outer space. The latter was commissioned by NASA, and its European counterparts, the European Space Agency and The Royal Astronomical Society.
"The first time we could prove Kaggle worked was with HIV research. In a week and a half, we were able to outdo four years of research," says Kaggle CEO Anthony Goldbloom.
Some of the prizes are pretty unbelievable. Dr. Richard Murkin, the billionaire CEO of Heritage Provider Network, has personally put up $3 million and posted a problem based on data from his company.
The company's revenue currently comes from a platform fee, charged to those who sponsor each study. "Once we scale up, the idea is to take a percentage of each prize. We hope to do tens of thousands of these studies each year," says Kaggle's Chief Data Scientist Jeremy Howard.
Tuesday, November 1, 2011
Churchill event: THE BIG DATA EFFECT
An Open Forum Event:
THE BIG DATA EFFECT
December 7, 2011
SPEAKERS:
Keith Collins, SVP and CTO, SASGil Elbaz, Founder and CEO, Factual - New !
Ping Li, Partner, Accel Partners - New !
Luke Lonergan, CTO, VP and Co-Founder, GreenplumAnand Rajaraman, SVP, Walmart Global Ecommerce, Head of @WalmartLabs
Moderator: Michael Chui, Senior Fellow, McKinsey Global Insitute
Time: 5:30 PM Registration | 6:00 PM Buffet | 7:00 PM Program
Location: Computer History Museum, Mountain View
Gartner predicts that data will grow by 800% in five years, with 80% of it unstructured. The World Economic Forum recently declared big data as an asset class. We're only getting started to discover the implications of making better sense of large amounts of unstructured data to uncover business opportunities, strategies, and more. Is it really an emerging market with lots of innovation, startups, job creation? Why is it suddenly possible? What are the obstacles, and the most promising areas of opportunity? How do you make it real in your organization? Join this group of thought leaders from Accel Partners, Factual, Greenplum, SAS, @Walmartlabs, and McKinsey Global Institute for a conversation that gets beyond the hype about big data.
THE BIG DATA EFFECT
December 7, 2011
SPEAKERS:
Keith Collins, SVP and CTO, SASGil Elbaz, Founder and CEO, Factual - New !
Ping Li, Partner, Accel Partners - New !
Luke Lonergan, CTO, VP and Co-Founder, GreenplumAnand Rajaraman, SVP, Walmart Global Ecommerce, Head of @WalmartLabs
Moderator: Michael Chui, Senior Fellow, McKinsey Global Insitute
Time: 5:30 PM Registration | 6:00 PM Buffet | 7:00 PM Program
Location: Computer History Museum, Mountain View
Gartner predicts that data will grow by 800% in five years, with 80% of it unstructured. The World Economic Forum recently declared big data as an asset class. We're only getting started to discover the implications of making better sense of large amounts of unstructured data to uncover business opportunities, strategies, and more. Is it really an emerging market with lots of innovation, startups, job creation? Why is it suddenly possible? What are the obstacles, and the most promising areas of opportunity? How do you make it real in your organization? Join this group of thought leaders from Accel Partners, Factual, Greenplum, SAS, @Walmartlabs, and McKinsey Global Institute for a conversation that gets beyond the hype about big data.
Monday, October 31, 2011
Cloud computing's Silver Lining: Jobs
NEED TO KNOW: TECHONOLOGY
Silver Lining
Yes, a new data-storage technology could cost jobs. But it could add even more.
Sara Jerome National Journal November 21, 2011
Here is an information-technology plan for the era of austerity: cloud computing. This innovation—namely, systems that store data on remote servers operated by host companies rather than on hardware owned by your employer—does away with expensive equipment and the hassle of maintaining it. So it’s no surprise that the Obama White House declared a “cloud-first” policy two years ago. The Office of Management and Budget said that if each department and agency moved just three projects to the cloud, the government would save $5 billion.
The trade-off, however, seemed to be that, in emancipating themselves from hardware, employers emancipate themselves from staff to support it. “[Our client was] able to eliminate a whole bunch of actually U.S.-based jobs and kind of replace them with two folks out of India to serve a 1,200-person engineering organization,” gloated Richard Marcello, an executive at the IT firm Unisys, at the Cloud Computing Conference & Expo in Santa Clara, Calif. A simple story of cutting spending at the expense of jobs, right?
Not so fast. If the story of cloud computing in the United States plays out as its backers promise, it could become one of the most successful recent job-creation trends. Cloud evangelists promise that a profitable new domestic industry will emerge from the ashes of the traditional IT model. “This is going to be a second version of the rise of the Internet. It’s about to explode,” says David LeDuc, senior director for public policy at the Software & Information Industry Association.
It’s true that, in this revolution, some tech professionals—particularly those focused on buying and running hardware—will lose their jobs. But the cloud isn’t decimating IT departments. Recruiting firm Robert Half Technology found that, in 2009, 43 percent of chief information officers said that their departments are either “very” or “somewhat” understaffed. Unemployment for IT professions was just 5 percent in September, far less than the national average, and in a Microsoft study, 54 percent of IT decision makers said they are “hiring as a result of the cloud.” At any rate, according to a report last year by the consultancy McKinsey, savings from cloud computing don’t come from labor. An average business spends some $107 on labor each month for traditional storage, compared with $207 for Amazon’s cloud, the report said. Lower costs come from less hardware, not fewer people.
Meanwhile, new jobs are sprouting up to service the cloud industry—and not just in the developing world. This new market, valued at $40.7 billion worldwide last year, is expected to reach $241 billion by 2020, according to a report this year by Forrester Research. Sure, some tech companies will base their servers overseas in low-cost environments, but the top cloud companies are American; all told, U.S. firms control 60 percent of the market, according to the latest data from 2009. “The massive computing infrastructure in the United States gives us an edge,” LeDuc says. His trade association predicts that other countries will increasingly outsource their data to the United States, where Google and Microsoft keep many of their cloud servers.
The cloud is already putting Americans to work. Google’s team has more than 1,000 employees, Texas cloud company RackSpace employs 3,700 people, and California-based provider Saleforces.com has 235 open positions, according to The Wall Street Journal. U.S. businesses paid almost $22 billion to move to the cloud last year, and that figure is expected to rise to $80 billion by 2015, according to a study by IT consulting firm IDC.
Economists haven’t yet studied how this will all play out in the United States, but the Center for Economics and Business Research, a British think tank, predicts that the cloud market will create 2.4 million jobs over the next four years in Europe, the Middle East, and Asia, with 300,000 alone in the United Kingdom. “Public and private organizations that preserve the status quo of wasteful spending [on IT] will be punished, while those that embrace the cloud will be rewarded with substantial savings and 21st-century jobs,” Vivek Kundra, the former U.S. chief information officer who pushed the government into the cloud, wrote in The New York Times in August.
The biggest hitch could be protectionism. Cloud providers such as Microsoft and Google are already working hard to prevent foreign governments from enacting laws banning “cross-border data flows.” Such laws could force cloud companies to keep servers in the country where information originates rather than in the storage provider’s country of choice. Kundra supports a global cloud policy “that forces nations to work together and resolve” cross-border issues. “The United States, along with leading nations in Europe and Asia, has an opportunity to announce such an initiative at the World Economic Forum meeting in January,” he wrote in The Times.
Many consumers are already familiar with cloud technology (Gmail stores users’ information on Google’s huge servers rather than eating up the finite space on their laptops), and the federal government’s buy-in has signaled that the cloud—once considered too vulnerable to cyberthreats—is now sufficiently secure for most offices. Washington spends $80 billion on IT each year, and tech officials hope to eventually move a quarter of that sum to the cloud. The payoff may take years, meaning that President Obama’s policy won’t affect today’s unemployment rate. But if the cloud sector catches fire as promised, it will add, not subtract, American jobs.
This article appeared in the Saturday, October 29, 2011 edition of National Journal.
Silver Lining
Yes, a new data-storage technology could cost jobs. But it could add even more.
Sara Jerome National Journal November 21, 2011
Here is an information-technology plan for the era of austerity: cloud computing. This innovation—namely, systems that store data on remote servers operated by host companies rather than on hardware owned by your employer—does away with expensive equipment and the hassle of maintaining it. So it’s no surprise that the Obama White House declared a “cloud-first” policy two years ago. The Office of Management and Budget said that if each department and agency moved just three projects to the cloud, the government would save $5 billion.
The trade-off, however, seemed to be that, in emancipating themselves from hardware, employers emancipate themselves from staff to support it. “[Our client was] able to eliminate a whole bunch of actually U.S.-based jobs and kind of replace them with two folks out of India to serve a 1,200-person engineering organization,” gloated Richard Marcello, an executive at the IT firm Unisys, at the Cloud Computing Conference & Expo in Santa Clara, Calif. A simple story of cutting spending at the expense of jobs, right?
Not so fast. If the story of cloud computing in the United States plays out as its backers promise, it could become one of the most successful recent job-creation trends. Cloud evangelists promise that a profitable new domestic industry will emerge from the ashes of the traditional IT model. “This is going to be a second version of the rise of the Internet. It’s about to explode,” says David LeDuc, senior director for public policy at the Software & Information Industry Association.
It’s true that, in this revolution, some tech professionals—particularly those focused on buying and running hardware—will lose their jobs. But the cloud isn’t decimating IT departments. Recruiting firm Robert Half Technology found that, in 2009, 43 percent of chief information officers said that their departments are either “very” or “somewhat” understaffed. Unemployment for IT professions was just 5 percent in September, far less than the national average, and in a Microsoft study, 54 percent of IT decision makers said they are “hiring as a result of the cloud.” At any rate, according to a report last year by the consultancy McKinsey, savings from cloud computing don’t come from labor. An average business spends some $107 on labor each month for traditional storage, compared with $207 for Amazon’s cloud, the report said. Lower costs come from less hardware, not fewer people.
Meanwhile, new jobs are sprouting up to service the cloud industry—and not just in the developing world. This new market, valued at $40.7 billion worldwide last year, is expected to reach $241 billion by 2020, according to a report this year by Forrester Research. Sure, some tech companies will base their servers overseas in low-cost environments, but the top cloud companies are American; all told, U.S. firms control 60 percent of the market, according to the latest data from 2009. “The massive computing infrastructure in the United States gives us an edge,” LeDuc says. His trade association predicts that other countries will increasingly outsource their data to the United States, where Google and Microsoft keep many of their cloud servers.
The cloud is already putting Americans to work. Google’s team has more than 1,000 employees, Texas cloud company RackSpace employs 3,700 people, and California-based provider Saleforces.com has 235 open positions, according to The Wall Street Journal. U.S. businesses paid almost $22 billion to move to the cloud last year, and that figure is expected to rise to $80 billion by 2015, according to a study by IT consulting firm IDC.
Economists haven’t yet studied how this will all play out in the United States, but the Center for Economics and Business Research, a British think tank, predicts that the cloud market will create 2.4 million jobs over the next four years in Europe, the Middle East, and Asia, with 300,000 alone in the United Kingdom. “Public and private organizations that preserve the status quo of wasteful spending [on IT] will be punished, while those that embrace the cloud will be rewarded with substantial savings and 21st-century jobs,” Vivek Kundra, the former U.S. chief information officer who pushed the government into the cloud, wrote in The New York Times in August.
The biggest hitch could be protectionism. Cloud providers such as Microsoft and Google are already working hard to prevent foreign governments from enacting laws banning “cross-border data flows.” Such laws could force cloud companies to keep servers in the country where information originates rather than in the storage provider’s country of choice. Kundra supports a global cloud policy “that forces nations to work together and resolve” cross-border issues. “The United States, along with leading nations in Europe and Asia, has an opportunity to announce such an initiative at the World Economic Forum meeting in January,” he wrote in The Times.
Many consumers are already familiar with cloud technology (Gmail stores users’ information on Google’s huge servers rather than eating up the finite space on their laptops), and the federal government’s buy-in has signaled that the cloud—once considered too vulnerable to cyberthreats—is now sufficiently secure for most offices. Washington spends $80 billion on IT each year, and tech officials hope to eventually move a quarter of that sum to the cloud. The payoff may take years, meaning that President Obama’s policy won’t affect today’s unemployment rate. But if the cloud sector catches fire as promised, it will add, not subtract, American jobs.
This article appeared in the Saturday, October 29, 2011 edition of National Journal.
Subscribe to:
Posts (Atom)