Plenty of smart companies and organizations have put together their own online communities. Some are pulling this off brilliantly; others, not so much.
By John Hagel and John Seely Brown, guest contributors, Fortune
January 18, 2012: 10:27 AM ET
FORTUNE -- As anyone who has started their own blog certainly knows, it takes all of an hour (if that) to create your very own space to share your most brilliant ponderings for the entire world to see. But when it comes to building a space online that people want to visit regularly and contribute to, well, most of us never get there, and for good reason. It's really hard.
Plenty of smart companies and organizations have put together their own online communities. Some are pulling this off brilliantly; others, not so much.
Take AARP, for example. True, the organization's website has a sizable amount of free information on member benefits and its advocacy work. But AARP also offers members the opportunity to participate in an online community where they can connect with others who have common interests. At first, participants share ideas and information with each other. Anyone, member or not, can register to join topic-specific groups such as "retirement planning" and "seniors as entrepreneurs" or start a new group.
As participants' relationships deepen, they begin to move from a conversation to actually working together. For example, single travelers in one AARP group have moved from exchanging tips to planning trips together. At this stage, they are not randomly connecting but collaborating.
The term "online community" is often used loosely and includes any sites that aggregate customers or an audience even though there is very little contribution by the participants or interaction among them. We use the term in a much more specific way, to highlight a potentially very powerful form of ecosystem that businesses can organize and nurture. These communities require extensive and sustained interaction among a growing number of participants to function well.
Why should companies care about these online communities? For one, they can increase the value derived from, and the lifespan of, a company's customer relationships. But more broadly, online communities can become valuable tools to understand how customers use a company's products or services and what kinds of improvements or changes might make sense. Companies pay a lot for focus groups and yet virtual communities represent a kind of 24-7 focus group.
Building an effective virtual community is no simple task. Most importantly, it requires a deep understanding of the unmet needs of potential community members rather than simply approaching it as a marketing opportunity for the company. It is no wonder that so many have tried to create these communities and yet so few have succeeded.
Many companies hire consultants like LiveWorld to build and moderate their communities. EBay (EBAY), for example, started a group a decade ago by bringing together online vendors and encouraging them to exchange tips on running auctions and how to build successful online businesses. It has since grown into a full-blown community where people talk personally about their families, work, and other interests.
One member, a skilled quilt maker, was so moved by how her work and friendships had grown as part of the group that she invited members to mail her cloth to stitch a quilt representing the eBay bond. A total of 114 members sent specially designed squares. With agreement from the group, she then sold the quilt on eBay, with proceeds going to a special children's program at the Canadian Diabetes Association. It was bought by an active member of the group who then lent it to eBay to display.
In another example, members of the eBay community launched a discussion board in response to Hurricane Katrina to give each other advance warning of potentially dangerous weather conditions. Each October, a seller group self-organizes to raise funds for the Susan G. Komen for the Cure. "Hundreds of thousands of members are active in the groups, discussion boards, and forums where this deep engagement happens," says Liveworld CEO Peter Friedman, and "tens of millions participate in the larger eBay network that uses the trusted feedback mechanism through which buyers and sellers rate and review each other."
Campbell Soup Company's (CPB) community started out as a recipe exchange, which was prompted by moderators who shared personal stories about how they cook for their families. The site now covers a broad range of interests around raising families.
As participants in these online communities develop relationships, they can begin to take on projects that often benefit the company that sponsors the community, also referred to as "co-creation." As one example, Lego has increasingly relied on its virtual community not only for ideas on new products but actually for help designing these products.
To design your online communities, as a starting point you need to understand the broader context of how your customers use your services. What are their unmet needs? How might online interaction among customers themselves help to address these unmet needs or opportunities?
Thursday, January 19, 2012
Wednesday, January 18, 2012
WSJ: Health Care Is Next Frontier for Big Data
Ben Rooney, The Wall Street Journal, January 18, 2012
Big Data—the ability to collect, process and interpret massive amounts of information—is one of today's most important technological drivers. While companies see it as a way of detecting weak market signals, one of the biggest potential areas of application for society is health care.
Historically, health care has been delivered by one doctor looking at one patient with only the information the doctor has at that time. But how much better if the doctor had access to information about thousands, or even tens of thousands, of people?
Acquiring medical data has, historically, been problematic. It is wrapped in layers of regulations and stringent safeguards and is expensive to collect.
It is also not representative of the general population, for the problem with health care is that only ill people use it. If you want to know what is going on in the general population ill people aren't terribly useful; they are, after all, ill.
So while the health services can already collect information from laboratories, hospitals and front-line doctors, the nature of the data is problematic, says Shamus Husheer, CEO of Cambridge Temperature Concepts. "It is massively skewed. It is highly biased by the people who go to doctors; the ill, the hypochondriacs and the elderly."
Mr. Husheer's company makes a product designed to help women with fertility problems conceive. Women wear a sensor 24 hours a day which records movement and changes in body temperature (an indication of ovulation), up to 20,000 readings a day. This is valuable medical data.
Because women using his product report anything that might affect their body temperature, "it allows us to build up profiles of illnesses without ever setting foot in a hospital," Mr. Husheer says. "More importantly we can see what those illnesses look like in the general population who are otherwise normally healthy. People who have a cold don't go to the doctor. Where else are you going to get that data?"
Because the recordings continue while the women are asleep, "for the first time, we also have extensive data on what normal sleep looks like," he said.
What this gives researchers is a huge control group. "You can compare sleep patterns from normal people with, say, pain sufferers. If you don't know what normal sleep looks like, how do you tease out the data?"
Getting patients to share data isn't always that easy, says Chris Edmonds, managing director of Replay Digital, a company specializing in medical smartphone apps which has recently built an app targeting asthma sufferers.
In return for sharing information, sufferers get advice on how to control their condition. The more people who use the app, the more anonymized, aggregated data can be built up about asthma sufferers whose condition isn't bad enough to require medical intervention.
But there is an implicit deal that has to be struck. If a sufferer is going to use an app, they have to get something back in return. "Most patients are pretty noncompliant," Mr. Edmonds says. "You can't just collect data from them."
Mr. Edmonds was dismissive of the idea of "gamification"—so that suffers of chronic conditions would get some kind of reward; say, measure your temperature four times in a week and unlock a badge to post on social networks.
"This idea about people wanting to play a game about their condition makes no sense whatsoever," he said. "If I am not going to measure my symptoms because it is good for me and makes my bowel cancer better, am I really going to do it to unlock a badge for Facebook?"
Kaveh Safavi, who leads Accenture's health practice for North America, says big data gives two benefits to clinicians. First is "the ability to see information across time in ways that aren't possible.
"The second is to begin the process of pattern recognition, particularly when you are looking for low frequency events, or things that where the signal is very small and might not be discernible when looking at very small groups," he says. A well known example of this is the ability for Google to track the progress of flu through looking at search terms.
There is a problem. Medical data are subject to very stringent conditions on sharing and storage. It took Mr. Husheer nearly two years to get clearance from the U.S. Food and Drug Administration. Is there a danger that the regulatory framework is going to impede the benefits that can be accrued?
It is a risk to which the U.K. government is alert. As one of the few countries to provide a "cradle to grave" health system, the U.K. has access to some of the most detailed and complete patient information. Some of that patient data are to be made available. "We will consider how best to achieve an appropriate balance between the protection of patient information and the use and sharing of information to improve care," said a U.K. Department of Health spokesperson.
A huge upside of technology has been its democratization, giving ordinary access to information and tools that had previously been the preserve of the few. Industry after industry has seen the creative destruction wreaked upon it as Internet technologies pull down walls.
Health care has, so far, remained relatively unscathed. For how much longer?
--------------------------------------------------
Please consider the environment before printing this e-mail.
Big Data—the ability to collect, process and interpret massive amounts of information—is one of today's most important technological drivers. While companies see it as a way of detecting weak market signals, one of the biggest potential areas of application for society is health care.
Historically, health care has been delivered by one doctor looking at one patient with only the information the doctor has at that time. But how much better if the doctor had access to information about thousands, or even tens of thousands, of people?
Acquiring medical data has, historically, been problematic. It is wrapped in layers of regulations and stringent safeguards and is expensive to collect.
It is also not representative of the general population, for the problem with health care is that only ill people use it. If you want to know what is going on in the general population ill people aren't terribly useful; they are, after all, ill.
So while the health services can already collect information from laboratories, hospitals and front-line doctors, the nature of the data is problematic, says Shamus Husheer, CEO of Cambridge Temperature Concepts. "It is massively skewed. It is highly biased by the people who go to doctors; the ill, the hypochondriacs and the elderly."
Mr. Husheer's company makes a product designed to help women with fertility problems conceive. Women wear a sensor 24 hours a day which records movement and changes in body temperature (an indication of ovulation), up to 20,000 readings a day. This is valuable medical data.
Because women using his product report anything that might affect their body temperature, "it allows us to build up profiles of illnesses without ever setting foot in a hospital," Mr. Husheer says. "More importantly we can see what those illnesses look like in the general population who are otherwise normally healthy. People who have a cold don't go to the doctor. Where else are you going to get that data?"
Because the recordings continue while the women are asleep, "for the first time, we also have extensive data on what normal sleep looks like," he said.
What this gives researchers is a huge control group. "You can compare sleep patterns from normal people with, say, pain sufferers. If you don't know what normal sleep looks like, how do you tease out the data?"
Getting patients to share data isn't always that easy, says Chris Edmonds, managing director of Replay Digital, a company specializing in medical smartphone apps which has recently built an app targeting asthma sufferers.
In return for sharing information, sufferers get advice on how to control their condition. The more people who use the app, the more anonymized, aggregated data can be built up about asthma sufferers whose condition isn't bad enough to require medical intervention.
But there is an implicit deal that has to be struck. If a sufferer is going to use an app, they have to get something back in return. "Most patients are pretty noncompliant," Mr. Edmonds says. "You can't just collect data from them."
Mr. Edmonds was dismissive of the idea of "gamification"—so that suffers of chronic conditions would get some kind of reward; say, measure your temperature four times in a week and unlock a badge to post on social networks.
"This idea about people wanting to play a game about their condition makes no sense whatsoever," he said. "If I am not going to measure my symptoms because it is good for me and makes my bowel cancer better, am I really going to do it to unlock a badge for Facebook?"
Kaveh Safavi, who leads Accenture's health practice for North America, says big data gives two benefits to clinicians. First is "the ability to see information across time in ways that aren't possible.
"The second is to begin the process of pattern recognition, particularly when you are looking for low frequency events, or things that where the signal is very small and might not be discernible when looking at very small groups," he says. A well known example of this is the ability for Google to track the progress of flu through looking at search terms.
There is a problem. Medical data are subject to very stringent conditions on sharing and storage. It took Mr. Husheer nearly two years to get clearance from the U.S. Food and Drug Administration. Is there a danger that the regulatory framework is going to impede the benefits that can be accrued?
It is a risk to which the U.K. government is alert. As one of the few countries to provide a "cradle to grave" health system, the U.K. has access to some of the most detailed and complete patient information. Some of that patient data are to be made available. "We will consider how best to achieve an appropriate balance between the protection of patient information and the use and sharing of information to improve care," said a U.K. Department of Health spokesperson.
A huge upside of technology has been its democratization, giving ordinary access to information and tools that had previously been the preserve of the few. Industry after industry has seen the creative destruction wreaked upon it as Internet technologies pull down walls.
Health care has, so far, remained relatively unscathed. For how much longer?
--------------------------------------------------
Please consider the environment before printing this e-mail.
Tuesday, January 17, 2012
WSJ: Big Firms Try Crowdsourcing
Rachel Emma Silverman, The Wall Street Journal, January 17, 2012
When AOL Inc. set out to determine whether it was getting the best use of its video library last year, the task required an inventory to measure which of the thousands of Web pages it publishes daily contained videos.
Daniel Maloney, the AOL executive in charge of the project, considered building video-detecting software, but determined that it would take too long to meet the project's deadline. He also thought about hiring temps to accomplish the task, but realized that, too, would drag on.
So he turned to another option that is gaining traction among large firms: crowdsourcing.
Crowdsourced labor usually involves breaking a project into tiny component tasks and farming those tasks out to the general public by posting the requests on a website.
Many firms that use crowdsourcing pay pennies per microtask to complete projects such as tagging or verifying data, digitizing handwritten forms and database entry. Dozens of services, such as Amazon.com Inc.'s Mechanical Turk and CrowdFlower, have cropped up in recent years to help companies cheaply crowdsource tasks.
Revenues of business-focused crowdsourcing firms grew 74% between 2010 and 2011, and 53% a year earlier, according to preliminary data from 14 large crowdsourcing firms with total revenues of about $50 million, by Crowdsourcing.org, a research firm which tracks the industry.
Companies that have assigned work to the crowd say it is generally cheaper and faster than hiring temps or traditional outsourcing firms. Crowdsourced labor can cost companies less than half as much as typical outsourcing, says Panagiotis G. Ipeirotis, an associate professor at New York University's Stern School of Business, who studies crowdsourcing.
Some individual microtasks can take just a few seconds and pay a few cents per task. More complex writing or transcription tasks might pay $10 or $20 per job, while some highly skilled work, such as writing programming code, commands higher rates.
For its video project, AOL ended up using Mechanical Turk, a crowdsourcing service that has over 500,000 registered workers from 190 different countries. Amazon says firms such as Microsoft Corp. and LinkedIn Corp. have used its service, for which Amazon charges clients 10% of the total labor cost. The company also collects tax identification information to track "Turkers'" earnings, but hiring companies are responsible for distributing tax forms to independent contractors.
AOL asked crowd workers to determine whether Web pages contained a video and to identify both its source and location on the page. With the help of Statera, a technology-services firm, AOL posted the tasks on Mechanical Turk and made sure that the quality of the many thousands of answers was up to par.
The project was up and running within a week, Mr. Maloney says, and took only a couple of months to complete, far less than it would have taken the company to vet and hire a temp staff or build a software program to do the job. AOL declined to release the exact cost of the effort, but Mr. Maloney says the whole project cost about as much as two temp workers would have been paid in the same time.
"We had a very high number of pages we need to process. Being able to tap into a scaled work force was massively helpful," Mr. Maloney says.
Crowdsourced labor raises some concerns. Workers may sign up for tasks unaware of what their labor may be used for. One 2010 study by Dr. Ipeirotis of NYU estimated that some 40% of microtask requests from new employers who joined Amazon's Turk in a two-month period were actually used to create spam. Amazon disputes the study results, and says it has very extensive anti-spam measures. The company has since been more aggressive in policing spam and email marketers on Turk, according to Dr. Ipeirotis.
For firms, vetting the quality of individual contributors is a tall order—work is done in bits and pieces by a large group. Some crowdsourcing companies say they embed various tests into the tasks to help track workers' accuracy.
InsideView Inc., a San Francisco company that filters data for sales forces, tapped CrowdFlower for an online data-collection project last year. CrowdFlower broke down the tasks into simple steps and to ensure quality, the firm quizzes workers and rates them with computer-generated accuracy scores.
"You have no idea who is doing the work," says Gordon Anderson, InsideView's vice president for content. "It could be a housewife in Iowa or someone in a refugee camp in India. They could be anywhere…. It puts the onus on us to make the project as simple as it can be."
The whole project, which processed about 300,000 company names, took seven weeks and was about 50% cheaper than buying the data from a data provider, another option the company was considering, says Mr. Anderson.
Another concern is that crowdsourced labor risks creating what Harvard Law School professor Jonathan Zittrain calls "digital sweatshops," where workers who may be underage work long hours on mind-numbing tasks for very little money—or, if the work is structured as a game, for no money at all. Crowdsourcing sites often pitch their work to stay-at-home moms or students who can pick up a few tasks to do during short breaks.
A recent survey of 1,046 active U.S. based Mechanical Turk workers found that about 25% of respondents said that their crowdsourcing work comprised more than 10% of their annual income, according to research conducted by CrowdControl, which manages crowdsourcing projects. More than half of these workers are women, and most are between age 26 and 35, the survey found.
Christopher Berry, a special-education teacher with four children, began doing jobs on Mechanical Turk a little over a year ago, to earn extra income to support his family. He sometimes does up to 1,000 tasks a day, some paying as little as two cents a pop, ranging from tagging Twitter posts to copy editing. Last year, he earned almost $10,000 from his Turk work. "You kind of get hooked on it," says Mr. Berry, 39, of Roseville, Calif.
Projects outsourced to the crowd can be used to help support workers with few other options. Crowdsourcing firm Samasource, a San Francisco nonprofit, specializes in assigning microtasks, such as verifying business addresses, to workers in the developing world. Samasource says it promises to pay workers a living wage.
Write to Rachel Emma Silverman at rachel.silverman@wsj.com
When AOL Inc. set out to determine whether it was getting the best use of its video library last year, the task required an inventory to measure which of the thousands of Web pages it publishes daily contained videos.
Daniel Maloney, the AOL executive in charge of the project, considered building video-detecting software, but determined that it would take too long to meet the project's deadline. He also thought about hiring temps to accomplish the task, but realized that, too, would drag on.
So he turned to another option that is gaining traction among large firms: crowdsourcing.
Crowdsourced labor usually involves breaking a project into tiny component tasks and farming those tasks out to the general public by posting the requests on a website.
Many firms that use crowdsourcing pay pennies per microtask to complete projects such as tagging or verifying data, digitizing handwritten forms and database entry. Dozens of services, such as Amazon.com Inc.'s Mechanical Turk and CrowdFlower, have cropped up in recent years to help companies cheaply crowdsource tasks.
Revenues of business-focused crowdsourcing firms grew 74% between 2010 and 2011, and 53% a year earlier, according to preliminary data from 14 large crowdsourcing firms with total revenues of about $50 million, by Crowdsourcing.org, a research firm which tracks the industry.
Companies that have assigned work to the crowd say it is generally cheaper and faster than hiring temps or traditional outsourcing firms. Crowdsourced labor can cost companies less than half as much as typical outsourcing, says Panagiotis G. Ipeirotis, an associate professor at New York University's Stern School of Business, who studies crowdsourcing.
Some individual microtasks can take just a few seconds and pay a few cents per task. More complex writing or transcription tasks might pay $10 or $20 per job, while some highly skilled work, such as writing programming code, commands higher rates.
For its video project, AOL ended up using Mechanical Turk, a crowdsourcing service that has over 500,000 registered workers from 190 different countries. Amazon says firms such as Microsoft Corp. and LinkedIn Corp. have used its service, for which Amazon charges clients 10% of the total labor cost. The company also collects tax identification information to track "Turkers'" earnings, but hiring companies are responsible for distributing tax forms to independent contractors.
AOL asked crowd workers to determine whether Web pages contained a video and to identify both its source and location on the page. With the help of Statera, a technology-services firm, AOL posted the tasks on Mechanical Turk and made sure that the quality of the many thousands of answers was up to par.
The project was up and running within a week, Mr. Maloney says, and took only a couple of months to complete, far less than it would have taken the company to vet and hire a temp staff or build a software program to do the job. AOL declined to release the exact cost of the effort, but Mr. Maloney says the whole project cost about as much as two temp workers would have been paid in the same time.
"We had a very high number of pages we need to process. Being able to tap into a scaled work force was massively helpful," Mr. Maloney says.
Crowdsourced labor raises some concerns. Workers may sign up for tasks unaware of what their labor may be used for. One 2010 study by Dr. Ipeirotis of NYU estimated that some 40% of microtask requests from new employers who joined Amazon's Turk in a two-month period were actually used to create spam. Amazon disputes the study results, and says it has very extensive anti-spam measures. The company has since been more aggressive in policing spam and email marketers on Turk, according to Dr. Ipeirotis.
For firms, vetting the quality of individual contributors is a tall order—work is done in bits and pieces by a large group. Some crowdsourcing companies say they embed various tests into the tasks to help track workers' accuracy.
InsideView Inc., a San Francisco company that filters data for sales forces, tapped CrowdFlower for an online data-collection project last year. CrowdFlower broke down the tasks into simple steps and to ensure quality, the firm quizzes workers and rates them with computer-generated accuracy scores.
"You have no idea who is doing the work," says Gordon Anderson, InsideView's vice president for content. "It could be a housewife in Iowa or someone in a refugee camp in India. They could be anywhere…. It puts the onus on us to make the project as simple as it can be."
The whole project, which processed about 300,000 company names, took seven weeks and was about 50% cheaper than buying the data from a data provider, another option the company was considering, says Mr. Anderson.
Another concern is that crowdsourced labor risks creating what Harvard Law School professor Jonathan Zittrain calls "digital sweatshops," where workers who may be underage work long hours on mind-numbing tasks for very little money—or, if the work is structured as a game, for no money at all. Crowdsourcing sites often pitch their work to stay-at-home moms or students who can pick up a few tasks to do during short breaks.
A recent survey of 1,046 active U.S. based Mechanical Turk workers found that about 25% of respondents said that their crowdsourcing work comprised more than 10% of their annual income, according to research conducted by CrowdControl, which manages crowdsourcing projects. More than half of these workers are women, and most are between age 26 and 35, the survey found.
Christopher Berry, a special-education teacher with four children, began doing jobs on Mechanical Turk a little over a year ago, to earn extra income to support his family. He sometimes does up to 1,000 tasks a day, some paying as little as two cents a pop, ranging from tagging Twitter posts to copy editing. Last year, he earned almost $10,000 from his Turk work. "You kind of get hooked on it," says Mr. Berry, 39, of Roseville, Calif.
Projects outsourced to the crowd can be used to help support workers with few other options. Crowdsourcing firm Samasource, a San Francisco nonprofit, specializes in assigning microtasks, such as verifying business addresses, to workers in the developing world. Samasource says it promises to pay workers a living wage.
Write to Rachel Emma Silverman at rachel.silverman@wsj.com
Cracking Open the Scientific Process
Thomas Lin, The New York Times, January 16, 2012
The New England Journal of Medicine marks its 200th anniversary this year with a timeline celebrating the scientific advances first described in its pages: the stethoscope (1816), the use of ether for anesthesia (1846), and disinfecting hands and instruments before surgery (1867), among others.
For centuries, this is how science has operated — through research done in private, then submitted to science and medical journals to be reviewed by peers and published for the benefit of other researchers and the public at large. But to many scientists, the longevity of that process is nothing to celebrate.
The system is hidebound, expensive and elitist, they say. Peer review can take months, journal subscriptions can be prohibitively costly, and a handful of gatekeepers limit the flow of information. It is an ideal system for sharing knowledge, said the quantum physicist Michael Nielsen, only “if you’re stuck with 17th-century technology.”
Dr. Nielsen and other advocates for “open science” say science can accomplish much more, much faster, in an environment of friction-free collaboration over the Internet.
And despite a host of obstacles, including the skepticism of many established scientists, their ideas are gaining traction.
Open-access archives and journals like arXiv and the Public Library of Science (PLoS) have sprung up in recent years. GalaxyZoo, a citizen-science site, has classified millions of objects in space, discovering characteristics that have led to a raft of scientific papers.
On the collaborative blog MathOverflow, mathematicians earn reputation points for contributing to solutions; in another math experiment dubbed the Polymath Project, mathematicians commenting on the Fields medalist Timothy Gower’s blog in 2009 found a new proof for a particularly complicated theorem in just six weeks.
And a social networking site called ResearchGate — where scientists can answer one another’s questions, share papers and find collaborators — is rapidly gaining popularity.
Editors of traditional journals say open science sounds good, in theory. In practice, “the scientific community itself is quite conservative,” said Maxine Clarke, executive editor of the commercial journal Nature, who added that the traditional published paper is still viewed as “a unit to award grants or assess jobs and tenure.”
Dr. Nielsen, 38, who left a successful science career to write “Reinventing Discovery: The New Era of Networked Science,” agreed that scientists have been “very inhibited and slow to adopt a lot of online tools.” But he added that open science was coalescing into “a bit of a movement.”
On Thursday, 450 bloggers, journalists, students, scientists, librarians and programmers will converge on North Carolina State University (and thousands more will join in online) for the sixth annual ScienceOnline conference. Science is moving to a collaborative model, said Bora Zivkovic, a chronobiology blogger who is a founder of the conference, “because it works better in the current ecosystem, in the Web-connected world.”
Indeed, he said, scientists who attend the conference should not be seen as competing with one another. “Lindsay Lohan is our competitor,” he continued. “We have to get her off the screen and get science there instead.”
Facebook for Scientists?
“I want to make science more open. I want to change this,” said Ijad Madisch, 31, the Harvard-trained virologist and computer scientist behind ResearchGate, the social networking site for scientists.
Started in 2008 with few features, it was reshaped with feedback from scientists. Its membership has mushroomed to more than 1.3 million, Dr. Madisch said, and it has attracted several million dollars in venture capital from some of the original investors of Twitter, eBay and Facebook.
A year ago, ResearchGate had 12 employees. Now it has 70 and is hiring. The company, based in Berlin, is modeled after Silicon Valley startups. Lunch, drinks and fruit are free, and every employee owns part of the company.
The Web site is a sort of mash-up of Facebook, Twitter and LinkedIn, with profile pages, comments, groups, job listings, and “like” and “follow” buttons (but without baby photos, cat videos and thinly veiled self-praise). Only scientists are invited to pose and answer questions — a rule that should not be hard to enforce, with discussion threads about topics like polymerase chain reactions that only a scientist could love.
Scientists populate their ResearchGate profiles with their real names, professional details and publications — data that the site uses to suggest connections with other members. Users can create public or private discussion groups, and share papers and lecture materials. ResearchGate is also developing a “reputation score” to reward members for online contributions.
ResearchGate offers a simple yet effective end run around restrictive journal access with its “self-archiving repository.” Since most journals allow scientists to link to their submitted papers on their own Web sites, Dr. Madisch encourages his users to do so on their ResearchGate profiles. In addition to housing 350,000 papers (and counting), the platform provides a way to search 40 million abstracts and papers from other science databases.
In 2011, ResearchGate reports, 1,620,849 connections were made, 12,342 questions answered and 842,179 publications shared. Greg Phelan, chairman of the chemistry department at the State University of New York, Cortland, used it to find new collaborators, get expert advice and read journal articles not available through his small university. Now he spends up to two hours a day, five days a week, on the site.
Dr. Rajiv Gupta, a radiology instructor who supervised Dr. Madisch at Harvard and was one of ResearchGate’s first investors, called it “a great site for serious research and research collaboration,” adding that he hoped it would never be contaminated “with pop culture and chit-chat.”
Dr. Gupta called Dr. Madisch the “quintessential networking guy — if there’s a Bill Clinton of the science world, it would be him.”
The Paper Trade
Dr. Sönke H. Bartling, a researcher at the German Cancer Research Center who is editing a book on “Science 2.0,” wrote that for scientists to move away from what is currently “a highly integrated and controlled process,” a new system for assessing the value of research is needed. If open access is to be achieved through blogs, what good is it, he asked, “if one does not get reputation and money from them?”
Changing the status quo — opening data, papers, research ideas and partial solutions to anyone and everyone — is still far more idea than reality. As the established journals argue, they provide a critical service that does not come cheap.
“I would love for it to be free,” said Alan Leshner, executive publisher of the journal Science, but “we have to cover the costs.” Those costs hover around $40 million a year to produce his nonprofit flagship journal, with its more than 25 editors and writers, sales and production staff, and offices in North America, Europe and Asia, not to mention print and distribution expenses. (Like other media organizations, Science has responded to the decline in advertising revenue by enhancing its Web offerings, and most of its growth comes from online subscriptions.)
Similarly, Nature employs a large editorial staff to manage the peer-review process and to select and polish “startling and new” papers for publication, said Dr. Clarke, its editor. And it costs money to screen for plagiarism and spot-check data “to make sure they haven’t been manipulated.”
Peer-reviewed open-access journals, like Nature Communications and PLoS One, charge their authors publication fees — $5,000 and $1,350, respectively — to defray their more modest expenses.
The largest journal publisher, Elsevier, whose products include The Lancet, Cell and the subscription-based online archive ScienceDirect, has drawn considerable criticism from open-access advocates and librarians, who are especially incensed by its support for the Research Works Act, introduced in Congress last month, which seeks to protect publishers’ rights by effectively restricting access to research papers and data.
In an Op-Ed article in The New York Times last week, Michael B. Eisen, a molecular biologist at the University of California, Berkeley, and a founder of the Public Library of Science, wrote that if the bill passes, “taxpayers who already paid for the research would have to pay again to read the results.”
In an e-mail interview, Alicia Wise, director of universal access at Elsevier, wrote that “professional curation and preservation of data is, like professional publishing, neither easy nor inexpensive.” And Tom Reller, a spokesman for Elsevier, commented on Dr. Eisen’s blog, “Government mandates that require private-sector information products to be made freely available undermine the industry’s ability to recoup these investments.”
Mr. Zivkovic, the ScienceOnline co-founder and a blog editor for Scientific American, which is owned by Nature, was somewhat sympathetic to the big journals’ plight. “They have shareholders,” he said. “They have to move the ship slowly.”
Still, he added: “Nature is not digging in. They know it’s happening. They’re preparing for it.”
Science 2.0
Scott Aaronson, a quantum computing theorist at the Massachusetts Institute of Technology, has refused to conduct peer review for or submit papers to commercial journals. “I got tired of giving free labor,” he said, to “these very rich for-profit companies.”
Dr. Aaronson is also an active member of online science communities like MathOverflow, where he has earned enough reputation points to edit others’ posts. “We’re not talking about new technologies that have to be invented,” he said. “Things are moving in that direction. Journals seem noticeably less important than 10 years ago.”
Dr. Leshner, the publisher of Science, agrees that things are moving. “Will the model of science magazines be the same 10 years from now? I highly doubt it,” he said. “I believe in evolution.
“When a better system comes into being that has quality and trustability, it will happen. That’s how science progresses, by doing scientific experiments. We should be doing that with scientific publishing as well.”
Matt Cohler, the former vice president of product management at Facebook who now represents Benchmark Capital on ResearchGate’s board, sees a vast untapped market in online science.
“It’s one of the last areas on the Internet where there really isn’t anything yet that addresses core needs for this group of people,” he said, adding that “trillions” are spent each year on global scientific research. Investors are betting that a successful site catering to scientists could shave at least a sliver off that enormous pie.
Dr. Madisch, of ResearchGate, acknowledged that he might never reach many of the established scientists for whom social networking can seem like a foreign language or a waste of time. But wait, he said, until younger scientists weaned on social media and open-source collaboration start running their own labs.
“If you said years ago, ‘One day you will be on Facebook sharing all your photos and personal information with people,’ they wouldn’t believe you,” he said. “We’re just at the beginning. The change is coming.”
The New England Journal of Medicine marks its 200th anniversary this year with a timeline celebrating the scientific advances first described in its pages: the stethoscope (1816), the use of ether for anesthesia (1846), and disinfecting hands and instruments before surgery (1867), among others.
For centuries, this is how science has operated — through research done in private, then submitted to science and medical journals to be reviewed by peers and published for the benefit of other researchers and the public at large. But to many scientists, the longevity of that process is nothing to celebrate.
The system is hidebound, expensive and elitist, they say. Peer review can take months, journal subscriptions can be prohibitively costly, and a handful of gatekeepers limit the flow of information. It is an ideal system for sharing knowledge, said the quantum physicist Michael Nielsen, only “if you’re stuck with 17th-century technology.”
Dr. Nielsen and other advocates for “open science” say science can accomplish much more, much faster, in an environment of friction-free collaboration over the Internet.
And despite a host of obstacles, including the skepticism of many established scientists, their ideas are gaining traction.
Open-access archives and journals like arXiv and the Public Library of Science (PLoS) have sprung up in recent years. GalaxyZoo, a citizen-science site, has classified millions of objects in space, discovering characteristics that have led to a raft of scientific papers.
On the collaborative blog MathOverflow, mathematicians earn reputation points for contributing to solutions; in another math experiment dubbed the Polymath Project, mathematicians commenting on the Fields medalist Timothy Gower’s blog in 2009 found a new proof for a particularly complicated theorem in just six weeks.
And a social networking site called ResearchGate — where scientists can answer one another’s questions, share papers and find collaborators — is rapidly gaining popularity.
Editors of traditional journals say open science sounds good, in theory. In practice, “the scientific community itself is quite conservative,” said Maxine Clarke, executive editor of the commercial journal Nature, who added that the traditional published paper is still viewed as “a unit to award grants or assess jobs and tenure.”
Dr. Nielsen, 38, who left a successful science career to write “Reinventing Discovery: The New Era of Networked Science,” agreed that scientists have been “very inhibited and slow to adopt a lot of online tools.” But he added that open science was coalescing into “a bit of a movement.”
On Thursday, 450 bloggers, journalists, students, scientists, librarians and programmers will converge on North Carolina State University (and thousands more will join in online) for the sixth annual ScienceOnline conference. Science is moving to a collaborative model, said Bora Zivkovic, a chronobiology blogger who is a founder of the conference, “because it works better in the current ecosystem, in the Web-connected world.”
Indeed, he said, scientists who attend the conference should not be seen as competing with one another. “Lindsay Lohan is our competitor,” he continued. “We have to get her off the screen and get science there instead.”
Facebook for Scientists?
“I want to make science more open. I want to change this,” said Ijad Madisch, 31, the Harvard-trained virologist and computer scientist behind ResearchGate, the social networking site for scientists.
Started in 2008 with few features, it was reshaped with feedback from scientists. Its membership has mushroomed to more than 1.3 million, Dr. Madisch said, and it has attracted several million dollars in venture capital from some of the original investors of Twitter, eBay and Facebook.
A year ago, ResearchGate had 12 employees. Now it has 70 and is hiring. The company, based in Berlin, is modeled after Silicon Valley startups. Lunch, drinks and fruit are free, and every employee owns part of the company.
The Web site is a sort of mash-up of Facebook, Twitter and LinkedIn, with profile pages, comments, groups, job listings, and “like” and “follow” buttons (but without baby photos, cat videos and thinly veiled self-praise). Only scientists are invited to pose and answer questions — a rule that should not be hard to enforce, with discussion threads about topics like polymerase chain reactions that only a scientist could love.
Scientists populate their ResearchGate profiles with their real names, professional details and publications — data that the site uses to suggest connections with other members. Users can create public or private discussion groups, and share papers and lecture materials. ResearchGate is also developing a “reputation score” to reward members for online contributions.
ResearchGate offers a simple yet effective end run around restrictive journal access with its “self-archiving repository.” Since most journals allow scientists to link to their submitted papers on their own Web sites, Dr. Madisch encourages his users to do so on their ResearchGate profiles. In addition to housing 350,000 papers (and counting), the platform provides a way to search 40 million abstracts and papers from other science databases.
In 2011, ResearchGate reports, 1,620,849 connections were made, 12,342 questions answered and 842,179 publications shared. Greg Phelan, chairman of the chemistry department at the State University of New York, Cortland, used it to find new collaborators, get expert advice and read journal articles not available through his small university. Now he spends up to two hours a day, five days a week, on the site.
Dr. Rajiv Gupta, a radiology instructor who supervised Dr. Madisch at Harvard and was one of ResearchGate’s first investors, called it “a great site for serious research and research collaboration,” adding that he hoped it would never be contaminated “with pop culture and chit-chat.”
Dr. Gupta called Dr. Madisch the “quintessential networking guy — if there’s a Bill Clinton of the science world, it would be him.”
The Paper Trade
Dr. Sönke H. Bartling, a researcher at the German Cancer Research Center who is editing a book on “Science 2.0,” wrote that for scientists to move away from what is currently “a highly integrated and controlled process,” a new system for assessing the value of research is needed. If open access is to be achieved through blogs, what good is it, he asked, “if one does not get reputation and money from them?”
Changing the status quo — opening data, papers, research ideas and partial solutions to anyone and everyone — is still far more idea than reality. As the established journals argue, they provide a critical service that does not come cheap.
“I would love for it to be free,” said Alan Leshner, executive publisher of the journal Science, but “we have to cover the costs.” Those costs hover around $40 million a year to produce his nonprofit flagship journal, with its more than 25 editors and writers, sales and production staff, and offices in North America, Europe and Asia, not to mention print and distribution expenses. (Like other media organizations, Science has responded to the decline in advertising revenue by enhancing its Web offerings, and most of its growth comes from online subscriptions.)
Similarly, Nature employs a large editorial staff to manage the peer-review process and to select and polish “startling and new” papers for publication, said Dr. Clarke, its editor. And it costs money to screen for plagiarism and spot-check data “to make sure they haven’t been manipulated.”
Peer-reviewed open-access journals, like Nature Communications and PLoS One, charge their authors publication fees — $5,000 and $1,350, respectively — to defray their more modest expenses.
The largest journal publisher, Elsevier, whose products include The Lancet, Cell and the subscription-based online archive ScienceDirect, has drawn considerable criticism from open-access advocates and librarians, who are especially incensed by its support for the Research Works Act, introduced in Congress last month, which seeks to protect publishers’ rights by effectively restricting access to research papers and data.
In an Op-Ed article in The New York Times last week, Michael B. Eisen, a molecular biologist at the University of California, Berkeley, and a founder of the Public Library of Science, wrote that if the bill passes, “taxpayers who already paid for the research would have to pay again to read the results.”
In an e-mail interview, Alicia Wise, director of universal access at Elsevier, wrote that “professional curation and preservation of data is, like professional publishing, neither easy nor inexpensive.” And Tom Reller, a spokesman for Elsevier, commented on Dr. Eisen’s blog, “Government mandates that require private-sector information products to be made freely available undermine the industry’s ability to recoup these investments.”
Mr. Zivkovic, the ScienceOnline co-founder and a blog editor for Scientific American, which is owned by Nature, was somewhat sympathetic to the big journals’ plight. “They have shareholders,” he said. “They have to move the ship slowly.”
Still, he added: “Nature is not digging in. They know it’s happening. They’re preparing for it.”
Science 2.0
Scott Aaronson, a quantum computing theorist at the Massachusetts Institute of Technology, has refused to conduct peer review for or submit papers to commercial journals. “I got tired of giving free labor,” he said, to “these very rich for-profit companies.”
Dr. Aaronson is also an active member of online science communities like MathOverflow, where he has earned enough reputation points to edit others’ posts. “We’re not talking about new technologies that have to be invented,” he said. “Things are moving in that direction. Journals seem noticeably less important than 10 years ago.”
Dr. Leshner, the publisher of Science, agrees that things are moving. “Will the model of science magazines be the same 10 years from now? I highly doubt it,” he said. “I believe in evolution.
“When a better system comes into being that has quality and trustability, it will happen. That’s how science progresses, by doing scientific experiments. We should be doing that with scientific publishing as well.”
Matt Cohler, the former vice president of product management at Facebook who now represents Benchmark Capital on ResearchGate’s board, sees a vast untapped market in online science.
“It’s one of the last areas on the Internet where there really isn’t anything yet that addresses core needs for this group of people,” he said, adding that “trillions” are spent each year on global scientific research. Investors are betting that a successful site catering to scientists could shave at least a sliver off that enormous pie.
Dr. Madisch, of ResearchGate, acknowledged that he might never reach many of the established scientists for whom social networking can seem like a foreign language or a waste of time. But wait, he said, until younger scientists weaned on social media and open-source collaboration start running their own labs.
“If you said years ago, ‘One day you will be on Facebook sharing all your photos and personal information with people,’ they wouldn’t believe you,” he said. “We’re just at the beginning. The change is coming.”
Big Data Needs Data Scientists, Or Quants, Or Excel Jockeys
Tom Groenfeldt, Forbes, January 17, 2012
To get the greatest business value from Big Data, companies are looking for multi-skilled experts who understand programming, large-scale mathematics, statistics and business. They call this new role a Data Scientist. Like Big Data, the term Data Scientist is both catching on and generating skepticism.
Also like Big Data, the Data Scientist role is required right now in some businesses and not so immediately in others. Commercial demand is apt to rise sharply within the next few years as social media, sensors, and other new data generators make petabytes a routine part of business analytics. Universities are busy creating Data Scientist programs while existing data-intensive firms are creating their own, even if they don’t use the label, by training their current analysts, quants, Excel jockeys and computer-savvy MBAs in some Big Data skills like Hadoop.
Keith Collins, vice president and chief technology office at SAS, told an audience at the Churchill Club in Silicon Valley that his firm sees a huge gap in the number of Data Scientists available.
In a May report on Big Data, the McKinsey Global Institute put some numbers on the demand: “By 2018, the United States alone could face a shortage of 140,000 to 190,000 people with deep analytical skills as well as 1.5 million managers and analysts with the know-how to use the analysis of big data to make effective decisions.”
Both Big Data believers and skeptics think large firms will need teams of three to ten analysts to bring together the expertise needed to understand their data, think up innovative ways to use it, select the hardware and software and write the programs to get the information business users need.
Luke Lonergan, CTO and co-founder at Greenplum, now part of EMC, said people underestimate the degree to which Big Data can fundamentally change their business.
“They are going to miss the opportunity or get overwhelmed. Those with data science teams begin to understand; others don’t see how much it can do.”
Some companies think they can buy the application to make Big Data happen, but and it doesn’t work that way, he added. “There should be an explosion of new activity around an explosion of big data.”
Randy Lea at Teradata’s Aster Center of Innovation, defines a Data Scientist as a person with mathematical and statistical skills, an investigative mind, an understanding of computer languages like C++ and Java and ability to write code. They are hard to find coming out of universities, he admits, but people with a combination of those skills have been around for 20 years and are called business analysts. They can learn a parallel programming framework like MapReduce, but he thinks much of the analytics runs just fine on SQL language databases; Teradata Aster has more than 25 clients using more than a petabyte of data. Aster can do SQL and SQL MapReduce codes and data mining using R in SAS.
“So we allow you to pick the best of both.”
No need to wait for universities to churn out Data Scientists in his opinion.
“The experienced quants have been using data mining tools. This is a great opportunity to educate them on some of the new tools and on MapReduce. They already have all of the background.”
Steve Hillion chief product officer at Alpine Data Labs, which promises the fastest path to insight from Big Data, said a few companies and government agencies have been using Big Data for years — Coca-Cola, Proctor and Gamble, telcos and intelligence services for example.
While pedigreed Data Scientists may be rare, people with most of the skills are working in companies today.
“At Alpine we are finding that you can take existing analysts such as a smart MBA or an Excel jockey and give them access to the tools of Big Data analytics. A lot of the tools, including our own, make this more accessible. You can get them trained and have them start providing insights into your business based on data you are already collecting.”
Data has become bigger, more pervasive and people are more aware of it, he said. Aggregation and sampling allow users to reduce the size of their data and run it in memory.
However, keeping the data in its raw form can provide insights into long tails, allowing firms to identify and target relatively small cohorts that fall outside the normal distribution.
To get the greatest business value from Big Data, companies are looking for multi-skilled experts who understand programming, large-scale mathematics, statistics and business. They call this new role a Data Scientist. Like Big Data, the term Data Scientist is both catching on and generating skepticism.
Also like Big Data, the Data Scientist role is required right now in some businesses and not so immediately in others. Commercial demand is apt to rise sharply within the next few years as social media, sensors, and other new data generators make petabytes a routine part of business analytics. Universities are busy creating Data Scientist programs while existing data-intensive firms are creating their own, even if they don’t use the label, by training their current analysts, quants, Excel jockeys and computer-savvy MBAs in some Big Data skills like Hadoop.
Keith Collins, vice president and chief technology office at SAS, told an audience at the Churchill Club in Silicon Valley that his firm sees a huge gap in the number of Data Scientists available.
In a May report on Big Data, the McKinsey Global Institute put some numbers on the demand: “By 2018, the United States alone could face a shortage of 140,000 to 190,000 people with deep analytical skills as well as 1.5 million managers and analysts with the know-how to use the analysis of big data to make effective decisions.”
Both Big Data believers and skeptics think large firms will need teams of three to ten analysts to bring together the expertise needed to understand their data, think up innovative ways to use it, select the hardware and software and write the programs to get the information business users need.
Luke Lonergan, CTO and co-founder at Greenplum, now part of EMC, said people underestimate the degree to which Big Data can fundamentally change their business.
“They are going to miss the opportunity or get overwhelmed. Those with data science teams begin to understand; others don’t see how much it can do.”
Some companies think they can buy the application to make Big Data happen, but and it doesn’t work that way, he added. “There should be an explosion of new activity around an explosion of big data.”
Randy Lea at Teradata’s Aster Center of Innovation, defines a Data Scientist as a person with mathematical and statistical skills, an investigative mind, an understanding of computer languages like C++ and Java and ability to write code. They are hard to find coming out of universities, he admits, but people with a combination of those skills have been around for 20 years and are called business analysts. They can learn a parallel programming framework like MapReduce, but he thinks much of the analytics runs just fine on SQL language databases; Teradata Aster has more than 25 clients using more than a petabyte of data. Aster can do SQL and SQL MapReduce codes and data mining using R in SAS.
“So we allow you to pick the best of both.”
No need to wait for universities to churn out Data Scientists in his opinion.
“The experienced quants have been using data mining tools. This is a great opportunity to educate them on some of the new tools and on MapReduce. They already have all of the background.”
Steve Hillion chief product officer at Alpine Data Labs, which promises the fastest path to insight from Big Data, said a few companies and government agencies have been using Big Data for years — Coca-Cola, Proctor and Gamble, telcos and intelligence services for example.
While pedigreed Data Scientists may be rare, people with most of the skills are working in companies today.
“At Alpine we are finding that you can take existing analysts such as a smart MBA or an Excel jockey and give them access to the tools of Big Data analytics. A lot of the tools, including our own, make this more accessible. You can get them trained and have them start providing insights into your business based on data you are already collecting.”
Data has become bigger, more pervasive and people are more aware of it, he said. Aggregation and sampling allow users to reduce the size of their data and run it in memory.
However, keeping the data in its raw form can provide insights into long tails, allowing firms to identify and target relatively small cohorts that fall outside the normal distribution.
Zions Bank, headquartered in Salt Lake City, Utah, used Alpine for a study of its customers. Alpine did some clustering and segmenting but noticed some anomalous views at the high end. The data showed a small cluster of high end customers who owned a small business and generated a lot of value. The individuals within the group showed similar behavior, but the group itself was so small that the individuals wouldn’t have appeared in an aggregated view.
But once they were identified the bank could target them with small business products.
Big Data has generated some hype, concluded Hillion.
“But when you get down to the everyday work of data scientists and analysts, in a very quiet average work-a-day way, they are finding a lot of insights that are becoming increasingly critical to the way companies are doing their business.”
Friday, January 13, 2012
Big Data: The Top 10 Data-Mining Links of 2011
Jonathan Stray, PSB.org, January 10, 2012
Overview is a project to create an open-source document-mining system for investigative journalists and other curious people. We've written before about the goals of the project, and we're developing some new technology, but mostly we're stealing it from other fields.
The following are some of the best ideas we saw in 2011, the data-mining work that we found most inspirational. Many of these links are educational resources for learning about specific technology. Some of this work illuminates how algorithms and humans treat information differently. Other are just amazing, mind-bending work.
1. What do your connections say about you? A lot. It is possible to accurately predict your political orientation solely on the basis of your network on Twitter. You can also work out gender and other things from public information.
2. Free textbooks from Stanford University. "Introduction to Information Retrieval" teaches you how a search engine works, in great detail. "Mining Massive Data Sets" covers a variety of big-data principles that apply to different types of information.
3. We're not above having a list of lists. Here's the Data Mining Blog's top 5 articles. Most of these are foundational, covering basic philosophy and technique such as choosing variables, finding clusters, and deciding what you're looking for.
4. The MINE technique looks for patterns between hundreds or thousands of variables -- say, patterns of gene expression inside a single cell. It's very general, and finds not only individual relationships but networks of cause and effect. Here's a nifty video, here's the original paper, and here's one statistician's review.
5. This is one of those papers that really changed the way I look at things. How do we know when a data visualization shows us something that is "actually there," as opposed to an artifact of the numbers? "Graphical Inference for Infovis" provides one excellent answer, based on a clever analogy with numerical statistics.
6. Lots of text-mining work uses "clustering" or "classification" techniques to sort documents into topics. But doesn't a categorization algorithm impose its own preconceptions? This is a deep issue, which you might think of as "framing" in code. To explore this question Justin Grimmer and Gary King went meta with a system that visualizes all possible categorizations of a document set, and how they relate.
7. A few years ago Google showed that the number of searches for "flu" was a great predictor of the actual number of outbreaks in a given location -- faster and more specific than the Center for Disease Control's own surveillance data. The team has now expanded the technique into Google Correlate, which instantly scans through petabytes of data to find search terms which follow any user-supplied time series.
Here's New Scientist taking it for a test drive.
8. Not content with free professional textbooks, Stanford has created two free online courses for machine learning and natural language processing. Both are live-streamed lecture series taught by experts, with homework. Learning these intricate technologies has never been easier.
9. Lots of people have speculated about the role of social media in protest movements. A team of researchers looked at the data, analyzing a huge set of tweets from the "May 20" protests in Spain last year. How do protests spread from social media? Now we have at least one solid answer.
10. And the craziest data-mining link we ran across in 2011: IBM's DeepQA project, which beat human Jeopardy champions. This project looks into an unstructured database to correctly answer about 80% of all general questions posed to it, in just a few seconds. Here's a TED talk, and here's the technical paper that explains how it works. I can't tell you how badly I want one of these in the newsroom. If enough journalist hackers build on each other's work, maybe one day ...
Happy data mining! We'll be releasing our own prototype document-mining system, and the source, at the NICAR conference next month. If these are the sorts of algorithms you like to play with, we're also hiring programmers who want to bring these sorts of advanced techniques within everyone's reach.
Tuesday, January 10, 2012
Big Data: Machine to read individual's DNA for $1,000
Clive Cookson, The Financial Times, January 10, 2012
A US biotechnology company will on Tuesday announce the first machine that can read all 3bn letters of an individual’s DNA for as little as $1,000 – a development that will greatly accelerate medical treatment tailored to a patient’s genes but also raises ethical questions.
Life Technologies says its new Ion Proton sequencer – a $149,000 instrument about the size of a laser printer – can read a whole human genome in less than a day for $1,000 including all chemicals, running costs and preliminary data analysis.
The landmark development, expected to be matched by other companies soon, will greatly increase knowledge about the links between genes and disease, while guiding patients – particularly those with cancer – to receive the treatments most likely to work with their individual genetic profile.
However, some fear that scientific enthusiasm for mass decoding of personal genomes could lead into an ethical minefield, raising problems such as access to DNA data by insurers – especially if most babies have their genome read at birth – and by employers.
For a decade since the completion of the $3bn international research project to decode the first human genome, the cost of DNA sequencing has been falling faster than almost any other field of technology, as new methods are introduced to read the genetic code shared by all life on Earth.
“A genome sequence for $1,000 was a pipedream just a few years ago,” said Richard Gibbs, director of the human genome sequencing centre at Baylor College of Medicine in Houston. “[It] will transform the clinical applications of sequencing.”
Baylor is one of three large US medical centres, along with Yale School of Medicine and the Broad Institute, that will receive the first Ion Proton sequencers at the end of January, said Jonathan Rothberg of Life Technologies, who invented the technology used. Deliveries to other academic and commercial customers will follow over the next few months.
Sequencing a human genome on most of the instruments working today costs $5,000 to $10,000 and takes up to a week, using optical technology to read the individual letters of DNA that are tagged with fluorescent marker. The Ion Proton machine cuts that substantially, by using semiconductor technology to read DNA directly through its chemistry.
Life Technologies will not have the $1,000 genome field to itself for long. Other gene sequencing companies, such as Illumina of the US and Oxford Nanopore of the UK, are rapidly developing competing systems – and the cost is expected to plummet further, leading some to speculate that it will become routine for every baby to have its genome read at birth.
Mr Rothberg estimates that between 5,000 and 10,000 people have had their full genome sequenced so far, almost all for research rather than medical treatment. “I believe millions or even tens of millions of people will have their personal genome read over the next decade,” he said.
P Please consider the environment before printing this e-mail.
A US biotechnology company will on Tuesday announce the first machine that can read all 3bn letters of an individual’s DNA for as little as $1,000 – a development that will greatly accelerate medical treatment tailored to a patient’s genes but also raises ethical questions.
Life Technologies says its new Ion Proton sequencer – a $149,000 instrument about the size of a laser printer – can read a whole human genome in less than a day for $1,000 including all chemicals, running costs and preliminary data analysis.
The landmark development, expected to be matched by other companies soon, will greatly increase knowledge about the links between genes and disease, while guiding patients – particularly those with cancer – to receive the treatments most likely to work with their individual genetic profile.
However, some fear that scientific enthusiasm for mass decoding of personal genomes could lead into an ethical minefield, raising problems such as access to DNA data by insurers – especially if most babies have their genome read at birth – and by employers.
For a decade since the completion of the $3bn international research project to decode the first human genome, the cost of DNA sequencing has been falling faster than almost any other field of technology, as new methods are introduced to read the genetic code shared by all life on Earth.
“A genome sequence for $1,000 was a pipedream just a few years ago,” said Richard Gibbs, director of the human genome sequencing centre at Baylor College of Medicine in Houston. “[It] will transform the clinical applications of sequencing.”
Baylor is one of three large US medical centres, along with Yale School of Medicine and the Broad Institute, that will receive the first Ion Proton sequencers at the end of January, said Jonathan Rothberg of Life Technologies, who invented the technology used. Deliveries to other academic and commercial customers will follow over the next few months.
Sequencing a human genome on most of the instruments working today costs $5,000 to $10,000 and takes up to a week, using optical technology to read the individual letters of DNA that are tagged with fluorescent marker. The Ion Proton machine cuts that substantially, by using semiconductor technology to read DNA directly through its chemistry.
Life Technologies will not have the $1,000 genome field to itself for long. Other gene sequencing companies, such as Illumina of the US and Oxford Nanopore of the UK, are rapidly developing competing systems – and the cost is expected to plummet further, leading some to speculate that it will become routine for every baby to have its genome read at birth.
Mr Rothberg estimates that between 5,000 and 10,000 people have had their full genome sequenced so far, almost all for research rather than medical treatment. “I believe millions or even tens of millions of people will have their personal genome read over the next decade,” he said.
P Please consider the environment before printing this e-mail.
Monday, January 9, 2012
Big Data: MIT researchers use smart phones to monitor health
John Tozzi, Bloomberg Businessweek Monday, January 9, 2012
In 2009, researchers at the Massachusetts Institute of Technology gave a dorm full of students smart phones and tracked where they went, whom they called and texted, and at what times they communicated. The researchers found that the data pouring out of the phones could reliably tell when a student was ill: Those stricken with the flu moved around much less, and those who were depressed had fewer calls and interactions with others.
Anmol Madan, the doctoral student who led the study, concluded that the findings might be useful outside of dorms. There are now more than 60 million smart phones in the United States, and they're "incredibly powerful diaries of a person's life," he says.
So in November 2010, Madan and his classmate Karan Singh, both 29, started Ginger.io to mine those diaries and provide the kind of detailed, persistent health monitoring that doctors and researchers have only dreamed of. "There hasn't been large-scale, real-world data about how people behave" before now, he says.
The seven-employee company raised $1.7 million in venture capital in October from True Ventures and Kapor Capital. They'll use the money to build a series of apps that health care providers, drug companies and insurers can offer to their patients.
As in the MIT study, Ginger.io (the name is a nod to the health benefits associated with ginger) relies on a branch of computer science known as "machine learning" to sort through the tens of thousands of data points coming out of a smart phone each month, identifying a user's typical pattern of behavior. When someone deviates from that pattern, Ginger.io can trigger a response that Singh likens to a car's "check engine" light, alerting friends or doctors that they may need to intervene.
The Cincinnati Children's Hospital Medical Center is currently conducting the first test of Ginger.io's technology, a study of teens and young adults suffering from inflammatory bowel disease. The chronic condition causes occasional diarrhea, stomach pain, and fever, and is often treated with a two-week course of steroids, which can be harmful when used for too long.
The hope is that Ginger.io can detect exactly when an attack ends by, say, noticing when someone leaves the house for the first time after several days at home, and reduce the time patients are on steroids.
A handful of patients started using a Ginger.io app in the fall, and Michael Seid, the researcher who's leading the study, expects to have 50 enrolled by early 2012. Ginger.io should be "an effortless way to get a much finer-grained continuous measure of health status," he says.
Ginger.io also won $100,000 in a November contest sponsored by Sanofi-Aventis to come up with ways to help diabetics. Diabetes patients are more susceptible to depression, which in turn increases the chance they'll stop taking their medication.
Ginger.io's prototype smart phone app raises an alert when a diabetic starts behaving in a way that signals depression, so that doctors or family can offer help.
"People don't self-report that isolation," says Singh. "They don't necessarily say, 'Yes, I'm talking to less people this week,' or they don't necessarily realize that they've actually started to close themselves off."
Other groups are working on similar, passive health monitoring: Researchers at the National Institutes of Health are using mobile phones to track recovering drug addicts in Baltimore to understand what triggers relapses, and companies including WellAware Systems place motion sensors in the homes of the elderly to detect if they're in trouble.
The challenge for tech startups in health care is that it's "traditionally a conservative market, and it's a regulated market," says Jonathan Collins, analyst at London tech researcher ABI Research.
Ginger.io will also have to reassure users that their data won't be misused. Madan and Singh say Ginger.io doesn't read the content of conversations or text messages and limits the information that goes to organizations such as insurance companies, so they can't use Ginger.io to tell how often someone goes out for a cigarette, for example.
The company expects its monitoring algorithms to get better as phones and accessories evolve to collect more data. One hint of the future: Some earbuds already measure heart rates for athletes, and Apple has a patent on earbuds that take body temperature readings.
"There's just so much information on (smart phones), and there's so many new sensors," says Singh, that the potential to improve people's health with better data "is only going to continue to grow."
http://sfgate.com/cgi-bin/article.cgi?f=/c/a/2012/01/09/BUTN1MLE7P.DTL
This article appeared on page D - 3 of the San Francisco Chronicle
P Please consider the environment before printing this e-mail.
In 2009, researchers at the Massachusetts Institute of Technology gave a dorm full of students smart phones and tracked where they went, whom they called and texted, and at what times they communicated. The researchers found that the data pouring out of the phones could reliably tell when a student was ill: Those stricken with the flu moved around much less, and those who were depressed had fewer calls and interactions with others.
Anmol Madan, the doctoral student who led the study, concluded that the findings might be useful outside of dorms. There are now more than 60 million smart phones in the United States, and they're "incredibly powerful diaries of a person's life," he says.
So in November 2010, Madan and his classmate Karan Singh, both 29, started Ginger.io to mine those diaries and provide the kind of detailed, persistent health monitoring that doctors and researchers have only dreamed of. "There hasn't been large-scale, real-world data about how people behave" before now, he says.
The seven-employee company raised $1.7 million in venture capital in October from True Ventures and Kapor Capital. They'll use the money to build a series of apps that health care providers, drug companies and insurers can offer to their patients.
As in the MIT study, Ginger.io (the name is a nod to the health benefits associated with ginger) relies on a branch of computer science known as "machine learning" to sort through the tens of thousands of data points coming out of a smart phone each month, identifying a user's typical pattern of behavior. When someone deviates from that pattern, Ginger.io can trigger a response that Singh likens to a car's "check engine" light, alerting friends or doctors that they may need to intervene.
The Cincinnati Children's Hospital Medical Center is currently conducting the first test of Ginger.io's technology, a study of teens and young adults suffering from inflammatory bowel disease. The chronic condition causes occasional diarrhea, stomach pain, and fever, and is often treated with a two-week course of steroids, which can be harmful when used for too long.
The hope is that Ginger.io can detect exactly when an attack ends by, say, noticing when someone leaves the house for the first time after several days at home, and reduce the time patients are on steroids.
A handful of patients started using a Ginger.io app in the fall, and Michael Seid, the researcher who's leading the study, expects to have 50 enrolled by early 2012. Ginger.io should be "an effortless way to get a much finer-grained continuous measure of health status," he says.
Ginger.io also won $100,000 in a November contest sponsored by Sanofi-Aventis to come up with ways to help diabetics. Diabetes patients are more susceptible to depression, which in turn increases the chance they'll stop taking their medication.
Ginger.io's prototype smart phone app raises an alert when a diabetic starts behaving in a way that signals depression, so that doctors or family can offer help.
"People don't self-report that isolation," says Singh. "They don't necessarily say, 'Yes, I'm talking to less people this week,' or they don't necessarily realize that they've actually started to close themselves off."
Other groups are working on similar, passive health monitoring: Researchers at the National Institutes of Health are using mobile phones to track recovering drug addicts in Baltimore to understand what triggers relapses, and companies including WellAware Systems place motion sensors in the homes of the elderly to detect if they're in trouble.
The challenge for tech startups in health care is that it's "traditionally a conservative market, and it's a regulated market," says Jonathan Collins, analyst at London tech researcher ABI Research.
Ginger.io will also have to reassure users that their data won't be misused. Madan and Singh say Ginger.io doesn't read the content of conversations or text messages and limits the information that goes to organizations such as insurance companies, so they can't use Ginger.io to tell how often someone goes out for a cigarette, for example.
The company expects its monitoring algorithms to get better as phones and accessories evolve to collect more data. One hint of the future: Some earbuds already measure heart rates for athletes, and Apple has a patent on earbuds that take body temperature readings.
"There's just so much information on (smart phones), and there's so many new sensors," says Singh, that the potential to improve people's health with better data "is only going to continue to grow."
http://sfgate.com/cgi-bin/article.cgi?f=/c/a/2012/01/09/BUTN1MLE7P.DTL
This article appeared on page D - 3 of the San Francisco Chronicle
P Please consider the environment before printing this e-mail.
Wednesday, January 4, 2012
WSJ: So, What's Your Algorithm?
Dennis K. Berman The Wall Street Journal January 4, 2012
We are ruined by our own biases. When making decisions, we see what we want, ignore probabilities, and minimize risks that uproot our hopes.
What's worse, "we are often confident even when we are wrong," writes Daniel Kahneman, in his masterful new book on psychology and economics called "Thinking, Fast and Slow."
An objective observer, he writes, "is more likely to detect our errors than we are."
The new year will bring plenty of splashy stories about iPads and IPOs. There is a more important theme gathering around us: How analytics harvested from massive databases will begin to inform our day-to-day business decisions. Call it Big Data, analytics, or decision science. Over time, this will change your world more than the iPad 3.
Computer systems are now becoming powerful enough, and subtle enough, to help us reduce human biases from our decision-making. And this is a key: They can do it in real-time. Inevitably, that "objective observer" will be a kind of organic, evolving database.
These systems can now chew through billions of bits of data, analyze them via self-learning algorithms, and package the insights for immediate use. Neither we nor the computers are perfect, but in tandem, we might neutralize our biased, intuitive failings when we price a car, prescribe a medicine, or deploy a sales force. This is playing "Moneyball" at life.
It means fewer hunches and more facts. Think you know something about mortgage bonds? These systems are now of such scale that they can analyze the value of tens of thousands of mortgage-backed securities by picking apart the ongoing, dynamic creditworthiness of tens of millions of individual homeowners. Just such a system has already been built for Wall Street traders.
Crunching millions of data points about traffic flows, an analytics system might find that on Fridays a delivery fleet should stick to the highways— despite your devout belief in surface-road shortcuts.
You probably hate the idea that human judgment can be improved or even replaced by machines, but you probably hate hurricanes and earthquakes too. The rise of machines is just as inevitable and just as indifferent to your hatred.
Business people have been having such fantasies of rationalism for decades. Until the last few years, they have been stymied by the cost of storage, slower processing speeds and the flood of data itself, spread sloppily across scores of different databases inside one company. These problems are now being solved.
"We've just got to the point where the technology really starts to work," says Michael Lynch, chief executive of Autonomy Corp. Hewlett-Packard Co. just spent $11 billion to buy Autonomy, which vacuums up "unstructured data" then applies it to these analytic approaches.
Of course, the hype is growing fast, too. Company valuations in this space have pushed higher, and surely some will falter along the way. That won't matter much in the long run. The story of 2012 is how these technologies are inching closer to each one of us.
For a glimpse, look inside The Schwan Food Co., whose 6,000 roving sales people deliver frozen products to homes of three million customers across the country.
Schwan home sales were listless for four straight years, beset by high customer churn and inventory pileups. Over 10 months, the venerable Minnesota company began a program with the aid of Opera Solutions Inc. of New York, an eight-year-old analytics firm.
Schwan already had a crude recommendation program. Its sales people could look at six weeks of orders, and suggest purchases from that list.
The new project took it into more sophisticated territory: Matching seemingly disparate customers with similar purchase patterns in their past. Opera calls them finding "genetic twins." It also added ways to track whether customers' spending was fading from certain categories—say, breakfast foods—and offered product suggestions and discounts to keep the spending intact.
Schwan's database is now pushing out more than 1.2 million dynamically-generated customer recommendations every day, sent directly to drivers' handheld devices. Opera says Schwan's revenues are up 3% to 4% because of it.
"There is a whole class of things that couldn't be done five years ago," says Opera CEO Arnab Gupta, who just landed an $84 million venture investment from investors including Accel-KKR and Silver Lake Sumeru. His company is now valued at around $500 million. "A few years ago it might take a month to run a project involving 30 billion separate calculations. Today it can be done in two to three hours."
The big goal is to push all the heavy back-end work forward to front-line workers, often as a "dashboard" on a handheld device.
Soon, a drug saleswoman will have real-time analytics that tell her to focus on the doctors who spent time on social networks that morning, and who are thus more apt to influence colleagues, says Dhiraj C. Rajaram, founder of analytics company Mu Sigma, of Northbrook, Ill. Last week Mu Sigma raised $108 million in venture funding from General Atlantic and Sequoia Capital.
A warning awaits, of course. As Mr. Rajaram explains, analytics will eventually become the norm, which will push adaptation and business cycles even faster than they are today. "As computers become better and better, our lives are becoming more and more complex. They create new problems as much as they solve old ones."
Until then, we should take some comfort—however difficult it may feel—that machines will help us eliminate our worst human tendencies. Mr. Kahneman reminds us best: "We often fail to allow for the possibility that evidence that should be critical to our judgment is missing. What we see is all there is."
The Game is a regular column covering the future of business. Follow on Twitter @dkberman or write to dennis.berman@wsj.com
Read more: http://online.wsj.com/article/SB10001424052970203462304577138961342097348.html#ixzz1iUsKKLqi
Tuesday, January 3, 2012
Big Data and Science: To know but Not Understand
To Know, but Not Understand: David Weinberger on Science and Big Data
David Weinberger The Atlantic January 3, 2012
In an edited excerpt from his new book, Too Big to Know, David Weinberger explains how the massive amounts of data necessary to deal with complex phenomena exceed any single brain's ability to grasp, yet networked science rolls on.
Thomas Jefferson and George Washington recorded daily weather observations, but they didn't record them hourly or by the minute. Not only did they have other things to do, such data didn't seem useful. Even after the invention of the telegraph enabled the centralization of weather data, the 150 volunteers who received weather instruments from the Smithsonian Institution in 1849 still reported only once a day. Now there is a literally immeasurable, continuous stream of climate data from satellites circling the earth, buoys bobbing in the ocean, and Wi-Fi-enabled sensors in the rain forest. We are measuring temperatures, rainfall, wind speeds, C02 levels, and pressure pulses of solar wind. All this data and much, much more became worth recording once we could record it, once we could process it with computers, and once we could connect the data streams and the data processors with a network.
How will we ever make sense of scientific topics that are too big to know? The short answer: by transforming what it means to know something scientifically.
This would not be the first time. For example, when Sir Francis Bacon said that knowledge of the world should be grounded in carefully verified facts about the world, he wasn't just giving us a new method to achieve old-fashioned knowledge. He was redefining knowledge as theories that are grounded in facts. The Age of the Net is bringing about a redefinition at the same scale. Scientific knowledge is taking on properties of its new medium, becoming like the network in which it lives.
In this excerpt from my new book, Too Big To Know, we'll look at a key property of the networking of knowledge: hugeness.
In 1963, Bernard K. Forscher of the Mayo Clinic complained in a now famous letter printed in the prestigious journal Science that scientists were generating too many facts. Titled Chaos in the Brickyard, the letter warned that the new generation of scientists was too busy churning out bricks -- facts -- without regard to how they go together. Brickmaking, Forscher feared, had become an end in itself. "And so it happened that the land became flooded with bricks. ... It became difficult to find the proper bricks for a task because one had to hunt among so many. ... It became difficult to complete a useful edifice because, as soon as the foundations were discernible, they were buried under an avalanche of random bricks."
If science looked like a chaotic brickyard in 1963, Dr. Forscher would have sat down and wailed if he were shown the Global Biodiversity Information Facility at GBIF.org. Over the past few years, GBIF has collected thousands of collections of fact-bricks about the distribution of life over our planet, from the bacteria collection of the Polish National Institute of Public Health to the Weddell Seal Census of the Vestfold Hills of Antarctica. GBIF.org is designed to be just the sort of brickyard Dr. Forscher deplored -- information presented without hypothesis, theory, or edifice -- except far larger because the good doctor could not have foreseen the networking of brickyards.
- Scientific knowledge is taking on properties of its new medium, becoming like the network in which it lives.
There are three basic reasons scientific data has increased to the point that the brickyard metaphor now looks 19th century. First, the economics of deletion have changed. We used to throw out most of the photos we took with our pathetic old film cameras because, even though they were far more expensive to create than today's digital images, photo albums were expensive, took up space, and required us to invest considerable time in deciding which photos would make the cut. Now, it's often less expensive to store them all on our hard drive (or at some website) than it is to weed through them.
Second, the economics of sharing have changed. The Library of Congress has tens of millions of items in storage because physics makes it hard to display and preserve, much less to share, physical objects. The Internet makes it far easier to share what's in our digital basements. When the datasets are so large that they become unwieldy even for the Internet, innovators are spurred to invent new forms of sharing. For example, Tranche, the system behind ProteomeCommons, created its own technical protocol for sharing terabytes of data over the Net, so that a single source isn't responsible for pumping out all the information; the process of sharing is itself shared across the network. And the new Linked Data format makes it easier than ever to package data into small chunks that can be found and reused. The ability to access and share over the Net further enhances the new economics of deletion; data that otherwise would not have been worth storing have new potential value because people can find and share them.
Third, computers have become exponentially smarter. John Wilbanks, vice president for Science at Creative Commons (formerly called Science Commons), notes that "[i]t used to take a year to map a gene. Now you can do thirty thousand on your desktop computer in a day. A $2,000 machine -- a microarray -- now lets you look at the human genome reacting over time." Within days of the first human being diagnosed with the H1N1 swine flu virus, the H1 sequence of 1,699 bases had been analyzed and submitted to a global repository. The processing power available even on desktops adds yet more potential value to the data being stored and shared.
The brickyard has grown to galactic size, but the news gets even worse for Dr. Forscher. It's not simply that there are too many brickfacts and not enough edifice-theories. Rather, the creation of data galaxies has led us to science that sometimes is too rich and complex for reduction into theories. As science has gotten too big to know, we've adopted different ideas about what it means to know at all.
For example, the biological system of an organism is complex beyond imagining. Even the simplest element of life, a cell, is itself a system. A new science called systems biology studies the ways in which external stimuli send signals across the cell membrane. Some stimuli provoke relatively simple responses, but others cause cascades of reactions. These signals cannot be understood in isolation from one another. The overall picture of interactions even of a single cell is more than a human being made out of those cells can understand. In 2002, when Hiroaki Kitano wrote a cover story on systems biology for Science magazine -- a formal recognition of the growing importance of this young field -- he said: "The major reason it is gaining renewed interest today is that progress in molecular biology ... enables us to collect comprehensive datasets on system performance and gain information on the underlying molecules." Of course, the only reason we're able to collect comprehensive datasets is that computers have gotten so big and powerful. Systems biology simply was not possible in the Age of Books.
The result of having access to all this data is a new science that is able to study not just "the characteristics of isolated parts of a cell or organism" (to quote Kitano) but properties that don't show up at the parts level. For example, one of the most remarkable characteristics of living organisms is that we're robust -- our bodies bounce back time and time again, until, of course, they don't. Robustness is a property of a system, not of its individual elements, some of which may be nonrobust and, like ants protecting their queen, may "sacrifice themselves" so that the system overall can survive. In fact, life itself is a property of a system.
The problem -- or at least the change -- is that we humans cannot understand systems even as complex as that of a simple cell. It's not that were awaiting some elegant theory that will snap all the details into place. The theory is well established already: Cellular systems consist of a set of detailed interactions that can be thought of as signals and responses. But those interactions surpass in quantity and complexity the human brains ability to comprehend them. The science of such systems requires computers to store all the details and to see how they interact. Systems biologists build computer models that replicate in software what happens when the millions of pieces interact. It's a bit like predicting the weather, but with far more dependency on particular events and fewer general principles.
Models this complex -- whether of cellular biology, the weather, the economy, even highway traffic -- often fail us, because the world is more complex than our models can capture. But sometimes they can predict accurately how the system will behave. At their most complex these are sciences of emergence and complexity, studying properties of systems that cannot be seen by looking only at the parts, and cannot be well predicted except by looking at what happens.
This marks quite a turn in science's path. For Sir Francis Bacon 400 years ago, for Darwin 150 years ago, for Bernard Forscher 50 years ago, the aim of science was to construct theories that are both supported by and explain the facts. Facts are about particular things, whereas knowledge (it was thought) should be of universals. Every advance of knowledge of universals brought us closer to fulfilling the destiny our Creator set for us.
This strategy also had a practical side, of course. There are many fewer universals than particulars, and you can often figure out the particulars if you know the universals: If you know the universal theorems that explain the orbits of planets, you can figure out where Mars will be in the sky on any particular day on Earth. Aiming at universals is a simplifying tactic within our broader traditional strategy for dealing with a world that is too big to know by reducing knowledge to what our brains and our technology enable us to deal with.
We therefore stared at tables of numbers until their simple patterns became obvious to us. Johannes Kepler examined the star charts carefully constructed by his boss, Tycho Brahe, until he realized in 1605 that if the planets orbit the Sun in ellipses rather than perfect circles, it all makes simple sense. Three hundred fifty years later, James Watson and Francis Crick stared at x-rays of DNA until they realized that if the molecule were a double helix, the data about the distances among its atoms made simple sense. With these discoveries, the data went from being confoundingly random to revealing an order that we understand: Oh, the orbits are elliptical! Oh, the molecule is a double helix!
- They are so complex that only our artificial brains can manage the amount of data and the number of interactions involved.
The same holds true for models of purely physical interactions, whether they're of cells, weather patterns, or dust motes. For example, Hod Lipson and Michael Schmidt at Cornell University designed the Eureqa computer program to find equations that make sense of large quantities of data that have stumped mere humans, including cellular signaling and the effect of cocaine on white blood cells. Eureqa looks for possible equations that explain the relation of some likely pieces of data, and then tweaks and tests those equations to see if the results more accurately fit the data. It keeps iterating until it has an equation that works.
Dr. Gurol Suel at the University of Texas Southwestern Medical Center used Eureqa to try to figure out what causes fluctuations among all of the thousands of different elements of a single bacterium. After chewing over the brickyard of data that Suel had given it, Eureqa came out with two equations that expressed constants within the cell. Suel had his answer. He just doesn't understand it and doesn't think any person could. It's a bit as if Einstein dreamed e = mc2, and we confirmed that it worked, but no one could figure out what the c stands for.
No one says that having an answer that humans cannot understand is very satisfying. We want Eureka and not just Eureqa. In some instances well undoubtedly come to understand the oracular equations our software produces. On the other hand, one of the scientists using Eureqa, biophysicist John Wikswo, told a reporter for Wired: "Biology is complicated beyond belief, too complicated for people to comprehend the solutions to its complexity. And the solution to this problem is the Eureqa project." The world's complexity may simply outrun our brains capacity to understand it.
Model-based knowing has many well-documented difficulties, especially when we are attempting to predict real-world events subject to the vagaries of history; a Cretaceous-era model of that eras ecology would not have included the arrival of a giant asteroid in its data, and no one expects a black swan. Nevertheless, models can have the predictive power demanded of scientific hypotheses. We have a new form of knowing.
This new knowledge requires not just giant computers but a network to connect them, to feed them, and to make their work accessible. It exists at the network level, not in the heads of individual human beings.
This article available online at:
http://www.theatlantic.com/technology/archive/2012/01/to-know-but-not-understand-david-weinberger-on-science-and-big-data/250820/
Monday, January 2, 2012
6 Big HealthTech Ideas That Will Change Medicine In 2012
Josh Constine TechCrunch January 1, 2012
“In the future we might not prescribe drugs all the time, we might prescribe apps.” Singularity University‘s executive director of FutureMed Daniel Kraft M.D. sat down with me to discuss the biggest emerging trends in HealthTech.
Here we’ll look at how A.I, big data, 3D printing, social health networks and other new technologies will help you get better medical care. Kraft believes that by analyzing where the field is going, we have the ability to reinvent medicine and build important new business models.
For background, HYPERLINKPractice Fusion conference
Artificial Intelligence
Siri and IBM’s Watson are starting to be applied to medical questions. They’ll assist with diagnostics and decision support for both patients and clinicians. Through the cloud, any device will be able to access powerful medical AI.
For example, an X-ray gun in remote africa could send shots to the cloud where an artificial intelligence augmented physician could analyze them. Pap smears and some mammograms are already read with some AI or elements of pattern recognition.
This has the potential to disintermediate some fields of medicine like dermatology which is a pattern based field — I look at the rash and I know what it is. Soon every primary care doctor is going to have an app on their phone that can send photos to the cloud. They’ll be analyzed by AI and determine “oh that mole looks like a dangerous melanoma” or “it’s normal”. So the referral pattern to the dermatologist will slow down.
On the plus side, there are consumer apps likeSkin Scan where for $5 you can take a picture of lesion and send it to the cloud, and it will at least give you an idea if it’s dangerous or not. If it is, it can help you find a nearby doctor, which could help dermatologists get more business. Many fields are going to change because of artificial intelligence, pattern recognition, and cheaper tests.
Big Data
We’re gaining the ability to get more and more data at lower and lower price points. The primary example is the human genome and genomic sequencing. It cost a billion dollars or more 10 years ago to get a complete human sequence. However, the cost and speed of getting that data has dropped faster than Moore’s law to the point where it’s less than $5,000 when ordered online. From 23andMeyou can now get a cheap snip test, and it has a pilot program for $999 for a whole exam.
Maybe there were 10,000 patients sequenced last year. Next year it could be 100,000 and soon millions. A genome sequence could be the cost of a blood count today. When that information becomes queryable in an a crowdsourced and cloudsourced way we can be more predictive about what you’re more likely to get based on your genomics. You can then take preventative steps or get screened more often.
So we’re pulling in huge data sets from low-cost genomics to proteomics (analyzing the proteins in the blood) to quantifiable self. The challenge is to make sense of that data and make it actionable information without making the patient or doctor overwhelmed.
I think we need to make smart dashboards like they have for fighter pilots. They would piece together data from ubiquitous sensors, like those made by GreenGoose, and Microsoft Kinect that can measure your activity around the house. It would be like the OnStar for your body that could give you clues about when you’re about to get in trouble, and it could call for help or guide you to appropriate therapy.
3D Printing
3D printing has been around for a while but now it’s being applied to medicine in ways such as being able to scan the remaining leg of a patient that’s missing one from an accident. It can then build a prosthetic leg with skin and size that matches. 3D printing is integrating with the fast-moving world of stem cells and regenerative medicine with 3D ink being replaced by stem cells. In the future we’ll probably use 3D printing and stem cells to make libraries of replacement parts. It will start with simple tissues and eventually maybe we’ll be printing organs.
Social Health Network
Social networks have the ability to change our behavior. When you wireless weight scale shares metrics with your friends, you get praised for success and pressured if you’re not maintaining your diet. Social networks are also quite powerful for tracking and predicting disease. James Fowler, co-author of the book Connected is now working with Facebook to look at health data. Not surprisingly, the more friends you have, the earlier in the flu season you’ll get influenza. This could help predict when you’ll get the flu and let you take steps to avoid it.
We’re in the Facebook era, and are more open to sharing information in the healthcare spectrum. Individuals will share their whole history through services including PatientsLikeMe andCureTogether where patients with similar problems from migraines to Lou Geghrig’s disease will consolidate health information. This will enable improvements in clinical trials.
Genomera is trying allow for low-cost web-based clinical trial around any question. Practice Fusioncan also crowdsource that data from its electronic medical records. By collecting data from all the patients within a hospital or a region you can see trends and almost run clinical studies on the fly. For example you could see all the patients that have this gene and that are taking this drug, and determine if that drug is effective for them or not.
Communication With Doctors
New communication platforms similar to a Skype or FaceTime will help you communicate differently with your clinician. Many of these things are basically already here. The challenge is often not the technology but the regulatory and reimbursement markets around them. If you’re going to be talking with your clinician on your iPhone you may need to do that in a HIPAA privacy protected way. The physician is also going to want to be paid for that in some way. They’re not going to want to get all your data every time you have a hiccup or look at your iPhone pictures of your rash unless there’s a way to get paid.
The regulatory system needs to adapt towards to becoming Accountable Care Organizations, which reward clinicians and healthcare plans for keeping patients healthy opposed to paying them to do extra procedures. This contrasts with a model of paying them for service like putting in stents and doing things after a problem has already progressed. Incentives need to be aligned and reimbursement needs to change to enable some of these new technologies to actually enter the clinic.
Mobile
The ability to have your phone tie to your healthcare record and track medical metrics will have vast repercussions. Though some aren’t cleared for sale in US yet, devices like the Alivecor electrocardiogram can monitor your heart in realtime, send the data to the cloud, and allow your cardiologist to look at it instantly. Other devices are turning phones into otoscopes for looking in your ears, or glucometers for monitoring blood sugar.
Quantified self devices like the Fitbit, Jawbone Up, and more medically themed devices will take what you used to do dsin a clinic or hospital and bring it home. This will allow therapies to be tuned much more effectively than scribbling data on a piece of paper and bringing it in to your doctor months later.
Eventually these devices will converge into the equivalent of Star Trek tricorder that can perform a wide variety of medical functions. There’s even an $10 million X Prize proposed to reward the inventor of the first functional tricorder.
Unfortunately, the strict regulatory system and entrenched, interested of the United States are pushing innovation offshore. A lot of the work for using mobile phones for health care is happening in Africa and India. Since there are few physicians in some of these areas mobile health and telemedicine are taking off. For example, microfluidics allows multiple tests to be done on a small chip at pennies per test, with the ability to connect to the web for analysis. The US will need to find a way to solve these regulatory problems while keeping patients safe, otherwise jobs and revenue could slip abroad.
To learn more about what’s happening next in healthtech, check out Singularity University’sFutureMed 2020 program, watch Daniel Kraft’s Ted Talk, and browse our healthtech channel.
“In the future we might not prescribe drugs all the time, we might prescribe apps.” Singularity University‘s executive director of FutureMed Daniel Kraft M.D. sat down with me to discuss the biggest emerging trends in HealthTech.
Here we’ll look at how A.I, big data, 3D printing, social health networks and other new technologies will help you get better medical care. Kraft believes that by analyzing where the field is going, we have the ability to reinvent medicine and build important new business models.
For background, HYPERLINKPractice Fusion conference
Artificial Intelligence
Siri and IBM’s Watson are starting to be applied to medical questions. They’ll assist with diagnostics and decision support for both patients and clinicians. Through the cloud, any device will be able to access powerful medical AI.
For example, an X-ray gun in remote africa could send shots to the cloud where an artificial intelligence augmented physician could analyze them. Pap smears and some mammograms are already read with some AI or elements of pattern recognition.
This has the potential to disintermediate some fields of medicine like dermatology which is a pattern based field — I look at the rash and I know what it is. Soon every primary care doctor is going to have an app on their phone that can send photos to the cloud. They’ll be analyzed by AI and determine “oh that mole looks like a dangerous melanoma” or “it’s normal”. So the referral pattern to the dermatologist will slow down.
On the plus side, there are consumer apps likeSkin Scan where for $5 you can take a picture of lesion and send it to the cloud, and it will at least give you an idea if it’s dangerous or not. If it is, it can help you find a nearby doctor, which could help dermatologists get more business. Many fields are going to change because of artificial intelligence, pattern recognition, and cheaper tests.
Big Data
We’re gaining the ability to get more and more data at lower and lower price points. The primary example is the human genome and genomic sequencing. It cost a billion dollars or more 10 years ago to get a complete human sequence. However, the cost and speed of getting that data has dropped faster than Moore’s law to the point where it’s less than $5,000 when ordered online. From 23andMeyou can now get a cheap snip test, and it has a pilot program for $999 for a whole exam.
Maybe there were 10,000 patients sequenced last year. Next year it could be 100,000 and soon millions. A genome sequence could be the cost of a blood count today. When that information becomes queryable in an a crowdsourced and cloudsourced way we can be more predictive about what you’re more likely to get based on your genomics. You can then take preventative steps or get screened more often.
So we’re pulling in huge data sets from low-cost genomics to proteomics (analyzing the proteins in the blood) to quantifiable self. The challenge is to make sense of that data and make it actionable information without making the patient or doctor overwhelmed.
I think we need to make smart dashboards like they have for fighter pilots. They would piece together data from ubiquitous sensors, like those made by GreenGoose, and Microsoft Kinect that can measure your activity around the house. It would be like the OnStar for your body that could give you clues about when you’re about to get in trouble, and it could call for help or guide you to appropriate therapy.
3D Printing
3D printing has been around for a while but now it’s being applied to medicine in ways such as being able to scan the remaining leg of a patient that’s missing one from an accident. It can then build a prosthetic leg with skin and size that matches. 3D printing is integrating with the fast-moving world of stem cells and regenerative medicine with 3D ink being replaced by stem cells. In the future we’ll probably use 3D printing and stem cells to make libraries of replacement parts. It will start with simple tissues and eventually maybe we’ll be printing organs.
Social Health Network
Social networks have the ability to change our behavior. When you wireless weight scale shares metrics with your friends, you get praised for success and pressured if you’re not maintaining your diet. Social networks are also quite powerful for tracking and predicting disease. James Fowler, co-author of the book Connected is now working with Facebook to look at health data. Not surprisingly, the more friends you have, the earlier in the flu season you’ll get influenza. This could help predict when you’ll get the flu and let you take steps to avoid it.
We’re in the Facebook era, and are more open to sharing information in the healthcare spectrum. Individuals will share their whole history through services including PatientsLikeMe andCureTogether where patients with similar problems from migraines to Lou Geghrig’s disease will consolidate health information. This will enable improvements in clinical trials.
Genomera is trying allow for low-cost web-based clinical trial around any question. Practice Fusioncan also crowdsource that data from its electronic medical records. By collecting data from all the patients within a hospital or a region you can see trends and almost run clinical studies on the fly. For example you could see all the patients that have this gene and that are taking this drug, and determine if that drug is effective for them or not.
Communication With Doctors
New communication platforms similar to a Skype or FaceTime will help you communicate differently with your clinician. Many of these things are basically already here. The challenge is often not the technology but the regulatory and reimbursement markets around them. If you’re going to be talking with your clinician on your iPhone you may need to do that in a HIPAA privacy protected way. The physician is also going to want to be paid for that in some way. They’re not going to want to get all your data every time you have a hiccup or look at your iPhone pictures of your rash unless there’s a way to get paid.
The regulatory system needs to adapt towards to becoming Accountable Care Organizations, which reward clinicians and healthcare plans for keeping patients healthy opposed to paying them to do extra procedures. This contrasts with a model of paying them for service like putting in stents and doing things after a problem has already progressed. Incentives need to be aligned and reimbursement needs to change to enable some of these new technologies to actually enter the clinic.
Mobile
The ability to have your phone tie to your healthcare record and track medical metrics will have vast repercussions. Though some aren’t cleared for sale in US yet, devices like the Alivecor electrocardiogram can monitor your heart in realtime, send the data to the cloud, and allow your cardiologist to look at it instantly. Other devices are turning phones into otoscopes for looking in your ears, or glucometers for monitoring blood sugar.
Quantified self devices like the Fitbit, Jawbone Up, and more medically themed devices will take what you used to do dsin a clinic or hospital and bring it home. This will allow therapies to be tuned much more effectively than scribbling data on a piece of paper and bringing it in to your doctor months later.
Eventually these devices will converge into the equivalent of Star Trek tricorder that can perform a wide variety of medical functions. There’s even an $10 million X Prize proposed to reward the inventor of the first functional tricorder.
Unfortunately, the strict regulatory system and entrenched, interested of the United States are pushing innovation offshore. A lot of the work for using mobile phones for health care is happening in Africa and India. Since there are few physicians in some of these areas mobile health and telemedicine are taking off. For example, microfluidics allows multiple tests to be done on a small chip at pennies per test, with the ability to connect to the web for analysis. The US will need to find a way to solve these regulatory problems while keeping patients safe, otherwise jobs and revenue could slip abroad.
To learn more about what’s happening next in healthtech, check out Singularity University’sFutureMed 2020 program, watch Daniel Kraft’s Ted Talk, and browse our healthtech channel.
Subscribe to:
Posts (Atom)