Wednesday, October 31, 2012

Economist: Cities Are Turning into Vast Data Factories


The Economist, October 27, 2012

Cheap and easy electronic communication has probably helped rather than hindered this. First, connectivity is usually better in cities than in the countryside, because it is more lucrative to build telecoms networks for dense populations than for sparse ones. Second, electronic chatter may reinforce rather than replace the face-to-face kind. In his 2011 book, “Triumph of the City”, Mr Glaeser theorises that this may be an example of what economists call “Jevons’s paradox”. In the 19th century the invention of more efficient steam engines boosted rather than cut the consumption of coal, because they made energy cheaper across the board. In the same way, cheap electronic communication may have made modern economies more “relationship-intensive”, requiring more contact of all kinds.

Recent research by Carlo Ratti, director of the SENSEable City Laboratory at the Massachusetts Institute of Technology, and colleagues, suggests there is something to this. The study, based on the geographical pattern of 1m mobile-phone calls in Portugal, found that calls between phones far apart (a first contact, perhaps) are often followed by a flurry within a small area (just before a meeting).

Data deluge
A third factor is becoming increasingly important: the production of huge quantities of data by connected devices, including smartphones. These are densely concentrated in cities, because that is where the people, machines, buildings and infrastructures that carry and contain them are packed together. They are turning cities into vast data factories. “That kind of merger between physical and digital environments presents an opportunity for us to think about the city almost like a computer in the open air,” says Assaf Biderman of the SENSEable lab. As those data are collected and analysed, and the results are recycled into urban life, they may turn cities into even more productive and attractive places.

Some of these “open-air computers” are being designed from scratch, most of them in Asia. At Songdo, a South Korean city built on reclaimed land, Cisco has fitted every home and business with video screens and supplied clever systems to manage transport and the use of energy and water. But most cities are stuck with the infrastructure they have, at least in the short term. Exploiting the data they generate gives them a chance to upgrade it. Potholes in Boston, for instance, are reported automatically if the drivers of the cars that hit them have an app called Street Bump on their smartphones. And, particularly in poorer countries, places without a well-planned infrastructure have the chance of a leap forward. Researchers from the SENSEable lab have been working with informal waste-collecting co-operatives in São Paulo whose members sift the city’s rubbish for things to sell or recycle. By attaching tags to the trash, the researchers have been able to help the co-operatives work out the best routes through the city so they can raise more money and save time and expense.

Exploiting data may also mean fewer traffic jams. A few years ago Alexandre Bayen, of the University of California, Berkeley, and his colleagues ran a project (with Nokia, then the leader of the mobile-phone world) to collect signals from participating drivers’ smartphones, showing where the busiest roads were, and feed the information back to the phones, with congested routes glowing red. These days this feature is common on smartphones. Mr Bayen’s group and IBM Research are now moving on to controlling traffic and thus easing jams rather than just telling drivers about them. Within the next three years the team is due to build a prototype traffic-management system for California’s Department of Transportation.

Cleverer cars should help, too, by communicating with each other and warning drivers of unexpected changes in road conditions. Eventually they may not even have drivers at all. And thanks to all those data they may be cleaner, too. At the Fraunhofer FOKUS Institute in Berlin, Ilja Radusch and his colleagues show how hybrid cars can be automatically instructed to switch from petrol to electric power if local air quality is poor, say, or if they are going past a school.

Enforcing the law may also become easier. Andrew Hudson-Smith, director of the Centre for Advanced Spatial Analysis at University College London, thinks that within five years or so police forces will be able to predict and prevent some crimes by watching Twitter and other social media. The thought may give civil libertarians the creeps, but some Londoners, recalling the part played by instant messaging in last year’s riots in their city, may wish the police already had such foresight.

More mundanely, existing data about crime can be analysed more systematically. “Law enforcement’s main problem is the fragmentation of information,” says Mark Cleverley of IBM. Local policemen may know that street robberies are likelier on certain days of the week; combining information like this with other data (such as weather, time of day and so forth) makes crimes easier to prevent. In Memphis predictive-analytics software helped to reduce serious crime by 30% and violent crime by 15% between 2006 and 2010.

However, the real prize, says John Day of IBM Research, lies not in single areas such as traffic or policing but in making whole cities better by drawing on data from multiple sources for multiple purposes. 

Smartphones and cameras, say, can track the flow of people as well as that of cars. A social-media flurry may show that more people than expected are going to turn up at a rock concert, suggesting that traffic should be redirected, more public transport laid on or more police deployed.

Even brilliant technology is not much use if cities are badly run or their politics are dysfunctional. Different departments or local authorities have to work together. In Rio de Janeiro’s “control centre”, officials from several departments watch screens side by side. When a thunderstorm strikes, the airport and schools can be closed and traffic redirected from this single centre. In newly built places institutions have to be designed along with the infrastructure. Simon Giles of Accenture, a consulting firm, explains that when his firm worked on a “creative digital city” in Guadalajara, Mexico, a structure of trusts was created to oversee the development and running of the place, helping to ensure that the benefits of development are shared equally within the community.

Another problem is how to pay for all the analytics and new infrastructure in cities. Private provision is one possibility. For example, German insurers are providing their customers with weather-warning systems developed at FOKUS, says Ulrich Meissen, head of the institute’s electronic-safety department. Cisco’s Mr Elfrink points out that cities themselves could charge for many smart services. Residents might pay a few dollars a month for online medical consultations, alerts that their children have reached school or internet access on the bus.

But quite a lot of things to make city life better can be done inexpensively by residents themselves. Many governments and cities are encouraging this by making public data available. The European Union is sponsoring a project called CitySDK, involving eight cities from Manchester to Istanbul, to give developers data and tools to create digital urban services. One pilot, in Helsinki, is meant to make it easier for citizens to report problems. Another, in Amsterdam, will use real-time traffic data to allow people to find the best way around town and avoid traffic jams. A third, in Lisbon, will guide tourists.

Mr Biderman says there is an even richer seam to be mined as people find and create their own data in real time. For example, air quality can be continually monitored by cyclists and cross-checked with time, place and weather. In fact, just about any object can tell a story. An earlier SENSEable project tracked 3,000 bits of rubbish from Seattle to nearby dumps, to Portland, via Chicago to Florida and California, and to landfills all over America. Information like this may make people think about what they throw away. Some may even do something about it.

Many ideas are brewing in the world’s cities, from grand projects to single apps. Some will be dead ends; others will rely on the enthusiasm of citizens, which will not always be in plentiful supply. But in lots of imperceptible ways, from better traffic management to bins that tweet when they are ready to be emptied, city life is getting better.


Thursday, October 25, 2012

Pew: Mobile is the Needle; Social is the Thread


Kathryn Zickuhr, Pew Research Center’s Internet & American Life Project, October 18, 2012

"Examining more than a decade of data on the social impact of technology in America, Pew Internet Research Analyst Kathryn Zickuhr discussed the patterns and trends shaping the new messaging realities of the digital age at the WSU Elliott School of Communications’ annual Comm Week conference."



Tuesday, October 23, 2012

Building A Culture Around Big Data


Deanna Glick, AOL Government, October 16, 2012

A report released today by the Partnership for Public Service aims to educate federal managers on how agencies can do just that. The report, From Data to Decisions II: Building an Analytics Culture, examines how to best use data – not anecdotes – to base decisions.

Building on an original report released last November that examined how several federal agencies use data, the new report identifies strategies for how to develop and grow an analytics culture within agencies and incorporate it into how federal workers perform the mission. It profiles seven agencies using analytics to achieve better results and the strategies used in a budget-cutting climate.

Both reports were joint efforts between the partnership and the IBM Center for The Business of Government.

"By sharing compelling stories of how agencies are developing, growing and sustaining their analytics and performance-management approaches, we hope to shed light on key steps and processes that are transferable to other agencies," the report states.

To complete the report, the organizations studies how agencies are using analytics; how they got started; what conditions helped to grow their approaches; what challenges arose and why; and what success looks like.

"We found many parallels in approach across agencies and programs," according to the report. "Driven by budget realities and the push for more data-driven actions, agency managers were examining their programs in a disciplined, comprehensive way to determine how they conduct their business."

The report features details of analytics efforts at agencies within the departments of Homeland Security, Health and Human Services, Interior, Defense and Treasury.

A common successful first step in creating a culture around analytics, researchers found, was agencies tying specific activities directly to what they are intended to achieve and linking them to goals. Focusing on these details help agencies employ a data-driven approach to managing programs, identify critical information to gauge progress and results, and ensure that only those activities that are key or essential to meeting desired results are performed.

To improve airport security, for example, a federal security director with the Transportation Security Administration worked with a team to break down the job of a transportation security officer at checkpoint and baggage areas. After analyzing and brainstorming around specific tasks related to the job, his team identified more than 1,300 knowledge areas, values and skills for a transportation security officer. Based on this analysis, they identified vulnerabilities in security screening and uncovered weaknesses in training, procedures or technology. They then pinpointed what could be improved through training and better application of procedures or policy and where technology could support improved performance.

"By instituting these types of systematic processes, agencies start building analytic cultures so they can look critically at what they do and thoroughly understand how their activities can lead to better results," the report states. "The reward for their meticulous appraisal is the enhanced ability to serve the American public cost-effectively and efficiently."

For more news and insights on innovations at work in government, please sign up for the AOL Gov newsletter. For the quickest updates, like us on Facebook.

Gartner Says Big Data Creates Big Jobs: 4.4 Million IT Jobs Globally to Support Big Data By 2015



Analysts Discuss Key Issues Facing the IT Industry During Gartner Symposium/ITxpo 2012, October 21-25, in Orlando

ORLANDO, Fla.--(BUSINESS WIRE)-- Worldwide IT spending is forecast to surpass $3.7 trillion in 2013, a 3.8 percent increase from 2012 projected spending of $3.6 trillion, but it's the outlook for big data that is creating much excitement, according to Gartner, Inc.

"By 2015, 4.4 million IT jobs globally will be created to support big data, generating 1.9 million IT jobs in the United States," said Peter Sondergaard, senior vice president at Gartner and global head of Research. "In addition, every big data-related role in the U.S. will create employment for three people outside of IT, so over the next four years a total of 6 million jobs in the U.S. will be generated by the information economy."

"But there is a challenge. There is not enough talent in the industry. Our public and private education systems are failing us. Therefore, only one-third of the IT jobs will be filled. Data experts will be a scarce, valuable commodity," Mr. Sondergaard said. "IT leaders will need immediate focus on how their organization develops and attracts the skills required. These jobs will be needed to grow your business. These jobs are the future of the new information economy."

Mr. Sondergaard provided the latest outlook for the IT industry today to an audience of more than 8,000 CIOs and IT leaders at Gartner Symposium/ITxpo, which is taking place here through October 25. He said the IT industry is entering the Nexus of Forces, which includes a confluence and integration of cloud, social collaboration, mobile and information.

"This is a time of accelerating change, where your current IT architecture will be rendered obsolete," Mr. Sondergaard said. "You must lead through this change, selectively destroy low impact systems, and aggressively change your IT cost structure. This is the New World of the Nexus, the next age of computing."

Cloud
The cloud is the carrier for the three other Forces: mobile is personal cloud, social media is only possible via the cloud, and big data is the killer app for the cloud. Cloud will be the permanent fixture, the foundation.

"Cloud is not merely about cost-cutting, the end game is not just cheap on-demand services. In fact, 90 percent of these services are still subscription based, not pay-as-you-go," Mr. Sondergaard said. "We are just at the beginning of realizing the cost benefits of cloud, but organizations moving to the cloud are also attracted by the new capabilities they do not get today. It is bringing new approaches to designing applications, specifically for the cloud, and providing more resilience by architecturing failure as a design concept. Cloud also teaches us about services and service levels, and the contrast between what the business wants for outcomes versus IT's old methods of getting there."

Mobile
In 2016, more than 1.6 billion smart mobile devices will be purchased globally. Two-thirds of the mobile workforce will own a smartphone, and 40 percent of the workforce will be mobile. The challenge for IT leaders is determining what to do with this new channel to their customers and employees.

"Mobile is about computing at the right time, in the moment. It is the point of entry for all applications, delivering personalized, contextual experiences," Mr. Sondergaard said. "It means: marketing gets more time with the customer; employees become more productive; and process flows get dramatically cut."

In less than two years, iPads will be more common in business than Blackberries. Mr. Sondergaard said some CIOs are now placing orders for tens of thousands of iPads at a time. Productivity is the driver. Two years from now, 20 percent of sales organizations will use tablets as the primary mobile platform for their field sales force. As a result, by 2018, 70 percent of mobile workers will use a tablet or a hybrid device that has tablet-like characteristics.

Gartner forecasts that in 2016, half of all non-PC devices will be purchased by employees. By the end of the decade, half of all devices in business will be purchased by employees.

Social Computing
In the next three years, the dominant consumer social networks will the limits of their growth. However, social computing will become even more important. Companies are establishing social media as a discipline. Gartner predicts that in three years, 10 organizations will each spend more than $1 billion on social media.

"Social computing is moving from being just on the outside of the organization to being at the core of business operations," Mr. Sondergaard said. "It is changing the fundamentals of management: how you establish a sense of purpose and motivate people to act. Social computing will move organizations from hierarchical structures and defined teams to communities that can cross any organizational boundary."

Big Data
By tapping a continual stream of information from internal and external sources, businesses today have an endless array of new opportunities for: transforming decision-making; discovering new insights; optimizing the business; and innovating their industries.

Big data creates a new layer in the economy which is all about information, turning information, or data, into revenue. This will accelerate growth in the global economy and create jobs.

"Big data is about looking ahead, beyond what everybody else sees," Mr. Sondergaard said. "You need to understand how to deal with hybrid data, meaning the combination of structured and unstructured data, and how you shine a light on 'dark data.' Dark data is the data being collected, but going unused despite its value. Leading organizations of the future will be distinguished by the quality of their predictive algorithms. This is the CIO challenge, and opportunity."

About Gartner Symposium/ITxpo
Gartner Symposium/ITxpo is the world's most important gathering of CIOs and senior IT executives. This event delivers independent and objective content with the authority and weight of the world's leading IT research and advisory organization, and provides access to the latest solutions from key technology providers. Gartner's annual Symposium/ITxpo events are key components of attendees' annual planning efforts. IT executives rely on Gartner Symposium/ITxpo to gain insight into how their organizations can use IT to address business challenges and improve operational efficiency.

Additional information about Gartner Symposium/ITxpo in Orlando, is available at www.gartner.com/symposium/us. Video replays of keynotes and sessions are available on Gartner Events on Demand at www.gartnerondemand.com. Follow news, photos and video coming from Gartner Symposium/ITxpo on Facebook at http://www.facebook.com/GartnerSymposium, and on Twitter at http://twitter.com/Gartner_inc and using #GartnerSym.

Friday, October 19, 2012

Will Big Data decide the election?


A new book traces the recent history of data mining in political campaigns. A review of The Victory Lab, by Sasha Issenberg.

Chip Lebovitz, Fortune, October 19, 2012

FORTUNE -- There's a powerful vignette in Sasha Issenberg's The Victory Lab in which political consultant Alexander Gage presents his new data targeting system to Mitt Romney's 2002 gubernatorial campaign.

Gage has combined consumer records with political voting history to identify potential Romney supporters among nontraditional Republican voting blocks. Gage sees his work as revolutionary -- a first in politics, and potentially a first anywhere. Yet just as he completes his presentation, Romney's deputy campaign manager Alex Dunn raises his hand and deadpans, "You mean you don't do this in politics."

Dunn's surprise was not out of place in 2001. And while much has changed since then, politics remains an analog art in many ways. The Victory Lab charts the recent history of political data mining designed to identify persuadable voters and swing elections.

Politico has called Issenberg's book "Moneyball for politics." There are obvious parallels between political operatives like Gage and the protagonist of Moneyball, Oakland Athletics general manager Billy Beane. 

Both pioneered data-centric techniques that gave their respective organizations a leg up over the competition.

Yet in his 14 years as GM, Billy Beane has yet to make it to the World Series. By contrast, the work of Gage and other data-mining political consultants has translated directly into electoral success. These operatives may not have singlehandedly won the past two presidential elections, but the side that had the most advanced data-driven mobilization efforts went a perfect 2-0.

The George W. Bush campaign's sophisticated data mining helped the president win reelection in 2004. Four years later, Barack Obama defeated John McCain in part because the Obama campaign outclassed McCain's operatives on the data front.

Issenberg seems to have interviewed everyone who's anyone in the field of political statistics, but his subject matter doesn't necessarily lend itself to prose. Numbers have power, but too many of them can addle rather than inform. After reading the book, you'll probably remember the outcome of a few experiments, but be stuck scouring your brain for the results of the rest. Readers are advised to keep a pen and pencil handy when reading The Victory Lab, so they can jot down key stats.

Yet Issenberg's material is sufficiently gripping that you'll want to keep turning the pages, even if it means deciphering the results of yet another randomized trial. For example, micro-targeting efforts on behalf of Senator Michael Bennett (D-Colo.)'s 2010 Senate reelection campaign likely spurred 25,000 additional Bennett votes. That might not seem like a big number, until you consider that the race was decided by 15,000 votes.

Issenberg avoids sweeping predictions about the future of political data mining. That's probably wise, given that political professionals tend to define the future as Election Day. Campaigns at their simplest are a one-day snapshot of personal preference. The Victory Lab reminds us, however, that every individual choice is the result of a thousand factors, many of them subject to manipulation by political operatives.


Thursday, October 18, 2012

Data from Health Care Reviews Could Power "Yelp for Health Care" Startups


Data-driven decision engines will need patient experience to complete the feedback loop.

Alex Howard, O'Reilly Radar, October 17, 2012

Given where my work and health has taken me this year, I’ve been thinking much more about the relationship of the Internet and health data to accountability and patient-driven health care.

When I was looking for a place in Maine to go for care this summer, I went online to look at my options. I consulted hospital data from the government at HospitalCompare.HHS.gov and patient feedback data on Yelp, and then made a decision based upon proximity and those ratings. If I had been closer to where I live in Washington D.C., I would also have consulted friends, peers or neighbors for their recommendations of local medical establishments.

My brush with needing to find health care when I was far from home reminded me of the prism that collective intelligence can now provide for the treatment choices we make, if we have access to the Internet.

Patients today are sharing more of their health data and experiences online voluntarily, which in turn means that the Internet is shaping health care. There’s a growing phenomenon of “e-patients” and caregivers going online to find communities and information about illness and disability.

Aided by search engines and social media, newly empowered patients are discussing health conditions with others suffering from disease and sickness — and they’re taking that peer-to-peer health care knowledge into their doctors’ offices with them, frequently on mobile devices. E-patients are sharing their health data of their own volition because they have a serious health condition, want to get healthy, and are willing.

From the perspective of practicing physicians and hospitals, the trend of patients contributing to and consulting on online forums adds the potential for errors, fraud, or misunderstanding. And yet, I don’t think there’s any going back from a networked future of peer-to-peer health care, anymore than we can turn back the dial on networked politics or disaster response.

What’s needed in all three of these areas is better data that informs better data-driven decisions. Some of that data will come from industry, some from government, and some from citizens.

This fall, the Obama administration proposed a system for patients to report medical mistakes. The system would create a new “consumer reporting system for patient safety” that would enable patients to tell the federal government about unsafe practices or errors. This kind of review data, if validated by government, could be baked into the next generation of consumer “choice engines,” adding another layer for people, like me, searching for care online.

There are precedents for the collection and publishing of consumer data, including the Consumer Product Safety Commission’s public complaint database at SaferProducts.gov and the Consumer Financial Protection Bureau’s complaint database. Each met with initial resistance by industry but have successfully gone online without massive abuse or misuse, at least to date.

It will be interesting to see how medical associations, hospitals and doctors react. Given that such data could amount to government collecting data relevant to thousands of “Yelps for health care,” there’s both potential and reason for caution. Health care is a bit different than product safety or consumer finance, particularly with respect to how a patient experiences or understands his or her treatment or outcomes for a given injury or illness. For those that support or oppose this approach, there is an opportunity for public comment on proposed data collection at the Federal Register.

The power of performance data
Combining patients review data with government-collected performance data could be quite powerful in helping to drive better decisions and adding more transparency to health care.
In the United Kingdom, officials are keen to find the right balance between open data, transparency and prosperity.

“David Cameron, the Prime Minister, has made open data a top priority because of the evidence that this public asset can transform outcomes and effectiveness, as well as accountability,” said Tim Kelsey, in an interview this year. He used to head up the United Kingdom’s transparency and open data efforts and now works at its National Health Service.

“There is a good evidence base to support this,” said Kelsey. “Probably the most famous example is how, in cardiac surgery, surgeons on both sides of the Atlantic have reduced the number of patient deaths through comparative analysis of their outcomes.”

More data collected by patients, advocates, governments and industry could help to shed light on the performance of more physicians and clinics engaged in other expensive and lifesaving surgeries and associated outcomes.

Should that be extrapolated across the medical industry, it’s a safe bet that some medical practices or physicians will use whatever tools or legislative influence they have to fight or discredit websites, services or data that puts them in a poor light. This might parallel the reception that BrightScope’s profiles of financial advisors have received in industry.

When I talked recently with Dr. Atul Gawande about health data and care givers, he said more transparency in these areas is crucial:

“As long as we are not willing to open up data to let people see what the results are, we will never actually learn. The experience of what happens in fields where the data is open is that it’s the practitioners themselves that use it.”

In that context, health data will be the backbone of the disruption in health care ahead. Part of that change will necessarily have to come from health care entrepreneurs and watchdogs connecting code to research. In the future, a move to open science and perhaps establish a health data commons could accelerate that change.

The ability of caregivers and patients alike to make better data-driven decisions is limited by access to data. To make a difference, that data will also need to be meaningful to both the patient and the clinician, said Dr. Gawande. He continued:

“[Health data] needs to be able to connect the abstract world of data to the physical world of what really happens, which means it has to be timely data. A six-month turnaround on data is not great. Part of what has made Wal-Mart powerful, for example, is they took retail operations from checking their inventory once a month to checking it once a week and then once a day and then in real-time, knowing exactly what’s on the shelves and what’s not. That equivalent is what we’ll have to arrive at if we’re to make our systems work. Timeliness, I think, is one of the under-recognized but fundamentally powerful aspects because we sometimes over prioritize the comprehensiveness of data and then it’s a year old, which doesn’t make it all that useful. Having data that tells you something that happened this week, that’s transformative.”

Health data, in other words, will need to be open, interoperable, timely, higher quality, baked into the services that people use, and put at the fingertips of caregivers, as US CTO Todd Park explains in the video below:

There is more that needs to be done than simply putting “how to live better” information online or into an app. To borrow a phrase from Robert Kirkpatrick, for data to change health care, we’ll need to apply the wisdom of the crowds, the power of algorithms and the intuition of experts to find meaning in health data and help patients and caregivers alike make better decisions.

That isn’t to say that health data, once published, can’t be removed or filtered. Witness the furor over the removal of a malpractice database from the Internet last year, along with its restoration.

But as more data about doctors, services, drugs, hospitals and insurance companies goes online, the ability of those institutions to control public perception of the institutions will shift, just as it has with government and media. Given
flaws in devices or poor outcomes, patients deserve such access, accountability and insight.


Enabling better health-data-driven decisions to happen across the world will be far from easy. It is, however, a future worth building toward.

IBM's Watson Is Learning Its Way To Saving Lives


A few years ago, IBM’s new computer was a game-playing curiosity. Now Watson is poised to change the way human beings make decisions about medicine, finance, and work.

Jon Gertner, Fast Company, October 15, 2012.

The woman was gravely ill. Her name was Ms. Yamato. Thirty-seven years old, born in Osaka, Japan, she had never smoked, and yet there it was anyway: a spot on her lung.
A doctor had already performed a bronchoscopy and had made the diagnosis of cancer. Then he referred the patient to Mark Kris, an oncologist at Memorial Sloan-Kettering Cancer Center in New York. Seated alongside me in his office on the Upper East Side of Manhattan, Kris is showing me Ms. Yamato's electronic medical record on an iPad. "I'm preparing for the first visit," he explains, swiping the screen to show what that entails. He's interested in running at least two tests on the patient. The first is an MRI, to find out if the cancer has spread to her brain. The second involves a deeper diagnostic regimen. Lung cancer tumors are not all the same; there are thousands of variations. So a test that examines the mutations within a tumor will be crucial, he says. It so happens that cancer patients born in East Asia who have never smoked often have a particular mutation that responds well to a medication by the name of Erlotinib. That may be the case here. One can hope.

Over the past year, IBM executives have come to believe that Watson represents the first machine of the third computer age.

The woman is not real. She happens to be a character within an app that IBM has created for Watson, its new computer. Watson's special talent, its reason for being, is a singular ability to grasp the intricacies of human language and answer exceedingly difficult questions. You may have heard about Watson already. Back in 2007, a group of computer engineers at IBM's research labs in upstate New York began building the machine--named for IBM's founder, Thomas J. Watson--with the goal of creating a question-and-answer technology that would be more authoritative and powerful than anything on the planet. The initial objective of the Watson group was simple: to win in the game show Jeopardy!, something Watson famously achieved in February 2011. Yet the group had a far more important goal: to turn Watson into a business, hopefully one of some scale. So starting in late 2009, a business development team at IBM began holding meetings outside the company in an effort to understand the ultimate worth of this new technology. No doubt it could be a business one day. But what kind of business?

"The first thing that hit us about Watson," recalls John Kelly, IBM's chief of research, "was that this thing could be applied almost anywhere." Early on, IBM executives decided to focus on a field in which Watson could have a notable social impact while also proving its ability to master a complex body of knowledge. The team chose medicine. They believed Watson could help doctors make diagnoses and, even more important, select treatments. Specifically, they thought Watson could be the perfect tool to chart the complex decision trees that cancer specialists like Kris negotiate every day as they weigh treatment options that might involve radiation, surgery, and any of countless chemotherapy drugs. Watson can ingest more data in a day than any human could in a lifetime. It can read all of the world's medical journals in less time than it takes a physician to drink a cup of coffee. All at once, it can peruse patient histories; keep an eye on the latest drug trials; stay apprised of the potency of new therapies; and hew closely to state-of-the-art guidelines that help doctors choose the best treatments. Watson never goes on vacation. And it never forgets a fact. On the contrary, it keeps learning.

This fall, after six months of teaching their treatment guidelines to Watson, the doctors at Sloan-Kettering will begin testing the IBM machine on real patients. The Ms. Yamato app shows how it will work. After Kris inputs the results of her medical tests, Watson begins deliberating. "It's going through its algorithms," Kris says as we stare at the iPad. "It's seeing where the data sends it today." On the screen, a colorful globe spins. In a few seconds, Watson offers three possible courses of chemotherapy, charted as bars with varying levels of confidence--one choice above 90% and two above 80%. "Watson doesn't give you the answer," Kris says. "It gives you a range of answers." Then it's up to Kris to make the call. He regards the options on the screen and wonders how they might change if Ms. Yamato happened to develop a common symptom: hemoptysis, or coughing up blood.

"Let's try that," he says. He inputs the information and shows me the result approvingly. Watson has dropped one drug from the top chemo regimen. That's just what Kris would have done.

To make sense of all this--that is, to gauge both the value of Watson to a hospital like Sloan-Kettering and its potential to change forever the worlds of medicine and business--you could follow two different paths. You might consider Watson's evolutionary promise. Watson can almost certainly generate huge administrative benefits. Already, one large health insurer--Indiana-based Wellpoint--has begun using a Watson computer in its Virginia data center to speed along the authorization for medical procedures. Usually, authorizations are evaluated by a team of trained nurses and can sometimes take weeks to come through. Watsonizing the process would speed it up--a boon for a doctor like Kris, who now must wait while assistants exchange faxes with insurers before he can get clearance for any expensive tests.

Kris shows me what happens when Watson's treatment plan calls for an MRI. A button pops up on his screen to ask for preauthorization. "I just click that," he says, and it's done instantly.

I ask him what if Watson's request is denied.

Kris seems amused by the question. Watson has already consulted the latest medical literature, and it's been trained by the best cancer doctors in the world. "Who is the authority that is going to trump that?" he asks. Insurers balk at paying for unnecessary procedures; Watson's expert opinion essentially guarantees the necessity.

But the more intriguing path is the second one--a consideration of Watson's potential to do something revolutionary. This is the trail that captivates Kris. Eventually, he thinks, Watson could provide any doctor anywhere with the world's best second opinion. A physician in a community hospital in the Midwest, or at a remote medical center in China, could have instant access to everything that the medical field's best oncologists--people like Kris and his colleagues at Sloan-Kettering--have taught Watson. What is more, Watson will be able to excavate facts beyond the ken of Sloan-Kettering's current lineup of specialists. As Kris says, "We could ask Watson: What is the best treatment for this rare condition based on all of Sloan-Kettering's records?" It could then go through several years of cancer cases looking for the most successful outcomes. In time, it could even look at hospital records from around the world. As Manoj Saxena, the IBM executive now in charge of commercializing Watson, tells me: "It's like being able to take a knowledge worker--cancer specialist, nurse, bond trader, portfolio manager, whatever--and equip that person with the best knowledge, and have it available at their fingertips." As Watson evolves, Saxena believes, these knowledge banks will significantly alter how, and how well, humans make decisions.

Within a few years, for instance, Watson may be reaching well beyond oncology to assist patients suffering from any chronic disease and help general practitioners make diagnoses in their offices. Ultimately, Saxena believes, Watson could play an essential role in the diagnosis and treatment of mental health; in the financial services industry, where Citibank is testing it now; and in education. It could become the world's smartest dietitian.

How Watson Works
IBM's Watson computer begins trials in the health care industry this fall. The initial goal is to help oncologists make better decisions for cancer treatment; eventually, the computer will also aid in the diagnosis and treatment of other chronic diseases.

1. For well over a year, the Watson computers have been "trained" in science and medicine. Technicians feed Watson medical textbooks and journals, patient histories, and treatment guidelines.

2. At Memorial Sloan-Kettering Cancer Center in New York, doctors have begun using a Watson appon a tablet to access the computer through the cloud. The doctor logs in to Watson and begins to input data and ask questions.

3. When the oncologist queries Watson about a course of treatment for a lung or breast cancer patient, the computer--with its ability to understand natural language--notes keywords in the query, such as the particular type of cancer and the genomic variant of the tumor.

4. Watson then springs into action, using its massively parallel processors to review millions of pages of text in seconds. It explores the patient's medical history, medications, and other existing conditions. It then combines this information with recent data from the patient's medical tests and may comb through studies of patient groups at Sloan-Kettering who have had similar types of cancer. It also reviews doctors' and nurses' notes, recent medical research, journal articles, and treatment guidelines.

5. Watson then generates hypotheses for treatment. On the tablet app, these appear as separate options with varying levels of confidence. For instance, Watson might score one treatment option--a combination of chemotherapy drugs--with a 95% confidence level, suggesting it would be the most sensible path. It might also highlight options with lower scores as alternative treatment courses. The doctor then weighs the options and makes the call.

Saxena now commands a team of about 200 people who are working to adapt Watson's skills for various IBM clients. He and I are discussing his progress over lunch one day near IBM's upstate New York headquarters when he leans back and tells me that after creating two successful tech startups, both of which he sold (the second to IBM), his current job is far and away the most meaningful endeavor of his life. Those startups, he confides, were exciting, important. "But this," he says of the Watson rollout, "this is stuff that is going to change the course of history."

Over the past year, IBM executives have come to believe that Watson represents the first machine of the third computer age, a category now referred to within the company as cognitive computing. As Kelly describes it, the first generation of computers were tabulating machines that added up figures. "The second generation," he says, "were the programmable systems--the mainframe, the first IBM 360, PCs, all the computers we have today." Now, Kelly believes, we've arrived at the cognitive moment--a moment of true artificial intelligence. These computers, such as Watson, can recognize important content within language, both written and spoken. They do not ask us to communicate with them in their coded language; they speak ours. And perhaps most important, they can learn, so they improve without constant human instruction.

Siri, on the iPhone, might be considered an elementary example. Watson is industrial strength. "Computers do numerical calculations, they move data around, and they've been doing that forever," David Ferrucci, the IBM researcher who commanded the team that built the first Watson computer, tells me one day at IBM's research labs. "When I think about Watson, it's interpreting the information in human terms. It's saying: What does this mean to me? And that's a big deal." Also significant is how Watson renders an answer. Unlike its responses in Jeopardy!, in the real world it will perform as it did for Kris at Sloan-Kettering--by giving not a single solution but a range of probable solutions, each backed up by Watson's evidence and ranked by its level of confidence. In the lingo of computer science, that makes the machine probabilistic rather than deterministic. One might say this trait gives Watson a humanizing glow of humility and diminishes concerns that it marks a stride toward a computer-led dystopia. Watson, in IBM's marketing schema, is here to help with our questions, rather than solve them. In the case of medicine, it--for Watson is not really a he--is here to support doctors, not replace them.


The Watson of today is not precisely the same machine that won in Jeopardy! IBM has fine-tuned its software and algorithms for medical applications (or, in the case of Citibank, financial services applications). Watson has shrunk, too, from a row of about a dozen server racks that would have filled a small bedroom to an assemblage about the size of a double-door refrigerator. But for all the concentrated power, it doesn't look like anything special. Its sleek black servers are standard IBM Power 750s. You could wander around Watson and regard its blinking lights, as I did on a quiet midsummer afternoon at IBM's research labs, and not think something unusual is happening inside it. But there is. The way Watson solves problems--or, rather, the way it looks for answers, simultaneously sending out thousands of inquiries in all directions and then scoring the evidence it collects--is different from how other computers work. One person at IBM likens Watson's process to (1) gathering hundreds or thousands of possible solutions from a vast data bank, (2) pouring them into a giant funnel, (3) stirring with a dash of algorithms, and (4) letting only the best drip out of the bottom.

At the moment, a half-dozen Watsons are scattered around the country. Some are on the premises of IBM clients, as with the insurer Wellpoint, while others are cloud based, which is how hospitals such as Sloan-Kettering will access Watson. "Effectively, there's no limit to how many Watsons there can be," Bernie Meyerson, IBM's VP of innovation, tells me. Watson is a creation of software, not hardware. "That's the beauty of it," he says.

Watson is different from big servers and mainframes in other ways, too. The best computers of today have the extraordinary processing power needed to create, say, complex supply chains for building a new automobile or planning a satellite launch. These machines are good at manipulating the vast amounts of clearly defined data--numbers and facts--known as structured information. But most of the world's information is more ambiguous and less precise and lies beyond their reckoning. "We now have this proliferation of what we call Big Data," Saxena, Watson's business manager, tells me, referring to the flood of information created by our computers, our electronic sensors, and ourselves. "Ninety percent of the world's information was created in the last two years," he says. "But 80% of that 90% is unstructured or semistructured information, like doctor's notes or product reviews on Amazon." This near infinitude also includes tweets, blogs, emails--all the noise and scribble of modern life. So any company that aspired to manage the data of all the world's businesses would today be able to analyze only a small part of it. Watson, though, is a genius at reading unstructured information. And it's precisely this facility that explains why IBM sees such a rich business opportunity here.

It likewise explains why medicine is a logical first choice. While some health information is indeed structured--think of blood-pressure readings or cholesterol counts--the vast majority is unstructured. This cache includes textbooks, medical journals, patient records, and nurse and doctor evaluations. In fact, medicine embodies so much unstructured information that its proliferation has, by the account of many medical professionals, far outstripped the ability of doctors to keep up. Neither better training nor continuing education could ever wholly remedy this problem. When I meet with Herbert Chase, a professor of clinical medicine at Columbia University who consulted with IBM during the early stages of the Watson project, he says it is "not humanly possible" for a busy doctor to keep abreast of the current literature.

One result of information overload is a high rate of misdiagnosis and consequently incorrect treatment. By some estimates, Saxena tells me, 20% of initial diagnoses of cancer are eventually altered. "Imagine the implications of cancer care if there is a one in five chance that for the next six months whatever therapy they're giving you is wrong," he says.
Deciding on a course of treatment is even tougher than making a diagnosis. "It's still possible for a doctor to know the ways that people get sick," says Chase, who is also a kidney specialist. "But what is unmanageable, and what has been for decades, is knowing what the best option is today." Some applications now available to doctors are meant to alleviate this problem; one popular web-based tool is named Isabel. But Watson, in Chase's view, reaches a different level of sophistication. "I'll give you an example of a test we thought up for Watson," he tells me one day in his Manhattan office. "A patient was pregnant, had Lyme disease, and was also allergic to penicillin. And Watson came up with a drug. The first thing I thought was, Watson made a mistake. That drug can't be given to someone allergic to penicillin." But Chase was wrong, not Watson. "My knowledge was about five years old," he says. "And in the past couple of years, all the muckety-mucks had reviewed all the studies and had concluded yes, you can give that drug to someone who's allergic to penicillin."
To Chase, this proves a point: If you're a patient, you don't want to believe your doctor doesn't know everything. But he or she doesn't, and can't. At its best, the dispensation of treatment is inefficient today. "At its worst," Chase says, "it's subpar, incorrect, wrong therapy," and doesn't reach the standard of care to which his profession aspires. "As you can imagine," he adds, "this is not something we like talking about."

Last year, IBM turned 100 years old, which sets it apart from West Coast counterparts like Amazon, Apple, Google, HP, and Microsoft--all younger and ostensibly the tech world's leading innovators. To delve into IBM's recent research, though, is to wonder if our perception of technological leadership sometimes suffers from the distortions of branding and familiarity. We use iPhones and search engines and laser printers every day. But IBM's technologies are lodged deeper within the infrastructure of daily life; you're tapping into them whenever you send an email, for instance, or log on to a website. IBM has been granted more patents than any other company in the world for 19 years in a row. Yet since getting out of the laptop business in 2004, it has not produced a single product that it sells directly to the consumer.

If you're a patient, you don't want to believe your doctor doesn't know everything. But he or she doesn't, and can't.

To understand how Watson figures into the company's culture of ideas, or to see how it represents the kind of large-scale innovation that arguably lies beyond the capabilities of any startup, it helps to understand what the company actually does these days. IBM has operations in 172 countries and an organizational chart that resembles a vast Soviet bureaucracy. It employs about 433,000 men and women. Though IBM still sells hardware--big mainframe computers, silicon chips, and supercomputers--mainly it makes money selling software and consulting services to businesses and governments. The company's strategy has been validated of late by its performance: IBM's stock price has been on an upward trek for the past five years, and its winning streak has attracted the likes of Warren Buffett, who last year decided the company merited an investment of $10.7 billion. Meanwhile, as one of the few global titans to invest staggering sums on R&D ($6 billion to $7 billion a year), IBM maintains one of the world's last great industrial laboratories. At its main research center in Yorktown Heights, New York, a jet-age dream of glass curtain walls and rusticated stone designed by the Finnish-American architect Eero Saarinen, IBM employs the bulk of what is likely the world's largest mathematics department, with 300 members. If you're looking for a new PC design, you're out of luck here. But if you're shopping around for a new or better algorithm, IBM can build you one.

Not everyone is impressed by the direction of IBM's management. A relentless focus on earnings and cost cutting has led to a significant offshoring of domestic jobs, and a vocal corps of disillusioned or laid-off IBMers regularly take to the web to lament that the company's best days are behind it. IBM has also had its share of technological stumbles, apparently bungling several high-profile government contracts in recent years (in Texas and Indiana, for example) that left the company embroiled in disagreements with unhappy clients. And though these flare-ups may be uncommon, the company otherwise rarely quickens the pulse, with a long-standing reputation for being slow, steady, reliable, and maybe a little dull. IBM doesn't have big growth spikes or ballyhooed product launches; rather, it has plodding, long-term client contracts built around its ability to help optimize, say, a company's global IT services or a public utility's electrical grid. The corporation moves along like a supertanker. "IBM's annual revenue base is huge--$100 billion," says Toni Sacconaghi, a technology analyst for Sanford C. Bernstein. "So to move the needle is tough. It's hard to find big new products."


The managers and engineers keep looking anyway. One way IBM tries to infuse the troops with a sense of mission is through its periodic attempts to create for itself a Grand Challenge, such as the construction of Deep Blue, a chess-playing computer, or, more recently, Watson. The Grand Challenges are focused and expensive efforts--IBM will not verify Watson's cost, but estimates put the sum between $100 million and $1 billion--to push the company beyond the competition.

Watson's origins can arguably be traced back some years to a more modest annual initiative IBM calls the Global Technology Outlook, or GTO. Anyone at IBM can contribute to the outlook, and most of the results are eventually made public. The GTO tries to identify future business opportunities by putting a spotlight on various technology trends. A while ago, the IBM outlook pointed to analytics as a potentially huge field. Not long after, then-CEO (and current chairman) Sam Palmisano green-lighted IBM's acquisition of about $16 billion in smaller companies that had computer technologies to do this kind of work--essentially, to comb through vast stores of data, both structured and unstructured, and help extract nuggets from the global corporate babel.

Like Big Data or cloud computing, analytics is one of those contemporary catchphrases that everyone talks about but no one pauses to define. Bernie Meyerson, IBM's VP of innovation, argues that the great promise of analytics is not just to spot trends or glean information for boosting sales but to use computers and software to change the future. "Analytics is the capability to see what no human can," he says. Recently, at a public event, Meyerson was asked if IBM missed out by not building a tablet to compete with the iPad. He responded that as part of its Smarter Cities Initiative, IBM had just spent several years gathering all of the data on car transportation in Singapore; it then fed the data into a model it had built to predict the time and location of traffic jams. "We know from history what happens in Singapore if you slow the lights down in one direction by three seconds, and how to tweak the model so the jam never happens," he told his questioner. "And so there will be a traffic jam that never occurs because we can predict what happens 20 minutes from now, because we can take enough Big Data and crunch it, and do analytics on it. So we're predicting the future, and changing it. And you're asking me if I'm worried about a tablet?"

Watson, too, fits into Meyerson's conception of analytics, though it aims to change not the future of a traffic jam but of illness and investing. And by all indications, that tantalizing promise is not lost on the business community. "I have my shoulder against the door," Saxena tells me. He means he is turning clients away--something I heard from several other sources, too--until IBM executives feel confident Watson has proved its credibility at places like Wellpoint and Sloan-Kettering. Saxena seems certain that Watson will be a multibillion-dollar business, though he will only go so far as to say that by 2015, IBM will have annual revenues of about $16 billion from its analytics portfolio, of which Watson will be a part. When I put the question of Watson's potential to John Kelly, IBM's chief of research, he says: "It's like asking, at the very beginning, How big will the PC industry be?"

Kelly notes that the business model for Watson is still to be determined. He isn't sure whether selling Watson as a computer or marketing it as a service will make the most sense. But he feels he has time to decide. None of IBM's competitors, more than a year after the Jeopardy! victory, has announced a Q&A technology like Watson. "I think we have a huge lead," Kelly tells me. "When people realize this is not a one-off game machine but a new era of computing, then you'll see other companies tripling down to catch up."

I asked a number of people, both within IBM and outside of it, whether other organizations could have built this machine first. The consensus was probably not. The reasons did not precisely connect to IBM's technological capabilities--Google and Microsoft have plenty of computer prodigies in their ranks too. Rather, it was the combination of assets at IBM that made the difference. The company had its vast corporate lab, huge sums it was ready to invest, a profound expertise in hardware as well as software, and a collaborative culture that brought in lots of help from academia. And crucially, it had its business clients. In this respect, being a company that doesn't cater to consumers has advantages. Watson is only as bright as its teachers. Without the staff at Sloan-Kettering, where doctors like Mark Kris teach it oncology, Watson would not be nearly so smart. In fact, it might be kinda dumb. Or it might get all sorts of things wrong, like Siri does, except you'll be looking not for a pizza parlor but for a tumor.


From the start, the team that originally built Watson under David Ferrucci has worked out of a big room on the second floor of IBM's Hawthorne Labs in Westchester County, New York. Hawthorne is a large glass cube of a building situated about 30 miles north of New York City. Inside the Watson work space are five fake wood-grained tables, each home to a group of computer engineers who sit around and alternately immerse themselves in their screens or break to discuss coding with a neighbor. The mood here is sober. The staffers bring water bottles, not junk food. These aren't the unlined faces you'll see at a startup. Indeed, Ferrucci, who sits off to the side, is a suburban dad who looks like he'd be just as comfortable standing in front of a grill with a basting brush as he is overseeing his team. The walls here are covered with huge whiteboards crammed with the hieroglyphics of computer science. Overhead lights cast the room in gloomy fluorescence. The place has the neglected feel of a finished basement in a 1970s-era subdivision.

In early fall, the Watson team, now about 45 strong, began moving its work to a gleaming new space in IBM's main Yorktown Heights research laboratory--a promotion that reflects their importance as they support Saxena's much larger business development group while simultaneously working on the next iteration of Watson, known as Watson 2.0. One of the team's goals is to make Watson adaptable enough so that it doesn't require several dozen people spending a year to get it ready for every new application, such as medicine or financial services. But a more immediate project is to help Watson through the U.S. Medical Licensing Examination, the complex test all med-school graduates must take before practicing. If it passes, says Ferrucci, "that doesn't mean I can have a computer be a doctor." But IBM would gain what he calls "a crisp metric" that proves Watson has a real proficiency in medicine. The credential would no doubt help Watson's standing with health insurers, doctors, and patients, too. Passing the licensing exam is a difficult task--far harder than winning at Jeopardy!--but in early September, Ferrucci seemed pleased by the results. The computer is doing "interestingly well," he said. He sounded confident that Dr. Watson will ace the test by year's end.

Harder to intuit is how soon afterward Watson will infiltrate society. When I ask Jaime Carbonell, a computer science professor at Carnegie Mellon, he says he has no doubt the impact of Watson will be significant. "But I don't think there will be one moment of, 'Now we have it and yesterday we didn't,'" Carbonell remarks. "It will take time to permeate. Like cell phones, which were big, clumsy things you could barely carry at first." Was there a year, or month, or day, he asks, when cell phones began to change the world? "I can't think of when that was," he says. "But now we can't do without them."

Such is the course of technology: Electronic tools initially available only to the elite grow ever faster, smaller, cheaper. Kelly tells me he believes that eventually Watson will shrink to the size of a handheld device. Randy Katz, a computer science professor at UC Berkeley, sees a more approachable Watson, too. "Can the person in the street ask Watson a question now? No, he can't," says Katz. "But in five or 10 years, will there be systems like that--like Siri, but much better? I think the answer is yes."

In many of my conversations at IBM, the talk often drifts to applications of Watson. All sorts of intriguing scenarios are presented to me--for instance, that Watson will soon analyze not just words but images, such as MRIs and EKGs. Or it will diagnose a spider bite on a child's arm in a crop field in Africa, transmitted via smartphone by his worried father to a U.S. hospital. One afternoon, Saxena suggests this one: When you think you're coming down with the flu, Watson will be able to discern, before you even arrive at the doctor's office, that it might be a ragweed allergy, based on your medical record (you've had the same symptoms twice before at this time of year); your symptoms (gleaned from the insurance claim and diagnostic information in journals); and recent news (it just read an article in the Austin-American Statesman on a ragweed outbreak near your hometown).

It all sounds amazing. It's also speculative. Watson has not yet saved a life or a dollar of medical costs, or added anything, really, to IBM's bottom line. It has not yet faced its resistors--doctors who may find the technology objectionable and slow its adoption. It has not yet, as Saxena believes it will, changed the course of history. It has only won a television game show.

Still, Saxena predicts the computer will begin to scale up dramatically late next year. "By then," he says, "we will have built the technology, demonstrated it, built the tooling and methods around it. We will have the recipe book, and then we'll just push it out." But he will only have reached the end of Watson's beginning.

A version of this article appears in the November 2012 issue of Fast Company.

Wednesday, October 17, 2012

Tim O'Reilly: Open Health Data in Practice: Increase Your Access to Lab Results Voice Your Support for a Proposed Federal Rule that Expands Patients' Access to Test Results


Tim O'Reilly, O'Reilly Radar, October 16, 2012

I’m convinced that there’s a wave of innovation coming in healthcare, driven by new kinds of data, new ways of extracting meaning from that data, and new business models that data can enable.  That’s one of the reasons why we launched our StrataRx Conference, which focuses on the importance of data science to the future of health care.

Unfortunately, much of the data that will enable an entrepreneurial explosion is still locked up — in paper records, in proprietary data formats, and by well-intentioned but conflicting privacy regulations.

We’re making progress towards open data in healthcare, but there are still so many obstacles!  Ann Waldo recently introduced me to one of these.

A 2009 law modernized patient access rights by allowing individuals to get copies of their medical records in electronic format. Unfortunately, however, these patients’ access rights surprisingly do not include lab test results – one of the types of medical records that people are most likely to find urgent and useful. Due to the interaction of HIPAA (the Federal medical privacy law), CLIA (a Federal laboratory regulatory law), and state laws, patients can only get direct access to their their test results from labs in a handful of states.
A recent New York Times story highlighted just how much pain and suffering can be caused by this inability to get access to your own lab results.

In 2011, the Department of Health and Human Services put forward a proposed Rule that would give patients the right to get their test results directly from laboratories. This Rule is still waiting to be finalized. In hopes of breaking the logjam, O’Reilly Media and a variety of other players have written a consensus letter that voices our whole-hearted support for that proposed Rule and encourages the Federal government to finalize it promptly.
We’d love to invite you to join us in signing this letter.

Patients’ rights should include direct access to their lab results, just like all their other medical records!



Should High Schools Teach Big Data?



Given the anticipated shortage of data scientists, some high school educators have jumped in to expose students to big data concepts.


Changing when advanced database technology is taught has real-world implications, given the realities of today's job market. Both data analytics and big data skills are in high demand in private industry and government.

But there is a looming shortage of workers with these abilities. McKinsey & Co. sounded this alarm back in 2011 with its seminal report that predicted the U.S. would face ashortage of 140,000 to 190,000 workers with the skills to manage and analyze big data.

The popular technology job board Dice.com has seen a spike in listings for "data scientist," up from just a handful a year ago to more than 35 at the start of October. While still an imprecise job designation, "data scientists" command high salaries compared to other IT job titles. (Separately, the unemployment rate for technology professionals dropped in the third quarter to 3.3%, as compared to 4.2% in the same quarter a year ago, according to the Bureau of Labor Statistics.)

These listings--many of which request a PhD in fields like mathematics, economics, or statistics--today cluster in financial services, retail, and e-commerce. Job listings using the phrase "big data" have increased from around 200 in January to nearly 800 in October.

"But increasingly, every industry is dealing with big data questions," said Alice Hill, managing director of Dice.com and president of Dice Labs.

Preparing the U.S. for future high-tech jobs, specifically ones oriented around data, was also a focus of TechAmerica Foundation's Big Data Commission.

Given the anticipated shortage of data scientists, should students start learning the precepts of big data in high school?Analytics and data science are central to making "business, the global economy, and our society work better," Steve Mills, senior VP and group executive at IBM, and co-chair of the Big Data Commission, said in a statement. "That's why it's critical that our country prepares a new generation of experts who know how to corral today's data deluge for world-changing insights."

The commission's new report, "Demystifying Big Data: A Practical Guide to Transforming the Business of Government," included the following recommendations for skills development: Strengthen and expand public-private partnerships to invest in skills-building initiatives for the federal workforce in the area of big data. These should include formal career tracks for IT managers; an IT leadership academy to provide big data and related training and certification; data-intensive degree programs; and scholarships to prepare a new generation of data scientists.

Big Data In High School?
Among those trying to push big data classes down to the high school level is Alex Philp, PhD. Philp is founder and CTO of TerraEchos, a developer of advanced intelligence and surveillance security systems. Philp has been working with the high schools in Missoula, Mont., and has even created a scholarship program at one--Sentinel High School--to introduce these topics to computer science students.

"[We] cooked up the idea of promoting a merit-based, micro-challenge grant at Sentinel," Philp explains. Currently, four student teams have been awarded grants. "Ultimately, I hope these high school students feed into opportunities at the University of Montana and then ultimately into the most exciting businesses and markets involving big data," Philp said. "This also relates to aspects of the overall economic competitiveness of our country and a renewed commitment to science, technology, engineering, arts, and mathematics (STEAM) competency in our country."

Meanwhile, at the university level, Philp has been working with Eric Tangedahl, IT director for the school of business administration at The University of Montana, to create a multidisciplinary course, now in its first year.

"The course draws students from computer science, management information systems, and math," Tangedahl said. "We felt that to tackle these problems, students have to work together in teams, teams with different skill sets," he said.

The three-hour class, now with 18 students, is half lecture about big data and half lab--specifically around IBM Infostream, a programming language that takes advantage of IBM's DB2 and WebSphere platforms.

Another high school on this path is the science and engineering magnet (SEM) school at Yvonne A. Ewell Townview Center in Dallas. SEM, ranked by Newsweek this year as one of the best high schools in the U.S., recently participated in IBM's annual Master the Mainframe Contest for high school students in the U.S. and Canada.

Marilyn Cadenhead, who teaches Advanced Placement computer science at SEM, has many students participating in the contest.

Cadenhead says her students use technology very effectively, and already have a computer-mediated learning style. "This is a digital age, very different from when I was in school," she wrote in an email. "I learned to type on a manual typewriter. If we as teachers do not engage [students] with the latest technology and teaching styles, we will be 'boring' and the students will not be motivated to learn."

"I want my students to know what is going on in the real world," Cadenhead concludes, explaining her enthusiasm for the IBM mainframe contest, as well the annual IBM Innovation summer camp that some SEM students attended this year.

IBM's Innovation summer camp Facebook page describes the program this way: "The students will gain hands-on experience with visualization, mobile application development, Linux, and DB2 coupled with demos of research underway at UTD in mind control, gaming, and robotics."

Other observers, however, point out that big data analysis is powerful precisely because it is about more than raw technology. The highly valued professionals in this space are those who ask the right questions, who can see business-relevant answers inside the data. That kind of maturity and domain expertise is beyond the capacity of high school students, they say.

The University of Montana's Tangedahl concurs. "I wouldn't throw high school student in and expect them to come up with a [big data] program for the Defense department," he said. But, he adds, "you've absolutely got to put the building blocks in place." He recommends adding data statistics and programming classes in high school to better prepare students for college-level classes, like his, that push teams of students to answer "real-world problems."

Like Tangedahl, Dice.com's Hill isn't sold on the idea of pushing big data instruction to high school students. She thinks students should learn relevant technologies, such as data interpretation, data extraction, and data modeling. "It's never too soon to learn those skills," she said.

But TerraEchos' Philp disagrees, arguing it is essential that educators encourage excellence in students at all ages, and not set limits.

"I've spent many years of my life working with thousands of students at various ages, attempting to raise the bar," he said. "High school kids are proposing [projects] that are as good as I'm seeing in college." With motivation, inspiration, and passion, he said, "my experience is, students rise to the occasion."