Showing posts with label Big data. Show all posts
Showing posts with label Big data. Show all posts

Monday, March 11, 2013

WSJ: Public Data…at Your Fingertips

Federal and state governments are a trove of information. Now a lot of it is just an app away.


Most government data have been available to the public for decades. But trying to find information quickly and in a useful format has been another story altogether.


All that is starting to change. A growing open-government movement has pushed successfully for state and federal data to be made more readily available online, in user-friendly formats. And the federal government has begun to release its data in a digital format that allows software engineers to more easily use it to develop interactive applications. App developers are now mixing and matching government data to help consumers, businesses and researchers harvest new insights.

Here's a closer look at some of these apps and websites.
Labor and Health Violations
DATA: Records of labor-law and health-and-safety violations for hotels, restaurants and shops
SOURCES:U.S. Labor Department; Occupational Safety and Health Administration
DESCRIPTION: The Eat Shop Sleep app by Rachel Moore maps health and labor-law violations by restaurants, hotels and shops, and combines that information with Yelp reviews. Now you can find out that the hamburger joint down the street owes employees $5,653 in back wages—but the sweet-potato fries are supposed to be divine.
-
Salaries
DATA: Average salaries for hundreds of jobs listed by state
SOURCE: U.S. Bureau of Labor Statistics
DESCRIPTION: Aequitas, an app from Fuzion Applications Inc., gathers salary data for more than 800 occupations from all around the U.S. Job seekers can see the range of salaries in any area and find out what they might command in, say, Nevada versus Illinois.
Business-Owner Addresses
DATA: Addresses of business-license holders
SOURCES: City authorities
DESCRIPTION: The Your Mapper app from Metro Mapper LLC shows information for business-license holders in Chicago; Portland, Ore.; and Louisville, Ky., on city maps. Salespeople can use it to target potential customers.
Hospital Records
DATA: Hospital safety records and quality indicators; patient surveys
SOURCE:U.S. Centers for Medicare and Medicaid Services
DESCRIPTION: Hospital Quality, an app from MicroStrategy Inc., combines hospital safety and quality scores with patient-satisfaction scores gathered by CMS and uses the data to rank hospitals.
Flu Incidence
DATA: Weekly state influenza levels; crowdsourced reports
SOURCES:U.S. Centers for Disease Control and Prevention; Foursquare

DESCRIPTION: The Flusquare app from Floodlight Software LLC updates users on flu levels in their states and asks them to indicate when they may be sick.
If a user who also has downloaded the Foursquare app indicates enough symptoms to suggest he has the flu, the app then pulls his recent location data from Foursquare to find out where he visited when he was contagious. The app then warns other Flusquare users who were in his vicinity at that time to watch for symptoms.
Energy Prices
DATA: Consumption and price information for electricity and fossil fuels; power plants' fuel use
SOURCE:U.S. Energy Information Administration
DESCRIPTION: EIA's Electricity Data Browser (at eia.gov/electricity/data/browser ) collects nationwide data on power consumption and prices for the past decade, helping companies project future energy costs. Plant-level data reveals what fuel types are used, allowing companies interested in reducing their carbon footprints to estimate the environmental impact of operating in one location versus another.
What's Offshore
DATA: Locations of offshore oil and gas platforms; traffic patterns of marine vessels and locations of shipwrecks; offshore areas designated for wind energy and protected wildlife areas
SOURCES: Include U.S. Defense Department; Bureau of Ocean Energy Management; National Oceanic and Atmospheric Administration
Description: MarineCadastre.gov pulls data from 16 government and nonprofit sources into an interactive online map of U.S. oceans. It is primarily intended to help developers of wind and wave energy, but it also can be used by fishing and drilling companies.
Where the Sun Shines
DATA: Average annual global horizontal irradiance; photovoltaic plant locations; transmission line locations
SOURCES: Include National Renewable Energy Laboratory; Argonne National Laboratory; U.S. Interior Department
DESCRIPTION: Designed to help developers find sites for large solar power plants, the National Renewable Energy Laboratory's Solar Prospector (at maps.nrel.gov/prospector) pulls data on sunlight patterns, routes of transmission lines and locations of existing power plants into a single interactive map.

USDA
Food and Fitness
DATA: Americans' access and proximity to grocery stores; obesity rates; diabetes rates; physical fitness levels; proximity to fast food
SOURCE:U.S. Department of Agriculture
DESCRIPTION: The USDA's Food Environment Atlas (at ers.usda.gov) pulls food, health and socioeconomic data into various maps, allowing researchers to search for correlations between the availability of healthy food and obesity-related illnesses like diabetes. The maps also allow food-service companies to look for possible sites to expand.

National Nanotechnology Initiative
Nanotechnology
DATA: Nationwide locations of nanotechnology facilities, initiatives and degree programs
SOURCE:National Nanotechnology Initiative
DESCRIPTION: The Nanotechnology Resource Map (at nano.gov) shows the location of degree programs, facilities and research-and-development projects in the field of nanotechnology, giving businesses, investors and students interested in the technology a comprehensive overview of what's out there.
Mr. Schectman is a Wall Street Journal reporter in New York. He can be reached at joel.schectman@wsj.com.

Algorithms Get a Human Hand in Steering Web

Trading stocks, targeting ads, steering political campaigns, arranging dates, besting people on "Jeopardy" and even choosing bra sizes: computer algorithms are doing all this work and more.

But increasingly, behind the curtain there is a decidedly retro helper — a human being.

Although algorithms are growing ever more powerful, fast and precise, the computers themselves are literal-minded, and context and nuance often elude them. Capable as these machines are, they are not always up to deciphering the ambiguity of human language and the mystery of reasoning. Yet these days they are being asked to be more humanlike in what they figure out.

"For all their brilliance, computers can be thick as a brick," said Tom M. Mitchell, a computer scientist at Carnegie Mellon University.

And so, while programming experts still write the step-by-step instructions of computer code, additional people are needed to make more subtle contributions as the work the computers do has become more involved. People evaluate, edit or correct an algorithm's work. Or they assemble online databases of knowledge and check and verify them — creating, essentially, a crib sheet the computer can call on for a quick answer. Humans can interpret and tweak information in ways that are understandable to both computers and other humans.

Question-answering technologies like Apple's Siri and I.B.M.'s Watson rely particularly on the emerging machine-man collaboration. Algorithms alone are not enough.

Twitter uses a far-flung army of contract workers, whom it calls judges, to interpret the meaning and context of search terms that suddenly spike in frequency on the service.

For example, when Mitt Romney talked of cutting government money for public broadcasting in a presidential debate last fall and mentioned Big Bird, messages with that phrase surged. Human judges recognized instantly that "Big Bird," in that context and at that moment, was mainly a political comment, not a reference to "Sesame Street," and that politics-related messages should pop up when someone searched for "Big Bird." People can understand such references more accurately and quickly than software can, and their judgments are fed immediately into Twitter's search algorithm.

"Humans are core to this system," two Twitter engineers wrote in a blog post in January.

Even at Google, where algorithms and engineers reign supreme in the company's business and culture, the human contribution to search results is increasing. Google uses human helpers in two ways. Several months ago, it began presenting summaries of information on the right side of a search page when a user typed in the name of a well-known person or place, like "Barack Obama" or "New York City." These summaries draw from databases of knowledge like Wikipedia, the C.I.A. World Factbook and Freebase, whose parent company, Metaweb, Google acquired in 2010. These databases are edited by humans.

When Google's algorithm detects a search term for which this distilled information is available, the search engine is trained to go fetch it rather than merely present links to Web pages.
 "There has been a shift in our thinking," said Scott Huffman, an engineering director in charge of search quality at Google. "A part of our resources are now more human curated."

Other human helpers, known as evaluators or raters, help Google develop tweaks to its search algorithm, a powerhouse of automation, fielding 100 billion queries a month. "Our engineers evolve the algorithm, and humans help us see if a suggested change is really an improvement," Mr. Huffman said.

Katherine Young, 23, is a Google rater — a contract worker and a college student in Macon, Ga. She is shown an ambiguous search query like "what does king hold," presented with two sets of Google search results and asked to rate their relevance, accuracy and quality. The current search result for that imprecise phrase starts with links to Web pages saying that kings typically hold ceremonial scepters, a reasonable inference.

Her judgments, Ms. Young said, are "not completely black and white; some of it is subjective." She added, "You try to put yourself in the shoes of the person who typed in the query."

I.B.M.'s Watson, the powerful question-answering computer that defeated "Jeopardy" champions two years ago, is in training these days to help doctors make diagnoses. But it, too, is turning to humans for help.

To prepare for its role in assisting doctors, Watson is being fed medical texts, scientific papers and digital patient records stripped of personal identifying information. Instead of answering questions, however, Watson is asking them of clinicians at the Cleveland Clinic and medical school students. They are giving answers and correcting the computer's mistakes, using a "Teach Watson" feature.

Watson, for example, might come across this question in a medical text: "What neurological condition contraindicates the use of bupropion?" The software may have bupropion, an antidepressant, in its database, but stumble on "contraindicates." A human helper will confirm that the word means "do not use," and Watson returns to its data trove to reason that the neurological condition is seizure disorder.
 "We're using medical experts to help Watson learn, make it smarter going forward," said Eric Brown, a scientist on I.B.M.'s Watson team.

Ben Taylor, 25, is a product manager at FindTheBest, a fast-growing start-up in Santa Barbara, Calif. The company calls itself a "comparison engine" for finding and comparing more than 100 topics and products, from universities to nursing homes, smartphones to dog breeds. Its Web site went up in 2010, and the company now has 60 full-time employees.

Mr. Taylor helps design and edit the site's education pages. He is not an engineer, but an English major who has become a self-taught expert in the arcane data found in Education Department studies and elsewhere. His research methods include talking to and e-mailing educators. He is an information sleuth.

On FindTheBest, more than 8,500 colleges can be searched quickly according to geography, programs and tuition costs, among other criteria. Go to the page for a university, and a wealth of information appears in summaries, charts and graphics — down to the gender and race breakdowns of the student body and faculty.

Mr. Taylor and his team write the summaries and design the initial charts and graphs. From hundreds of data points on college costs, for example, they select the ones most relevant to college students and their parents. But much of their information is prepared in templates and tagged with code a computer can read. So the process has become more automated, with Mr. Taylor and others essentially giving "go fetch" commands that the computer algorithm obeys.

The algorithms are getting better. But they cannot do it alone.

Sunday, March 10, 2013

Data Brokers Know More About Us than We Know

Data brokers, workplace sensor studies, unreported drug side effects revealed in search data, and the dark side of big data.

ProPublica's Lois Beckett takes a look this week at data brokers. She says that though Congress is making moves to make such companies give consumers more control over their data and what happens to it, many people not only don't know these data brokers exist, but they also don't know the extent of the data gathered and how it's used.


Beckett takes a step-by-step look at who these companies are, how much they know, how they get our data and from where, what kind of data they're allowed to collect, how they use the data, and much, much more. She says most of the time, consumers have no idea their data has been purchased. For instance, "When you're checking out at a store and a cashier asks you for your Zip code, the store isn't just getting that single piece of information," she writes. "Acxiom and other data companies offer services that allow stores to use your Zip code and the name on your credit card to pinpoint your home address — without asking you for it directly."

It's possible but very, very difficult for consumers to stop the companies from collecting and sharing their data, Beckett notes. Though most data brokers have an "opt-out" policy, consumers would "need to know about all the different data brokers and where to find their opt-outs" — information that most consumers don't have and don't know how to find. You can find Beckett's full report at ProPublica — it's this week's recommended read.

In related news, Rachel Emma Silverman at the Wall Street Journal takes a look at the use of sensors and data gathering practices in the workplace. "As big data becomes a fixture of office life, companies are turning to tracking devices to gather real-time information on how teams of employees work and interact," she writes. "Sensors, worn on lanyards or placed on office furniture, record how often staffers get up from their desks, consult other teams and hold meetings."

Though there's a fine line between big data and big brother, Silverman says, "[s]ensor proponents … argue that smartphones and corporate ID badges already can transmit their owner's location" and most companies will allow workers to opt out of the sensor studies. Silverman reviews a few real-world sensor study cases, along with the results and insights gleaned. You can read her full report at the Wall Street Journal.

Study shows search data reveals unreported drug side effects

John Markoff reports at the New York Times on a study published this week that shows by data mining Internet search data, scientists from Microsoft, Stanford and Columbia University have been able to discover unreported prescription drug side effects before the FDA's warning system flagged them. Markoff writes:
"Using automated software tools to examine queries by six million Internet users taken from web search logs in 2010, the researchers looked for searches relating to an antidepressant, paroxetine, and a cholesterol lowering drug, pravastatin. They were able to find evidence that the combination of the two drugs caused high blood sugar."
Users who opted to participate in the study installed a browser toolbar that gathered anonymized data, Markoff reports. Using data from 82 million drug-, symptom- and condition-related searches in 2010, researchers were able to cross-reference searches for "paroxetine" and "pravastatin" with the number of times users would also search for "hyperglycemia" or any of its 80 or so symptoms.

"They determined that people who searched for both drugs during the 12-month period were significantly more likely to search for terms related to hyperglycemia than were those who searched for just one of the drugs," Markoff reports. He also notes that the searches for symptoms relating to both drugs occurred within a short period of time — 30% the same day, 40% the same week, and 50% the same month. You can read Markoff's full report at The New York Times.

Keeping an eye on big data's dark side

Viktor Mayer-Schönberger's and Kenneth Cukier's new book Big Data: A Revolution That Will Transform How We Live, Work, and Think was released this week. The duo addressed one of the topics in their book — predicting and punishing crime before it happens, ala Minority Report — in a post at PopSci. They warn that along with all the benefits we're reaping from big data, we need to be conscious of big data's dark side, too:
"Already we see the seedlings of Minority Report-style predictions penalizing people. Parole boards in more than half of all U.S. states use predictions founded on data analysis as a factor in deciding whether to release somebody from prison or to keep him incarcerated. A growing number of places in the United States — from precincts in Los Angeles to cities like Richmond, Virginia — employ 'predictive policing': using big-data analysis to select what streets, groups, and individuals to subject to extra scrutiny, simply because an algorithm pointed to them as more likely to commit crime."
They warn further that it won't stop there — law enforcement will eventually attempt to predict crime on individual levels and, ultimately, to use big data to prevent the crime in the first place. You can read more from Mayer-Schönberger's and Cukier's piece at PopSci.

Cukier also sat down with O'Reilly Radar online managing editor Mac Slocum at the recent Strata Conference in Santa Clara to talk about government's use of big data, regulations and restrictions, and what needs to be done to keep data open. You can watch their interview in the following video:

Thursday, March 7, 2013

Big Data Helps Kaiser Close Healthcare Gaps

 Analytics from massive clinical data repository are central to closing gaps in care, HIMSS attendees told.


One benefit of Kaiser Permanente spending an estimated $6 billion for an integrated electronic health records (EHR) system to serve 9 million people across eight regions from coast to coast is it that has amassed a vast repository of clinical data. That storehouse also contains information from a patient portal, ancillary systems, smart medical devices and even home-based patient monitoring systems.

All those terabytes of electronic data now are helping to fuel a massive analytics operation, part of an overall organizational goal of improving care and reining in costs. "It's all about the data and information, not the electronic health record," Carol Cain, senior director of clinical information services for the Kaiser Permanente Care Management Institute, said this week at the Healthcare Information and Management Systems Society (HIMSS) annual conference in New Orleans.

Kaiser has embraced a concept of "complete care," which one Southern California Permanente Medical Group described as "giving my patients everything they need, whether they know it or not," according to Cain's presentation.

"We need to incorporate so much more data that is available," Cain said. Data needs to be "synthesized in a meaningful way" and delivered to primary care physicians at the point of care to help suggest appropriate interventions.

Cain said Kaiser views big data as being characterized by "volume, variety and velocity." The term "refers to datasets whose size is beyond the ability of typical database software tools to capture, store, manage and analyze," she said.

"Our ability to monitor our members' health is greater than our members' ability to know what needs to be monitored," Cain explained.

[ Are your patients taking leadership for their own health? See 7 Portals Powering Patient Engagement. ]

Kaiser Permanente has developed several modules of population management, all designed to identify and close gaps in care. If a patient shows up with knee pain, for example, management tools suggest doctors ask about a cancer screening, in an effort to make office visits "proactive" and organize care around the concept of the patient-centered medical home, Cain said.

The analytics also has to be done in a way that won't make patients feel like Big Brother is watching over them, Cain said. Instead, Kaiser wants people to think that the integrated delivery system is helping to prevent illness and find health problems early. If patients allow Kaiser to access information linked to their supermarket loyalty cards, the organization will not send warnings every time they purchase a candy bar, Cain said.

What Kaiser can do is rely on its platform to combine patient-specific knowledge, such as whether an individual has filled a prescription. This can help with medication adherence, according to Cain. Analytics are helpful for developing care plans before patients are discharged from hospitals, too.

Kaiser also can advise patients to telephone or schedule e-visits if a primary care physician determines a problem is not worth an in-person appointment. "That is something that is often appreciated by our members," Cain noted.

Cain said that patient needs are not always clinical, either. During a 12-hour hackathon in the analytics department, Kaiser IT professionals were able to correlate access to parks with rates of obesity in Oakland, Calif. "In some of our communities, we are investing in building parks," Cain said. Kaiser also has partnerships with YMCA and schools in some areas to address lifestyle issues that can affect health.

Wednesday, March 6, 2013

NYU to Train Data Scientists for E-Commerce, Pharma, and Other Markets

João-Pierre S. Ruth, Xconomy, February 20, 2013

With all the talk of big data these days, many businesses are still looking for talent to turn massive amounts of information into something they can use. New York University announced Tuesday it is creating a Center for Data Science, along with a master's-level degree program to train big-data professionals for commercial and public sectors.

"There's a continuous flow of data coming from customers and sensors," says Yann LeCun, director of the new center. "Companies have to make sense out of it."

Making more efficient use of big data is growing more critical to business, research, and government entities. NYU sees increasing demand for experts to better harness big data to work in industries that include sensors, e-commerce, and media. "It is very difficult to find data scientists right now on the job market," LeCun says. "It's because there is no place to learn data science."

The theme for the new center, he says, is to bring together people from different academic disciplines such as mathematics, computer science, and statistics. He explains that computer scientists might not know enough math to fill data science jobs, while mathematicians might not know enough computer science. Statisticians also might not be fluent with the computational tools of computer science.

"We're also bringing together people who have data problems—data they want to extract knowledge from," LeCun says. The center will be housed at the university's Courant Institute of Mathematical Sciences and will strive to collaborate with companies such as Google and Facebook that want to leverage information from big data as part of their businesses.

Smaller companies with access to large amounts of data may also want to collaborate with the center, according to LeCun. "Sometimes they have data that can be pooled with other people in their industries," he says.

Why the business interest? Data scientists can help e-commerce companies better understand and target their customers. "If people are buying stuff from your website, you want to propose things they are likely to buy," LeCun says. "There are tools to do this at a simple level, [but] to do this to scale is kind of difficult."

For example, pharmaceutical companies, he says, with vast amounts of data on clinical trials and health insurance companies with information on patients could use data science to help develop suggestions on treatment or courses of action. Utility companies might use big data to better predict system failures or increases in power consumption.

NYU wants to establish partnerships with companies to serve as mentors for students in the program, who must complete capstone projects at the end of their studies. Each project would typically be proposed by a partner company and partially advised by someone from that company. "Students will be confronted with the real world, and the company gets a preview of the students' skills—and perhaps some useful developments on their problem," LeCun says.

Companies might send scientists in residence to work with researchers and faculty at the center to gain new expertise. The companies would also be able take on students as summer interns. And there may be opportunities for companies to fund research projects they are interested in.

The data science center will start taking applications this month, with the master's program expected to launch in the fall with about 25 students. LeCun says 50 to 60 students will annually be accepted going forward into the two-year master's program. Plans for a doctoral program are also in the works.

Thursday, February 28, 2013

Privacy and WEF: Use, Not Collection, Should Be Focus of Data Rules, Report Says

Steve Loh, The New York Times, February 28, 2013

Personal data is a valuable asset that ought to be put to work.

Fluid data markets will benefit economies, societies and individuals.

Privacy rules should focus on how data is used rather than on the widespread collection of personal data.

That is the gist of a new report from World Economic Forum's Personal Data project, "Unlocking the Value of Personal Data: From Collection to Usage."

The modern digital world, with its explosion of data, has made the traditional approach to privacy based on "notice and consent" typically between two parties — a marketer and a consumer — obsolete, in the view of the report's authors.

"The technology has overrun the classical model," said Craig Mundie, a senior adviser to Microsoft's chief executive, Steven A. Ballmer.

Mr. Mundie was on the five-member steering board for the report. All five people represent corporations that stand to gain from tapping personal data.

Privacy advocates and regulators in Europe and the United States have been reluctant to give up on efforts to control the collection of data. Their concern is that once personal data is collected, its use is very difficult to monitor and control. Information brokers that consumers never see — and few know about — market personal data to advertisers, retailers, financial institutions and others. That problem prompted the Federal Trade Commission in a report last year to recommend that Congress enact legislation "to provide greater transparency for, and control over, the practices of information brokers."

But while recognizing the privacy challenges, the companies participating in the World Economic Forum project say what was needed was a careful balance. In a blog post on Wednesday, Raymond J. Baxter, a senior vice president of Kaiser Permanente, a major health care provider and insurer, emphasized the value of personal data, when used properly. He cited Kaiser's use of personal medical data for research.

For example, mining family data and outcomes over years, Kaiser scientists found that the children of women who took anti-depressant drugs while pregnant had more than twice the risk of developing autism disorders. "By discovering this correlation and leveraging this data in new ways, lives are improved," Mr. Baxter wrote.
According to Mr. Mundie of Microsoft, technology can help strike the right balance between individuals' concerns about privacy and the benefit of a fluid market in personal data. He said independent organizations, most likely nonprofits, would develop automated privacy preference services that individuals could subscribe to. A person would check off what he or she wanted his data to be used for and not. Those preferences, he explained, would then be encoded as software tags that traveled with the person's data.

Those preferences, Mr. Mundie added, could vary depending on context. For example, a person might say he or she did not want personal medical data shared beyond a family doctor and one or two specialists — unless the person was taken to an emergency ward.

"You can intelligently use computing technology to provide the benefits and curtail abuse," Mr. Mundie said.

Monday, February 25, 2013

The Promise and Peril of the 'Data-Driven Society'


A small group of academics, business executives and journalists gathered at the M.I.T. Media Lab last Thursday, and the purpose was to toss out ideas and discuss the concept of "Data-Driven Societies." A daunting topic, ambitious and vague at once, it seems.

Up to now, the focus on the power and implications of Big Data technology has been involved social media, business decision-making and online privacy. Those are big subjects in their own right. So it's not surprising that the notion of a data-driven society has not been much considered.

But someone who has was host of the meeting: Alex Pentland, a computational social scientist at the Media Lab. He put his intellectual stake in the ground last year in a presentation posted on Edge.org, "Reinventing Society in the Wake of Big Data."

Mr. Pentland's starting point is that the most important data that is becoming available on a vast new scale is information about people's behavior. For example, he cites location data from cellphones and evermore consumption data as people increasingly use credit cards for even the smallest purchases. He distinguishes this behavioral data from less-telling data — about people's beliefs like Facebook communications or Google searches.

The fine-grained behavioral data, according to Mr. Pentland, opens the way to changing how we think about society and how a society is governed. Adam Smith and Karl Marx, he explains, thought about markets and classes, respectively. "But those are aggregates," he said. "They're averages."

Yet now, Mr. Pentland says, it becomes possible to track social phenomena down to the individual level and the social and economic connections among individuals. The ability to monitor these "micro-patterns," Mr. Pentland said, means "we're entering a new era of social physics."

What might that mean in practice? Reed Hundt, the chairman of the Federal Communications Commission in the Clinton administration, observed at the meeting that Big Data played a major role in the last election — a reference to the Obama campaign's deft use of data analysis to identify potential Obama voters and encourage them to cast their ballots.

"You get elected with Big Data, but you govern without it," Mr. Hundt said. "How much sense does that make?"

Mr. Hundt, chief executive for the Coalition for Green Capital, a nonprofit organization, pointed to the waste in a range of government incentive and benefit programs, from tax credits for solar panels to Social Security, that results from the across-the-board approach — or policy by averages, as Mr. Pentland might put it.

Instead, a data-driven approach to solar-energy incentives would concentrate government incentives to where the payoff is greatest in terms of efficiently generating alternative energy — larger buildings with a lot of roof space instead of small houses, Mr. Hundt said. A by-the-data model for benefits programs, he added, would suggest means-testing Social Security payments as well as adjusting payments locally for differences in costs of living.

"So all men are created equal, but are subsidized individually," Mr. Hundt quipped.

Intriguing, and perhaps wise policy, but it would also seem to be a redefinition of fairness as it applies to broad benefit programs, like Social Security and Medicare, which typically make standard payments and avoid means testing.

What are the chances such a data-driven course would be politically acceptable? If the data points the way to greater efficiency, why not, Mr. Hundt replied. After all, he said, a major role of government is to transfer income to people who would benefit most — better data, closely analyzed, means government can perform that role more effectively, Mr. Hundt said.

An underlying assumption of tilting toward a data-driven society is that, as one participant put it, "information over time wins out." That is, data will change attitudes and policy, combating bias and causing policy-making to be more of a science. To data optimists, then, the endless political squabbling and stalemate in Washington points to all the room there is for improvement.

In a Big Data world, the data-mining for patterns and insights to guide policy will be done automatically — by software algorithms. Of course, algorithms are created by people and they contain inferences and assumptions coded in. Those coded-in values shape the output — computer-generated predictions, recommendations and simulations.

That raises question of the human design and control of the computerized helpers in policy-making, as in other realms of decision-making. "At some point, you're in the hands of the algorithm," observed John Henry Clippinger, chief executive of the Institute for Data Driven Design, a nonprofit research and educational organization. "You're whistling in the dark if you don't think that day is coming."

Sunday, February 24, 2013

CUSP in NYT: SimCity, for Real: Measuring an Untidy Metropolis

Steve Lohr, The New York Times, February 23, 2013

THE notion of a "science of cities" seems contradictory. Science is a realm of grand theory and precise measurement, while cities are messy agglomerations of people and human foible. But science is precisely the ambition of New York University's Center for Urban Science and Progress. Founded last year, the center has been getting under way in recent weeks, moving into new office space and firing off its first project proposal to the National Science Foundation.

The center's director is Steven E. Koonin, a Brooklyn native and graduate of Stuyvesant High School, who came to N.Y.U. after a stint in the Obama administration as the under secretary for science in the Department of Energy. He is both a theoretical physicist and science policy expert. The center shouldn't lack for intellectual rigor.

The initiative at N.Y.U. is part of a broader trend: the global drive to apply modern sensor, computing and data-sifting technologies to urban environments, in what has become known as "smart city" technology. The goals are big gains in efficiency and quality of life by using digital technology to better manage traffic and curb the consumption of water and electricity, for example. By some estimates, water and electricity use can be cut by 30 to 50 percent over the course of a decade.

Cities from Stockholm to Singapore are deep into smart city projects. The market looms as big, lucrative business for technology companies. "The Smart City movement," according to a report this month from IDC, a technology research firm, "is emerging and growing as a significant force of innovation and investment at all levels of government." The N.Y.U. center's partners include technology companies like I.B.M., Cisco Systems and Xerox, as well as universities and the New York City government.

City governments, like other institutions, have collected data for years to try to become more efficient. There have been some notable achievements, like CompStat, the New York Police Department's system for identifying crime patterns, introduced in the mid-1990s and later widely adopted elsewhere.

What is different today, says Dr. Koonin, is that digital technologies — sensors, wireless communication, storage and clever software algorithms — are advancing so rapidly that it is becoming possible to see and measure activities in an urban environment as never before.

"We can build an observatory to be able to see the pulse of the city in detail and as a whole," Dr. Koonin explains.

Dr. Koonin's digital "observatory" of urban life raises questions about privacy. He is keenly aware of that issue, and vows that the center is engaged in science rather than surveillance. For example, individuals' names or tax identification numbers would be stripped from personal records.

The collected data, he says, will be the raw material for modeling outcomes — say, the steps required to reduce electricity consumption in a high-rise office building or in an individual apartment. Those modeled predictions, he adds, can guide policy or inform citizens.

"I'd like to create SimCity for real," Dr. Koonin says, referring to the classic computer simulation game.
To help, Dr. Koonin is forging partnerships with government laboratories to tap their expertise in building complex computer simulations, like climate models for weather prediction.

The path to SimCity will come step by step, through tackling specific projects. The first one is a program to monitor and analyze noise. The largest single cause of complaints to New York's 311 phone and online service is noise. It is a quality-of-life issue, Dr. Koonin says, and one related to health, especially when noise disrupts sleep.

The 10-member project team includes music professors, computer scientists and graduate students. The group will use the city's 311 data, but also plans to employ wireless sensors — tiny ones outside windows, noise meters on traffic lights and street corners, perhaps a smartphone app for crowdsourced data gathering. To inform policy choices, data on noise limits for vehicles and muffler costs might be added to the street-level noise readings. Then, computer simulations could predict the likely effect of enforcement steps, charges or incentives to buy properly working mufflers for vehicles without them.

The project, Dr. Koonin says, might also pull in data on traffic flows, garbage pickup times and building classifications. For example, he says, a 2 a.m. garbage pickup could be routed to a neighborhood with little residential housing.

The hope, he says, is that a problem many people view as an inevitable, if grating fact of urban life can be made less severe. "It's the beginning of what we want to do," he explains.

Another project on the drawing board is technology for capturing thermal images of buildings across much of the city, as a starting point for research on energy use.

The center will focus its research and resources on one city — New York, as "a living laboratory."

That may give the center a leg up, since New York, under Mayor Michael R. Bloomberg, is at the forefront of using data to guide operations. In 2010, the city even set up a team of data scientists for special projects in the mayor's office.

ONE problem the team tackled was illegal conversions, landlords packing far more people into an apartment building or house than its zoning permits. These locations are fire hazards. Data from 19 agencies — including late tax payments, repair permits, foreclosure records and ages of the buildings — were mined to predict where to send the city's 200 building inspectors, who field more than 20,000 complaints a year.
Inspectors responding to complaints usually find high-risk conditions 13 percent of the time. Guided by data predictions, inspectors greatly improved their odds when pursuing complaint reports, finding those risky conditions 70 percent of the time, says Michael P. Flowers, analytics director in the mayor's office.

The city government is committed to giving the N.Y.U. center access to all its public data. That is a rich asset not only for research, but also for its potential to change government operations and public behavior. In many "smart city" projects, "the single biggest impact is transparency — the effect of measurement and communicating the data," observes Jonathan R. Woetzel, a director of McKinsey & Company in Shanghai, who heads the firm's consulting work with cities.

Communicating effectively with data, experts say, requires skills beyond technology. Jurij R. Paraszczak, director of smarter cities research at I.B.M., pointed to a water-management pilot study in Dubuque, Iowa, in which 150 households were equipped with sensors to measure and analyze their water use. They had the data, but the households were also grouped into teams for an informal competition. Water use dropped by 7 percent in two months.

"People live in cities," Dr. Paraszczak says. "So much of the equation is not just the data but how you encourage people to change their behavior."

The social ingredients of motivation, habit and incentives, according to Dr. Koonin, will be part of the research agenda at the N.Y.U. center. "The approach we're taking here is from sensors to sociologists. This has got to be science with a social dimension."

Monday, January 28, 2013

GE to IBM: Watch your Data, We Are Coming


General Electric, the massive industrial conglomerate, will not be content to let IT leaders like IBM and Google hog all the glory in the internet of things era.


It sure looks like General Electric — the conglomerate that builds stuff ranging from appliances to jet engines — is spending a ton of time and resources to boost its profile in high (as opposed to “low”) tech. In fact it looks like it’s waging a massive PR campaign to show that it is not some grimy industrial relic but a force at the cutting edge of big data and “the internet of things.” If you don’t believe it, just download its November report on the industrial internet, which we covered here.

The latest evidence of this push? An interview with William Ruh, VP of software for GE Research, in ComputerWeekly.com. In the piece, Ruh appeared to take a veiled swipe IBM — which loves to portray itself as the thought leader in bleeding-edge tech and the kingpin in tech patents. (For the record, in 2012 GE came in ninth in patents with a total of 1,652 compared to IBM’s 6,478 — but who’s counting?)

Ruh said the airline industry has gathered tons of data about how jet engines have performed over the past two decades and that historical data should help guide predictive maintenance going forward. Ruh told ComputerWeekly:

“In emerging markets, we are seeing dirt and sandy environments … How are these affecting aero engines? [Business intelligence] cannot answer this. Nor can a supercomputer … Watson cannot tell me when this machine part will break.”

Watson is IBM’s much-hyped computer that boasts human-like thought processes and beat the human champion in Jeopardy a few years back.


GE is banking on the growing acknowledgement that machine data — information generated and collected by the types of industrial gear it makes — gives it an entry into the booming world of big data. That’s probably why GE CEO Jeff Immelt has been cropping up in a lot of interesting venues, including in an interview with Om Malik last month. And why GE came to San Francisco to announce its “Industrial Internet Quests” and tap into the wealth of software and data expertise there. As my colleague Katie Fehrenbacher put it at the time, the quest “calls on developers, data scientists and designers to make algorithms and applications that can increase productivity for the health and aviation sectors” — all sectors where GE plays.
It may be easy for folks in the valley to forget that GE has thousands of its own software developers on staff and builds sophisticated medical imaging and other high-tech gear: it does have credibility. And, at a time when the emphasis on making and building actual products is more valued, GE has lessons to teach.
The conglomerate obviously wants to be seen as a leader in this realm and won’t be content to let the likes of IBM hog all the glory in the internet of things era. After all, it builds an awful lot of those “things.”



Bill gates on Metrics and Big Data


From the fight against polio to fixing education, what's missing is often good measurement and a commitment to follow the data. We can do better. We have the tools at hand.

Bill Gates, The Wall Street Journal, January 25, 2013

By custom, many Ethiopian parents won't name a child for weeks, in case the baby dies. Sebsebila Nassir, pictured above with a health worker, named her newborn daughter Amira—'princess' in Arabic—on her immunization card the day she was born.

We can learn a lot about improving the 21st-century world from an icon of the industrial era: the steam engine.

Harnessing steam power required many innovations, as William Rosen chronicles in the book "The Most Powerful Idea in the World." Among the most important were a new way to measure the energy output of engines and a micrometer dubbed the "Lord Chancellor" that could gauge tiny distances.

Such measuring tools, Mr. Rosen writes, allowed inventors to see if their incremental design changes led to the improvements—such as higher power and less coal consumption—needed to build better engines. There's a larger lesson here: Without feedback from precise measurement, Mr. Rosen writes, invention is "doomed to be rare and erratic." With it, invention becomes "commonplace."

In the past year, I have been struck by how important measurement is to improving the human condition. You can achieve incredible progress if you set a clear goal and find a measure that will drive progress toward that goal—in a feedback loop similar to the one Mr. Rosen describes.

This may seem basic, but it is amazing how often it is not done and how hard it is to get right. Historically, foreign aid has been measured in terms of the total amount of money invested—and during the Cold War, by whether a country stayed on our side—but not by how well it performed in actually helping people. Closer to home, despite innovation in measuring teacher performance world-wide, more than 90% of educators in the U.S. still get zero feedback on how to improve.

An innovation—whether it's a new vaccine or an improved seed—can't have an impact unless it reaches the people who will benefit from it. We need innovations in measurement to find new, effective ways to deliver those tools and services to the clinics, family farms and classrooms that need them.

I've found many examples of how measurement is making a difference over the past year—from a school in Colorado to a health post in rural Ethiopia. Our foundation is supporting these efforts. But we and others need to do more. As budgets tighten for governments and foundations world-wide, we all need to take the lesson of the steam engine to heart and adapt it to solving the world's biggest problems.

One of the greatest successes in terms of using measurement to drive global change has been an agreement signed in 2000 by the United Nations. The Millennium Development Goals, supported by 189 nations, set 2015 as a deadline for making specific percentage improvements across a set of crucial areas—such as health, education and basic income. Many people assumed the pact would be filed away and forgotten like so many U.N. and government pronouncements. The decades before had brought many well-meaning declarations to combat problems from nutrition to human rights, but most lacked a road map for measuring progress. However, the Millennium goals were backed by a broad consensus, were clear and concrete, and brought focus to the highest priorities.

When Ethiopia signed on to the Millennium goals in 2000, the country put hard numbers to its ambition to bring primary health care to all of its citizens. The concrete goal of reducing child mortality by two-thirds created a clear target by which to measure success or failure. Ethiopia's commitment attracted a surge of donor money toward improving the country's primary health-care services.

With help from the Indian state of Kerala, which had built a successful network of community health-care posts, Ethiopia launched its own program in 2004 and today has more than 15,000 health posts staffed by 34,000 workers. (This is one of the greatest benefits of measurement—the ability it gives government leaders to make comparisons across countries and then learn from the best.)

Last March, I visited the Germana Gale Health Post in the Dalocha region of Ethiopia, where I saw charts of immunizations, malaria cases and other data plastered to its walls. This information goes into a system—part paper-based and part computerized—that helps government officials see where things are working and to take action in places where they aren't. In recent years, data from the field have helped the government respond more quickly to outbreaks of malaria and measles. Perhaps even more important, the government previously didn't have any official record of a child's birth or death in rural Ethiopia. It now tracks those metrics closely.

The health workers provide most services at the posts, though they also visit the homes of pregnant women and sick people. They ensure that each home has access to a bed net to protect the family from malaria, a pit toilet, first-aid training and other basic health and safety practices. All these interventions are quite simple, yet they've dramatically improved the lives of people in this country.

Consider the story of one young mother in Dalocha. Sebsebila Nassir was born in 1990, when about 20% of all children in Ethiopia did not survive to see their fifth birthdays. Two of Sebsebila's six siblings died as infants. But when a health post opened its doors in Dalocha, life started to change. Last year when Sebsebila became pregnant, she received regular checkups. On Nov. 28, Sebsebila traveled to a health center where a midwife was at her bedside during her seven-hour labor. Shortly after her daughter was born, a health worker gave the baby vaccines against polio and tuberculosis.

According to Ethiopian custom, parents wait to name a baby because children often die in the first weeks of life. When Sebsebila's first daughter was born three years ago, she followed tradition and waited a month to bestow a name. This time, with more confidence in her new baby's chances of survival, Sebsebila put "Amira"—"princess" in Arabic—in the blank at the top of her daughter's immunization card on the day she was born. Sebsebila isn't alone: Many parents in Ethiopia now have the confidence to do the same.

Ethiopia has lowered child mortality more than 60% since 1990, putting the country on track to achieve the Millennium goal of lowering child mortality two-thirds by 2015, compared with 1990. Though the world won't quite meet the goal, we've still made great progress: The number of children under 5 years old who die world-wide fell to 6.9 million in 2011, down from 12 million in 1990 (despite a growing global population).
Another story of success driven by better measurement is polio. Starting in 1988, global health organizations (along with many countries) established a goal of eradicating polio, which focused political will and opened purse strings to pay for large-scale immunization campaigns. By 2000, the virus had nearly been wiped out; there are now fewer than 1,000 cases world-wide.

But getting rid of the very last cases is the hardest part. In order to stop the spread of infections, health workers have to vaccinate nearly all children under the age of 5 multiple times a year in polio-affected countries. There are now just three countries that have not eliminated polio: Nigeria, Pakistan and Afghanistan. I visited northern Nigeria four years ago to try to understand why eradication is so difficult there. I saw that routine public health services were failing: Fewer than half the kids were getting vaccines regularly. One huge problem was that many small settlements in the region were missing from vaccinators' hand-drawn maps and lists documenting the locations of villages and numbers of children.

To fix this, the polio workers walked through all high-risk areas in the northern part of the country, which enabled them to add 3,000 previously overlooked communities to the immunization campaigns. The program is also using high-resolution satellite images to create even more detailed maps. As a result, managers can now allocate vaccinators efficiently.

What's more, the program is piloting the use of phones equipped with a GPS application for the vaccinators. Tracks are downloaded from the phone at the end of the day so managers can see the route the vaccinators followed and compare it to the route they were assigned. This helps ensure that areas that were missed can be revisited.

I believe these kinds of measurement systems will help us to finish the job of polio eradication within the next six years. And those systems can be used to help expand routine vaccination and other health activities, which means the legacy of polio eradication will live beyond the disease itself.

Another place where measurement is starting to lead to vast improvements is in education.

In October, Melinda and I sat among two dozen 12th-graders at Eagle Valley High School near Vail, Colo. Mary Ann Stavney, a language-arts teacher, was leading a lesson on how to write narrative nonfiction pieces. She engaged her students, walking among them and eliciting great participation. We could see why Mary Ann is a master teacher, a distinction given to the school's best teachers and an important component of a teacher-evaluation system in Eagle County.

Ms. Stavney's work as a master teacher is informed by a three-year project our foundation funded to better understand how to build an evaluation and feedback system for educators. Drawing input from 3,000 classroom teachers, the project highlighted several measures that schools should use to assess teacher performance, including test data, student surveys and assessments by trained evaluators. Over the course of a school year, each of Eagle County's 470 teachers is evaluated three times and is observed in class at least nine times by master teachers, their principal and peers called mentor teachers.

The Eagle County evaluations are used to give a teacher not only a score but also specific feedback on areas to improve and ways to build on their strengths. In addition to one-on-one coaching, mentors and masters lead weekly group meetings in which teachers collaborate to spread their skills. Teachers are eligible for annual salary increases and bonuses based on the classroom observations and student achievement.
The program faces challenges from tightening budgets, but Eagle County so far has been able to keep its evaluation and support system intact—likely one reason why student test scores have improved in Eagle County over the past five years.

I think the most critical change we can make in U.S. K–12 education, with America lagging countries in Asia and Northern Europe when it comes to turning out top students, is to create teacher-feedback systems that are properly funded, high quality and trusted by teachers.

And there are plenty of other areas where our ability to measure can improve people's lives in powerful ways—areas where we are falling short, unnecessarily.

In poor countries, we still need better ways to measure the effectiveness of the many government workers providing health services. They are the crucial link bringing tools such as vaccines and education to the people who need them most. How well trained are they? Are they showing up to work? How can measurement enable them to perform their jobs better?

In the U.S., we should be measuring the value being added by colleges. Currently, college rankings are focused on inputs—the scores and quality of students entering college—and on judgments and prejudices about a school's "reputation." Students would be better served by measures of which colleges were best preparing their graduates for the job market. They then could know where they would get the most for their tuition money.

In agriculture, creating a global productivity target would help countries focus on a key but neglected area: the efficiency and output of hundreds of millions of small farmers who live in poverty. It would go a long way toward reducing poverty if we had public scorecards showing how developing-country governments, donors and others are helping those farmers.

And if I could wave a wand, I'd love to have a way to measure how exposure to risks like disease, infection, malnutrition and problem pregnancies impact children's potential—their ability to learn and contribute to society. Measuring that could help us quantify the broader impact of those risks and help us tackle them.
The lives of the poorest have improved more rapidly in the past 15 years than ever before. And I am optimistic that we will do even better in the next 15 years. The process I have described—setting clear goals, choosing an approach, measuring results, and then using those measurements to continually refine our approach—helps us to deliver tools and services to everybody who will benefit, be they students in the U.S. or mothers in Africa. Following the path of the steam engine long ago, thanks to measurement, progress isn't "doomed to be rare and erratic." We can, in fact, make it commonplace.

Friday, January 25, 2013

Has Big Data Reached Its Moment of Disillusionment?


Aril Hasseldahl, The Wall Street Journal, January 24, 2013


Last year was a year when the phrase “Big Data” was all over the place. Dig through the troves of data your business generates, the thinking goes, and some useful business intelligence falls out. That, at least, is the idea, and there are numerous companies — some startups, some big and established players — trying to build business plans on different aspects of that idea.

But in the course of being reduced to a simple buzzy phrase, Big Data as a concept implies some expectations, some realistic, some undoubtedly not. The research house Gartner has a phrase for this tendency as well: The Hype Cycle.

The Hype Cycle goes like this: A new technology that promises to fundamentally “change everything” gets talked up incessantly in the press and at industry events and often also in research reports. At some point the chatter peaks, and expectations reach a fever pitch. Soon, maybe a year or two after it all started to build and some money has been spent and everything that was supposed to have changed for the better actually hasn’t, the narrative focus turns negative. What seemed so brilliant and earthshaking 18 months ago, seems in restrospect to have been an ill-advised waste of time, money and attention.

This is what Gartner calls the “Trough of Disillusionment” phase of the Hype Cycle.

Gartner analyst Svetlana Sicular argues in a blog post that Big Data may have reached that point. She has been “hearing from people in the center of the Hadoop movement,” the open-source technology central to companies like Cloudera, Hortonworks and MapR. She also presents a video of a Hadoop gathering called Elephant Riders where reps from these three companies are debating its current state. (It’s about 90 minutes and if you’re so inclined, you can see it here.)

By Gartner’s standards, the trough of disillusionment may indeed have arrived, though you certainly wouldn’t be able to tell from the level of investment interest in companies like Cloudera, which late last year raised a massive $65 million round of funding.

One source of that disillusionment, she writes, is that companies are struggling with a basic problem: What questions do you attempt to answer with your data in the first place? “Several days ago, a financial industry client told me that framing a right question to express a game-changing idea is extremely challenging,” Sicular wrote. “First, selecting a question from multiple candidates; second, breaking it down to many sub-questions; and, third, answering even one of them reliably. It is hard.”

Hadoop doesn’t exactly make that process any easier. Once you’ve decided to use it, getting anything useful out of it requires some pretty specialized knowledge and training, and finding the right people to do that isn’t easy. But the industry is beginning to respond to that need: Startups like Mortar Data have sought to make Hadoop more readily accessible to mainstream programmers, while another called Trifacta makes the resulting data easier to manipulate.

And versions of Hadoop itself are getting incrementally easier to work with. Hortonworks, for example, recently released HDP 1.2, a new version of the open-source platform, but also Sandbox, a set of training tools that lets developers play around with Hadoop and get a feel for its use.

I talked with Hortonworks CEO Rob Bearden recently, and he said that, in 2011, companies had no idea what Hadoop could be used for, then spent 2012 experimenting with it, and now want to get some real-world value out of it in 2013. “This year, all the technology is coming together in a way that is consumable,” he said. “In the last quarter of last year we saw a lot of interesting production environments. Now the objectives are becoming clear for getting useful in 2013.”

The next milestone in the Hype Cycle, Sicular writes, is negative press. Eventually it’s followed by a period called the “Slope of Enlightenment,” and finally the “Plateau of Productivity.” It’s nice to know there could be a positive conclusion to all this somewhere down the road.