Tuesday, April 30, 2013

Measuring the Benefits of Tech Tools

April 30, 2013

Measuring the Benefits of Tech Tools

By EDUARDO PORTER
http://www.nytimes.com/2013/05/01/business/statistics-miss-the-benefits-of-technology.html?pagewanted=all&_r=0&pagewanted=print

When I was a young reporter we could not afford cellphones. I remember waiting in line for a pay phone in downtown Mexico City one afternoon to call in the news about the auction to privatize the phone company Telmex, driving those behind me crazy while a copy editor on the other end patiently typed it into those glowing green letters of an earlier information age.

I traveled to Japan with a TRS-80 portable computer, which ran on AA batteries and had plastic cups to put over the phone receiver. It transmitted copy at the blistering speed of 300 bits per second. And I wrote about Mexico’s tequila crisis of 1994 without the benefit of a full set of Mexican financial statistics a few clicks away.

From my perspective, the evolution of the tools of journalism between then and now has been nothing less than breathtaking.

Articles are more thorough — informed by complementary data and analysis, enriched with links to things like interactive charts, videos and slide shows. They get to readers much more quickly. Most important, they reach many more of them.

For all its financial troubles, never has The New York Times been read by more people: 44 million unique viewers online in the United States every month. Yet if you were to rummage through American economic statistics you would find little evidence of journalism’s technological leaps. Measured by its contribution to gross domestic product, the most prominent indicator of the nation’s economic well-being, much of this new journalistic value enabled by information technology is not worth much.

This is true not only of journalism. The failure of I.T. to deliver measurable value has been a popular meme among economists for years. Back in 1987 Nobel laureate Robert Solow posed a now famous paradox: “We can see the computers everywhere except in the productivity statistics.”

The meme is back. The burst of productivity during the dot-com revolution of the 1990s gave skeptics pause. But as productivity has slowed substantially in recent years, doubts have re-emerged about whether information technology can power economic growth like the steam engine and the internal combustion engine did in the past.

Last year, Robert J. Gordon of Northwestern University proposed that the I.T. revolution has pretty much exhausted its promise. He asked, provocatively: “Is U.S. economic growth over?” And he forecast stagnating living standards for the vast majority of Americans for decades to come.

Government statistics lend support to his skepticism: Value added by the information technology and communications industries — mostly hardware and software — has remained stuck at around 4 percent of the nation’s economic output for the last quarter century.

But these statistics do not tell the whole story. Because they miss much of what technology does for people’s well-being.

News organizations that take advantage of computers to let go of journalists, secretaries and research assistants will show up in the economic statistics as more productive, making more with less. But statisticians have no way to value more thorough, useful, fact-dense articles.

What’s more, gross domestic product only values the goods and services people pay for. It does not capture the value to consumers of economic improvements that are given away free. And until recently this is what news media organizations like The New York Times were doing online.

The Commerce Department is in the process of revising the way it measures G.D.P. to take better account of the contributions of investment in research and development and artistic creation. But even though the revisions to be announced this summer are expected to make the economy look bigger, they are not devised to capture the value that Americans get from digital technologies.

“G.D.P. is not a measure of how much value is produced for consumers,” said Erik Brynjolfsson of the Massachusetts Institute of Technology. “Everybody should recognize that G.D.P. is not a welfare metric.”

G.D.P. misses what Americans gain from sharing information on Facebook or finding information on Google or Wikipedia. It misses how dating sites reduce the cost and increase the odds of finding a mate. It misses the time saved by drivers who use Google Maps and the time gained by consumers from shopping online. Measured in money — what it contributes to G.D.P. — the recording industry is shrinking. Yet never before have Americans had access to so much music.

“Pretty much every human on earth can access all human knowledge,” said Hal Varian, Google’s chief economist. That may not be as impressive as civilization’s 60-year technological jump from the horse-drawn buggy to the man on the moon. But it’s probably useful, and it is mostly ignored by our measures of progress.

So how to measure the Internet’s contribution to our lives? A few years ago, Austan Goolsbee of the University of Chicago and Peter J. Klenow of Stanford gave it a shot. They estimated that the value consumers gained from the Internet amounted to about 2 percent of their income — an order of magnitude larger than what they spent to go online. Their trick was to measure not only how much money users spent on access but also how much of their leisure time they spent online.

The approach makes intuitive sense. Time is not only a valuable asset. Its relative value increases with economic development, as workers’ incomes grow while their allotment of time remains stubbornly fixed.

It puts the Internet in a different light. Earlier this year, Yan Chen, Grace YoungJoo Jeon and Yong-Mi Kim of the University of Michigan published the result of an experiment that found that people who had access to a search engine took 15 minutes less to answer a question than those without online access.

Using the average wage of $22 an hour as the value of workers’ time, and assuming that people who could answer more questions would ask more of them, Mr. Varian estimated that a search engine might be worth about $500 annually to the average worker. Across the working population, this would add up to $65 billion a year.

Last year, Mr. Brynjolfsson and an M.I.T. postgraduate, JooHee Oh, used a similar accounting technique to that of Mr. Goolsbee and Mr. Klenow and concluded that the consumer surplus from free online services — the value derived by consumers from the experience above what they paid for it — has been growing by $34 billion a year, on average, since 2002. If it were tacked on as “economic output,” it would add about 0.26 of a percentage point to annual G.D.P. growth.

Technoskeptics may scoff at these calculations. The Internet is hardly the first technology to offer consumers valuable free goods. The consumer surplus from television is about five times as large as that delivered by free stuff online, according to Mr. Brynjolfsson’s calculations.

Gross domestic product has always failed to capture many things — from the costs of pollution and traffic jams to the gains of unpaid household work. As the economist Paul Samuelson once pointed out, if a man married his maid, G.D.P. would decline.

Notably, it almost inevitably misses some economic gains from new technologies. The missed consumer surplus from the Internet may be no bigger than the unmeasured gains in the production, for example, of electric light.

But there is a case to be made that the unmeasured benefits from the Internet deserve more attention.

The amount of time Americans devote to the Internet has doubled in the last five years. Information, encoded in bits, is bound to become a larger and larger share of our economic output. Much of its value will be delivered to each additional consumer at a marginal cost of nearly zero.

“We know less about the sources of value in the economy than we did 25 years ago,” wrote Mr. Brynjolfsson and Adam Saunders of the University of British Columbia. If we really want to understand the impact of information technology on our future well-being, we first need to find a consistent way to measure it.

 

Sunday, April 28, 2013

How Big Data Is Playing Recruiter for Specialized Workers

April 27, 2013

How Big Data Is Playing Recruiter for Specialized Workers

By MATT RICHTEL
http://www.nytimes.com/2013/04/28/technology/how-big-data-is-playing-recruiter-for-specialized-workers.html?curator=MediaReDEF&emc=rss&partner=rss&_r=0&pagewanted=print#h[]

WHEN the e-mail came out of the blue last summer, offering a shot as a programmer at a San Francisco start-up, Jade Dominguez, 26, was living off credit card debt in a rental in South Pasadena, Calif., while he taught himself programming. He had been an average student in high school and hadn’t bothered with college, but someone, somewhere out there in the cloud, thought that he might be brilliant, or at least a diamond in the rough.

That someone was Luca Bonmassar. He had discovered Mr. Dominguez by using a technology that raises important questions about how people are recruited and hired, and whether great talent is being overlooked along the way. The concept is to focus less than recruiters might on traditional talent markers — a degree from M.I.T., a previous job at Google, a recommendation from a friend or colleague — and more on simple notions: How well does the person perform? What can the person do? And can it be quantified?

The technology is the product of Gild, the 18-month-old start-up company of which Mr. Bonmassar is a co-founder. His is one of a handful of young businesses aiming to automate the discovery of talented programmers — a group that is in enormous demand. These efforts fall in the category of Big Data, using computers to gather and crunch all kinds of information to perform many tasks, whether recommending books, putting targeted ads onto Web sites or predicting health care outcomes or stock prices.

Of late, growing numbers of academics and entrepreneurs are applying Big Data to human resources and the search for talent, creating a field called work-force science. Gild is trying to see whether these technologies can also be used to predict how well a programmer will perform in a job. The company scours the Internet for clues: Is his or her code well-regarded by other programmers? Does it get reused? How does the programmer communicate ideas? How does he or she relate on social media sites?

Gild’s method is very much in its infancy, an unproven twinkle of an idea. There is healthy skepticism about this idea, but also excitement, especially in industries where good talent can be hard to find.

The company expects to have about $2 million to $3 million in revenue this year and has raised around $10 million, including a chunk from Mark Kvamme, a venture capitalist who invested early in LinkedIn. And Gild has big-name customers testing or using its technology to recruit, including Facebook, Amazon, Wal-Mart Stores, Google and Twitter.

Companies use Gild to mine for new candidates and to assess candidates they are already considering. Gild itself uses the technology, which was how the company, desperate for programming talent and unable to match the salaries offered by bigger tech concerns, found this guy named Jade outside of Los Angeles. Its algorithm had determined that he had the highest programming score in Southern California, a total that almost no one achieves. It was 100.

Who was Jade? Could he help the company? What does his story tell us about modern-day recruiting and hiring, about the concept of meritocracy?

PEOPLE in Silicon Valley tend to embrace certain assumptions: Progress, efficiency and speed are good. Technology can solve most things. Change is inevitable; disruption is not to be feared. And, maybe more than anything else, merit will prevail.

But Vivienne Ming, who since late in 2012 has been the chief scientist at Gild, says she doesn’t think Silicon Valley is as merit-based as people imagine. She thinks that talented people are ignored, misjudged or fall through the cracks all the time. She holds that belief in part because she has had some experience of it.

Dr. Ming was born male, christened Evan Campbell Smith. He was a good student and a great athlete — holding records at his high school in track and field in the triple jump and long jump. But he always felt a disconnect with his body. After high school, Evan experienced a full-blown identity crisis. He flopped at college, kicked around jobs, contemplated suicide, hit the proverbial bottom. But rather than getting stuck there, he bounced. At 27, he returned to school, got an undergraduate degree in cognitive neuroscience from the University of California, San Diego, and went on to receive a Ph.D. at Carnegie Mellon in psychology and computational neuroscience.

During a fellowship at Stanford, he began gender transition, becoming, fully, Dr. Vivienne Ming in 2008.

As a woman, Dr. Ming started noticing that people treated her differently. There were small things that seemed innocuous, like men opening the door for her. There were also troubling things, like the fact that her students asked her fewer questions about math then they had when she was a man, or that she was invited to fewer social events — a baseball game, for instance — by male colleagues and business connections.

Bias often takes forms that people may not recognize. One study that Dr. Ming cites, by researchers at Yale, found that faculty members at research universities described female applicants for a manager position as significantly less competent than male applicants with identical qualifications. Another study, published by the National Bureau of Economic Research, found that people who sent in résumés with “black-sounding” names had a considerably harder time getting called back from employers than did people who sent in résumés showing equal qualifications but with “white-sounding” names.

Everybody can pretty much agree that gender, or how people look, or the sound of a last name, shouldn’t influence hiring decisions. But Dr. Ming takes the idea of meritocracy further. She suggests that shortcuts accepted as a good proxy for talent — like where you went to school or previously worked — can also shortchange talented people and, ultimately, employers. “The traditional markers people use for hiring can be wrong, profoundly wrong,” she said.

Dr. Ming’s answer to what she calls “so much wasted talent” is to build machines that try to eliminate human bias. It’s not that traditional pedigrees should be ignored, just balanced with what she considers more sophisticated measures. In all, Gild’s algorithm crunches thousands of bits of information in calculating around 300 larger variables about an individual: the sites where a person hangs out; the types of language, positive or negative, that he or she uses to describe technology of various kinds; self-reported skills on LinkedIn; the projects a person has worked on, and for how long; and, yes, where he or she went to school, in what major, and how that school was ranked that year by U.S. News & World Report.

“Let’s put everything in and let the data speak for itself,” Dr. Ming said of the algorithms she is now building for Gild.

Gild is not the only company now scouring for information. TalentBin, another San Francisco start-up firm, searches the Internet for talented programmers, trawling sites where they gather, collecting “data exhaust,” according to the company Web site, and creating lists of potential hires for employers. Another competitor is RemarkableHire, which assesses a person’s talents by looking at how his or her online contributions are rated by others.

And there’s Entelo, which tries to figure out who might be looking for a job before they even start their exploration. According to its Web site, the company uses more than 70 variables to find indications of possible career change, such as how someone presents herself on social sites. The Web site reads: “We crunch the data so you don’t have to.”

This application of Big Data to recruiting is “is absolutely worth a try,” said Susan Etlinger, an analyst of the data and analytics industries at the Altimeter Group. But she questioned whether an algorithm would be an improvement over what employers already do: gathering résumés, or referrals, and using traditional markers associated with success.

“The big hole is actual outcomes,” she said. “What I’m not buying yet is that probability equals actuality.”

Sean Gourley, co-founder and chief technology officer at Quid, a Big Data company, said that data trawling could inform recruiting and hiring, but only if used with an understanding of what the data can’t reveal. “Big Data has its own bias,” he said. “You measure what you can measure,” and “you’re denigrating what can’t be measured, like gut instinct, charisma.”

He added: “When you remove humans from complex decision-making, you can optimize the hell out of the algorithm, but at what cost?”

Dr. Ming doesn’t suggest eliminating human judgment, but she does think that the computer should lead the way, acting as an automated vacuum and filter for talent. The company has amassed a database of seven million programmers, ranking them based on what it calls a Gild score — a measure, the company says, of what a person can do. Ultimately, Dr. Ming wants to expand the algorithm so it can search for and assess other kinds of workers, like Web site designers, financial analysts and even sales people at, say, retail outlets.

“We did our own internal gold strike,” Dr. Ming said. “We found this kid in Los Angeles just kicking around his computer.”

She’s talking about Jade.

MR. DOMINGUEZ grew up in Los Angeles, the middle child of five. His mother took care of the household; his dad installed telecommunications equipment — a blue-collar guy who prized education.

But Jade had a rebellious streak. Halfway through high school, Mr. Dominguez, previously a straight-A student, began wondering whether going to school was more about satisfying requirements than real learning. “The value proposition is to go to school to get a good job,” he told me. “Philosophically, shouldn’t you go to school to learn?” His grades fell sharply, and he said he graduated from Alhambra High School in 2004 with less than a 3.0 grade-point average.

Not only did he reject college, he also wanted to prove that he could succeed wildly without it. He devoured books on entrepreneurship. He started a company that printed custom T-shirts, first from his house, then from a 1,000-square-foot warehouse space he rented. He decided that he needed a Web site, so he taught himself programming.

“I was out to prove myself on my own merit,” he said. He concedes that he might have taken it a little far. “It’s a little immature to be motivated by proving people wrong,” he said.

He got a tattoo on his arm in flowery script that read “Believe.” He sort of laughs about it now, though he still feels that he can accomplish what he puts his mind to. “It’s the great thing about code,” he said of computer language. “It’s largely merit-driven. It’s not about what you’ve studied. It’s about what you’ve shipped.”

When Gild went looking for talent, it assumed that the San Francisco and Silicon Valley areas would be picked over. So it ran its algorithm in Southern California and came up with a list of programmers. At the top was Mr. Dominguez, who had a very solid reputation on GitHub — a place where software developers gather to share code, exchange ideas and build reputations. Gild combs through GitHub and a handful of other sites, including Bitbucket and Google Code, looking for bright people in the field.

Mr. Dominguez had made quite a contribution. His code for Jekyll-Bootstrap, a function used in building Web sites, was reused by an impressive 1,267 other developers. His language and habits showed a passion for product development and several programming tools, like Rails and JavaScript, which were interesting to Gild. His blogs and posts on Twitter suggested that he was opinionated, something that the company wanted on its initial team.

A recruiter from Gild sent him an e-mail and had him come to San Francisco for an interview. The company founders met a charismatic, confident person — poised, articulate, thoughtful, with an easy smile, a tad rougher around the edges than other interview candidates, said Sheeroy Desai, Mr. Bonmassar’s co-founder at Gild and the company’s chief executive.

Mr. Dominguez wore a vibrant green hoodie to the interview. He asked pointed questions, like this one: Did the company worry that it would be perceived as violating privacy by scoring engineers without their knowledge? (It didn’t believe so, and he didn’t, either. Gild says it uses only publicly available information.)

They asked him some pointed but gentle questions, too, like whether he could work in a structured environment. He said he could. The company made Mr. Dominguez a job offer right away, and he accepted a position that pays around $115,000 a year.

“He’s a symbol of someone who is smart, highly motivated and yet, for whatever reason, wasn’t motivated in high school and didn’t see value in college,” Mr. Desai said.

Mr. Desai did go to college, at M.I.T., one of those schools that recruiters value so highly. It was there, he said, that he learned how to cope with pressure and to work with brilliant people and sometimes feel humbled. But while one’s work at school isn’t inconsequential, he said,it’s not the whole story.” He asserts that despite his degree in computer science, “I’m a terrible developer.”

David Lewin, a professor at the University of California, Los Angeles, and an expert in management of human resources, said that asking what someone could do was an important question, but so was asking whether the person could accomplish it with other people. Of all the efforts to predict whether someone will perform well in an organization, the most proven method, Dr. Lewin said, is a referral from someone already working there. Current employees know the culture, he said, and have their reputations and their work environment on the line. A recent study from the Yale School of Management that uses Big Data offers a refinement to the notion, finding that employee referrals are a great way to find good hires but that the method tends to work much better if the employee making the referral is highly productive.

For his part, Dr. Lewin is skeptical that an algorithm would be a good substitute for a good referral from a trusted employee.

One of Gild’s customers is Square, a San Francisco-based mobile payment system. Like many other high-tech companies, Square is aggressively hiring, and it’s finding the competition for great talent as intense as it was during the dot-com boom, according to Bryan Power, the company’s director of talent and a Silicon Valley veteran. Mr. Power says Gild offers a potential leg up in finding programmers who aren’t the obvious catches.

“Getting out of Stanford or Google is a very good proxy” for talent, Mr. Power said. “They have reputations for a reason.” But those prospects have many choices, and they might not choose Square. “We need more pools to draw from,” he said, “and that’s what Gild represents.”

Gild’s technology has turned up some prospects for Square, but hasn’t led directly to a hire. Mr. Power says the Gild algorithm provides a generalized programming score that is not as specific as Square needs for its job slots. “Gild has an opinion of who is good but it’s not that simple,” he said, adding that Square was talking to Gild about refining the model.

Despite the limited usefulness thus far, Mr. Power says that what Gild is doing is the start of something powerful. Today’s young engineers are posting much more of their work online, and doing open-source work, providing more data to mine in search of the diamonds. “It’s all about finding unrecognized talent,” he said.

MR. DOMINGUEZ has worked at Gild for eight months and has proved himself a talented programmer, Mr. Desai said. But he also said that Mr. Dominguez “sometimes struggles to work in a structured environment.” His co-workers try not to bug him when he’s sitting at his computer, locked into that work zone.

In meetings, Mr. Dominguez speaks his mind. He’s happier, he said, “as long as I can have a say in how the system is built,” or it’s just another system he would have to conform to. He bristles slightly at the growth of the company, which has expanded to 40 people from 10 in the last six months, adding layers of management and bureaucracy.

“The truth is that’s in my nature to do stuff in my own way; inevitably I want to start my own company,” he said, but he’s quick to add: “I do appreciate and the respect the opportunity the company’s given me because I think it’s very clear they hired me on merit. I will always appreciate that.”

Dr. Ming says the young man is both a great find and still an unknown. Of course, he is just a single example, one heralded by the company, but who cannot alone either validate or disprove the method.

“He’s got the lone-wolf thing going on,” Dr. Ming said. “It’s going well early but it could get tougher later on.”

 

Saturday, April 27, 2013

How data is changing the car game for Ford

How data is changing the car game for Ford

By Derrick Harris

http://gigaom.com/2013/04/26/how-data-is-changing-the-car-game-for-ford/

Summary:

The advent of big data is affecting Ford Motor Co. in some significant ways, from how it analyzes its supply chain to the features it puts into its cars.

tweet this

When most people think about how cars are built, they probably think about assembly lines, manufacturing robots, and batteries of safety and performance simulations on massive supercomputers. But at Ford, big data is having a significant impact on the parts and features of those cars before they’re ever part of a design file. From the cars in stock at the dealership to the performance of the engine in a rainstorm, big data is infiltrating nearly every aspect of the Ford experience and the company itself.

Obviously, data is nothing new to the automotive industry — companies have been trying to optimize supply chains and analyze sales numbers for decades — but the advent of big data, as well as related technlogies such as sensors and smartphones, is changing how companies are thinking about data. Ford isn’t alone in its quest to take advantage of these new technologies, either. For example, General Motors collects data from its OnStar system to help lower drivers’ insurance premiums, and also collects lots of data on its Chevrolet Volt electric car that it feeds to drivers via a mobile app. We recently noted how a luxury automobile company used big data software from Aster Data Systems to determine the relationships between malfunctions so it could provide a more thorough and beneficial service-department experience.

But in an industry notoriously unwilling to talk about information technology, Ford’s experiences might shed a lot on what other companies are thinking and doing, as well.

Building a better experience through data

According to John Ginder, manager for systems analytics with Ford Research & Innovation, the company has been doing advanced business modeling for about 20 years, but big data is something else. Today’s technologies are allowing Ford to handle larger, more-diverse datasets than ever before possible, and its efforts are already beginning to bear fruit in numerous places — including in the cars themselves.

The most obvious example of data influencing the driving experience might be the types of data car companies are actually giving back to drivers. At Ford, its Energi line of plug-in hybrid cars generate 25 gigabytes of data per hour that’s then processed and given back to drivers via a mobile app. It tells them about battery life, the nearest charging stations and other data about the vehicle’s performance.

The MyFord mobile app architecture.

Ginder said all that data is the result of a “convergence of need and opportunity.” The opportunity is a way to experiment with collecting and presenting vehicle data on a group of early adopters that’s probably more interested in this type of advanced technology. The need has to do with what Ginder calls “range anxiety” — when drivers are getting used to electric vehicles, they need reassurance they’re not going to run out juice.

However, Ginder said, the company is just scratching the surface of what’s possible, because there aren’t that many of the electric vehicles on the road yet. The goal is to better understand how drivers are using the vehicles and use that information to continuously improve the vehicles and the overall experience. Ford’s Super Duty line of pickup trucks also offers a “crew chief” package that lets bosses monitor the fuel consumption, engine performance and other data about their fleets of vehicles.

Mike Cavaretta, technical leader for predictive analytics and data mining with Ford Research & Innovation, added that Ford is really interested in collecting more data from more vehicles, but noted there’s also a privacy concern that could come into play. The potential of someone knowing where and how you’re driving might not appeal to the mainstream just yet (just look at all that data Tesla collects about its cars and can present if it really wants to), but as with the Energi, data does present some opportunities to improve the customer experience.

The test cars in Ford’s research labs are collecting about 250 gigabytes of data per hour from high-resolution cameras and an array of sensors, Cavaretta noted, and the company is trying to find out what data is most useful and how it might be rolled into production vehicles.

Building betters cars through data

Of course, sometimes the best data isn’t the stuff you see, but the stuff that just makes your car better. Cavaretta said Ford analyzes a lot of social media and other external data in order to figure out, for example, what customers are saying about their vehicles compared with other makes and what problems they’re having.

Opens with the touch of a foot. Source: Ford

In one recent case, the product development team was curious as to whether the Ford Escape sport-utility vehicle should have a standard liftgate (i.e., it opens manually and the rear window can flip open) or a power liftgate in which the glass and the gate are one piece. In the latter option, the gate opens automatically by tapping under the rear bumper with your foot, but the window doesn’t open at all. Regular surveys hadn’t addressed the question, so Cavaretta and his team took to social media, where people were actually talking about it quite a bit and seemed to heavily favor the power liftgate in most cases. It’s now a feature.

Back in 2004, Ford built a self-learning neural network system for its Aston Martin luxury brand that maintains proper engine function by recognizing engine misfires and particular driving conditions and adjusting warnings and performance accordingly.

Ginder said his team has been improving on that technology ever since and actually expanded its use into a system, called Smart Inventory Management System, that lets dealers ensure they have the optimal stock of vehicles and features on their lots. Historically, he said, some dealers were very sophisticated about inventory management, while others were more reactionary (“They just sold a red Mustang,” he joked, “so they think they need to go order another red Mustang.”) With SIMS, all sorts of data about vehicle sales and other locally relevant data from across the country is aggregated in Ford’s big data platform, and the neural network algorithms learn the current patterns so Ford can make better recommendations — whether or not dealers choose to heed the advice.

Selling big data internally

Cavaretta characterizes the division in which he and Ginder work as “an Ernst & Young, but just for Ford,” an internal consultancy (as opposed to Ford’s more-traditional research and development division) in charge of solving business problems via analytics. About 80 percent of those problems come directly from those lines of business, while about 20 percent are the research division’s own ideas. However, although he’s excited about how big data can help his team answer these questions in novel ways, it’s not always an easy sell with other parts of the company.

Mashing up data sources such as social and sales in order to find insights is a pretty easy sell, Cavaretta explained, but getting people to put sensors in everything and collect data every second or with every transaction can still be a bit challenging. In part, this is just a lingering effect of the constraints that legacy technologies imposed on the company. It wasn’t possible to store all this data, so people just got accustomed to the status quo of summarizing data hourly, for example.

Source: Ford

Now, however, he’s pushing them to “dial it down” and collect data at the lowest level possible and as often as possible. In manufacturing alone, he explained, there are between 20,000 and 25,000 parts in any given vehicle, and there’s a supply chain that spans from parts suppliers all the way up to dealerships. Getting a complete view of this process could help drive serious efficiencies and, Cavaretta said, “We don’t see anything but big data technologies that can get us there.”

Other areas where Ford is collecting, or wants to collect, more real-time data is from websites, call centers and the company’s credit-processing arm, he added.

Building big data internally

In order to accomplish their lofty goals, the Research & Innovation analytics team relies heavily on open source technologies, most prominently Hadoop. However, Cavaretta said, they’ve been experimenting with a variety of natural-language processing tools, too, and even did a proof-of-concept with SAP’s HANA in-memory analytic database. The NLP tools were first turned on text analysis of internal surveys and dealer network documents, but now are used pretty heavily on social media and other web data.

Their team has some systems numbering in the dozens of nodes in its own building, but on weekends it’s able to borrow high-performance computing cycles from Ford’s Numerically Intensive Computing Center next door in order to model recommendation engines and other tasks that demand serious computing power.

But as a part of a specialized research division, the work that Ginder, Cavaretta and their team do on everything from Hadoop to visualization with tools like Tableau isn’t automatically ready for primetime. In fact, Cavaretta said, it looks at “what’s the art of the possible” and tries to show the value of it. It’s like a vanguard, he added, going out and seeing what’s ahead and then reporting back.

At that point, projects are often handed off to Ford’s central IT team that actually puts the technologies into production. A system that took the research team weeks to deploy and start deriving insights from might take IT months to make production-ready. However, Ginder added, his team can’t just throw stuff over the wall and abandon it — it has to collaborate with the IT team and individual departments throughout the project’s lifecycle.

An important part of this cross-company relationship — and something many CIOs have likely heard before — is having data scientists on board that can see the world through the eyes of both technologists and businesspeople, two groups that often have different concerns and goals in mind. “We look for people who can bridge those worlds,” Ginder said. “It’s hard to find these people, but they’re hugely important to organizations.”

Feature image courtesy of Shutterstock user PhotoSmart.

 

Tuesday, April 23, 2013

The Rise of Big Data (Foreign Affairs)

The Rise of Big Data

How It's Changing the Way We Think About the World

By Kenneth Neil Cukier and Viktor Mayer-Schoenberger

May/June 2013

 

http://www.foreignaffairs.com/articles/139104/kenneth-neil-cukier-and-viktor-mayer-schoenberger/the-rise-of-big-data?page=show

Everyone knows that the Internet has changed how businesses operate, governments function, and people live. But a new, less visible technological trend is just as transformative: “big data.” Big data starts with the fact that there is a lot more information floating around these days than ever before, and it is being put to extraordinary new uses. Big data is distinct from the Internet, although the Web makes it much easier to collect and share data. Big data is about more than just communication: the idea is that we can learn from a large body of information things that we could not comprehend when we used only smaller amounts.

In the third century BC, the Library of Alexandria was believed to house the sum of human knowledge. Today, there is enough information in the world to give every person alive 320 times as much of it as historians think was stored in Alexandria’s entire collection -- an estimated 1,200 exabytes’ worth. If all this information were placed on CDs and they were stacked up, the CDs would form five separate piles that would all reach to the moon.

This explosion of data is relatively new. As recently as the year 2000, only one-quarter of all the world’s stored information was digital. The rest was preserved on paper, film, and other analog media. But because the amount of digital data expands so quickly -- doubling around every three years -- that situation was swiftly inverted. Today, less than two percent of all stored information is nondigital.

We can learn from a large body of information things that we could not comprehend when we used only smaller amounts.

Given this massive scale, it is tempting to understand big data solely in terms of size. But that would be misleading. Big data is also characterized by the ability to render into data many aspects of the world that have never been quantified before; call it “datafication.” For example, location has been datafied, first with the invention of longitude and latitude, and more recently with GPS satellite systems. Words are treated as data when computers mine centuries’ worth of books. Even friendships and “likes” are datafied, via Facebook.

This kind of data is being put to incredible new uses with the assistance of inexpensive computer memory, powerful processors, smart algorithms, clever software, and math that borrows from basic statistics. Instead of trying to “teach” a computer how to do things, such as drive a car or translate between languages, which artificial-intelligence experts have tried unsuccessfully to do for decades, the new approach is to feed enough data into a computer so that it can infer the probability that, say, a traffic light is green and not red or that, in a certain context, lumière is a more appropriate substitute for “light” than léger.

Using great volumes of information in this way requires three profound changes in how we approach data. The first is to collect and use a lot of data rather than settle for small amounts or samples, as statisticians have done for well over a century. The second is to shed our preference for highly curated and pristine data and instead accept messiness: in an increasing number of situations, a bit of inaccuracy can be tolerated, because the benefits of using vastly more data of variable quality outweigh the costs of using smaller amounts of very exact data. Third, in many instances, we will need to give up our quest to discover the cause of things, in return for accepting correlations. With big data, instead of trying to understand precisely why an engine breaks down or why a drug’s side effect disappears, researchers can instead collect and analyze massive quantities of information about such events and everything that is associated with them, looking for patterns that might help predict future occurrences. Big data helps answer what, not why, and often that’s good enough.

The Internet has reshaped how humanity communicates. Big data is different: it marks a transformation in how society processes information. In time, big data might change our way of thinking about the world. As we tap ever more data to understand events and make decisions, we are likely to discover that many aspects of life are probabilistic, rather than certain.

APPROACHING "N=ALL"

For most of history, people have worked with relatively small amounts of data because the tools for collecting, organizing, storing, and analyzing information were poor. People winnowed the information they relied on to the barest minimum so that they could examine it more easily. This was the genius of modern-day statistics, which first came to the fore in the late nineteenth century and enabled society to understand complex realities even when little data existed. Today, the technical environment has shifted 179 degrees. There still is, and always will be, a constraint on how much data we can manage, but it is far less limiting than it used to be and will become even less so as time goes on.

The way people handled the problem of capturing information in the past was through sampling. When collecting data was costly and processing it was difficult and time consuming, the sample was a savior. Modern sampling is based on the idea that, within a certain margin of error, one can infer something about the total population from a small subset, as long the sample is chosen at random. Hence, exit polls on election night query a randomly selected group of several hundred people to predict the voting behavior of an entire state. For straightforward questions, this process works well. But it falls apart when we want to drill down into subgroups within the sample. What if a pollster wants to know which candidate single women under 30 are most likely to vote for? How about university-educated, single Asian American women under 30? Suddenly, the random sample is largely useless, since there may be only a couple of people with those characteristics in the sample, too few to make a meaningful assessment of how the entire subpopulation will vote. But if we collect all the data -- “n = all,” to use the terminology of statistics -- the problem disappears.

This example raises another shortcoming of using some data rather than all of it. In the past, when people collected only a little data, they often had to decide at the outset what to collect and how it would be used. Today, when we gather all the data, we do not need to know beforehand what we plan to use it for. Of course, it might not always be possible to collect all the data, but it is getting much more feasible to capture vastly more of a phenomenon than simply a sample and to aim for all of it. Big data is a matter not just of creating somewhat larger samples but of harnessing as much of the existing data as possible about what is being studied. We still need statistics; we just no longer need to rely on small samples.

There is a tradeoff to make, however. When we increase the scale by orders of magnitude, we might have to give up on clean, carefully curated data and tolerate some messiness. This idea runs counter to how people have tried to work with data for centuries. Yet the obsession with accuracy and precision is in some ways an artifact of an information-constrained environment. When there was not that much data around, researchers had to make sure that the figures they bothered to collect were as exact as possible. Tapping vastly more data means that we can now allow some inaccuracies to slip in (provided the data set is not completely incorrect), in return for benefiting from the insights that a massive body of data provides.

Consider language translation. It might seem obvious that computers would translate well, since they can store lots of information and retrieve it quickly. But if one were to simply substitute words from a French-English dictionary, the translation would be atrocious. Language is complex. A breakthrough came in the 1990s, when IBM delved into statistical machine translation. It fed Canadian parliamentary transcripts in both French and English into a computer and programmed it to infer which word in one language is the best alternative for another. This process changed the task of translation into a giant problem of probability and math. But after this initial improvement, progress stalled.

Using big data will sometimes mean forgoing the quest for why in return for knowing what.

Then Google barged in. Instead of using a relatively small number of high-quality translations, the search giant harnessed more data, but from the less orderly Internet -- “data in the wild,” so to speak. Google inhaled translations from corporate websites, documents in every language from the European Union, even translations from its giant book-scanning project. Instead of millions of pages of texts, Google analyzed billions. The result is that its translations are quite good -- better than IBM’s were--and cover 65 languages. Large amounts of messy data trumped small amounts of cleaner data.

FROM CAUSATION TO CORRELATION

These two shifts in how we think about data -- from some to all and from clean to messy -- give rise to a third change: from causation to correlation. This represents a move away from always trying to understand the deeper reasons behind how the world works to simply learning about an association among phenomena and using that to get things done.

Of course, knowing the causes behind things is desirable. The problem is that causes are often extremely hard to figure out, and many times, when we think we have identified them, it is nothing more than a self-congratulatory illusion. Behavioral economics has shown that humans are conditioned to see causes even where none exist. So we need to be particularly on guard to prevent our cognitive biases from deluding us; sometimes, we just have to let the data speak.

Take UPS, the delivery company. It places sensors on vehicle parts to identify certain heat or vibrational patterns that in the past have been associated with failures in those parts. In this way, the company can predict a breakdown before it happens and replace the part when it is convenient, instead of on the side of the road. The data do not reveal the exact relationship between the heat or the vibrational patterns and the part’s failure. They do not tell UPS why the part is in trouble. But they reveal enough for the company to know what to do in the near term and guide its investigation into any underlying problem that might exist with the part in question or with the vehicle.

A similar approach is being used to treat breakdowns of the human machine. Researchers in Canada are developing a big-data approach to spot infections in premature babies before overt symptoms appear. By converting 16 vital signs, including heartbeat, blood pressure, respiration, and blood-oxygen levels, into an information flow of more than 1,000 data points per second, they have been able to find correlations between very minor changes and more serious problems. Eventually, this technique will enable doctors to act earlier to save lives. Over time, recording these observations might also allow doctors to understand what actually causes such problems. But when a newborn’s health is at risk, simply knowing that something is likely to occur can be far more important than understanding exactly why.

Medicine provides another good example of why, with big data, seeing correlations can be enormously valuable, even when the underlying causes remain obscure. In February 2009, Google created a stir in health-care circles. Researchers at the company published a paper in Nature that showed how it was possible to track outbreaks of the seasonal flu using nothing more than the archived records of Google searches. Google handles more than a billion searches in the United States every day and stores them all. The company took the 50 million most commonly searched terms between 2003 and 2008 and compared them against historical influenza data from the Centers for Disease Control and Prevention. The idea was to discover whether the incidence of certain searches coincided with outbreaks of the flu -- in other words, to see whether an increase in the frequency of certain Google searches conducted in a particular geographic area correlated with the CDC’s data on outbreaks of flu there. The CDC tracks actual patient visits to hospitals and clinics across the country, but the information it releases suffers from a reporting lag of a week or two -- an eternity in the case of a pandemic. Google’s system, by contrast, would work in near-real time.

Google did not presume to know which queries would prove to be the best indicators. Instead, it ran all the terms through an algorithm that ranked how well they correlated with flu outbreaks. Then, the system tried combining the terms to see if that improved the model. Finally, after running nearly half a billion calculations against the data, Google identified 45 terms -- words such as “headache” and “runny nose” -- that had a strong correlation with the CDC’s data on flu outbreaks. All 45 terms related in some way to influenza. But with a billion searches a day, it would have been impossible for a person to guess which ones might work best and test only those.

Moreover, the data were imperfect. Since the data were never intended to be used in this way, misspellings and incomplete phrases were common. But the sheer size of the data set more than compensated for its messiness. The result, of course, was simply a correlation. It said nothing about the reasons why someone performed any particular search. Was it because the person felt ill, or heard sneezing in the next cubicle, or felt anxious after reading the news? Google’s system doesn’t know, and it doesn’t care. Indeed, last December, it seems that Google’s system may have overestimated the number of flu cases in the United States. This serves as a reminder that predictions are only probabilities and are not always correct, especially when the basis for the prediction -- Internet searches -- is in a constant state of change and vulnerable to outside influences, such as media reports. Still, big data can hint at the general direction of an ongoing development, and Google’s system did just that.

BACK-END OPERATIONS

There will be a special need to carve out a place for the human: to reserve space for intuition, common sense, and serendipity.

Many technologists believe that big data traces its lineage back to the digital revolution of the 1980s, when advances in microprocessors and computer memory made it possible to analyze and store ever more information. That is only superficially the case. Computers and the Internet certainly aid big data by lowering the cost of collecting, storing, processing, and sharing information. But at its heart, big data is only the latest step in humanity’s quest to understand and quantify the world. To appreciate how this is the case, it helps to take a quick look behind us.

Appreciating people’s posteriors is the art and science of Shigeomi Koshimizu, a professor at the Advanced Institute of Industrial Technology in Tokyo. Few would think that the way a person sits constitutes information, but it can. When a person is seated, the contours of the body, its posture, and its weight distribution can all be quantified and tabulated. Koshimizu and his team of engineers convert backsides into data by measuring the pressure they exert at 360 different points with sensors placed in a car seat and by indexing each point on a scale of zero to 256. The result is a digital code that is unique to each individual. In a trial, the system was able to distinguish among a handful of people with 98 percent accuracy.

The research is not asinine. Koshimizu’s plan is to adapt the technology as an antitheft system for cars. A vehicle equipped with it could recognize when someone other than an approved driver sat down behind the wheel and could demand a password to allow the car to function. Transforming sitting positions into data creates a viable service and a potentially lucrative business. And its usefulness may go far beyond deterring auto theft. For instance, the aggregated data might reveal clues about a relationship between drivers’ posture and road safety, such as telltale shifts in position prior to accidents. The system might also be able to sense when a driver slumps slightly from fatigue and send an alert or automatically apply the brakes.

Koshimizu took something that had never been treated as data -- or even imagined to have an informational quality -- and transformed it into a numerically quantified format. There is no good term yet for this sort of transformation, but “datafication” seems apt. Datafication is not the same as digitization, which takes analog content -- books, films, photographs -- and converts it into digital information, a sequence of ones and zeros that computers can read. Datafication is a far broader activity: taking all aspects of life and turning them into data. Google’s augmented-reality glasses datafy the gaze. Twitter datafies stray thoughts. LinkedIn datafies professional networks.

Once we datafy things, we can transform their purpose and turn the information into new forms of value. For example, IBM was granted a U.S. patent in 2012 for “securing premises using surface-based computing technology” -- a technical way of describing a touch-sensitive floor covering, somewhat like a giant smartphone screen. Datafying the floor can open up all kinds of possibilities. The floor could be able to identify the objects on it, so that it might know to turn on lights in a room or open doors when a person entered. Moreover, it might identify individuals by their weight or by the way they stand and walk. It could tell if someone fell and did not get back up, an important feature for the elderly. Retailers could track the flow of customers through their stores. Once it becomes possible to turn activities of this kind into data that can be stored and analyzed, we can learn more about the world -- things we could never know before because we could not measure them easily and cheaply.

BIG DATA IN THE BIG APPLE

Big data will have implications far beyond medicine and consumer goods: it will profoundly change how governments work and alter the nature of politics. When it comes to generating economic growth, providing public services, or fighting wars, those who can harness big data effectively will enjoy a significant edge over others. So far, the most exciting work is happening at the municipal level, where it is easier to access data and to experiment with the information. In an effort spearheaded by New York City Mayor Michael Bloomberg (who made a fortune in the data business), the city is using big data to improve public services and lower costs. One example is a new fire-prevention strategy.

Illegally subdivided buildings are far more likely than other buildings to go up in flames. The city gets 25,000 complaints about overcrowded buildings a year, but it has only 200 inspectors to respond. A small team of analytics specialists in the mayor’s office reckoned that big data could help resolve this imbalance between needs and resources. The team created a database of all 900,000 buildings in the city and augmented it with troves of data collected by 19 city agencies: records of tax liens, anomalies in utility usage, service cuts, missed payments, ambulance visits, local crime rates, rodent complaints, and more. Then, they compared this database to records of building fires from the past five years, ranked by severity, hoping to uncover correlations. Not surprisingly, among the predictors of a fire were the type of building and the year it was built. Less expected, however, was the finding that buildings obtaining permits for exterior brickwork correlated with lower risks of severe fire.

Using all this data allowed the team to create a system that could help them determine which overcrowding complaints needed urgent attention. None of the buildings’ characteristics they recorded caused fires; rather, they correlated with an increased or decreased risk of fire. That knowledge has proved immensely valuable: in the past, building inspectors issued vacate orders in 13 percent of their visits; using the new method, that figure rose to 70 percent -- a huge efficiency gain.

Of course, insurance companies have long used similar methods to estimate fire risks, but they mainly rely on only a handful of attributes and usually ones that intuitively correspond with fires. By contrast, New York City’s big-data approach was able to examine many more variables, including ones that would not at first seem to have any relation to fire risk. And the city’s model was cheaper and faster, since it made use of existing data. Most important, the big-data predictions are probably more on target, too.

Big data is also helping increase the transparency of democratic governance. A movement has grown up around the idea of “open data,” which goes beyond the freedom-of-information laws that are now commonplace in developed democracies. Supporters call on governments to make the vast amounts of innocuous data that they hold easily available to the public. The United States has been at the forefront, with its Data.gov website, and many other countries have followed.

At the same time as governments promote the use of big data, they will also need to protect citizens against unhealthy market dominance. Companies such as Google, Amazon, and Facebook -- as well as lesser-known “data brokers,” such as Acxiom and Experian -- are amassing vast amounts of information on everyone and everything. Antitrust laws protect against the monopolization of markets for goods and services such as software or media outlets, because the sizes of the markets for those goods are relatively easy to estimate. But how should governments apply antitrust rules to big data, a market that is hard to define and that is constantly changing form? Meanwhile, privacy will become an even bigger worry, since more data will almost certainly lead to more compromised private information, a downside of big data that current technologies and laws seem unlikely to prevent.

Regulations governing big data might even emerge as a battleground among countries. European governments are already scrutinizing Google over a raft of antitrust and privacy concerns, in a scenario reminiscent of the antitrust enforcement actions the European Commission took against Microsoft beginning a decade ago. Facebook might become a target for similar actions all over the world, because it holds so much data about individuals. Diplomats should brace for fights over whether to treat information flows as similar to free trade: in the future, when China censors Internet searches, it might face complaints not only about unjustly muzzling speech but also about unfairly restraining commerce.

BIG DATA OR BIG BROTHER?

States will need to help protect their citizens and their markets from new vulnerabilities caused by big data. But there is another potential dark side: big data could become Big Brother. In all countries, but particularly in nondemocratic ones, big data exacerbates the existing asymmetry of power between the state and the people.

The asymmetry could well become so great that it leads to big-data authoritarianism, a possibility vividly imagined in science-fiction movies such as Minority Report. That 2002 film took place in a near-future dystopia in which the character played by Tom Cruise headed a “Precrime” police unit that relied on clairvoyants whose visions identified people who were about to commit crimes. The plot revolves around the system’s obvious potential for error and, worse yet, its denial of free will.

Although the idea of identifying potential wrongdoers before they have committed a crime seems fanciful, big data has allowed some authorities to take it seriously. In 2007, the Department of Homeland Security launched a research project called FAST (Future Attribute Screening Technology), aimed at identifying potential terrorists by analyzing data about individuals’ vital signs, body language, and other physiological patterns. Police forces in many cities, including Los Angeles, Memphis, Richmond, and Santa Cruz, have adopted “predictive policing” software, which analyzes data on previous crimes to identify where and when the next ones might be committed.

For the moment, these systems do not identify specific individuals as suspects. But that is the direction in which things seem to be heading. Perhaps such systems would identify which young people are most likely to shoplift. There might be decent reasons to get so specific, especially when it comes to preventing negative social outcomes other than crime. For example, if social workers could tell with 95 percent accuracy which teenage girls would get pregnant or which high school boys would drop out of school, wouldn’t they be remiss if they did not step in to help? It sounds tempting. Prevention is better than punishment, after all. But even an intervention that did not admonish and instead provided assistance could be construed as a penalty -- at the very least, one might be stigmatized in the eyes of others. In this case, the state’s actions would take the form of a penalty before any act were committed, obliterating the sanctity of free will.

Another worry is what could happen when governments put too much trust in the power of data. In his 1999 book, Seeing Like a State, the anthropologist James Scott documented the ways in which governments, in their zeal for quantification and data collection, sometimes end up making people’s lives miserable. They use maps to determine how to reorganize communities without first learning anything about the people who live there. They use long tables of data about harvests to decide to collectivize agriculture without knowing a whit about farming. They take all the imperfect, organic ways in which people have interacted over time and bend them to their needs, sometimes just to satisfy a desire for quantifiable order.

This misplaced trust in data can come back to bite. Organizations can be beguiled by data’s false charms and endow more meaning to the numbers than they deserve. That is one of the lessons of the Vietnam War. U.S. Secretary of Defense Robert McNamara became obsessed with using statistics as a way to measure the war’s progress. He and his colleagues fixated on the number of enemy fighters killed. Relied on by commanders and published daily in newspapers, the body count became the data point that defined an era. To the war’s supporters, it was proof of progress; to critics, it was evidence of the war’s immorality. Yet the statistics revealed very little about the complex reality of the conflict. The figures were frequently inaccurate and were of little value as a way to measure success. Although it is important to learn from data to improve lives, common sense must be permitted to override the spreadsheets.

HUMAN TOUCH

Big data is poised to reshape the way we live, work, and think. A worldview built on the importance of causation is being challenged by a preponderance of correlations. The possession of knowledge, which once meant an understanding of the past, is coming to mean an ability to predict the future. The challenges posed by big data will not be easy to resolve. Rather, they are simply the next step in the timeless debate over how to best understand the world.

Still, big data will become integral to addressing many of the world’s pressing problems. Tackling climate change will require analyzing pollution data to understand where best to focus efforts and find ways to mitigate problems. The sensors being placed all over the world, including those embedded in smartphones, provide a wealth of data that will allow climatologists to more accurately model global warming. Meanwhile, improving and lowering the cost of health care, especially for the world’s poor, will make it necessary to automate some tasks that currently require human judgment but could be done by a computer, such as examining biopsies for cancerous cells or detecting infections before symptoms fully emerge.

Ultimately, big data marks the moment when the “information society” finally fulfills the promise implied by its name. The data take center stage. All those digital bits that have been gathered can now be harnessed in novel ways to serve new purposes and unlock new forms of value. But this requires a new way of thinking and will challenge institutions and identities. In a world where data shape decisions more and more, what purpose will remain for people, or for intuition, or for going against the facts? If everyone appeals to the data and harnesses big-data tools, perhaps what will become the central point of differentiation is unpredictability: the human element of instinct, risk taking, accidents, and even error. If so, then there will be a special need to carve out a place for the human: to reserve space for intuition, common sense, and serendipity to ensure that they are not crowded out by data and machine-made answers.

This has important implications for the notion of progress in society. Big data enables us to experiment faster and explore more leads. These advantages should produce more innovation. But at times, the spark of invention becomes what the data do not say. That is something that no amount of data can ever confirm or corroborate, since it has yet to exist. If Henry Ford had queried big-data algorithms to discover what his customers wanted, they would have come back with “a faster horse,” to recast his famous line. In a world of big data, it is the most human traits that will need to be fostered -- creativity, intuition, and intellectual ambition -- since human ingenuity is the source of progress.

Big data is a resource and a tool. It is meant to inform, rather than explain; it points toward understanding, but it can still lead to misunderstanding, depending on how well it is wielded. And however dazzling the power of big data appears, its seductive glimmer must never blind us to its inherent imperfections. Rather, we must adopt this technology with an appreciation not just of its power but also of its limitations.

 

Wednesday, April 17, 2013

The New Digital Age: Reshaping the Future of People, Nations and Business

The New Digital Age: Reshaping the Future of People, Nations and Business by Eric Schmidt and Jared Cohen, Knopf, 2013

Scientific American: "Schmidt, executive chairman of Google, and Cohen, director of Google Ideas and a foreign policy wonk who has advised Hillary Clinton, deliver their vision of the future in this ambitious, fascinating account. For gadget geeks, the book is filled with tantalizing examples of futuristic goods and services: robotic plumbers; automated haircuts; computers that read body language; and 3-D holographs of weddings projected into the living rooms of relatives who couldn't attend. Not surprisingly, the authors are bullish on how connectivity—access to the Internet that will soon be nearly universal—will transform education, terrorism, journalism, government, privacy and war. The result, they argue, though not perfect, will be “more egalitarian, more transparent and more interesting than we can even imagine.”

 

David Brooks: What You'll Do Next

April 15, 2013

What You’ll Do Next

By DAVID BROOKS
http://www.nytimes.com/2013/04/16/opinion/brooks-what-youll-do-next.html?pagewanted=print

Over the past few centuries, there have been many efforts to come up with methods to help predict human behavior — what Leon Wieseltier of The New Republic calls mathematizing the subjective. The current one is the effort to understand the world by using big data.

Other efforts to predict behavior were based on models of human nature. The people using big data don’t presume to peer deeply into people’s souls. They don’t try to explain why people are doing things. They just want to observe what they are doing.

The theory of big data is to have no theory, at least about human nature. You just gather huge amounts of information, observe the patterns and estimate probabilities about how people will act in the future.

As Viktor Mayer-Schönberger and Kenneth Cukier write in their book, “Big Data,” this movement asks us to move from causation to correlation. People using big data are not like novelists, ministers, psychologists, memoirists or gossips, coming up with intuitive narratives to explain the causal chains of why things are happening. “Contrary to conventional wisdom, such human intuiting of causality does not deepen our understanding of the world,” they write.

Instead, they aim to stand back nonjudgmentally and observe linkages: “Correlations are powerful not only because they offer insights, but also because the insights they offer are relatively clear. These insights often get obscured when we bring causality back into the picture.”

This method has yielded some impressive observations. Analysts can look at Google search terms and pick up where flu outbreaks are occurring. In doctor’s offices, statistical predictions often make better diagnoses than clinical predictions. Wal-Mart executives looked at the data and noticed that, as hurricanes approach, people buy large quantities of Strawberry Pop-Tarts. They began to put Pop-Tarts at the front of the stores with storm supplies.

In my columns, I’m trying to appreciate the big data revolution, but also probe its limits. One limit is that correlations are actually not all that clear. A zillion things can correlate with each other, depending on how you structure the data and what you compare. To discern meaningful correlations from meaningless ones, you often have to rely on some causal hypothesis about what is leading to what. You wind up back in the land of human theorizing.

Another obvious problem is that unlike physical objects and even animals, people are discontinuous. We have multiple selves. We are ambiguous and ambivalent. We get bored, and we self-deceive. We learn and mislearn from experience. Thus, the passing of time can produce gigantic and unpredictable changes in taste and behavior, changes that are poorly anticipated by looking at patterns of data on what just happened.

Another limit is that the world is error-prone and dynamic. I recently interviewed George Soros about his financial decision-making. While big data looks for patterns of preferences, Soros often looks for patterns of error. People will misinterpret reality, and those misinterpretations will sometimes create a self-reinforcing feedback loop. Housing prices skyrocket to unsustainable levels.

If you are relying just on data, you will have a tendency to trust preferences and anticipate a continuation of what is happening right now. Soros makes money by exploiting other people’s misinterpretations and anticipating when they will become unsustainable.

Then there is the distinction between commodity decisions and flourishing decisions. Some decisions are straightforward commodities: what route to work is likely to be fastest. Big data can help. Flourishing decisions are things like who to marry, who to befriend, what career calling to pursue and what college to choose. These decisions involve trying to find people, places and things that harmonize with your subjective self. It’s a mistake to take subjective intuition out of this decision because subjectivity is the whole point.

One of my take-aways is that big data is really good at telling you what to pay attention to. It can tell you what sort of student is likely to fall behind. But then to actually intervene to help that student, you have to get back in the world of causality, back into the world of responsibility, back in the world of advising someone to do x because it will cause y.

Big data is like the offensive coordinator up in the booth at a football game who, with altitude, can see patterns others miss. But the head coach and players still need to be on the field of subjectivity.

Most of the advocates understand data is a tool, not a worldview. My worries mostly concentrate on the cultural impact of the big data vogue. If you adopt a mind-set that replaces the narrative with the empirical, you have problems thinking about personal responsibility and morality, which are based on causation. You wind up with a demoralized society. But that’s a subject for another day.

 

Tuesday, April 16, 2013

Open Science, Open Data, Open Access

Open Science, Open Data, Open Access

 

By Matt Luchette 

http://www.bio-itworld.com/2013/4/16/open-science-open-data-open-access.html

 

April 16, 2013 | BOSTON–In two compelling presentations at the Bio-IT World Conference* last week, Atul Butte and Steven Salzberg provided formidable advocacy for the virtues of open data and open science.

 

Salzberg, a computer scientist at Johns Hopkins University, accepted the 2013 Benjamin Franklin Award for Open Access in the Life Sciences for his work promoting “free and open access to the materials and methods used in the life sciences.” 

 

Salzberg is perhaps best-known for developing a series of popular open-source software platforms, including Glimmer (a bacterial gene finder) and the Tuxedo software suite of next generation sequencing tools (see, "Steven Salzberg on Microbial Genomes, Open Access, Flu Shots, and Gene Patents") . Salzberg insists the software stay open-source because “free software gets used.”

 

Salzberg has also been a fervent advocate for a more open atmosphere in science for over a decade. In a 2003 letter to the editor of Nature, he asserted that “genome data-collection projects should be freely available to the entire scientific community, immediately and with no restrictions or conditions.” In a paper last year on “the perils of gene patents,” he argued that “gene patents are antithetical to scientific process.”

 

After accepting the award from Jeff Bizzaro, president of Bioinformatics.org, Salzberg delivered a fast-paced talk on three components of open-science he feels are essential: free software, open data, and open access publication. He discussed some of his lab’s accomplishments in encouraging researchers to be more open with their experimental data. 

 

For the Influenza Genome Sequencing Project he co-founded with NCBI’s David Lipman, for example, Salzberg set out to sequence strains of the flu, a project he hoped would help researchers develop therapies and “improve understanding of the overall molecular evolution of influenza.” The project helped sequence more than 10,000 flu genomes. Even more surprisingly, though, the group published the genomes in real time to assist collaborating research teams. Prior to his project, “the community wasn’t doing that at all,” he said.

 

Salzberg ended on the importance of open-access publications in disseminating knowledge. 

 

“We already write the papers, we already review the papers, why can’t we be the ones who publish?” he asked. The Public Library of Science (PLoS) was one of the first open-access publishing projects to employ this model—and earned co-founder Michael Eisen the first Benjamin Franklin award in 2002—but Salzberg hopes this method will become ubiquitous. He added that it was important for researchers to make the raw data behind their published work available as well. Such an approach could provide scientists a much wider data pool than what they alone can produce in their lab. 

 

“Open science makes us free!” Salzberg remarked in closing. “It allows us to do our work and not worry about all the restrictions on it.”

 

Open Science vs Free Tools 

 

Salzberg’s talk echoed many of the themes in the preceding keynote from Stanford University’s Atul Butte, who highlighted how open access to experimental data could democratize science (see, "Bits and SNPs: Atul Butte and Medicine in the Era of Big Data"). 

 

His talk came just two months after the White House’s Office of Science and Technology released a memorandum directing federal agencies with more than $100 million in R&D expenditures to develop plans for making their experimental data publicly available. With vast, publicly available data libraries, perhaps one day, Butte mused, students could create biotech startups “out of their garage,” the way many technology startups began in the past few decades. 

 

“If you’re not going to do this, try and get your kids interested,” Butte challenged the packed audience. He sprinkled his talk with examples of commercial services such as Assay Depot, offering easy access to cell lines, animal models and so on, greatly expediting new experimental ideas.

 

While the plenary speakers suggested the halcyon days of open access in biomedical research are on the near horizon, the tone was a little different in the exhibit hall downstairs.  

 

While many company representatives at the conference agreed that open-source software and open access to experimental data provide enormous benefits for researchers, they argue a "pay for profit" model of resource distribution provides advantages that open-access platforms aren't prepared to address.

 

Some companies stressed that while publicly available experimental data sets will be invaluable for researchers, not all of the data is created equally. Thomas Reuters’ program MetaCore, for example, providers its clients access to a “high-quality, manually-curated database” formed from “2,700 scientific journals, reviewed by PhD and M.D. level research professionals.” 

 

“We don't just look at the data,” said one Reuters representative. “We look at the experimental design for appropriate controls and that the conclusions are valid."

 

Source Code 

 

Other companies emphasized the on-demand support network users can turn to when they encounter problems. “When I have a problem with [an open-access coding program like] Python, I can look at the source code or send an email and hope somebody responds,” a representative from Wolfram explained. But when software is purchased, he continued, there’s a team of paid developers you can turn to if things go wrong. 

 

Like MetaCore, in addition to a support staff, Wolfram’s Mathematica provides users with “gigabytes of carefully curated and continually updated data” from multiple academic fields, in addition to the program’s computational capabilities.

 

“It’s fine if you’re just starting out and don’t have a lot of money,” one Pekin Elmer representative said about open-software, “but it isn’t scalable to a larger company with lots of researchers,” where lost time due to software issues means lost money.

 

But for many researchers, the rate that the curated databases are updated or the software is upgraded just isn’t fast enough. 

 

“If I’m a researcher, open-source software is always going to be better,” said one representative from Seven Bridges Genomics. Open-source genome analysis programs provide scientists with state-of-the-art computational power, he explained, and are updated faster than many for-profit products as distributors manage licensing and patent logistics before release. Scientists can adapt the software to fit their evolving needs whenever they like. What companies like Seven Bridges provide instead is the ability for researchers to customize their experiments and analysis with a team of genomics experts, as well as the infrastructure to run their analysis and store their data.

 

The advantages and drawbacks of open-science versus for-profit research resources have been a popular conversation topic in academic and industry circles recently. And as seen in the Myriad Genetics patent hearing at the US Supreme Court earlier this week, voices on both sides have grown more resolved. 

 

While open-software has taken up strong roots in many research fields, and experts like Salzberg or Butte are confident open-science could revolutionize the way scientists conduct their research, companies still argue that the cost of for-profit resources is not likely to exceed its value anytime soon.

 

*Bio-IT World Conference & Expo, Boston, April 9-11, 2013.