Monday, May 6, 2013

Life in the City Is Essentially One Giant Math Problem

Life in the City Is Essentially One Giant Math Problem

Experts in the emerging field of quantitative urbanism believe that many aspects of modern cities can be reduced to mathematical formulas

  • By Jerry Adler
  • Photographs by Jordan Hollender
  • Smithsonian magazine, May 2013,

 

Glen Whitney stands at a point on the surface of the Earth, north latitude 40.742087, west longitude 73.988242, which is near the center of Madison Square Park, in New York City. Behind him is the city’s newest museum, the Museum of Mathematics, which Whitney, a former Wall Street trader, founded and now runs as executive director. He is facing one of New York’s landmarks, the Flatiron Building, which got its name because its wedge- like shape reminded people of a clothes iron. Whitney observes that from this perspective you can’t tell that the building, following the shape of its block, is actually a right triangle—a shape that would be useless for pressing clothes—although the models sold in souvenir shops represent it in idealized form as an isosceles, with equal angles at the base. People want to see things as symmetrical, he muses. He points to the building’s narrow prow, whose outline corresponds to the acute angle at which Broadway crosses Fifth Avenue.

“The cross street here is 23rd Street,” Whitney says, “and if you measure the angle at the building’s point, it is close to 23 degrees, which also happens to be approximately the angle of inclination of the Earth’s axis of rotation.”

“That’s remarkable,” he is told.

“Not really. It’s coincidence.” He adds that, twice each year, a few weeks on either side of the summer solstice, the setting sun shines directly down the rows of Manhattan’s numbered streets, a phenomenon sometimes called “Manhattanhenge.” Those particular dates don’t have any special significance, either, except as one more example of how the very bricks and stones of the city illustrate the principles of the highest product of the human intellect, which is math.

Cities are particular: You would never mistake a favela in Rio de Janeiro for downtown Los Angeles. They are shaped by their histories and accidents of geography and climate. Thus the “east-west” streets of Midtown Manhattan actually run northwest-southeast, to meet the Hudson and East rivers at roughly 90 degrees, whereas in Chicago the street grid aligns closely with true north, while medieval cities such as London don’t have right-angled grids. But cities are also, at a deep level, universal: the products of social, economic and physical principles that transcend space and time. A new science—so new it doesn’t have its own journal, or even an agreed-upon name—is exploring these laws. We will call it “quantitative urbanism.” It’s an effort to reduce to mathematical formulas the chaotic, exuberant, extravagant nature of one of humanity’s oldest and most important inventions, the city.

The systematic study of cities dates back at least to the Greek historian Herodotus. In the early 20th century, scientific disciplines emerged around specific aspects of urban development: zoning theory, public health and sanitation, transit and traffic engineering. By the 1960s, the urban-planning writers Jane Jacobs and William H. Whyte used New York as their laboratory to study the street life of neighborhoods, the walking patterns of Midtown pedestrians, the way people gathered and sat in open spaces. But their judgments were generally aesthetic and intuitive (although Whyte, photographing the plaza of the Seagram Building, derived the seat-of-the-pants formula for bench space in public spaces: one linear foot per 30 square feet of open area). “They had fascinating ideas,” says Luís Bettencourt, a researcher at the Santa Fe Institute, a think tank better known for its contributions to theoretical physics, “but where is the science? What is the empirical basis for deciding what kind of cities we want?” Bettencourt, a physicist, practices a discipline that shares a deep affinity with quantitative urbanism. Both require understanding complex interactions among large numbers of entities: the 20 million people in the New York metropolitan area, or the countless subatomic particles in a nuclear reaction.

The birth of this new field can be dated to 2003, when researchers at SFI convened a workshop on ways to “model”—in the scientific sense of reducing to equations—aspects of human society. One of the leaders was Geoffrey West, who sports a neatly trimmed gray beard and retains a trace of the accent of his native Somerset. He was also a theoretical physicist, but had strayed into biology, exploring how the properties of organisms relate to their mass. An elephant is not just a bigger version of a mouse, but many of its measurable characteristics, such as metabolism and life span, are governed by mathematical laws that apply all up and down the scale of sizes. The bigger the animal, the longer but the slower it lives: A mouse heart rate is around 500 beats per minute; an elephant’s pulse is 28. If you plotted those points on a logarithmic graph, comparing size with pulse, every mammal would fall on or near the same line. West suggested that the same principles might be at work in human institutions. From the back of the room, Bettencourt (then at Los Alamos National Laboratory) and José Lobo, an economist at Arizona State University (who majored in physics as an undergraduate), chimed in with the motto of physicists since Galileo: “Why don’t we get the data to test it?”

Out of that meeting emerged a collaboration that produced the seminal paper in the field: “Growth, Innovation, Scaling, and the Pace of Life in Cities.” In six pages dense with equations and graphs, West, Lobo and Bettencourt, along with two researchers from the Dresden University of Technology, laid out a theory about how cities vary according to size. “What people do in cities—create wealth, or murder each other—shows a relationship to the size of the city, one that isn’t tied just to one era or nation,” says Lobo. The relationship is captured by an equation in which a given parameter—employment, say—varies exponentially with population. In some cases, the exponent is 1, meaning whatever is being measured increases linearly, at the same rate as population. Household water or electrical use, for example, shows this pattern; as a city grows bigger its residents don’t use their appliances more. Some exponents are greater than 1, a relationship described as “superlinear scaling.” Most measures of economic activity fall into this category; among the highest exponents the scholars found were for “private [research and development] employment,” 1.34; “new patents,” 1.27; and gross domestic product, in a range of 1.13 to 1.26. If the population of a city doubles over time, or comparing one big city with two cities each half the size, gross domestic product more than doubles. Each individual becomes, on average, 15 percent more productive. Bettencourt describes the effect as “slightly magical,” although he and his colleagues are beginning to understand the synergies that make it possible. Physical proximity promotes collaboration and innovation, which is one reason the new CEO of Yahoo recently reversed the company’s policy of letting almost anyone work from home. The Wright brothers could build their first flying machines by themselves in a garage, but you can’t design a jet airliner that way.

Unfortunately, new AIDS cases also scale superlinearly, at 1.23, as does serious crime, 1.16. Lastly, some measures show an exponent of less than 1, meaning they increase more slowly than population. These are typically measures of infrastructure, characterized by economies of scale that result from increasing size and density. New York doesn’t need four times as many gas stations as Houston, for instance; gas stations scale at 0.77; total surface area of roads, 0.83; and total length of wiring in the electrical grid, 0.87.

Remarkably, this phenomenon applies to cities all over the world, of different sizes, regardless of their particular history, culture or geography. Mumbai is different from Shanghai is different from Houston, obviously, but in relation to their own pasts, and to other cities in India, China or the U.S., they follow these laws. “Give me the size of a city in the United States and I can tell you how many police it has, how many patents, how many AIDS cases,” says West, “just as you can calculate the life span of a mammal from its body mass.”

One implication is that, like the elephant and the mouse, “big cities are not just bigger small cities,” says Michael Batty, who runs the Centre for Advanced Spatial Analysis at University College London. “If you think of cities in terms of potential interactions [among individuals], as they get bigger you get more opportunities for that, which amounts to a qualitative change.” Consider the New York Stock Exchange as a microcosm of a metropolis. In its early years, investors were few and trades sporadic, Whitney says. Hence “specialists” were needed, intermediaries who kept an inventory of stock in certain companies, and would “make a market” in the shares, pocketing the margin between their selling and buying price. But over time, as more participants joined the market, buyers and sellers could find one another more easily, and the need for specialists—and their profits, which amounted to a small tax on everyone else—diminished. There is a point, Whitney says, at which a system—a market, or a city—undergoes a phase shift and reorganizes itself in a more efficient and productive way.

Whitney, who has a slight build and a meticulous manner, walks swiftly through Madison Square Park to the Shake Shack, a hamburger stand famous for its food and its lines. He points out the two service windows, one for customers who can be served quickly, the other for more complicated orders. This distinction is supported by a branch of mathematics called queuing theory, whose fundamental principle can be stated as “the shortest aggregate waiting time for all customers is achieved when the person with the shortest expected wait time is served first, provided the guy who wants four hamburgers with different toppings doesn’t go berserk when he keeps getting sent to the back of the line.” (This assumes that the line closes at a certain time so everyone gets served eventually. The equations can’t handle the concept of an infinite wait.) That idea “seems intuitive,” says Whitney, “but it had to be proved.” In the real world, queuing theory is used for designing communications networks, in deciding which packet of data gets sent first.

At the Times Square subway station, Whitney buys a fare card, in an amount he has calculated to take advantage of the bonus for paying in advance and come out with an even number of rides, with no money left unspent. On the platform, as passengers rush back and forth between trains, he talks about the mathematics of running a transit system. You might think, he says, that an express should always leave as soon as it’s ready, but there are times when it makes sense to hold it in the station—to make a connection with an incoming local. The calculation, simplified, is this: Multiply the number of people on the express train by the number of seconds they will be kept waiting while it idles in the station. Now estimate how many people on the arriving local will transfer, and multiply that by the average time they will save by taking the express to their destination rather than the local. (You’ll have to model how far passengers who bother to switch are going.) This can lead to the potential savings, in person-seconds, for comparison. The principle is the same at any scale, but it is only above a certain size of population that the investment in dual-track subway lines or two-window hamburger stands makes sense. Whitney boards the local, heading downtown toward the museum.

***

It also can be readily seen that the more data you have on transit usage (or hamburger orders), the more detailed and accurate you can make these calculations. If Bettencourt and West are building a theoretical science of urbanism, then Steven Koonin, the first director of New York University’s newly created Center for Urban Science and Progress, intends to be in the forefront of applying it to real-world problems. Koonin, as it happens, is also a physicist, a former Cal Tech professor and assistant secretary of the Department of Energy. He describes his ideal student, when CUSP begins its first academic year this fall, as “someone who helped find the Higgs boson and now wants to do something with her life that will make society better.” Koonin is a believer in what is sometimes called Big Data, the bigger the better. Only in the past decade has the ability to collect and analyze information about the movement of people begun to catch up to the size and complexity of the modern metropolis itself. Around the time he took the job at CUSP, Koonin read a paper about the ebb and flow of population in Manhattan’s business district, based on an exhaustive analysis of published data on employment, transit and traffic patterns. It was a great piece of research, Koonin says, but in the future, that’s not how it will be done. “People carry tracking devices in their pockets all day long,” he says. “They’re called cellphones. You don’t need to wait for some agency to publish statistics from two years ago. You can get this data almost in real time, block by block, hour by hour.

“We have acquired the technology to know virtually anything that goes on in an urban society,” he adds, “so the question is, how can we leverage that to do good? Make the city run better, enhance security and safety and promote the private sector?” Here’s a simple example of what Koonin envisions, in the near future. If you are, say, deciding whether to drive or take the subway from Brooklyn to Yankee Stadium, you can consult a website for real-time transit data, and another for traffic. Then you can make a choice based on intuition, and your personal feelings about the trade-offs among speed, economy and convenience. This by itself would have seemed miraculous even a few years ago. Now imagine a single app that would have access to that data (plus GPS locations of taxis and buses along the route, cameras surveying the stadium’s parking lots and Twitter feeds from people stuck on FDR Drive), factor in your preferences and tell you instantly: Stay home and watch the game on TV.

Or some slightly less simple examples of how Big Data can be used. At a lecture last year Koonin presented an image of a large swath of Lower Manhattan, showing the windows of some 50,000 offices and apartments. It was taken with an infrared camera, and so could be used for environmental surveillance, identifying buildings, or even individual units, that were leaking heat and wasting energy. Another example: As you move around the city, your cellphone tracks your location and that of everyone you come into contact with. Koonin asks: How would you like to get a text message telling you that yesterday you were in a room with someone who just checked into the emergency room with the flu?

***

Inside the Museum of Mathematics, kids and the occasional adult manipulate various solids on a series of screens, rotating them, extending or compressing or twisting them into fantastical shapes, then extruding them in plastic on a 3-D printer. They sit inside a tall cylinder whose base is a rotating platform and whose sides are defined by vertical strings; as they twist the platform, the cylinder deforms into a hyperboloid, a curved surface that somehow is created out of straight lines. Or they demonstrate how it is possible to have a smooth ride on a square-wheeled tricycle, if you contour the track beneath it to keep the axle level. Geometry, unlike formal logic, which was Whitney’s field before he went to Wall Street, lends itself particularly well to hands-on experiment and demonstration—although there are also exhibits touching on fields he identifies as “calculus, calculus of variations, differential equations, combinatorics, graph theory, mathematical optics, symmetry and group theory, statistics and probability, algebra, matrix analysis—and arithmetic.” It troubled Whitney that in a world with museums devoted to ramen noodles, ventriloquism, lawn mowers and pencils, “most of the world has never seen the raw beauty and adventure that is the world of mathematics.” That’s what he set out to remedy.

As Whitney points out on the popular math tours he runs, the city has a distinctive geometry, which can be described as occupying two-and-a-half dimensions. Two of these are those you see on the map. He describes the half-dimension as the network of elevated and underground walkways, roads and tunnels that can only be accessed at specific points, like the High Line, an abandoned railroad trestle that has been turned into an elevated linear park. This space is analogous to an electronic printed-circuit board, in which, as mathematicians have shown, certain configurations cannot be achieved in a single plane. The proof is in the famous “three-utilities puzzle,” a demonstration of the impossibility of routing gas, water and electric service to three houses without any of the lines crossing. (You can see this for yourself by drawing three boxes and three circles, and trying to connect each circle to each box with nine lines that do not intersect.) In a circuit board, for conductors to cross without touching, one of them sometimes must leave the plane. Just so, in the city, sometimes you have to climb up or down to get to where you are going.

Whitney heads uptown, to Central Park, where he walks on a path that for the most part skirts the hills and declivities created by the most recent glaciation and improved by Olmsted and Vaux. On a certain class of continuous surfaces—of which parkland is one—you can always find a path that stays on one level. From various points in Midtown, the Empire State Building appears and disappears behind the interposing structures. This brings to mind a theory Whitney has about the height of skyscrapers. Obviously big cities have more tall buildings than small cities, but the height of the tallest building in a metropolis doesn’t bear a strong relationship to its population; based on a sample of 46 metropolitan areas around the world, Whitney has found that it tracks the economy of the region, approximating the equation H=134 + 0.5(G), where H is the height of the tallest building in meters, and G is the Gross Regional Product, in billions of dollars. But building heights are constrained by engineering , while there’s no limit to how big a pile you can make out of money, so there are two very rich cities whose tallest towers are lower than the formula would predict. They are New York and Tokyo. Also, his equation has no term for “national pride,” so there are a few outliers in the other direction, cities whose reach toward the sky exceeds their grasp of GDP: Dubai, Kuala Lumpur.

No city exists in pure Euclidean space; geometry always interacts with geography and climate, and with social, economic and political factors. In Sunbelt metropolises such as Phoenix, other things being equal the more desirable suburbs are to the east of downtown, where you can commute both ways with the sun behind you as you drive. But where there is a prevailing wind, the best place to live is (or was, in the era before pollution controls) upwind of the city center, which in London means to the west. Deep mathematical principles underlie even such seemingly random and historically contingent facts as the distribution of the sizes of cities within a country. There is, typically, one largest city, whose population is twice that of the second-largest, and three times the third-largest, and increasing numbers of smaller cities whose sizes also fall into a predictable pattern. This principle is known as Zipf’s law, which applies across a wide range of phenomena. (Among other unrelated phenomena, it predicts how incomes are distributed across the economy and the frequency of the appearance of words in a book.) And the rule holds true even though individual cities move up and down in the rankings all the time—St. Louis, Cleveland and Baltimore, all in the top 10 a century ago, making way for San Diego, Houston and Phoenix.

As West and his colleagues are well aware, this research takes place against the background of a huge demographic shift, the predicted movement of literally billions of people to cities in the developing world over the next half century. Many of them are going to end up in slums—a word that describes, without judgment, informal settlements on the outskirts of cities, generally inhabited by squatters with limited or no government services. “No one has done a serious scientific study of these communities,” West says. “How many people live in how many structures of how many square feet? What is their economy? The data we do have, from governments, is often worthless. In the first set we got from China, they reported no murders. So you throw that out, but what are you left with?”

To answer those questions, the Santa Fe Institute, with backing from the Gates Foundation, has begun a partnership with Slum Dwellers International, a network of community organizations based in Cape Town, South Africa. The plan is to analyze the data gathered from 7,000 settlements in cities such as Mumbai, Nairobi and Bangalore, and begin the work of developing a mathematical model for these places, and a path toward integrating them into the modern economy. “For a long time, policy makers have assumed it’s a bad thing for cities to keep getting larger,” says Lobo. “You hear things like, ‘Mexico City has grown like a cancer.’ A lot of money and effort has been devoted to stemming this, and by and large it has failed miserably. Mexico City is bigger than it was ten years ago. So we think policy makers should worry instead about making those cities more livable. Without glorifying the conditions in these places, we think they’re here to stay and we think they hold opportunities for the people who live there.”

And one had better hope he is right, if Batty is correct in predicting that by the end of the century, practically the entire population of the world will live in what amounts to “a completely global entity...in which it will be impossible to consider any individual city separately from its neighbors...indeed perhaps from any other city.” We are seeing now, in Bettencourt’s words, “the last big wave of urbanization that we will experience on Earth.” Urbanization gave the world Athens and Paris, but also the chaos of Mumbai and the poverty of Dickens’ London. If there’s a formula for assuring that we are headed for one rather than the other, West, Koonin, Batty and their colleagues are hoping to be the ones to find it.

 

 

 

 

Find this article at:
http://www.smithsonianmag.com/ideas-innovations/Life-in-the-City-Is-Essentially-One-Giant-Math-Problem-204138731.html?c=y&story=fullstory

 

Friday, May 3, 2013

The Big Data Debate: Correlation vs. Causation

The Big Data Debate: Correlation vs. Causation

 

43

http://smartdatacollective.com/gilpress/121531/big-data-debate-correlation-vs-causation?utm_source=feedburner&utm_medium=feed&utm_campaign=Smart+Data+Collective+%28all+posts%29

In the first quarter of 2013, the stock of big data has experienced sudden declines followed by sporadic bouts of enthusiasm. The volatility—a new big data “V”—continues and Ted Cuzzillo summed up the recent negative sentiment in “Big data, big hype, big danger” on SmartDataCollective:

"A remarkable thing happened in Big Data last week. One of Big Data’s best friends poked fun at one of its cornerstones: the Three V’s. The well-networked and alert observer Shawn Rogers, vice president of research at Enterprise Management Associates, tweeted his eight V’s: '…Vast, Volumes of Vigorously, Verified, Vexingly Variable Verbose yet Valuable Visualized high Velocity Data.' He was quick to explain to me that this is no comment on Gartner analyst Doug Laney’s three-V definition. Shawn’s just tired of people getting stuck on V’s.”

Indeed, all the people who “got stuck” on Laney’s “definition,” conveniently forgot that he first used the “three-Vs” to describe data management challenges in 2001. Yes, 2001. If big data is a “revolution,” how come its widely-used “definition” is based on a dozen year-old analyst note?

Ranting about how “blogs and articles yammer on with the benefits of ‘big data,’” Cuzzillo correctly observes that they are simply “repeating promises made years ago about the benefits of small data and small analytics. This is old decision support super-sized and warmed over, the ‘new and improved’ that won’t satisfy any better than the original but which costs much, much more.”

Cuzzillo is joined by a growing chorus of critics that challenge some of the breathless pronouncements of big data enthusiasts. Specifically, it looks like the backlash theme-of-the-month is correlation vs. causation, possibly in reaction to the success of Viktor Mayer-Schönberger and Kenneth Cukier’s recent big data book in which they argued for dispensing “with a reliance on causation in favor of correlation” (see my discussion of the book and this argument).

In “Steamrolled by Big Data,” The New Yorker’s Gary Marcus declares that “Big Data isn’t nearly the boundless miracle that many people seem to think it is.” He concedes that “Big Data can be especially helpful in systems that are consistent over time, with straightforward and well-characterized properties, little unpredictable variation, and relatively little underlying complexity.” But Marcus warns that “not every problem fits those criteria; unpredictability, complexity, and abrupt shifts over time can lead even the largest data astray. Big Data is a powerful tool for inferring correlations, not a magic wand for inferring causality.” Calling for “a sensitivity to when humans should and should not remain in the loop,” Marcus quotes Alexei Efros, “one of the leaders in applying Big Data to machine vision,” who described big data as “a fickle, coy mistress.”

Matti Keltanen at The Guardian agrees, explaining “Why 'lean data' beats big data.” Writes Keltanen: “…the lightest, simplest way to achieve your data analysis goals is the best one…The dirty secret of big data is that no algorithm can tell you what's significant, or what it means. Data then becomes another problem for you to solve. A lean data approach suggests starting with questions relevant to your business and finding ways to answer them through data, rather than sifting through countless data sets. Furthermore, purely algorithmic extraction of rules from data is prone to creating spurious connections, such as false correlations… today's big data hype seems more concerned with indiscriminate hoarding than helping businesses make the right decisions.”

In “Data Skepticism,” O’Reilly Radar’s Mike Loukides adds this gem to the discussion: “The idea that there are limitations to data, even very big data, doesn’t contradict Google’s mantra that more data is better than smarter algorithms; it does mean that even when you have unlimited data, you have to be very careful about the conclusions you draw from that data. It is in conflict with the all-too-common idea that, if you have lots and lots of data, correlation is as good as causation.”

Isn’t more-data-is-better the same as correlation-is-as-good-as-causation? Or, in the words of Chris Andersen, "with enough data, the numbers speak for themselves."

That’s much more than a mantra. It’s the big data religion, its core mystical experience: The data speak (how prescient was Larry Ellison when he re-named his company in 1982?).

“Can numbers actually speak for themselves?” non-believer Kate Crawford asks in “The Hidden Biases in Big Data” on the Harvard Business Review blog and answers: “Sadly, they can't. Data and data sets are not objective; they are creations of human design. We give numbers their voice, draw inferences from them, and define their meaning through our interpretations. Hidden biases in both the collection and analysis stages present considerable risks, and are as important to the big-data equation as the numbers themselves. We get a much richer sense of the world when we ask people the why and the how not just the 'how many.'"

A NPR blogger notes that “while Big Data can uncover correlations between data, it doesn't reveal causation. Sometimes, that doesn't really matter, but other times, it might — in ways we're not always aware of.” He (or she) also quotes The New York Times Steve Lohr who quotes Albert Einstein: "Not everything that counts can be counted, and not everything that can be counted counts."

Speaking of Einstein (“imagination is more important than knowledge”), E. O. Wilson in The Wall Street Journal takes the discussion to a whole new level. While he doesn’t specifically mention big data, Wilson (in “great scientists don’t need math”) makes an important distinction between using only mathematics and using one’s imagination or intuition: “I have a professional secret to share: Many of the most successful scientists in the world today are mathematically no more than semiliterate… Fortunately, exceptional mathematical fluency is required in only a few disciplines, such as particle physics, astrophysics and information theory. Far more important throughout the rest of science is the ability to form concepts, during which the researcher conjures images and processes by intuition… The annals of theoretical biology are clogged with mathematical models that either can be safely ignored or, when tested, fail. Possibly no more than 10% have any lasting value. Only those linked solidly to knowledge of real living systems have much chance of being used.”

And David Brooks in The New York Times, while probing the limits of “the big data revolution,” takes the discussion to yet another level: “One limit is that correlations are actually not all that clear. A zillion things can correlate with each other, depending on how you structure the data and what you compare. To discern meaningful correlations from meaningless ones, you often have to rely on some causal hypothesis about what is leading to what. You wind up back in the land of human theorizing… Most of the advocates understand data is a tool, not a worldview. My worries mostly concentrate on the cultural impact of the big data vogue. If you adopt a mind-set that replaces the narrative with the empirical, you have problems thinking about personal responsibility and morality, which are based on causation. You wind up with a demoralized society.”

I don’t think that the big data mind-set replaces “the narrative” with the empirical. It replaces it with numbers and correlations. There is nothing wrong with a scientific mind-set, based on empirical observations, as long as people don’t mistake number-crunching for scientific inquiry or see cause-and-effect in correlations.

Kaiser Fung concludes in his summary of the recent Reinhart-Rogof kerfuffle (“Occupational hazards in data science”) that the problem of seeing (or implying) causation in correlations is found not just in economics but also in medical research and other fields using observational data: “The usual ploy is first acknowledge that the data could not prove causality (‘we found an association between sleeping less and snoring; our data does not allow us to prove causation.’), then quietly assume that the causal link is there, and wax on the implications (‘if you want to snore less, sleep less.’)” Or, as The Atlantic’s Matthew O’Brien puts it: “R-R whisper ‘correlation’ to other economists, but say ‘causation’ to everyone else.”

Whether you use small or big data, your imagination (developing theories) and integrity (following the scientific method) are what counts. Correlations can count, too, in certain situations. Just don’t expect them to explain anything.

[Originally published on Forbes.com]

 

Tuesday, April 30, 2013

Measuring the Benefits of Tech Tools

April 30, 2013

Measuring the Benefits of Tech Tools

By EDUARDO PORTER
http://www.nytimes.com/2013/05/01/business/statistics-miss-the-benefits-of-technology.html?pagewanted=all&_r=0&pagewanted=print

When I was a young reporter we could not afford cellphones. I remember waiting in line for a pay phone in downtown Mexico City one afternoon to call in the news about the auction to privatize the phone company Telmex, driving those behind me crazy while a copy editor on the other end patiently typed it into those glowing green letters of an earlier information age.

I traveled to Japan with a TRS-80 portable computer, which ran on AA batteries and had plastic cups to put over the phone receiver. It transmitted copy at the blistering speed of 300 bits per second. And I wrote about Mexico’s tequila crisis of 1994 without the benefit of a full set of Mexican financial statistics a few clicks away.

From my perspective, the evolution of the tools of journalism between then and now has been nothing less than breathtaking.

Articles are more thorough — informed by complementary data and analysis, enriched with links to things like interactive charts, videos and slide shows. They get to readers much more quickly. Most important, they reach many more of them.

For all its financial troubles, never has The New York Times been read by more people: 44 million unique viewers online in the United States every month. Yet if you were to rummage through American economic statistics you would find little evidence of journalism’s technological leaps. Measured by its contribution to gross domestic product, the most prominent indicator of the nation’s economic well-being, much of this new journalistic value enabled by information technology is not worth much.

This is true not only of journalism. The failure of I.T. to deliver measurable value has been a popular meme among economists for years. Back in 1987 Nobel laureate Robert Solow posed a now famous paradox: “We can see the computers everywhere except in the productivity statistics.”

The meme is back. The burst of productivity during the dot-com revolution of the 1990s gave skeptics pause. But as productivity has slowed substantially in recent years, doubts have re-emerged about whether information technology can power economic growth like the steam engine and the internal combustion engine did in the past.

Last year, Robert J. Gordon of Northwestern University proposed that the I.T. revolution has pretty much exhausted its promise. He asked, provocatively: “Is U.S. economic growth over?” And he forecast stagnating living standards for the vast majority of Americans for decades to come.

Government statistics lend support to his skepticism: Value added by the information technology and communications industries — mostly hardware and software — has remained stuck at around 4 percent of the nation’s economic output for the last quarter century.

But these statistics do not tell the whole story. Because they miss much of what technology does for people’s well-being.

News organizations that take advantage of computers to let go of journalists, secretaries and research assistants will show up in the economic statistics as more productive, making more with less. But statisticians have no way to value more thorough, useful, fact-dense articles.

What’s more, gross domestic product only values the goods and services people pay for. It does not capture the value to consumers of economic improvements that are given away free. And until recently this is what news media organizations like The New York Times were doing online.

The Commerce Department is in the process of revising the way it measures G.D.P. to take better account of the contributions of investment in research and development and artistic creation. But even though the revisions to be announced this summer are expected to make the economy look bigger, they are not devised to capture the value that Americans get from digital technologies.

“G.D.P. is not a measure of how much value is produced for consumers,” said Erik Brynjolfsson of the Massachusetts Institute of Technology. “Everybody should recognize that G.D.P. is not a welfare metric.”

G.D.P. misses what Americans gain from sharing information on Facebook or finding information on Google or Wikipedia. It misses how dating sites reduce the cost and increase the odds of finding a mate. It misses the time saved by drivers who use Google Maps and the time gained by consumers from shopping online. Measured in money — what it contributes to G.D.P. — the recording industry is shrinking. Yet never before have Americans had access to so much music.

“Pretty much every human on earth can access all human knowledge,” said Hal Varian, Google’s chief economist. That may not be as impressive as civilization’s 60-year technological jump from the horse-drawn buggy to the man on the moon. But it’s probably useful, and it is mostly ignored by our measures of progress.

So how to measure the Internet’s contribution to our lives? A few years ago, Austan Goolsbee of the University of Chicago and Peter J. Klenow of Stanford gave it a shot. They estimated that the value consumers gained from the Internet amounted to about 2 percent of their income — an order of magnitude larger than what they spent to go online. Their trick was to measure not only how much money users spent on access but also how much of their leisure time they spent online.

The approach makes intuitive sense. Time is not only a valuable asset. Its relative value increases with economic development, as workers’ incomes grow while their allotment of time remains stubbornly fixed.

It puts the Internet in a different light. Earlier this year, Yan Chen, Grace YoungJoo Jeon and Yong-Mi Kim of the University of Michigan published the result of an experiment that found that people who had access to a search engine took 15 minutes less to answer a question than those without online access.

Using the average wage of $22 an hour as the value of workers’ time, and assuming that people who could answer more questions would ask more of them, Mr. Varian estimated that a search engine might be worth about $500 annually to the average worker. Across the working population, this would add up to $65 billion a year.

Last year, Mr. Brynjolfsson and an M.I.T. postgraduate, JooHee Oh, used a similar accounting technique to that of Mr. Goolsbee and Mr. Klenow and concluded that the consumer surplus from free online services — the value derived by consumers from the experience above what they paid for it — has been growing by $34 billion a year, on average, since 2002. If it were tacked on as “economic output,” it would add about 0.26 of a percentage point to annual G.D.P. growth.

Technoskeptics may scoff at these calculations. The Internet is hardly the first technology to offer consumers valuable free goods. The consumer surplus from television is about five times as large as that delivered by free stuff online, according to Mr. Brynjolfsson’s calculations.

Gross domestic product has always failed to capture many things — from the costs of pollution and traffic jams to the gains of unpaid household work. As the economist Paul Samuelson once pointed out, if a man married his maid, G.D.P. would decline.

Notably, it almost inevitably misses some economic gains from new technologies. The missed consumer surplus from the Internet may be no bigger than the unmeasured gains in the production, for example, of electric light.

But there is a case to be made that the unmeasured benefits from the Internet deserve more attention.

The amount of time Americans devote to the Internet has doubled in the last five years. Information, encoded in bits, is bound to become a larger and larger share of our economic output. Much of its value will be delivered to each additional consumer at a marginal cost of nearly zero.

“We know less about the sources of value in the economy than we did 25 years ago,” wrote Mr. Brynjolfsson and Adam Saunders of the University of British Columbia. If we really want to understand the impact of information technology on our future well-being, we first need to find a consistent way to measure it.

 

Sunday, April 28, 2013

How Big Data Is Playing Recruiter for Specialized Workers

April 27, 2013

How Big Data Is Playing Recruiter for Specialized Workers

By MATT RICHTEL
http://www.nytimes.com/2013/04/28/technology/how-big-data-is-playing-recruiter-for-specialized-workers.html?curator=MediaReDEF&emc=rss&partner=rss&_r=0&pagewanted=print#h[]

WHEN the e-mail came out of the blue last summer, offering a shot as a programmer at a San Francisco start-up, Jade Dominguez, 26, was living off credit card debt in a rental in South Pasadena, Calif., while he taught himself programming. He had been an average student in high school and hadn’t bothered with college, but someone, somewhere out there in the cloud, thought that he might be brilliant, or at least a diamond in the rough.

That someone was Luca Bonmassar. He had discovered Mr. Dominguez by using a technology that raises important questions about how people are recruited and hired, and whether great talent is being overlooked along the way. The concept is to focus less than recruiters might on traditional talent markers — a degree from M.I.T., a previous job at Google, a recommendation from a friend or colleague — and more on simple notions: How well does the person perform? What can the person do? And can it be quantified?

The technology is the product of Gild, the 18-month-old start-up company of which Mr. Bonmassar is a co-founder. His is one of a handful of young businesses aiming to automate the discovery of talented programmers — a group that is in enormous demand. These efforts fall in the category of Big Data, using computers to gather and crunch all kinds of information to perform many tasks, whether recommending books, putting targeted ads onto Web sites or predicting health care outcomes or stock prices.

Of late, growing numbers of academics and entrepreneurs are applying Big Data to human resources and the search for talent, creating a field called work-force science. Gild is trying to see whether these technologies can also be used to predict how well a programmer will perform in a job. The company scours the Internet for clues: Is his or her code well-regarded by other programmers? Does it get reused? How does the programmer communicate ideas? How does he or she relate on social media sites?

Gild’s method is very much in its infancy, an unproven twinkle of an idea. There is healthy skepticism about this idea, but also excitement, especially in industries where good talent can be hard to find.

The company expects to have about $2 million to $3 million in revenue this year and has raised around $10 million, including a chunk from Mark Kvamme, a venture capitalist who invested early in LinkedIn. And Gild has big-name customers testing or using its technology to recruit, including Facebook, Amazon, Wal-Mart Stores, Google and Twitter.

Companies use Gild to mine for new candidates and to assess candidates they are already considering. Gild itself uses the technology, which was how the company, desperate for programming talent and unable to match the salaries offered by bigger tech concerns, found this guy named Jade outside of Los Angeles. Its algorithm had determined that he had the highest programming score in Southern California, a total that almost no one achieves. It was 100.

Who was Jade? Could he help the company? What does his story tell us about modern-day recruiting and hiring, about the concept of meritocracy?

PEOPLE in Silicon Valley tend to embrace certain assumptions: Progress, efficiency and speed are good. Technology can solve most things. Change is inevitable; disruption is not to be feared. And, maybe more than anything else, merit will prevail.

But Vivienne Ming, who since late in 2012 has been the chief scientist at Gild, says she doesn’t think Silicon Valley is as merit-based as people imagine. She thinks that talented people are ignored, misjudged or fall through the cracks all the time. She holds that belief in part because she has had some experience of it.

Dr. Ming was born male, christened Evan Campbell Smith. He was a good student and a great athlete — holding records at his high school in track and field in the triple jump and long jump. But he always felt a disconnect with his body. After high school, Evan experienced a full-blown identity crisis. He flopped at college, kicked around jobs, contemplated suicide, hit the proverbial bottom. But rather than getting stuck there, he bounced. At 27, he returned to school, got an undergraduate degree in cognitive neuroscience from the University of California, San Diego, and went on to receive a Ph.D. at Carnegie Mellon in psychology and computational neuroscience.

During a fellowship at Stanford, he began gender transition, becoming, fully, Dr. Vivienne Ming in 2008.

As a woman, Dr. Ming started noticing that people treated her differently. There were small things that seemed innocuous, like men opening the door for her. There were also troubling things, like the fact that her students asked her fewer questions about math then they had when she was a man, or that she was invited to fewer social events — a baseball game, for instance — by male colleagues and business connections.

Bias often takes forms that people may not recognize. One study that Dr. Ming cites, by researchers at Yale, found that faculty members at research universities described female applicants for a manager position as significantly less competent than male applicants with identical qualifications. Another study, published by the National Bureau of Economic Research, found that people who sent in résumés with “black-sounding” names had a considerably harder time getting called back from employers than did people who sent in résumés showing equal qualifications but with “white-sounding” names.

Everybody can pretty much agree that gender, or how people look, or the sound of a last name, shouldn’t influence hiring decisions. But Dr. Ming takes the idea of meritocracy further. She suggests that shortcuts accepted as a good proxy for talent — like where you went to school or previously worked — can also shortchange talented people and, ultimately, employers. “The traditional markers people use for hiring can be wrong, profoundly wrong,” she said.

Dr. Ming’s answer to what she calls “so much wasted talent” is to build machines that try to eliminate human bias. It’s not that traditional pedigrees should be ignored, just balanced with what she considers more sophisticated measures. In all, Gild’s algorithm crunches thousands of bits of information in calculating around 300 larger variables about an individual: the sites where a person hangs out; the types of language, positive or negative, that he or she uses to describe technology of various kinds; self-reported skills on LinkedIn; the projects a person has worked on, and for how long; and, yes, where he or she went to school, in what major, and how that school was ranked that year by U.S. News & World Report.

“Let’s put everything in and let the data speak for itself,” Dr. Ming said of the algorithms she is now building for Gild.

Gild is not the only company now scouring for information. TalentBin, another San Francisco start-up firm, searches the Internet for talented programmers, trawling sites where they gather, collecting “data exhaust,” according to the company Web site, and creating lists of potential hires for employers. Another competitor is RemarkableHire, which assesses a person’s talents by looking at how his or her online contributions are rated by others.

And there’s Entelo, which tries to figure out who might be looking for a job before they even start their exploration. According to its Web site, the company uses more than 70 variables to find indications of possible career change, such as how someone presents herself on social sites. The Web site reads: “We crunch the data so you don’t have to.”

This application of Big Data to recruiting is “is absolutely worth a try,” said Susan Etlinger, an analyst of the data and analytics industries at the Altimeter Group. But she questioned whether an algorithm would be an improvement over what employers already do: gathering résumés, or referrals, and using traditional markers associated with success.

“The big hole is actual outcomes,” she said. “What I’m not buying yet is that probability equals actuality.”

Sean Gourley, co-founder and chief technology officer at Quid, a Big Data company, said that data trawling could inform recruiting and hiring, but only if used with an understanding of what the data can’t reveal. “Big Data has its own bias,” he said. “You measure what you can measure,” and “you’re denigrating what can’t be measured, like gut instinct, charisma.”

He added: “When you remove humans from complex decision-making, you can optimize the hell out of the algorithm, but at what cost?”

Dr. Ming doesn’t suggest eliminating human judgment, but she does think that the computer should lead the way, acting as an automated vacuum and filter for talent. The company has amassed a database of seven million programmers, ranking them based on what it calls a Gild score — a measure, the company says, of what a person can do. Ultimately, Dr. Ming wants to expand the algorithm so it can search for and assess other kinds of workers, like Web site designers, financial analysts and even sales people at, say, retail outlets.

“We did our own internal gold strike,” Dr. Ming said. “We found this kid in Los Angeles just kicking around his computer.”

She’s talking about Jade.

MR. DOMINGUEZ grew up in Los Angeles, the middle child of five. His mother took care of the household; his dad installed telecommunications equipment — a blue-collar guy who prized education.

But Jade had a rebellious streak. Halfway through high school, Mr. Dominguez, previously a straight-A student, began wondering whether going to school was more about satisfying requirements than real learning. “The value proposition is to go to school to get a good job,” he told me. “Philosophically, shouldn’t you go to school to learn?” His grades fell sharply, and he said he graduated from Alhambra High School in 2004 with less than a 3.0 grade-point average.

Not only did he reject college, he also wanted to prove that he could succeed wildly without it. He devoured books on entrepreneurship. He started a company that printed custom T-shirts, first from his house, then from a 1,000-square-foot warehouse space he rented. He decided that he needed a Web site, so he taught himself programming.

“I was out to prove myself on my own merit,” he said. He concedes that he might have taken it a little far. “It’s a little immature to be motivated by proving people wrong,” he said.

He got a tattoo on his arm in flowery script that read “Believe.” He sort of laughs about it now, though he still feels that he can accomplish what he puts his mind to. “It’s the great thing about code,” he said of computer language. “It’s largely merit-driven. It’s not about what you’ve studied. It’s about what you’ve shipped.”

When Gild went looking for talent, it assumed that the San Francisco and Silicon Valley areas would be picked over. So it ran its algorithm in Southern California and came up with a list of programmers. At the top was Mr. Dominguez, who had a very solid reputation on GitHub — a place where software developers gather to share code, exchange ideas and build reputations. Gild combs through GitHub and a handful of other sites, including Bitbucket and Google Code, looking for bright people in the field.

Mr. Dominguez had made quite a contribution. His code for Jekyll-Bootstrap, a function used in building Web sites, was reused by an impressive 1,267 other developers. His language and habits showed a passion for product development and several programming tools, like Rails and JavaScript, which were interesting to Gild. His blogs and posts on Twitter suggested that he was opinionated, something that the company wanted on its initial team.

A recruiter from Gild sent him an e-mail and had him come to San Francisco for an interview. The company founders met a charismatic, confident person — poised, articulate, thoughtful, with an easy smile, a tad rougher around the edges than other interview candidates, said Sheeroy Desai, Mr. Bonmassar’s co-founder at Gild and the company’s chief executive.

Mr. Dominguez wore a vibrant green hoodie to the interview. He asked pointed questions, like this one: Did the company worry that it would be perceived as violating privacy by scoring engineers without their knowledge? (It didn’t believe so, and he didn’t, either. Gild says it uses only publicly available information.)

They asked him some pointed but gentle questions, too, like whether he could work in a structured environment. He said he could. The company made Mr. Dominguez a job offer right away, and he accepted a position that pays around $115,000 a year.

“He’s a symbol of someone who is smart, highly motivated and yet, for whatever reason, wasn’t motivated in high school and didn’t see value in college,” Mr. Desai said.

Mr. Desai did go to college, at M.I.T., one of those schools that recruiters value so highly. It was there, he said, that he learned how to cope with pressure and to work with brilliant people and sometimes feel humbled. But while one’s work at school isn’t inconsequential, he said,it’s not the whole story.” He asserts that despite his degree in computer science, “I’m a terrible developer.”

David Lewin, a professor at the University of California, Los Angeles, and an expert in management of human resources, said that asking what someone could do was an important question, but so was asking whether the person could accomplish it with other people. Of all the efforts to predict whether someone will perform well in an organization, the most proven method, Dr. Lewin said, is a referral from someone already working there. Current employees know the culture, he said, and have their reputations and their work environment on the line. A recent study from the Yale School of Management that uses Big Data offers a refinement to the notion, finding that employee referrals are a great way to find good hires but that the method tends to work much better if the employee making the referral is highly productive.

For his part, Dr. Lewin is skeptical that an algorithm would be a good substitute for a good referral from a trusted employee.

One of Gild’s customers is Square, a San Francisco-based mobile payment system. Like many other high-tech companies, Square is aggressively hiring, and it’s finding the competition for great talent as intense as it was during the dot-com boom, according to Bryan Power, the company’s director of talent and a Silicon Valley veteran. Mr. Power says Gild offers a potential leg up in finding programmers who aren’t the obvious catches.

“Getting out of Stanford or Google is a very good proxy” for talent, Mr. Power said. “They have reputations for a reason.” But those prospects have many choices, and they might not choose Square. “We need more pools to draw from,” he said, “and that’s what Gild represents.”

Gild’s technology has turned up some prospects for Square, but hasn’t led directly to a hire. Mr. Power says the Gild algorithm provides a generalized programming score that is not as specific as Square needs for its job slots. “Gild has an opinion of who is good but it’s not that simple,” he said, adding that Square was talking to Gild about refining the model.

Despite the limited usefulness thus far, Mr. Power says that what Gild is doing is the start of something powerful. Today’s young engineers are posting much more of their work online, and doing open-source work, providing more data to mine in search of the diamonds. “It’s all about finding unrecognized talent,” he said.

MR. DOMINGUEZ has worked at Gild for eight months and has proved himself a talented programmer, Mr. Desai said. But he also said that Mr. Dominguez “sometimes struggles to work in a structured environment.” His co-workers try not to bug him when he’s sitting at his computer, locked into that work zone.

In meetings, Mr. Dominguez speaks his mind. He’s happier, he said, “as long as I can have a say in how the system is built,” or it’s just another system he would have to conform to. He bristles slightly at the growth of the company, which has expanded to 40 people from 10 in the last six months, adding layers of management and bureaucracy.

“The truth is that’s in my nature to do stuff in my own way; inevitably I want to start my own company,” he said, but he’s quick to add: “I do appreciate and the respect the opportunity the company’s given me because I think it’s very clear they hired me on merit. I will always appreciate that.”

Dr. Ming says the young man is both a great find and still an unknown. Of course, he is just a single example, one heralded by the company, but who cannot alone either validate or disprove the method.

“He’s got the lone-wolf thing going on,” Dr. Ming said. “It’s going well early but it could get tougher later on.”

 

Saturday, April 27, 2013

How data is changing the car game for Ford

How data is changing the car game for Ford

By Derrick Harris

http://gigaom.com/2013/04/26/how-data-is-changing-the-car-game-for-ford/

Summary:

The advent of big data is affecting Ford Motor Co. in some significant ways, from how it analyzes its supply chain to the features it puts into its cars.

tweet this

When most people think about how cars are built, they probably think about assembly lines, manufacturing robots, and batteries of safety and performance simulations on massive supercomputers. But at Ford, big data is having a significant impact on the parts and features of those cars before they’re ever part of a design file. From the cars in stock at the dealership to the performance of the engine in a rainstorm, big data is infiltrating nearly every aspect of the Ford experience and the company itself.

Obviously, data is nothing new to the automotive industry — companies have been trying to optimize supply chains and analyze sales numbers for decades — but the advent of big data, as well as related technlogies such as sensors and smartphones, is changing how companies are thinking about data. Ford isn’t alone in its quest to take advantage of these new technologies, either. For example, General Motors collects data from its OnStar system to help lower drivers’ insurance premiums, and also collects lots of data on its Chevrolet Volt electric car that it feeds to drivers via a mobile app. We recently noted how a luxury automobile company used big data software from Aster Data Systems to determine the relationships between malfunctions so it could provide a more thorough and beneficial service-department experience.

But in an industry notoriously unwilling to talk about information technology, Ford’s experiences might shed a lot on what other companies are thinking and doing, as well.

Building a better experience through data

According to John Ginder, manager for systems analytics with Ford Research & Innovation, the company has been doing advanced business modeling for about 20 years, but big data is something else. Today’s technologies are allowing Ford to handle larger, more-diverse datasets than ever before possible, and its efforts are already beginning to bear fruit in numerous places — including in the cars themselves.

The most obvious example of data influencing the driving experience might be the types of data car companies are actually giving back to drivers. At Ford, its Energi line of plug-in hybrid cars generate 25 gigabytes of data per hour that’s then processed and given back to drivers via a mobile app. It tells them about battery life, the nearest charging stations and other data about the vehicle’s performance.

The MyFord mobile app architecture.

Ginder said all that data is the result of a “convergence of need and opportunity.” The opportunity is a way to experiment with collecting and presenting vehicle data on a group of early adopters that’s probably more interested in this type of advanced technology. The need has to do with what Ginder calls “range anxiety” — when drivers are getting used to electric vehicles, they need reassurance they’re not going to run out juice.

However, Ginder said, the company is just scratching the surface of what’s possible, because there aren’t that many of the electric vehicles on the road yet. The goal is to better understand how drivers are using the vehicles and use that information to continuously improve the vehicles and the overall experience. Ford’s Super Duty line of pickup trucks also offers a “crew chief” package that lets bosses monitor the fuel consumption, engine performance and other data about their fleets of vehicles.

Mike Cavaretta, technical leader for predictive analytics and data mining with Ford Research & Innovation, added that Ford is really interested in collecting more data from more vehicles, but noted there’s also a privacy concern that could come into play. The potential of someone knowing where and how you’re driving might not appeal to the mainstream just yet (just look at all that data Tesla collects about its cars and can present if it really wants to), but as with the Energi, data does present some opportunities to improve the customer experience.

The test cars in Ford’s research labs are collecting about 250 gigabytes of data per hour from high-resolution cameras and an array of sensors, Cavaretta noted, and the company is trying to find out what data is most useful and how it might be rolled into production vehicles.

Building betters cars through data

Of course, sometimes the best data isn’t the stuff you see, but the stuff that just makes your car better. Cavaretta said Ford analyzes a lot of social media and other external data in order to figure out, for example, what customers are saying about their vehicles compared with other makes and what problems they’re having.

Opens with the touch of a foot. Source: Ford

In one recent case, the product development team was curious as to whether the Ford Escape sport-utility vehicle should have a standard liftgate (i.e., it opens manually and the rear window can flip open) or a power liftgate in which the glass and the gate are one piece. In the latter option, the gate opens automatically by tapping under the rear bumper with your foot, but the window doesn’t open at all. Regular surveys hadn’t addressed the question, so Cavaretta and his team took to social media, where people were actually talking about it quite a bit and seemed to heavily favor the power liftgate in most cases. It’s now a feature.

Back in 2004, Ford built a self-learning neural network system for its Aston Martin luxury brand that maintains proper engine function by recognizing engine misfires and particular driving conditions and adjusting warnings and performance accordingly.

Ginder said his team has been improving on that technology ever since and actually expanded its use into a system, called Smart Inventory Management System, that lets dealers ensure they have the optimal stock of vehicles and features on their lots. Historically, he said, some dealers were very sophisticated about inventory management, while others were more reactionary (“They just sold a red Mustang,” he joked, “so they think they need to go order another red Mustang.”) With SIMS, all sorts of data about vehicle sales and other locally relevant data from across the country is aggregated in Ford’s big data platform, and the neural network algorithms learn the current patterns so Ford can make better recommendations — whether or not dealers choose to heed the advice.

Selling big data internally

Cavaretta characterizes the division in which he and Ginder work as “an Ernst & Young, but just for Ford,” an internal consultancy (as opposed to Ford’s more-traditional research and development division) in charge of solving business problems via analytics. About 80 percent of those problems come directly from those lines of business, while about 20 percent are the research division’s own ideas. However, although he’s excited about how big data can help his team answer these questions in novel ways, it’s not always an easy sell with other parts of the company.

Mashing up data sources such as social and sales in order to find insights is a pretty easy sell, Cavaretta explained, but getting people to put sensors in everything and collect data every second or with every transaction can still be a bit challenging. In part, this is just a lingering effect of the constraints that legacy technologies imposed on the company. It wasn’t possible to store all this data, so people just got accustomed to the status quo of summarizing data hourly, for example.

Source: Ford

Now, however, he’s pushing them to “dial it down” and collect data at the lowest level possible and as often as possible. In manufacturing alone, he explained, there are between 20,000 and 25,000 parts in any given vehicle, and there’s a supply chain that spans from parts suppliers all the way up to dealerships. Getting a complete view of this process could help drive serious efficiencies and, Cavaretta said, “We don’t see anything but big data technologies that can get us there.”

Other areas where Ford is collecting, or wants to collect, more real-time data is from websites, call centers and the company’s credit-processing arm, he added.

Building big data internally

In order to accomplish their lofty goals, the Research & Innovation analytics team relies heavily on open source technologies, most prominently Hadoop. However, Cavaretta said, they’ve been experimenting with a variety of natural-language processing tools, too, and even did a proof-of-concept with SAP’s HANA in-memory analytic database. The NLP tools were first turned on text analysis of internal surveys and dealer network documents, but now are used pretty heavily on social media and other web data.

Their team has some systems numbering in the dozens of nodes in its own building, but on weekends it’s able to borrow high-performance computing cycles from Ford’s Numerically Intensive Computing Center next door in order to model recommendation engines and other tasks that demand serious computing power.

But as a part of a specialized research division, the work that Ginder, Cavaretta and their team do on everything from Hadoop to visualization with tools like Tableau isn’t automatically ready for primetime. In fact, Cavaretta said, it looks at “what’s the art of the possible” and tries to show the value of it. It’s like a vanguard, he added, going out and seeing what’s ahead and then reporting back.

At that point, projects are often handed off to Ford’s central IT team that actually puts the technologies into production. A system that took the research team weeks to deploy and start deriving insights from might take IT months to make production-ready. However, Ginder added, his team can’t just throw stuff over the wall and abandon it — it has to collaborate with the IT team and individual departments throughout the project’s lifecycle.

An important part of this cross-company relationship — and something many CIOs have likely heard before — is having data scientists on board that can see the world through the eyes of both technologists and businesspeople, two groups that often have different concerns and goals in mind. “We look for people who can bridge those worlds,” Ginder said. “It’s hard to find these people, but they’re hugely important to organizations.”

Feature image courtesy of Shutterstock user PhotoSmart.