Friday, June 28, 2013

Siri's Creators Demonstrate an Assistant That Takes the Initiative

Siri’s Creators Demonstrate an Assistant That Takes the Initiative

An SRI project aims to build a powerful predictive assistant for office workers.

Why It Matters

Computer interfaces that improve worker productivity can have a huge impact.

In a small, dark, room off a long hallway within a sprawling complex of buildings in Silicon Valley, an array of massive flat-panel displays and video cameras track Grit Denker’s every move. Denker, a senior computer scientist at the nonprofit R&D institute SRI, is showing off Bright, an intelligent assistant that could someday know what information you need before you even ask.

Initially, Bright is meant to cut down on the cognitive overload faced by workers in high-stress, data-intensive jobs like emergency response and network security. Bright may, for instance, aid network administratorsin trying to stop the spread of a fast-moving virus by quickly providing crucial infection information, or help 911 operators send the right kind of assistance to the scene of an accident. But like many other technologies developed at SRI, such as the digital personal assistant Siri (now owned by Apple), Bright could eventually trickle down to laptops and smartphones. It might take the form of software that automatically brings up listings for your favorite shows when it thinks you’re about to sit down and watch TV, or searches the Web for information relevant to your latest research project without requiring you to lift a finger.

Already some assistant software, such as Google Now for Android smartphones, tries to predict what information a user may need and serve it up automatically. It does this by, for example, recognizing that the user is waiting at a bus stop and delivering bus timetables. The aim of Bright is to develop something even more sophisticated and capable in an office setting. But the big challenge for Bright and similar projects is: how do you learn from a relatively small amount of information?

Originally created by Stanford University as a research institution in 1946 (it’s been operating independently since 1970), SRI International, based in Menlo Park, California, has developed key technologies including the computer mouse, the LCD, and even the first twinklings of the Internet, called ARPAnet. In recent years, it has had success in the artificial-intelligence field with Siri, which was spun out of a project SRI did for the Department of Defense’s Defense Advanced Research Projects Agency, or DARPA, called CALO (that’s “cognitive agent that learns and organizes”).

Denker describes Bright as a “cognitive desktop” and “a desktop that really understands what you’re doing, and not just for you, but also in a collaborative setting for people.” In its current setup, three cameras stare out at her; a monitor shows where she’s looking and displays a real-time log of every action she takes, as well as a familiar-looking computer desktop of files and folders. When she uses the monitor in front of her to open an e-mail from Wells Fargo bank requesting a meeting, for example, Bright records all her actions on a monitor off to the left, noting that she opened the message, that she spent time looking at it (rather than just gazing elsewhere on the screen), and that she closed it.

As Denker demonstrates Bright’s nascent capabilities, it’s not hard to imagine the technology easing everything from scheduling tasks to searching the Web. She explains that her team is trying to adapt existing computer science techniques that try to increase efficiency by anticipating what information will be needed next and testing different actions in advance to speed up response time. Bright, she says, uses the same ideas to anticipate what the user will want to do, so it requires additional equipment to monitor the user. A touch-sensitive display can track finger touches, and hand motions—such as waving—are tracked too.

While it is being developed for cybersecurity and emergency response, Bright could be tailored for other types of users. In schools, for example, Bright might be able to determine that a student is struggling and adjust itself to better meet his or her needs.

There’s a long way to go, however. The system is currently focused on “cognitive indexing”—the mechanism that ties various clues together and then tries to predict what is important. The team behind Bright also needs to build its abilities to predict interests and automate tasks. And before it can be rolled out anywhere, Bright needs to learn how to study what you’re using your computer for.

Getting to know a user is difficult, says Bill Mark, vice president of information and computing sciences at SRI and one of the principal investigators behind CALO. Mark calls this the “small-data problem”; while “big data” efforts focus on gleaning insights from mountains of information, systems like Bright are looking for patterns in much smaller quantities, and this can be very tricky. The limited data set, combined with users’ tendency to change behavior, is very unfriendly to pattern-finding algorithms, he says: “We’re not putting in that much data. These machine-learning algorithms like to generalize over very large amounts of data.”

There are plenty of other challenges. Krzysztof Gajos, an assistant professor of computer science at Harvard who also spent a year working on CALO, notes that one of the difficulties in building intelligent interactive systems is figuring out how to distinguish mandatory tasks like office work from voluntary tasks like playing games. For office-related tasks, he says, it’s hard to design automation in a way that leaves the user feeling in control and seems worth using even though it will occasionally screw up.

“If you look back to systems like the Microsoft Clippy, you can see an example of a system that failed at that,” Gajos says. “The few times it failed were just so aggravating that it overshadowed any benefits the system might have provided for many users.”

 

Thursday, June 27, 2013

How Big Data can revolutionize health care

http://www.politico.com/story/2013/06/how-big-data-can-revolutionize-health-care-93449.html

 

How Big Data can revolutionize health care
By:
Eric Dishman
June 26, 2013 10:01 PM EDT

Twenty-four years ago when I was a sophomore in college, I began experiencing a series of unusual fainting spells. The spells eventually landed me in student health, where they ran lab work and uncovered some pretty serious kidney problems. After six months of tests with six doctors across two hospitals, I received my diagnosis: I had two rare diseases that would eventually destroy my kidneys. I had cancer-like cells in my immune system that needed immediate treatment. I would never be eligible for a kidney transplant, and I wasn’t likely to live more than two years.

That diagnosis changed my life forever. But after I found myself preparing to die according to their schedule, a fellow patient snapped me out of it. She dragged me to a medical library and dug up some research that showed the diagnoses didn’t fit me at all. She told me to wake up and take control of my health. And I did.

I became committed to creating a personal health system that wasn’t focused just on increasing my chances of survival but on improving my quality of life. I actively pursued access to technologies, data and cutting-edge treatments that helped maximize my time with friends and family. I strived to have my care at home as much as possible, away from hospitals and outpatient centers, which can be dangerous places for my compromised immune system.

Perhaps the most important step I took was having my genome sequenced. I’m lucky to be one of the 47,000 people on the planet to have this raw data on all 6 billion letters of my DNA. In my case, my physician said it changed everything we knew about my diseases and course of treatment — which had been wrong for two decades. The data they were able to uncover paved the way for a kidney transplant I was never supposed to have and a life I was never supposed to live.

I’m just one of millions of examples of how data can transform our lives. But there are still many challenges that stand in the way of us scaling truly “personal” health both nationally and internationally. Right now, these technologies are limited to top research hospitals and universities, not only because of how expensive and early they are but because there are still a number of privacy, security and education issues that must be addressed. Now is the time for leading experts to determine how to overcome these barriers and make personal health attainable for everyone.

That’s why this week, I joined Sen. Ron Wyden (D-Ore.) and a host of industry experts at a Bipartisan Policy Center forum to explore the promise, challenges and policies that are critical to encouraging innovation and improving health care through Big Data. We explored the ways Big Data is transforming health care today, its promise for the future and how we can get there.

Right now, our population-based health care system leads us to draw conclusions for patients based on what we know of others and how we care for them. But Big Data provides us an opportunity to transition to a personal care system. Rather than making assumptions based on what has worked for other people, this personal view would allow us to take data about a patient’s genomes, medical history and behaviors to construct a virtual model that would help predict which treatments will be most effective and customize them to an individual — improving quality of life for the patient and saving the delivery system money.

But how do we scale a data-driven, personal health system so access is afforded to everyone? One of the most important aspects is care coordination — ensuring your team of health care providers communicates regularly, has access to the same information and can make informed care decisions based on complete sets of data across medical disciplines. Data standardization is a crucial component to coordinated care. It allows us to ensure that data sets held by doctors, hospitals and health plans are consistent and can be shared efficiently.

Equally important to the conversation about Big Data and health care are issues relating to privacy and security. We have often debated the major challenges to securing our personal data and keeping it private. But rather than viewing Big Data as a roadblock to privacy or a burden to security — we should think about the different ways people are solving these challenges through Big Data, not in spite of it.

There’s no question that Big Data can transform our health care system by improving access and quality while driving down costs. My hope is that by talking about these issues with some of the brightest minds in the field, we can determine how best to overcome these challenges and make data-driven, personal health a reality for everyone.

Eric Dishman is an Intel Fellow and general manager of the Health & Life Sciences Group at Intel.

 

Tuesday, June 25, 2013

Connecting the Dots, Missing the Story

Connecting the Dots, Missing the Story

With Big Data, the government doesn’t need to know the “why” behind anything.

Could Big Data have prevented 9/11? Perhaps—Dick Cheney, for one, seems to think so. But let's consider another, far more provocative question: What if 9/11 happened today, in the era of Big Data, making it all but inevitable that all the 19 hijackers had extensive digital histories?

It used to be that one's propensity for terrorism was measured in books or sermons. Today, it's measured in clicks. It's not that books or sermons no longer matter—they still do—it's just today they are consumed digitally, in a way that leaves a trail. And that trail allows us to establish patterns. Are the books you bought on Amazon today more radical than the books you bought last month? If so, you might be a person of interest.

The Tsarnaev brothers, who allegedly bombed the Boston Marathon earlier this year, are of this new breed of terrorists. The brothers felt at home in the world of Twitter and YouTube. And some of the videos reportedly favorited by Tamerlan, the older brother, are clearly of extremist nature. Had someone been analyzing the brothers' viewing habits in real time, a great tragedy might have been averted.

Advertisement

The good news—at least to Big Data proponents—is that we don't need to understand what any of these clicks or videos mean. We just need to establish some relationship between the unknown terrorists of tomorrow and the established terrorists of today. If the terrorists we do know have a penchant for, say, hummus, then we might want to apply extra scrutiny to anyone who's ever bought it—without ever developing a hypothesis as to why the hummus is so beloved. (In fact, for a brief period of time in 2005 and 2006, the FBI, hoping to find some underground Iranian terrorist cells, did just that: They went through customer data collected by grocery stores in the San Francisco area searching for sales records of Middle Eastern food.)

The great temptation of Big Data is that we can stop worrying about comprehension and focus on preventive action instead. Instead of wasting precious public resources on understanding the “why”—i.e., exploring the reasons as to why terrorists become terrorists—one can focus on predicting the “when” so that a timely intervention could be made. And once someone has been identified as a suspect, it's wise to get to know everyone in his social network: Catching just one Tsarnaev brother early on may not have stopped the Boston bombing. Thus, one is simply better off recording everything—you never know when it might be useful.

Gus Hunt, the chief technology officer of the CIA, said as much earlier this year. "The value of any piece of information is only known when you can connect it with something else that arrives at a future point in time,” he said at a Big Data conference. Thus, “since you can't connect dots you don't have … we fundamentally try to collect everything and hang on to it forever." The end of theory, which Chris Anderson predicted in Wired a few years ago, has reached the intelligence community: Just like Google doesn't need to know why some sites get more links from other sites—securing a better place on its search results as a result—the spies do not need to know why some people behave like terrorists. Acting like a terrorist is good enough.

As the media academic Mark Andrejevic points out in Infoglut, his new book on the political implications of information overload, there is an immense—but mostly invisible—cost to the embrace of Big Data by the intelligence community (and by just about everyone else in both the public and private sectors). That cost is the devaluation of individual and institutional comprehension, epitomized by our reluctance to investigate the causes of actions and jump straight to dealing with their consequences. But, argues Andrejevic, while Google can afford to be ignorant, public institutions cannot.

"If the imperative of data mining is to continue to gather more data about everything," he writes, "its promise is to put this data to work, not necessarily to make sense of it. Indeed, the goal of both data mining and predictive analytics is to generate useful patterns that are far beyond the ability of the human mind to detect or even explain." In other words, we don't need to inquire why things are the way they are as long as we can affect them to be the way we want them to be. This is rather unfortunate. The abandonment of comprehension as a useful public policy goal would make serious political reforms impossible.

Forget terrorism for a moment. Take more mundane crime. Why does crime happen? Well, you might say that it’s because youths don't have jobs. Or you might say that's because the doors of our buildings are not fortified enough. Given some limited funds to spend, you can either create yet another national employment program or you can equip houses with even better cameras, sensors, and locks. What should you do?

If you're a technocratic manager, the answer is easy: Embrace the cheapest option. But what if you are that rare breed, a responsible politician? Just because some crimes have now become harder doesn't mean that the previously unemployed youths have finally found employment. Surveillance cameras might reduce crime—even though the evidence here is mixed—but no studies show that they result in greater happiness of everyone involved. The unemployed youths are still as stuck as they were before—only that now, perhaps, they displace anger onto one another. On this reading, fortifying our streets without inquiring into the root causes of crime is a self-defeating strategy, at least in the long run.

Big Data is very much like the surveillance camera in this analogy: Yes, it can help us avoid occasional jolts and disturbances and, perhaps, even stop the bad guys. But it can also blind us to the fact that the problem at hand requires a more radical approach. Big Data buys us time, but it also gives us a false illusion of mastery.

We can draw a distinction here between Big Data—the stuff of numbers that thrives on correlations—and Big Narrative—a story-driven, anthropological approach that seeks to explain why things are the way they are. Big Data is cheap where Big Narrative is expensive. Big Data is clean where Big Narrative is messy. Big Data is actionable where Big Narrative is paralyzing.
The promise of Big Data is that it allows us to avoid the pitfalls of Big Narrative. But this is also its greatest cost. With an extremely emotional issue such as terrorism, it's easy to believe that Big Data can do wonders. But once we move to more pedestrian issues, it becomes obvious that the supertool it's made out to be is a rather feeble instrument that tackles problems quite unimaginatively and unambitiously. Worse, it prevents us from having many important public debates.

As Band-Aids go, Big Data is excellent. But Band-Aids are useless when the patient needs surgery. In that case, trying to use a Band-Aid may result in amputation. This, at least, is the hunch I drew from Big Data.

This article arises from Future Tense, a collaboration among Arizona State University, the New America Foundation, and Slate. Future Tense explores the ways emerging technologies affect society, policy, and culture. To read more, visit the Future Tense blog and the Future Tense home page. You can also follow us on Twitter.

 

Thursday, June 20, 2013

European Parliament: EU citizens' data must be properly protected against US surveillance

PRISM: EU citizens' data must be properly protected against US surveillance

LIBE Citizens' rights − 20-06-2013 - 13:03

http://www.europarl.europa.eu/news/en/pressroom/content/20130617IPR12352/html/PRISM-EU-citizens%27-data-must-be-properly-protected-against-US-surveillance

 

The US PRISM internet surveillance case highlights the urgent need to pass legislation to protect EU citizens' personal data, most MEPs agreed in Wednesday's Civil Liberties Committee debate with Justice Commissioner Viviane Reding. MEPs also called for safeguards for personal data transferred outside the EU.

"The PRISM case was a wake-up call that shows how urgent it is to advance with a solid piece of legislation" on data protection, said Commissioner Reding in her opening remarks. Reporting back on her 14 June meeting with US Attorney General Eric Holder in Ireland, she said: "We agreed to set up a transatlantic group of experts to address concerns".

Group of experts to start work in July

"What is happening now is really shocking: (...) we cannot allow Americans to spy on EU citizens (...) even if it is a security matter", said Veronique Mathieu (EPP, FR). She also stressed the need to speed up work on the new EU data protection legislation and asked to be fully informed on the work of the above expert group.

"Are those experts known already? When are they meeting?" asked Judith Sargentini (Greens/EFA, NL). Timothy Kirkhope (ECR, UK) called for "a proper investigation" to "gather facts and details". He welcomed the use of IT tools to fight terrorism, provided it is always done in a "lawful way", and expressed support for the Commission.

Ms Reding confirmed that "not all the questions have been answered in Ireland". She stressed that EU citizens' data should have the same protection as those of US citizens and announced that the first meeting of the expert group should be held in July.

She also agreed that new data protection rules must be agreed quickly and "apply to all companies that operate in the EU", regardless of nationality or headquarters country.

Safeguarding data transferred outside the EU

"Our friends and partners go behind our backs and fish our citizens' data: this is dramatic" said Birgit Sippel (S&D, DE). "It is not true that this data is only used to fight terrorism. It is also used for immigration control", she continued, stressing that "We need to ensure that people's data are protected, whether or not they are suspected" of a crime".

"Our allies treat us not as friends but as suspects", said Sophia in't Veld (ALDE, NL). The EU needs to "show some backbone" and say where the limits are, she added.

Quizzing Ms Reding about a proposed data transfer safeguard, which would oblige third country authorities to request data through legal channels, she asked "Why between the first leaked draft and the official draft (...) was the jurisdiction deleted?" "Have Americans have been going through the draft with a red pen?"

Ms Reding replied that the red line is "never agree to go under the 1995 (data protection) directive standards". She added that the data transfer safeguard is currently just a recital in the draft legislation, but if Parliament "wants to make it an article I have no objection".

Committee on Civil Liberties, Justice and Home Affairs

In the chair: Juan Fernando López Aguilar (S&D, ES)

Procedure: debate

 

Tuesday, June 18, 2013

Review of Jim Manzi's book

Experiments in Democracy 

Jeremy Rozansky

http://www.thenewatlantis.com/publications/experiments-in-democracy

There is a timeworn joke about a man who, walking home after midnight, comes across a drunkard on his knees under a streetlamp, patting the pavement with his hands. “What are you looking for?” asks the passerby. “I can’t find my keys,” says the drunk. The passerby, being a kindly man, gets on his own knees and joins in the search — to no avail. Finally, after a quarter of an hour, he asks, “Are you sure this is the spot where you lost your keys?” “Oh, not at all,” the drunk answers, pointing: “I dropped them on my front stoop across the street.” “Then why are we looking on this side of the street!” “The light’s better over here.”

This dusty old story may be good for a laugh or a groan, but it also serves as a parable about social science and policymaking. The light shone by the methods of social science is limited, sometimes too dim, and may not illuminate what is important or useful for policymakers — but a social scientist, like the drunkard in the joke, is going to scour the ground he can see. For most public policy questions, the streetlamp under which social scientists search for useful and important knowledge is what is known as the “econometric method.” The econometric method takes a large sample of observed human interactions and creates a model to predict what a particular policy or treatment will do, controlling for a large number of variables like race, sex, income, habits, location, and attitudes.

In his debut book Uncontrolled, entrepreneur and policy analyst Jim Manzi argues that social scientists and policymakers should instead adopt the “experimental method.” The essential tool of this method is the randomized field trial (RFT), a technique that already informs many of our successful private enterprises. Perhaps the best known example of RFTs — one that Manzi uses to illustrate the concept — is the kind of clinical trial performed to test new medicines, wherein researchers “undertake a painstaking series of replicated controlled experiments to measure the effects of various interventions under various conditions,” as he puts it.

The central argument of Uncontrolled is that RFTs should be adopted more widely by businesses as well as government. The book is helpful and holds much wisdom — although the approach he recommends is ultimately just another streetlamp in the night, casting a pale light that tapers off after a few yards. Much still lies beyond its glow.

The relative merits of different social science methodologies may seem to be a subject of merely academic interest. But Manzi persuasively shows why it is important for policymakers to think about science, beginning with a recent example illustrating the reality that policymakers often don’t know whether or not their plans will work until they actually put them into place.

In early 2009, President Obama and congressional Democrats proposed to jumpstart the U.S. economy out of recession by passing a fiscal-stimulus bill. Whether this infusion of about $800 billion in deficit-financed government spending would improve the economy was not at all certain, but the White House could point to several economists, including a handful of Nobel laureates, who predicted that every dollar spent in stimulus would increase the national income to the tune of a dollar-and-a-half by “stimulating” demand. However, other Nobel laureates disagreed, arguing that the bill was not worth its price tag: The stimulative effects of the deficit spending would be diminished by (among other things) expectations of future taxation, and so the return on every dollar of deficit spending would be closer to zero. Paul Krugman, a pro-stimulus economist, lambasted anti-stimulus economists as priests from the “Dark Age of macroeconomics.” Other academics hit back, alleging that it was the Keynesian Krugman who was himself stuck in the Dark Ages.

A policymaker or citizen looking at both arguments could see that each side made models of fiscal spending and GDP growth, and each had a smattering of econometric analyses of past fiscal spending that they said validated their models. But as Manzi wryly notes, “the only thing an observer could say with high confidence before the stimulus program launched was that at least several Nobel laureates in economics would be directionally incorrect about its effects.” In such a policy stalemate — professor against professor, Nobelist against Nobelist, Princeton against Chicago, all appealing ultimately to their own authority — what can we lay citizens do but throw up our hands and concede ignorance? And what does this mean for democratic self-governance?

A similar policy stalemate arose during the 2012 presidential contest. Mitt Romney had promised both to cut tax rates by 20 percent and to maintain existing levels of revenue. He claimed he would do so by eliminating unspecified carve-outs and loopholes. The Obama campaign cited a Tax Policy Center study claiming that the only way Governor Romney could achieve both goals would be by eliminating a variety of tax deductions popular with the middle class. Unless Romney did so, the study argued, he would either have to renege on his 20 percent rate-cut promise or fail to meet his revenue goals. When challenged on this subject in the first debate against President Obama, Romney fought study with study:

Now, you cite a study. There are six other studies that looked at the study you described and say it’s completely wrong. I saw a study that came out today that said you’re going to raise taxes by $3,000 to $4,000 on middle-income families. There are all these studies out there.

Being a man for whom data analysis is perhaps a more intuitive craft than politics, Romney probably could have explained the differing assumptions behind the various studies and evaluated the validity of each. Perhaps he could have compared the studies in detail and shown the method of one to be obviously superior. Instead, he made two moves that well illustrate an all-too-common pathology of contemporary politics. First, he did not challenge the assumptions or methods used by Obama’s preferred study. Second, he merely cited other studies, presenting them as equal and opposite authorities to the Tax Policy Center’s. Romney knew that even though the electorate is disposed to believe what “studies” show, most voters lack the expertise, the time, and the inclination to compare the validity of contradictory studies. Thus, so long as social science is not unanimous, its authority will be inconclusive for the voter.

But a self-contradictory authority is, in effect, no authority at all. Voters tend to ignore much of the daily business of government as hopelessly complicated — but while understandable, this state of affairs is a recipe for both shallow political debates and rule by technocrats. Manzi’s book, which addresses the underlying assumptions of different social science methods, offers a solution to the problem of dueling Nobelists or think-tank studies — a solution that promises not only better policy but, more importantly, better democratic politics.

Uncontrolled is in many ways a book about the scientific method, and Manzi begins by staking out a position on an important question about science that most Americans rarely think to ask: What is science — in varieties both “hard” and “soft” — good for? Manzi’s answer, which may startle some readers, and may even offend some scientists, is not “finding truth.” Rather, the answer is utility. Science, Manzi argues, chiefly aims to discover effective means for reaching the ends that human beings choose through other forms of reflection, such as philosophy, theology, or the arts. While common sense can be useful for identifying the means to accomplish desired ends, “the key value of science is that it provides causal rules that are nonobvious, that is, that extend beyond common sense.”

Still, as Manzi describes, the causal rules that science gives us are not definitive. In fact, it is fundamentally impossible to know any causal rules with certainty: Even when we see one kind of event precede another kind of event under many conditions, we can never be completely certain that the first kind causes the second. Every time we see someone let go of a rock, we see it fall, and so we infer that dropping a rock causes it to fall. But although this apparent causal relationship has always served as a reliable rule, we cannot know with perfect certainty that the next time someone drops a rock it will fall, and even if it does fall, we cannot conclude that dropping it is what causes it to fall.

This limitation on both science and common sense, which philosophers call the problem of induction, can of course seem preposterous when it calls into question our ability to draw any conclusions about causation. But there are always hidden factors that can account for apparent cause-effect relationships. Any ancient thinkers who might have seen dropping and falling as inviolably linked would have been brought up short by the later discovery that a rock will float rather than fall when it is released in outer space. When we drop a rock here on the ground, the proximity to a large gravitational body is a hidden conditional. So too would the ancients have been unlikely to anticipate that some rocks let go here on earth sometimes will not fall, for example, when a large magnet is present.

So we cannot demonstrate causal rules to be true, and should not try to; we can only demonstrate causal rules to be false or in need of amendment, when we find a hidden conditional. The practice of science involves making useful assumptions and gradually and meticulously adding nuance to our assumptions in order to make them more useful. Manzi both narrows the focus of science and demystifies its methods, bringing it down to its rightful place among the useful arts.

As we will see, in lowering the sights of science, particularly social science, Manzi points toward why he prefers the experimental method over the econometric method. But the crucial reason for Manzi’s preference for RFTs over econometrics is more technical — an idea Manzi calls “causal density.” In seeking the cause of a given effect, the general approach of science is to isolate each potential variable that might play a causal role and manipulate it while holding everything else equal. By performing experiments and measuring the apparent effects of each isolated cause, scientists can make useful assumptions about cause and effect. While we can never control for everything that could possibly be causally significant — in part because we never know where hidden conditionals might be lurking — we can be satisfied with an assumption if the cause seems well isolated and the effect is reliably observed when we replicate the causal conditions we think are relevant. The model we create from these observed rules can never completely capture the actual system, and we can never know we have included every hidden conditional, but it can be a useful predictor.

Still, some systems are more easily modeled than others; it all depends on how easily the relevant conditionals are isolated. Creating models is appropriate, and often relatively straightforward, in a field like astrophysics, where objects are far away from each other and replicating an observed rule is easy given the vast expanse of data. Astrophysics is a science of low causal density.

Social science, by contrast, has very high causal density. The subjects — human beings and their institutions — are complexly intertwined. It is very difficult to isolate a conditional, and it is impossible (or at least terribly intrusive and often unethical in practice) to hold all other things equal. Think of our debates over education, and all the variables that can affect whether a child becomes well educated: the resources available to his school, whether his parents value education and how much they do, the skill of his teacher, the influence of his classmates, his I.Q., self-motivation, nutrition, access to school supplies, the particular textbooks he reads (or doesn’t read), and so forth. Now consider how all these variables interact with one another. Would anyone really expect two groups of students — one in, say, Moline, Illinois and the other across the river in Davenport, Iowa — to react exactly the same to a new fourth-grade math curriculum, even if virtually every feature of their lives that social scientists can measure were the same? Would we be reasonable to chalk up the observed differences to the differences between Illinois and Iowa? The answer to both these questions is most likely “no” — because human beings and institutions are just too complicated to justify such claims. Manzi vividly compares social science to medical research in which every test tube is poorly cleaned and contains a foreign residue — a hidden conditional.

The econometric method now dominates the social sciences because it helps to cope with the problem of high causal density. It begins with a large data set: economic records, election results, surveys, and other similar big pools of data. Then the social scientist uses statistical techniques to model the interactions of sundry independent variables (causes) and a dependent variable (the effect). But for this method to work properly, social scientists must know all the causally important variables beforehand, because a hidden conditional could easily yield a false positive.

The experimental method, which Manzi prefers, offers a different way of coping with high causal density: sidestepping the problem of isolating exact causes. To sort out whether a given treatment or policy works, a scientist or social scientist can try it out on a random section of a population, and compare the results to a different section of the population where the treatment or policy was not implemented. So while econometric models aim to identify which particular variables are responsible for different results, RFTs have more modest aims, as they do not seek to identify every hidden conditional. By using the RFT approach, we may not know precisely why we achieved a desired effect, since we do not model all possible variables. But we can gain some ability to know that we will achieve a desired effect, at least under certain conditions.

Strictly speaking, even a randomized field trial only tells us with certainty that some exact technique worked with some specific population on some specific date in the past when conducted by some specific experimenters. We cannot know whether a given treatment or policy will work again under the same conditions at a later date, much less on a different population, much less still on the population as a whole. But scientists must always be cautious about moving from particular results to general conclusions; this is why experiments need to be replicated. And the more we do replicate them, the more information we can gain from those particular results, and the more reliably they can build toward teaching us which treatments or policies might work or (more often) which probably won’t. The result is that the RFT approach is very well suited to the business of government, since policymakers usually only need to know whether a given policy will work — whether it will produce a desired outcome.

Manzi offers plenty of evidence of the efficacy of RFTs. In the business world, he himself has built a career by consulting with companies to help them run RFTs and improve their profits. One of his former consulting colleagues founded the credit card company Capital One; its rise in an industry with high barriers to entry is the result of following the RFT approach, conducting thousands of experiments. Outside of business, RFTs also contributed to some of the great broken-windows innovations in criminology that helped bring down the crime rate two decades ago. Some RFTs help by disproving theories. Milton Friedman’s idea that a negative income tax — essentially a guaranteed minimum income — could replace the welfare-state bureaucracy and eliminate welfare’s perverse incentives was tested in several massive programs. Policymakers discovered that the negative income tax actually exacerbated some of the perverse incentives Friedman was hoping to fix. In the end, the experiment was most useful in that the failure of Friedman’s hypothesis pointed to a different way to fix welfare: the work requirements in the 1996 reform bill. More recently, RFTs have made their way into electoral politics, first with Rick Perry’s successful gubernatorial campaign and later with Barack Obama’s reelection.

Manzi looks at the RFTs conducted for public policy questions — which are dwarfed by the number conducted in the business world — and draws a few general conclusions. First, most policy experiments don’t work, so policymakers should not be too enthralled with their own designs. Second, programs that focus on “raising skills or consciousness” tend to fail because people’s character is hard to change; the ones that do work tend to be the ones that focus on changing behavior by changing incentives. Third, grand counterintuitive or surprising causal effects that make up much of the pop-social-science literature are generally not true or are only half-true; although science is supposed to find non-obvious rules, social science mostly confirms (or refutes) ideas already held by common sense. And so Manzi does not anticipate that the greater use of RFTs will revolutionize policymaking. His expectations are more modest.

With these rules in mind, Manzi urges policymakers to embrace the use of randomized field trials. He recommends not just that we use RFTs to test specific policies but also, more broadly, that we adopt an experimental disposition. Such a disposition entails a general deference to decentralized systems — like federalism and the market — that encourage trial-and-error improvements. This preference for decentralization should not be taken to extremes: some interventions, like changes in interest rates, must be done at a large scale. Rather, a disposition toward trial and error will encourage experimentation by the most local competent political authority, and by firms of all sizes.

Manzi’s prescription is in many ways deeply conservative. History is often seen, incorrectly, as encompassing a few revolutionary moments when new truths are discovered. A more accurate view, Manzi posits, holds that history is simply a long record of trial and error that has built up and retained a reserve of “implicit knowledge.” Under this view, the arrangements that have withstood experiments over the generations must seem workable and wise. Thus a preference for the status quo is rational; Manzi places the burden of proof on those who advocate radical change.

There is, however, a tension between, on the one hand, the restless experimentation that Manzi recommends, and on the other, the conservative bias toward the present order. Manzi does not presume that he can cleanly resolve this tension, but he does offer some ideas for how it can — and has — been managed. For instance, the use of decentralized systems can offer an advantage because within them growth is more incremental and therefore less disruptive. Some government policies can also reduce the tension between innovation and social cohesion by improving the adaptability of vulnerable individuals to the jarring effects of economic growth. Public education is one way of improving adaptability, and redistributive policies can be another. Manzi warns, however, that the welfare state can stifle trial-and-error processes, especially given its tendency to be highly centralized. But since it is necessary to smooth over some of the harsh edges of innovation, Manzi writes that, as much as possible, the welfare state should be structured so as not to choke the very innovation that it is meant to make palatable.

There are of course some caveats that those disposed to trial and error should keep in mind. For one, policymakers cannot always conduct a satisfactory experiment — one randomized and replicable — so they will sometimes have to settle for implementing an idea without a real record of success. In these cases Manzi simply encourages policymakers to try out new ideas on a smaller scale so as to reduce the risk.

But naturally, there are situations where even such small-scale attempts are impossible. To revisit the example of the 2009 stimulus, the theory behind the policy was that deficit spending would improve GDP growth in the midst of a recession. But the states, our laboratories of democracy, could not conduct their own stimulus experiments because nearly all of them are legally required to have balanced budgets, preventing them from running the deficits that the stimulus required. Nor could the federal government experiment with different ideas on a smaller scale. If, say, the federal government selected one hundred counties for workers to have payroll-tax breaks, workers would flood into those counties. Policies like the stimulus bill are not made by social scientists testing theories, but by politicians facing a national crisis. Many policy decisions are made under circumstances that force large-scale and unpredictable — and therefore risky — ventures. And crises do not give policymakers enough time to learn from their mistakes.

In the final chapter of the book, Manzi offers a handful of specific policy recommendations for how the government can “embed a trial-and-error process within humane constraints.” First, and most obviously, we should decentralize and start conducting more and better experiments, especially at the state level. For example, we could have a better sense of the costs and benefits of universal preschool if several counties, states, or foundations were to undertake more rigorous experiments. Much of the hype about the benefits of universal preschool has been fueled by the intense Perry and Abecedarian preschool projects of the 1960s, but since so much of the success of those unusual projects was attributable to the singular talents of the individuals involved, it seems inappropriate to cite those projects as evidence for the proposition that preschool is generally a wise investment.

The federal government’s involvement in policy experimentation has received little public attention. The Centers for Medicare and Medicaid Services (CMS) runs a number of demonstration projects, although they are generally less useful and replicable than true RFTs. Nonetheless, the CMS experience with experimentation is an apt illustration of Manzi’s broader point: There seems to be a wall between the knowledge gained and broader policymaking. For instance, many of the savings in Medicare that Obamacare hopes to achieve are gained through the introduction of what is known as “bundled payments”: payments based on standardized rates for the treatment of specific clinical conditions, regardless of the services employed in treating that condition. Yet CMS has already run a demonstration project on this exact topic and learned that we should be skeptical of bundled payments saving even one percent of Medicare costs. As long as the experiments are conducted without fanfare and their results ignored, policymakers can go on hyping bad ideas as if they might work.

Manzi also wants to see a proliferation of different and creative policy experiments that allow state governments to better achieve the goals of federal mandates and programs. He therefore proposes that the states and the federal government agree to a simple trade: The federal government will broadly waive regulations governing program design and other federal mandates on a trial basis if the states accept certain standards of experimental rigor. Manzi also calls for the creation of an organization within the federal government to create and enforce standards for the design and interpretation of randomized public policy experiments, much like what the Food and Drug Administration does for clinical trials in the field of medicine.

Manzi also offers a set of several proposals to build human capital. He calls for a universal voucher program for public education, a bias in visa-granting toward highly skilled immigrants, and an increase in the share of the federal budget dedicated to research and development. He also suggests that policymakers involved in education and R&D be more “ruthless” in killing off ideas that do not lead to helpful results, while “pointing a fire hose of money at those that succeed.” Unfortunately, he has no ideas for how to craft such a bureaucracy.

He also urges a rethinking of the welfare state, in which policymakers list all its discrete tasks, creating separate programs for each that employ market mechanisms where feasible. For instance, Social Security is both a means of forcing workers to save their income and a safety net for the vulnerable elderly. These are separate tasks, and the first could be accomplished by a managed market of private accounts, while the second could be achieved through direct federal transfers.

What is striking about Manzi’s policy ideas is how commonsensical and unoriginal they are in the world of conservative policy analysis. Of course we would rather have a Ph.D. in chemistry than a high-school dropout join our workforce! Of course we should introduce vouchers and transparent transfer payments where possible. Yet Uncontrolled is billed as revealing the “surprising payoff of trial-and-error.” Nothing is too surprising in this last chapter. In fact, Manzi’s concrete policy proposals do not directly follow from knowledge gained through specific trial-and-error processes or experiments. With the exception of school vouchers, which Manzi discusses earlier in the book, these policy prescriptions have little experimental evidence to support them.

The lack of cited experiments should lead us to wonder if new data will find Manzi is wrong about the positive effects of high-skilled immigration or the unbundling of Social Security. Manzi’s method of choice does not lead directly to his policies of choice. Instead, his method follows from principles which also more or less guide him, in parallel, to his policies. This could be seen as the great error of Uncontrolled: The method he uses many pages to argue for is really not very proximate to the policies the reader is encouraged to prefer. Instead, this should be seen as a point in the book’s favor. Social science has always had a hubristic ambition to draw a straight line from its general findings to particular policy proposals. Even at its best, however, social science can offer us only partial knowledge of human affairs. But because we live in the world, we must still muddle along — and Manzi’s book offers a few sensible ways to muddle better, without claiming to do much more than that.

To Manzi’s great credit, he never describes the randomized field trial as a silver bullet, only a sensible alternative to the overrated econometric method. He summarizes the academic critiques of RFTs — most notably from economist James Heckman — and basically agrees with them. In practice, it is terribly difficult to truly randomize a field test for a public policy. Small-scale experiments will also not be useful in every policy context. As social scientist Robert A. Moffit wrote about experiments with welfare reform, “the RFT methodology is poorly suited to measuring the effects of structural, system-wide reforms.” And, again, the great strength of RFTs — that they can show whether a policy works — does not mean that they offer insight into why it does or does not work. The econometric method still has the advantage of breaking results down into the sort of piecemeal conclusions that can inform a theory of the actual reasons for the effectiveness of a treatment or policy. If researchers have a sense of the why, they can perhaps design programs that are even more effective, and more narrowly targeted.

Manzi concedes all of this. For him, however, the RFT is a second streetlamp to the econometric method. Together, the two barely illuminate a city block, but the RFT reveals some parts of the pavement that econometrics leave in shadow.

By arguing for a more commonsense approach to social science, Manzi denies the privileged place that experts with arcane econometric models occupy in contemporary politics. His appeal is rightly understood not merely as an endorsement of randomized field trials but rather as an argument for the empiricist disposition characterized by incrementalism, libertarianism, caution, and epistemic modesty.

This disposition is intuitively American. Alexis de Tocqueville observed nearly two centuries ago that although “there is no country in the civilized world where they are less occupied with philosophy than the United States,” Americans do have certain inclinations: the use of “tradition only as information, and current facts only as a useful study for doing otherwise and better,” and striving “for a result without letting themselves be chained to the means.” Tocqueville identifies this cast of mind with the skeptical and rationalistic French philosopher René Descartes, calling Americans natural Cartesians.

The greatest work of American political philosophy, The Federalist, joins practical wisdom gathered from the study of history — empirical evidence, of a sort — with philosophic reflections on human nature. Neither of these is especially helpful without the other, and Alexander Hamilton assails “the reveries of those political doctors whose sagacity disdains the admonitions of experimental instruction.” Hamilton even quotes the great Scottish empiricist David Hume in the final paragraph of the final paper of the Federalist. Hume writes that the craft of creating a constitution is such a difficult work “that no human genius, however comprehensive, is able by the mere dint of reason and reflection, to effect it. The judgments of many must unite in the work: experience must guide their labour: time must bring it to perfection: and the feeling of inconveniences must correct the mistakes which they inevitably fall into, in their first trials and experiments.”

All these refractions of American thinking seem to argue that, given the limitations of human reason, trial and error is very often the best means toward the practical improvement of any human undertaking. The experimental method lends itself to incremental progress, with experiment after experiment adding to the useful knowledge of mankind. The econometric method gathers the data of the past to create a predictive model of human behavior. But this predictive model too often takes the place of the accumulated experience of many generations and becomes a single, supremely-but-unjustifiably confident reference for policymaking. It uses history, but it uses history in order to wipe history away.

In our increasing deference to econometric studies and technocratic experts, we Americans have come a long way from being the “natural Cartesians” who so impressed Tocqueville. Experimental knowledge — the residue of trial and error — can also be haggled over and made esoteric, but, being more empirical and concrete, it is less liable to become mired in abstractions or to lead us down the dangerous alleyway of a false positive.

If Manzi gets his way and such experimental social science becomes more common, it will become the norm in political debates to say “show me.” Rather than be stuck in a stalemate of whether or not the “studies say” a proposed policy will work, experiments will be conducted, giving the public a sense of the costs, benefits, and tradeoffs associated with that policy. Then we can argue about whether the policy is worth implementing on a wider scale instead of engaging in rhetorical theatrics, with politicians bludgeoning one another with studies. We should return to making moral arguments — the sort of arguments that self-governing citizens ought to have basic competence to judge.

Jim Manzi says he offers a modest solution for some of the problems of poor policymaking. Whether or not his proposal for changing the way social science is done in America improves policy outcomes, by advancing the experimental method for social science, he may just help to revive the intellectually independent disposition of the American citizen.


Jeremy Rozansky is assistant editor of National Affairs.

 

Monday, June 17, 2013

Bill Keller From NYT: Living With the Surveillance State

Living With the Surveillance State

By BILL KELLER
Published: June 16, 2013 105 Comments

http://www.nytimes.com/2013/06/17/opinion/keller-living-with-the-surveillance-state.html?ref=opinion

MY colleague Thomas Friedman’s levelheaded take on the National Security Agency eavesdropping uproar needs no boost from me. His column soared to the top of the “most e-mailed” list and gathered a huge and mostly thoughtful galaxy of reader comments. Judging from the latest opinion polling, it also reflected the prevailing mood of the electorate. It reflected mine. But this is a discussion worth prolonging, with vigilant attention to real dangers answering overblown rhetoric about theoretical ones.

Enlarge This Image

Tony Cenicola/The New York Times

Bill Keller

Go to Columnist Page »
Bill Keller's Blog »

Connect With Us on Twitter

For Op-Ed, follow @nytopinion and to hear from the editorial page editor, Andrew Rosenthal, follow @andyrNYT.

Enlarge This Image

R.O. Blechman

Readers’ Comments

Share your thoughts.

Tom’s important point was that the gravest threat to our civil liberties is not the N.S.A. but another 9/11-scale catastrophe that could leave a panicky public willing to ratchet up the security state, even beyond the war-on-terror excesses that followed the last big attack. Reluctantly, he concludes that a well-regulated program to use technology in defense of liberty — even if it gives us the creeps — is a price we pay to avoid a much higher price, the shutdown of the world’s most open society. Hold onto that qualifier: “well regulated.”

The N.S.A. data-mining is part of something much larger. On many fronts, we are adjusting to life in a surveillance state, relinquishing bits of privacy in exchange for the promise of other rewards. We have a vague feeling of uneasiness about these transactions, but it rarely translates into serious thinking about where we set the limits.

Exhibit A: In last Thursday’s Times Joseph Goldstein reported that local law enforcement agencies, “largely under the radar,” are amassing their own DNA databanks, and they often do not play by the rules laid down for the databases compiled by the F.B.I. and state crime labs. As a society, we have accepted DNA evidence as a reliable tool both for bringing the guilty to justice and for exonerating the wrongly accused. But do we want police agencies to have complete license — say, to sample our DNA surreptitiously, or to collect DNA from people not accused of any wrongdoing, or to share our most private biological information? Barry Scheck, co-director of the Innocence Project and a member of the New York State Commission on Forensic Science, says regulators have been slow to respond to what he calls rogue databanks. And a recent Supreme Court ruling that defined DNA-gathering as a legitimate police practice comparable to fingerprinting is likely to encourage more freelancing. Scheck says his fear is that misuse will arouse public fears of government overreach and discredit one of the most valuable tools in our justice system. “If you ask the American people, do you support using DNA to catch criminals and exonerate the innocent, everybody says yes,” Scheck told me. “If you ask, do you trust the government to have your DNA, everybody says no.”

Exhibit B: Nothing quite says Big Brother like closed-circuit TV. In Orwell’s Britain, which is probably the democratic world’s leading practitioner of CCTV monitoring, the omnipresent pole-mounted cameras are being supplemented in some jurisdictions by wearable, night-vision cop-cams that police use to record every drunken driver, domestic violence call and restive crowd they encounter. New York last year joined with Microsoft to introduce the eerily named Domain Awareness System, which connects 3,000 CCTV cameras (and license-plate scanners and radiation detectors) around the city and allows police to cross-reference databases of stolen cars, wanted criminals and suspected terrorists. Fans of TV thrillers like “Homeland,” “24” and the British series “MI-5” (guilty, guilty and guilty) have come to think of the omnipresent camera as a crime-fighting godsend. But who watches the watchers? Announcing the New York system, the city assured us that no one would be monitored because of race, religion, citizenship status, political affiliation, etc., to which one skeptic replied, “But we’ve heard that one before.”

Exhibit C: Congress has told the F.A.A. to set rules for the use of spy drones in American air space by 2015. It is easy to imagine the value of this next frontier in surveillance: monitoring forest fires, chasing armed fugitives, search-and-rescue operations. Predator drones already patrol our Southern border for illegal immigrants and drug smugglers. Indeed, border surveillance may be critical in persuading Congress to pass immigration reform that would extend our precious liberty to millions living in the shadows. I for one would count that a fair trade. But where does it stop? Scientific American editorialized in March: “Privacy advocates rightly worry that drones, equipped with high-resolution video cameras, infrared detectors and even facial-recognition software, will let snoops into realms that have long been considered private.” Like your backyard. Or, with the sort of thermal imaging used to catch the Boston bombing fugitive hiding under a boat tarp, your bedroom.

And then there is the Internet. We seem pretty much at peace, verging on complacent, about the exploitation of our data for commercial, medical and scientific purposes — as trivial as the advertising algorithm that pitches us camping gear because we searched the Web for wilderness travel, as valuable as the digital record-sharing that makes sure all our doctors know what meds we’re on.

In an online debate about the N.S.A. eavesdropping story the other day, Eric Posner, a professor at the University of Chicago Law School, pointed out that we have grown comfortable with the Internal Revenue Service knowing our finances, employees of government hospitals knowing our medical histories, and public-school teachers knowing the abilities and personalities of our children.

“The information vacuumed up by the N.S.A. was already available to faceless bureaucrats in phone and Internet companies — not government employees but strangers just the same,” Posner added. “Many people write as though we make some great sacrifice by disclosing private information to others, but it is in fact simply the way that we obtain services we want — whether the market services of doctors, insurance companies, Internet service providers, employers, therapists and the rest or the nonmarket services of the government like welfare and security.”

Privacy advocates will retort that we surrender this information wittingly, but in reality most of us just let it slip away. We don’t pay much attention to privacy settings or the “terms of service” fine print. Our two most common passwords are “password” and “123456.”

From time to time we get worrisome evidence of data malfeasance, such as the last big revelation of N.S.A. eavesdropping, in 2005, which disclosed that the agency was tapping Americans without the legal nicety of a warrant, or the more recent I.R.S. targeting of right-wing political groups. But in most cases the advantages of intrusive technology are tangible and the abuses are largely potential. Edward Snowden’s leaks about N.S.A. data-mining have, so far, not included evidence of any specific abuse.

The danger, it seems to me, is not surveillance per se. We have already decided, most of us, that life on the grid entails a certain amount of intrusion. Nor is the danger secrecy, which, as Posner notes, “is ubiquitous in a range of uncontroversial settings,” a promise the government makes to protect “taxpayers, inventors, whistle-blowers, informers, hospital patients, foreign diplomats, entrepreneurs, contractors, data suppliers and many others.”

The danger is the absence of rigorous, independent regulation and vigilant oversight to keep potential abuses of power from becoming a real menace to our freedom. The founders created a system of checks and balances, but the safeguards have not kept up with technology. Instead, we have an executive branch in a leak-hunting frenzy, a Congress that treats oversight as a form of partisan combat, a political climate that has made “regulation” an expletive and a public that feels a generalized, impotent uneasiness. I don’t think we’re on a slippery slope to a police state, but I think if we are too complacent about our civil liberties we could wake up one day and find them gone — not in a flash of nuclear terror but in a gradual, incremental surrender.

 

Thursday, June 13, 2013

Your New Secretary: An Algorithm

Your New Secretary: An Algorithm

Startups Like RelateIQ Are Aiming to Help Improve Employees' Work Life With Software

By EVELYN M. RUSLI

The next frontier for data is improving your work relationships.

Can an algorithm improve your work life? Evelyn Rusli explains why the next frontier for data is improving your work relationships.

Jon Porter, the CEO of private wealth-management firm Three Bell Capital, used to keep track of clients by manually typing information about meetings and leads with software from Salesforce.com Inc. CRM +0.61%

Enlarge Image

RelateIQ employee Chip Camden prepares a demo presentation at the company's office in Palo Alto, Calif.

This spring, he got an algorithm to do the work.

Software from startup RelateIQ Inc. now looks at every digital scrap of Mr. Porter's work life—incoming emails, social-network contacts and phone calls—compares it with his colleagues' data, and figures out what and who is important. Two weeks ago, the algorithm prodded Mr. Porter to follow up on a time-sensitive question from a client.

"Had we not been on top of that, our client would have missed the window" for an investment, said Mr. Porter, who estimates he saves about two hours a week with the software.

Data scientists are beginning to peer into work relationships, trying to identify patterns that can improve how employees collaborate with peers, manage sales relationships, or see how they stack up against colleagues. It is a nascent market, but up-and-coming startups have their eyes set on upending established business-technology companies like Salesforce, which are also increasingly digging into data.

"We wanted to build an algorithm that could do what a highly trained relationship manager, with 20 years of experience, could do," said Adam Evans, a co-founder of RelateIQ, which has raised $29 million from investors including Accel Partners, Allen & Co., Battery Ventures, and Facebook Inc. FB -0.21% co-founder Dustin Moskovitz. The investment values the startup at $100 million.

How RelateIQ Works

  • Users sign up, connect their email and relevant social-media accounts.
  • Contacts and data are automatically pulled in from these sources, creating an address book.
  • Users create lists, or project pages, where they can define objectives and add companies or contacts to that page.
  • Users download mobile app to log calls and missed calls.
  • As users interact with contacts, communications are logged automatically.
  • Users decide whether to share their entire list, some of it or none of it with colleagues. In the context of a group, lists track items such as sales partners or recruits.
  • RelateIQ's algorithm constantly collects data signals to identify whether relationships are cooling and if list members should be prodded to take action, such as respond to an email.

Source: RelatedIQ

Elsewhere, Boston-based Sociometric Solutions Inc. uses physical sensors to collect data on employees' movements and the tone of their conversations to tell managers where interactions are dipping and where employees are congregating. In San Francisco, tenXer Inc., a program for computer engineers, tracks code modifications and hours spent in meetings to help them see how their productivity stacks up against colleagues. And Boston-based Yesware Inc. helps employees track emails, monitors how many times their emails are opened, what devices recipients are using, and provides analytic reports on the email traffic of colleagues.

Though harvesting of such employee data may raise eyebrows of some privacy advocates—particularly in light of the recent debate over technology companies' involvement in National Security Agency data-gathering programs—the startups are emerging as a hotbed for venture-capital dollars.

"It's a trend toward a more transparent workplace," said Michael Abbott, a general partner at Kleiner Perkins Caufield & Byers, which is currently on the hunt for investments on the theme.

The idea is that software can detect patterns that humans can't.

Angus Davis, the CEO of online payments service Swipely Inc., used Yesware during his last fundraising round to determine which venture capitalists were reading his emails, how many links they were clicking and if they forwarded it to others in the office. "When I saw an email opened 30 times, I thought, 'Wow' they are interested," Mr. Davis said.

RelateIQ, which has been operating in "stealth mode" underneath a home décor store in Palo Alto Calif., for two years, is one of the most ambitious of the big-data work apps. Its software offers a central hub for work groups to see what their co-workers are doing and keep track of relationships in real time.

RelateIQ absorbs massive amounts of data—it scans about 10,000 emails, calendar entries and other data points per minute at first run—but does offer privacy controls for employees, such as the option to hide the content of email messages from colleagues.

Of the company's roughly 30 employees, four make up the data science team, a group with Ph.D.'s in statistics, fluid dynamics and physics. The team constantly tweaks the software to better identify patterns, such as the average time it takes for a person to respond and what types of punctuation and phrases typically elicit responses.

It also tries to detect sarcasm and words that are usually associated with important questions. These data can reveal, for example, whether a relationship is stagnating or progressing.

The data startups will have to tackle entrenched business-software companies.

Salesforce, which has more than 100,000 customers using its customer relationship management software, is increasingly mining data to provide recommendations. For example, its communication tool Chatter recommends people and topics users should follow, based on patterns in their past activity.

"Everyone is chasing us," said Kendall Collins, an executive vice president of Salesforce. "I've never seen another player in the market, at scale, with the level of social integration and the level of data enrichment."

Oracle Inc., ORCL +2.18% also a major provider of such software, declined to comment. SAP AG, SAP +0.92% another competitor, says it is increasingly working with relationship management startups to provide some back-end tools to process huge data sets.

RelateIQ, which has about 100 clients, says its software requires companies to enter in less information to extract useful insight. "We automatically capture everything that is happening with your relationships and surface only what you should worry about," said Steve Loughlin, RelateIQ's CEO and co-founder.

Another big challenge is identifying the right communication data and organizing it in ways that make sense to humans who, unlike computers, aren't very predictable.

"We have very complex communications environments," said Tom Davenport, a professor in management and information technology at Babson College who studies technology in the workplace. Machine learning, he says, can result in recommendations that seem to come out of nowhere.

Mr. Loughlin of RelateIQ says he's aware of the limitations of an algorithm to read actual relationships. For instance, software could interpret a long delay from a client as a negative signal, even if he was just out with an illness.

Because of the squishiness of human relationships, the company says it is careful about language. "We don't say you 'need' to follow up, we say 'we suggest,'" Mr. Loughlin said.

Write to Evelyn M. Rusli at evelyn.rusli@wsj.com

 

Wednesday, June 12, 2013

John Gapper in FT: Big data has to show that it’s not like Big Brother

http://www.ft.com/cms/s/0/5af52e98-d2a5-11e2-88ed-00144feab7de.html#ixzz2W3YcmHGy

Big data has to show that it's not like Big Brother

We do not know yet what this new technology of data analysis and artificial intelligence means

©Ingram Pinn

Sales of George Orwell's Nineteen Eighty-Four have risen since Edward Snowden revealed how the National Security Agency of the US gains access to telephone records and data from technology companies. So far, if people do not exactly love Big Brother, they are prepared to accept some invasion of their privacy in return for security.

What about "big data"? Companies that hold rapidly expanding amounts of personal information are using new kinds of data analysis and artificial intelligence to shape products and services, and to predict what customers will want. Larry Page, Google's chief executive, describes his ideal form of technology as "a really smart assistant doing things for you so you don't have to think about it".

  •  
  •  
  •  
  •  

More

On this story

On this topic

John Gapper

The vision of living in a virtual Downton Abbey, with a computer to plan your day, suggest the best route to travel, the films you might want to watch and the best flight to catch – even to book it for you – has an allure. We are all pressed for time and want an easy life. Instead of being bombarded with information and forced to choose, it's nice to get personal service.

But just as the NSA disclosures have taken people by surprise, although it has existed for 60 years, I doubt whether many grasp either the size of the data trail they create daily, or the advances in technology that are permitting a select group of big data enterprises to exploit it. The technology is evolving so quickly that what was unthinkable two years ago is routine.

"It is both a wonderful and scary future. Companies with huge amounts of data will know more about you than yourself. They will be able to predict what you might do next," says Kai-Fu Lee, a Beijing-based investor and the former head of Google in China.

In a column last week I compared Google to General Electric in the late 19th century – an innovative industrial enterprise riding a wave of new technology. The flip side of that is that Google, Amazon, Microsoft and other technology giants are amassing powers that need to be controlled carefully.

The NSA and big data companies put their databases and computing power to different uses – one to identify spies and terrorists, and the others to match services to users. They have in common the use of very large databases and techniques such as pattern recognition and network analysis.

At the advanced end, this shades into artificial intelligence of the kind that, for example, intuits what you meant to search for even when you misspell the key words; can translate speech into another language in real time (as Microsoft demonstrated in China last year); or learns to recognise a photograph of a cat by viewing thousands of images.

The ability of computers to learn in a similar manner to humans is known as "deep learning" and it is notable that Google has hired several pioneers in the field, including the scientist and author Ray Kurzweil. Among the technology transfer offered by the NSA to private US companies are "cutting-edge machine learning technologies".

Such software can infer a lot from scraps of information, provided that it has enough of them, as shown by the NSA's effort to analyse phone call metadata from Verizon (and perhaps other operators). President Barack Obama assured Americans that "no one is listening to your phone calls", but this alone is a trove.

A study by Latanya Sweeney, a professor at Harvard University, found that 87 per cent of people can be identified simply by knowing their age, gender and postcode, if these are cross-checked against public databases. That is typical of the data collected by social networks and internet companies.

The extraordinary power of big data companies comes from being able to combine the personal data of customers with observations about them, from which products they buy to where (as measured by global positioning satellite data from mobile phones) they are. That produces a set of "inferred data" about what they probably want.

If I search on an Android phone for "Taj Mahal" while standing in India, for example, Google will prioritise results for the shrine in Uttar Pradesh. If I do the same in Brick Lane, east London, it will suggest local Bangladeshi restaurants. How long before it offers to book a restaurant based on how I rated others as I walk around a foreign city at dusk?

At one level, I would be pleased if it did (as long as it was a good one) since it would save me doing the work myself. At another, as a World Economic Forum report on personal data put it: "Inferred data can feel like an all-knowing Big Brother watching the security camera."

One of the concerns that springs from this is that big data companies with such software are very difficult to compete with. The more data that I and other users provide them with, the better they are at predicting what we want. The machine brain becomes cleverer with use.

Another is trust. Social networks have been poor at protecting users' data, and they hold only a fraction of the information on people's behaviour, habits and intentions on the new generation of services. It is no wonder that the NSA turns to them – it has computing power and they have swaths of material.

A third is ownership. We each have rights over our own information, but what happens when it gets mixed up with that of others and combined into a vast database of intentions? If I change my mind, how can it be unscrambled?

Above all, we don't know what this technology means because we are only at the beginning of the era of big data. There are plenty of aspects to admire but it will take some time to love.

john.gapper@ft.com