Monday, April 30, 2012

ACLU: Eight Problems With "Big Data"

Jay Stanley, ACLU, April 25, 2012
 
The idea of “Big Data” is in the air. At the South by Southwest Interactive conference last month, it was probably the hot topic, dominating or surfacing in numerous panels, including one on which I spoke, on “Big Data: Privacy Threat or Business Model?”

Let’s be clear on what we’re talking about. The term refers to something more specific than the general fact that companies and government agencies are collecting lots of personal information about people. What it refers to is the fact that once you store up huge amounts of information, you can mine those databases to discover subtle patterns, correlations, or relationships that our brains can’t perceive on their own because the scales involved are beyond our ability process (either the time scales at work, or the sheer number of data points). 

Such data mining has been called “the macroscope” – like a telescope or microscope, making things visible to us that have never been visible before.

In many ways Big Data is just a new buzzword for data mining, which we and others have been grappling with since not long after 9/11. The New York Times, for example, wrote about it using the term “data mining” in this 2007 piece.

A more recent (and much-discussed) article by Charles Duhigg in the New York Times offers a good example to keep in mind during discussions of the subject.  The piece described how Target identifies customers who are pregnant (sometimes before their own family members know) by tracking customers’ purchases and identifying patterns in their behavior. It then uses that insight to sell them baby-related goods.

And that may be only the beginning. What else might companies be identifying? Customers who are showing signs of Parkinson’s, or diabetes, or depression? I suppose that could actually be helpful if they notified you of their findings. But, there’s reason to think they won’t – Target found that it sort of freaks out their customers when they reveal what they know, so according to Duhigg the company has taken to hiding its ads to pregnant women among other “decoy” ads for things like lawnmowers so that the targets of the ads will think they’re just receiving the same flyers as everyone else.

Okay so that’s all pretty spooky.  But if we were to put our fingers on the precise problems with Big Data, what are they?

They include:
1.      It incentivizes more collection of data and longer retention of it. If any and all data sets might turn out to prove useful for discovering some obscure but valuable correlation, you might as well collect it and hold on to it. In long run, the more useful big data proves to be, the stronger this incentivizing effect will be – but in the short run it almost doesn’t matter; the current buzz over the idea is enough to do the trick.
2.      When you combine someone’s personal information with vast external data sets, you create new facts about that person (such as the fact that they’re pregnant, or are showing early signs of Parkinson’s disease, or are unconsciously drawn toward products that are colored red or purple). And when it comes to such facts, a person a) might not want the data owner to know b) might not want anyone to know c) might not even know themselves. The fact is, humans like to control what other people do and do not know about them – that’s the core of what privacy is, and data mining threatens to violate that principle.
3.      Many (perhaps most) people are not aware of how much information is being collected (for example, that stores are tracking their purchases over time), let alone how it is being used (scrutinized for insights into their lives).  The fact that Target goes to considerable trouble to hide its knowledge from its customers tells you all you need to know on that front.
4.      Big data can further tilt the playing field toward big institutions and away from individuals. In economic terms, it accentuates the information asymmetries of big companies over other economic actors and allows for people to be manipulated. If a store can gain insight into just how badly I want to buy something, just how much I can afford to pay for it, just how knowledgeable I am about the marketplace, or the best way to scare me into buying it, it can extract the maximum profit from me.
5.      It holds the potential to accentuate power differentials among individuals in society by amplifying existing advantages and disadvantages. Those who are savvy and well educated may get improved treatment from companies and government – while those who are poor, underprivileged, and perhaps already have some strikes against them in life (such as a criminal record) will be easily identified, and treated worse. In that way data mining may increase social stratification.
6.      Data mining can be used for so-called “risk analysis” in ways that treat people unfairly and often capriciously – for example, by insurance companies or banks to approve or deny applications. Credit card companies sometimes lower a customer’s credit limit based on the repayment history of the other customers of stores where a person shops. Such “behavioral scoring” is a form of economic guilt-by-association based on making statistical inferences about a person that go far beyond anything that person can control or be aware of.
7.      Its use by law enforcement raises even sharper issues – and when our national security agencies start using it to try to spot terrorists, those stakes can get even more serious. We know too little about how our security agencies are using Big Data, but such approaches have been discussed since the days of the Total Information Awareness program and before – and there is strong evidence that it’s being used by the NSA to sift through the vast volumes of communications that agency collects.  The threat here is that people will be tagged and suffer adverse consequences without due process, the ability to fight back, or even knowledge that they have been discriminated against. The threat of bad effects is magnified by the fact that data mining is so ineffective at spotting true terrorists.
8.      Over time such consequences will lead to chilling effects, as people become more reluctant to engage in any behaviors that will put them under the macroscope (more about that in a future post).

The Intersection of Information and Energy Technologies


Why I think computational power will solve the world's energy problems in this century.

Bill Gross, Technology Review, April 27, 2012

See the rest of our Business Impact report on Computers Storm the Grid.
 
Two talks at the TED conference this year formed, back to back, a sort of debate about the future of our planet. First, Paul Gilding gave a talk entitled "The Earth Is Full," about how we are using up all Earth's resources, with possibly devastating consequences. Next, X Prize creator Peter Diamandis gave a presentation entitled "Abundance," about how we will invent innovative ways to solve the challenges that loom before us. 

I believe that we will need great ingenuity to enable our planet to provide successfully for more than seven billion human beings, let alone the nine billion that will probably inhabit it by 2050, and I believe that information technology will make this ingenuity possible. Because of fluid marketplaces and an ever more globalized economy, nearly every important resource is becoming scarcer and more costly. Evidence of this is seen in the price not only of oil but also of aluminum, concrete, wood, water, rare-earth elements, and even common elements like copper. Everything is getting more expensive because billions of people are trying creatively to repackage and consume these materials. But there is one resource whose price has consistently has gone down: computation. 

The power, cost, and energy use involved in one unit of computation is declining at a more consistent, dependable rate than we have seen with any other commodity in human history. That declining cost curve must be tapped to lower energy prices—and I believe it will be. This will happen as people ask: To achieve my purpose (in designing whatever device or system), can I use more "atoms" or more "bits" (computation power)? The choice will have to be bits, because atoms are going up in price while bits are going down. 

Here are a few examples. When designing a car, one can put a bit more effort into stronger, lighter-weight materials, which will increase energy efficiency but possibly drive up cost; or one can put a lot more effort into using computational power to run simulations that optimize the use of materials. Today, computational fluid dynamics allow a designer to accurately design a new shape of car, put it in a computer wind tunnel instead of a physical one, and test 1,000,000 body designs to improve fuel mileage by significant amounts. This was never before possible for those constructing vehicles. 

In solar energy, large fields of mirrors or photovoltaic panels can be optimized to be lighter, more reliable, and more power-efficient by putting a $2 microprocessor in every panel. An onboard computer that lets each panel track the sun independently replaces previous systems that used more steel, bigger gears, and bigger gearboxes—basically, more materials. 

As little as 10 years ago, the computing power and sensors needed to build a closed-loop, sun-tracking solar panel might have cost $2,000, or more than the panel itself, and thus the system would not have been cost effective. But with computing costs coming down by a factor of 1,000 every 15 years, all kinds of new opportunities arise to improve system design. 

At eSolar, one of our companies, we designed and built a utility-scale solar-thermal power plant with a huge amount of computation embedded into the field of mirrors. We reduced the size of the components, cut the installation expense, and drove the cost of the system down to nearly half what had been achieved before. This experience proved to me the feasibility of replacing atoms with bits. 

The price reduction curve for computing is not over—it's continuing, and each year will open up further avenues for ingenuity. That is important because our current energy resources are not at all easy to compete with. Fuels that we dig out of the ground and burn are extremely cheap. They are, in effect, the concentrated storage of millions of years of sunlight falling on Earth. Ironically, the biggest component of energy costs is the expense of moving the fuel to consumers from where it's obtained—and transportation costs are mostly fuel, too. So we are in a kind of vicious cycle. The way to break free of fossil fuels is to introduce something new to our energy equation that isn't fuel. 

I believe ingenuity in the form of information technology is the only variable that offers sufficient leverage. We need to replace a cheap, unsustainable form of energy with sustainable forms of energy that are equally cheap. The only way to compete with cheap fuels is to be more clever with computation; that is, to use as little of anything else as possible. 

Bill Gross is a lifelong entrepreneur and CEO of Idealab, an incubator for ideas and prototypes that has spun off more than 75 operating companies.

Paul Krugman: Wasting Our Minds


Paul Krugman, The New York Times, April 29, 2012

In Spain, the unemployment rate among workers under 25 is more than 50 percent. In Ireland almost a third of the young are unemployed. Here in America, youth unemployment is “only” 16.5 percent, which is still terrible — but things could be worse. 

And sure enough, many politicians are doing all they can to guarantee that things will, in fact, get worse. We’ve been hearing a lot about the war on women, which is real enough. But there’s also a war on the young, which is just as real even if it’s better disguised. And it’s doing immense harm, not just to the young, but to the nation’s future. 

Let’s start with some advice Mitt Romney gave to college students during an appearance last week. After denouncing President Obama’s “divisiveness,” the candidate told his audience, “Take a shot, go for it, take a risk, get the education, borrow money if you have to from your parents, start a business.” 

The first thing you notice here is, of course, the Romney touch — the distinctive lack of empathy for those who weren’t born into affluent families, who can’t rely on the Bank of Mom and Dad to finance their ambitions. But the rest of the remark is just as bad in its own way.

I mean, “get the education”? And pay for it how? Tuition at public colleges and universities has soared, in part thanks to sharp reductions in state aid. Mr. Romney isn’t proposing anything that would fix that; he is, however, a strong supporter of the Ryan budget plan, which would drastically cut federal student aid, causing roughly a million students to lose their Pell grants. 

So how, exactly, are young people from cash-strapped families supposed to “get the education”? Back in March Mr. Romney had the answer: Find the college “that has a little lower price where you can get a good education.” Good luck with that. But I guess it’s divisive to point out that Mr. Romney’s prescriptions are useless for Americans who weren’t born with his advantages. 

There is, however, a larger issue: even if students do manage, somehow, to “get the education,” which they do all too often by incurring a lot of debt, they’ll be graduating into an economy that doesn’t seem to want them. 

You’ve probably heard lots about how workers with college degrees are faring better in this slump than those with only a high school education, which is true. But the story is far less encouraging if you focus not on middle-aged Americans with degrees but on recent graduates.  
Unemployment among recent graduates has soared; so has part-time work, presumably reflecting the inability of graduates to find full-time jobs. Perhaps most telling, earnings have plunged even among those graduates working full time — a sign that many have been forced to take jobs that make no use of their education. 

College graduates, then, are taking it on the chin thanks to the weak economy. And research tells us that the price isn’t temporary: students who graduate into a bad economy never recover the lost ground. Instead, their earnings are depressed for life. 

What the young need most of all, then, is a better job market. People like Mr. Romney claim that they have the recipe for job creation: slash taxes on corporations and the rich, slash spending on public services and the poor. But we now have plenty of evidence on how these policies actually work in a depressed economy — and they clearly destroy jobs rather than create them. 

For as you look at the economic devastation in Europe, you should bear in mind that some of the countries experiencing the worst devastation have been doing everything American conservatives say we should do here. Not long ago, conservatives gushed over Ireland’s economic policies, especially its low corporate tax rate; the Heritage Foundation used to give it higher marks for “economic freedom” than any other Western nation. When things went bad, Ireland once again received lavish praise, this time for its harsh spending cuts, which were supposed to inspire confidence and lead to quick recovery. 

And now, as I said, almost a third of Ireland’s young can’t find jobs. 

What should we do to help America’s young? Basically, the opposite of what Mr. Romney and his friends want. We should be expanding student aid, not slashing it. And we should reverse the de facto austerity policies that are holding back the U.S. economy — the unprecedented cutbacks at the state and local level, which have been hitting education especially hard. 

Yes, such a policy reversal would cost money. But refusing to spend that money is foolish and shortsighted even in purely fiscal terms. Remember, the young aren’t just America’s future; they’re the future of the tax base, too. 

A mind is a terrible thing to waste; wasting the minds of a whole generation is even more terrible. Let’s stop doing it.

Animated Introduction to Big Data [video]


Big Data is everywhere. Everyone knows it, everyone is talking about it. But what exactly is Big Data?

"How big, really, is Big Data? This is actually a very intriguing question whose answer seems to lack consensus at the moment but whose ambiguity has not stopped the use of the term. A common misconception, however, is that big data refers solely to the size of the data: if it is data and it is big then it must be big data. While size is certainly an element of the equation, there are other aspects or properties of big data not necessarily associated with size."

In this video, we learn interesting facts about Big Data, how it is being generated and at which pace. It also discusses misconceptions about it and provides a clear definition of Big Data. Below are some conclusions from the video:

·       Big Data is any attribute that challenges contraints of a system capability or business need.
·       Big Data refers to the sheer volume of data to be analyzed within a given time frame or within a geographical boundary.
·       Not all Big Data are the same from a structure perspective: some Big Data have its format well defined like transactions in a database, some Big Data may be only a collection of Blog entries that contain text, tables, images and video all kept in the same data storage.
In summary:

Despite of the size, speed or source of the data, Big Data drives the need to make sense out of the chaos, Big Data drives the need to find meaning on the data that is constantly changing and to find relationships between the data created. Understanding this interconnectedness and being able to harvest the information hidden in Big Data unlocks Big Data's value which can only be gained by being able to tackle our own Big Data challenges. 

Collecting, analyzing and understanding Big Data is becoming a differentiated strategy today but it will become a fact of life tomorrow. Running analyses at the finest granularities while you still have enough data for the results to be meaningful and accurate leads to more precise action and in turn more profits and savings for the company and customers. So when it comes to big data the question is not "Why should I care about Big Data?" but "How can I get closer to Big Data and how can I start taking advantage of it now?

Thursday, April 26, 2012

Big Data's Big Problem: Little Talent


Ben Rooney,The Wall Street Journal, April 26, 2012

It seems that the markets are as much in love with "Big Data"—the ability to acquire, process and sort vast quantities of data in real time—as the technology industry.

The first Big Data initial public offering hit the market last week to roaring approval. Splunk Inc., which helps businesses organize and make sense of all the information they gather, soared 109% on its first day of trading. Big Data, big price.

And this week, in cities in the U.S. and the U.K., Big Data Week events are being held to proselytize the unbelievers. 

Big Data refers to the idea that an enterprise can mine all the data it collects right across its operations to unlock golden nuggets of business intelligence. And whereas companies in the past have had to rely on sampling, Big Data, or so the promise goes, means you can use your entire corpus of digitized corporate knowledge. It is, by all accounts, the next big thing.

However, according to a report published last year by McKinsey, there is a problem. "A significant constraint on realizing value from Big Data will be a shortage of talent, particularly of people with deep expertise in statistics and machine learning, and the managers and analysts who know how to operate companies by using insights from Big Data," the report said. "We project a need for 1.5 million additional managers and analysts in the United States who can ask the right questions and consume the results of the analysis of Big Data effectively." What the industry needs is a new type of person: the data scientist. 

According to Pat Gelsinger, president and chief operating officer of EMC Corp., the giant U.S. data company, this isn't an unprecedented problem. "IBM started a generation of Cobol programmers," he said, referring to one of the first dominant programming languages. 

"Thirty years ago we didn't have computer-science departments; now every quality school on the planet has a CS department. Now nobody has a data-science department; in 30 years every school on the planet will have one."

Hilary Mason, chief scientist for the URL shortening service bit.ly, says a data scientist must have three key skills. "They can take a data set and model it mathematically and understand the math required to build those models; they can actually do that, which means they have the engineering skills…and finally they are someone who can find insights and tell stories from their data. That means asking the right questions, and that is usually the hardest piece."

It is this ability to turn data into information into action that presents the most challenges. It requires a deep understanding of the business to know the questions to ask. The problem that a lot of companies face is that they don't know what they don't know, as former U.S. Defense Secretary Donald Rumsfeld would say. The job of the data scientist isn't simply to uncover lost nuggets, but discover new ones and more importantly, turn them into actions. Providing ever-larger screeds of information doesn't help anyone.

One of the earliest tests for biggish data was applying it to the battlefield. The Pentagon ran a number of field exercises of its Force XXI—a device that allows commanders to track forces on the battlefield—around the turn of the century. The hope was that giving generals "exquisite situational awareness" (i.e. knowing everything about everyone on the battlefield) would turn the art of warfare into a science. What they found was that just giving bad generals more information didn't make them good generals; they were still bad generals, just better informed. 

At conference in London this week on the subject, the data scientist was called, only half-jokingly, "a caped superhero."

So where can companies find these superheros? Not from universities, it seems. Nigel Shadbolt, who doubles up as the professor of artificial intelligence at the University of Southampton as well as co-director (along with Tim Berners-Lee) of the U.K.'s Open Data Institute, said the courses don't yet exist. "Bits of it do exist in various departments around the country, and also in businesses, but as an integrated discipline it is only just starting to emerge."

Nor can they be found in recruitment agencies. Rob Grimsey, a director of IT recruitment agency Harvey Nash, said they had limited experience in recruiting data scientists—"which might be a statement in itself about how common these kind of roles are," he added.

One of the problems with Big Data is the fact that it has to deal with real data from the real world, which tends to be messy and difficult to represent. Conventional relational databases are excellent at handling stuff that comes in discreet packets, such as your social security number or a stock price. They are less useful when it comes to, say, the content of a phone call, a video, or an email. Out in the real world, most data is unstructured. Handling this sort of real, messy, scrappy data, isn't so simple.

"People have been doing data mining for years, but that was on the premise that the data was quite well behaved and lived in big relational databases," said Mr. Shadbolt. "How do you deal with data sets that might be very ragged, unreliable, with missing data?"

In the meantime, companies will have to be largely self-taught, said Nick Halstead, CEO of DataSift, one of the U.K. start-ups actually doing Big Data. When recruiting, he said that the ability to ask questions about the data is the key, not mathematical prowess. "You have to be confident at the math, but one of our top people used to be an architect".

But Fernando Lucini, chief architect for Autonomy Corp., a U.K. software maker recently acquired by Hewlett-Packard Co., is much more optimistic. Mr. Lucini said the industry is fretting unnecessarily and should have more confidence in its own abilities. Most of these problems can be tackled through algorithms, he said, which coincidentally is the promise of Autonomy. "The problem can be solved by better tools. The tools need to help you understand the data. They can do the heavy lifting for you so that anyone in a business can use them and ask the questions they need to answer."

The challenge of 'big data' for data protection

Kuner, Christopher,*  and Fred H. Cate,** Christopher Millard,** and Dan Jerker B. Svantesson.*** International Data Privacy Law , 2012 2 (2): 47-49.

Data protection, like almost everything else in our lives, is challenged by the advent of ‘big data’. The Economist reports in its 2012 Outlook that the quantity of global digital data expanded from 130 exabytes in 2005 to 1,227 in 2010, and is predicted to rise to 7,910 exabytes in 2015.1

An exabyte is a quintillion bytes. If you find that hard to visualize, consider this: someone has calculated that if you loaded an exabyte of data on to DVDs in slimline jewel cases, and then loaded them into Boeing 747 aircraft, it would take 13,513 planes to transport one exabyte of data. Using DVDs to move the data collected globally in 2010 would require a fleet of more than 16 million jumbo jets. 

And exabytes are rapidly becoming passé. The volume of stored information in the world is growing so fast that scientists have had to create new terms, including zettabyte and yottabyte, to describe the flood of data. 

The importance of big data is not just a result of its size or how fast it is growing (about 60 per cent a year), but also the reality that the data come from an amazing array of sources. The Internet captures lots of data. Facebook alone has more than 800 million active users, more than half of whom log in every day, where they generate more than 900 million web pages and upload more than 250 million photos every day. 

In 2010, a lifetime ago in Internet time, Google sites were used by more than 1 billion unique visitors every month who spent a collective 200 billion minutes on its sites. Google-owned YouTube passed 1 trillion video playbacks in 2011. Email, IM, VOIP calls, and other communications generate tens of trillions of recorded messages every year.
Credit and …
[Full Text of this Article]

Sunday, April 22, 2012

IBM Fellow Jeff Jonas on the evolution of Big Data

Recently named IBM Fellow Jeff Jonas is one of the most interesting big data thought-leaders. He spoke to CNET about the increasing value of data-driven decisions.


Dave Rosenberg, CNET News, April 21, 2012

Last week I reconnected with Jeff Jonas, chief scientist of the IBM Entity Analytics group and a recently named IBM Fellow, about what's going on in the realm of big data.

When I first met Jonas, back in
June of 2010, he was focused on how companies are dealing with the deluge of information associated with Big Data. His focus hasn't changed, but he told me his perspective on how we make sense of data continues to evolve -- especially as we move in and out of demand for real-time versus batch data processing.

New Big Data tools make it much more affordable to gather and organize large sets of data that can be analyzed in its raw form. As advanced analytics applications get applied against that data, it becomes dramatically easier to identify the direct cause-and-effect relationship between business events, regardless of what department is nominally in charge of that event or the process associated with it.


According to Jonas, the three V's -- volume, velocity, and variety -- are the essential characteristics of "Big Data" that will grow exponentially, rather than in a linear fashion. Accordingly, you have to plan for data growth in conjunction with any projects you plan to undertake.

But planning is just one aspect of the Big Data situation. More important is knowing what you want to get out of the data analysis you're performing. Trying to make more sense of data is growing in importance for businesses of all kinds, but the techniques employed are relative to the problem you're trying to solve.

Jonas found that by default most organizations go with a batch approach, using tools like Hadoop and other MapReduce implementations. This approach works well for "thinking apps," where you are looking for information and context to inform a bigger notion or data amalgamation as opposed to a real-time decision. But there are many things that can and should be done in real-time primarily because batch/MapReduce processes are too late with decisions and suggestions.

Based on his experience at IBM, Jonas suggested that as the value of the analysis rises, business users will demand that data be delivered sooner -- even if it's not totally practical. "I need to know now, not later, what this data means to my business."

Ultimately, this all feeds into the work Jonas has been doing around puzzles and problem solving, changing the way to think about the context of the data. There may be more than one aspect to the problem. Moneygram, for example, is using IBM Identity Insight for context accumulation -- weaving data together to address fraud complaints, which have dropped 72 percent since they started using it, proving the value of Big Data analysis in real business terms.

As a side note, Jonas is a big triathlete. He told me that if completes all the rest of the races he's planned this year and with two more next year he's been given the impression that he will be the third human being to have done every international Ironman event. Cheer for him as he swims/bikes/runs by.

Big Data age puts privacy in question as information becomes currency

Exploiting Big Data's opportunities will need a delicate balance between the right to knowledge and the right of the individual

Aleks Krotoski,  guardian.co.uk,  April 22, 2012 


When Social Calendar users give personal details about themselves or their friends, the data ends up in Walmart's hands. Photograph: Marc F Henning/Alamy

This month, the US chain Walmart bought the startup Social Calendar, one of the most popular calendar apps on Facebook, which lets users record special events, birthdays and anniversaries. More than 15 million registered users have posted over 110m personal notifications, and users receive email reminders totalling over 10m a month.

Of course, when a Social Calendar user listed a friend's birthday or details of a holiday to Malaga, she or he probably had no idea the information would end up in the hands of a US supermarket. But now it will be cross-referenced with Walmart's own data, plus any other databases that are available, to generate a compelling profile of individual Social Calendar users and their non-Social Calendar-using friends.

The second decade of the 21st century is epitomised by Big Data. From the status updates, friendship connections and preferences generated by Facebook and Twitter to search strings on Google, locations on mobile phones and purchasing history on store cards, this is data that's too big to compute easily, yet is so rich that it is being used by institutions in the public and private sectors to identify what people want before they are even aware they want it.

The most important thing for data holders in the Big Data age is the kind of information they have access to. Facebook's projected $100bn value is based on the data it offers people who want to exploit its social graph. Its holdings include more than 800m records about who's in a user's social circle, relationship information, likes, dislikes, public and private messages and even physiological characteristics.

Google's recent privacy policy change has integrated the various accounts an individual maintains, creating a single profile that includes intentions from its search engine and the connections identified from its social network Google+; preferences and interests from mail, documents or YouTube; and location from its maps and mobile phone operating system.

Aggregated, this data can prove powerful. "Given enough data, intelligence and power, corporations and government can connect dots in ways that only previously existed in science fiction," said Alexander Howard, government 2.0 correspondent at the technology publisher O'Reilly Media.

In a trend that is remarkably similar to the plotline of Philip K Dick's Minority Report, Big Data is being used to predict social unrest or criminal intent. For example, Pax, an experimental system developed by the documentary maker and historian Brian Lapping, predicts the conditions for uprisings using aggregated search terms in different regions of the world. The analysed intelligence is then sold to governments, which can act accordingly.

The systems used to parse, synthesise, assimilate and make sense of the information are starting to make sophisticated connections and learn patterns. Big Data proponents view this as an opportunity to observe behaviours in real time, draw real-time conclusions and affect real-time change. Yet their conclusions can trip into areas that require human sensibilities to truly understand their implications.

In one recent high-profile example, a Minneapolis man discovered his teenage daughter was pregnant because coupons for baby food and clothing were arriving at his address from the US superstore Target. The girl, who had not registered her pregnancy with the chain, had been identified by a system that looked for pregnancy patterns in her purchase behaviour. "Data can say quite a lot," said Howard. "Though one has to be very careful to verify quality and balance it with human expertise and intuition."

In an infamous case in 2006, anonymised search terms released into the public domain by AOL were quickly de-anonymised, identifying individual searchers. And last month, police in New York used a photo from Facebook in combination with their own photo files and facial recognition software to arrest a man for attempted murder.

"People give out their data often without thinking about it," said the European commission vice-president Viviane Reding. "They have no idea that it will be sold to third parties." So users continue to populate databases such as Social Calendar with increasingly valuable personal information that, as commercial property, can be transferred to a new company with a different privacy ethos.

Privacy is not about control over personal data, according to the web theorist Danah Boyd, but the control individuals think they have. "People seek privacy so that they can make themselves vulnerable in order to gain something: personal support, knowledge, friendship," she said at the WWW conference in 2010. Increasingly, people are gaining services that deliver value, relevance and connection – as Google and Facebook do – in exchange for their personal information.

Expectations of privacy are being renegotiated. "When I grew up in Greensborough, Alabama, the population was 1,200," said Jim Adler, chief privacy officer and general manager of data systems at the information commerce firm Intelius. "If you cut school, everyone knew it by dinner. The expectation of privacy was low.

"Now, the expectation of privacy that we've had before Big Data, and our parents had, has been pulled away."

To some degree, this is happening because web users and web developers may not share a universal sense of what is and what is not private. As Boyd put it, privacy is contextual. An individual may be willing to share what they had for breakfast on Twitter, divulge where they are via FourSquare or record every keystroke made on their computers since 1998, but they wouldn't want information about their health or their children's whereabouts made public.

It becomes even more complicated when the users of software systems and architectures are a global population but the privacy expectations have been put in place by primarily US services. "Our expectations of privacy in the US versus Europe are very different," said Adler. "We are currently negotiating which is more important: the rights of the individual or the rights of knowledge."

In the EU, Reding has campaigned for the "right to be forgotten", already part of the 1995 data protection directive, which establishes by law that private data is the property of the individual and must be deleted from a system on request at any time. "More and more people feel uncomfortable about being traced everywhere, about a brave new world," she said. Information held by public bodies, however, remains exempt.

Reding's motivation is primarily to maintain a business ecosystem friendly for foreign investment. "This isn't about the reputation of the individual," she explained. "It's about the reputation of the companies. Data is their currency.

"What we're aiming for is privacy by design," she said. Companies should initiate a hallmark system that informs users that the privacy policy adheres to the guidelines. This, she argued, would ensure that people continue to share their data.

Sceptics like Adler argue that the right to be forgotten is flawed because it ignores how social boundaries are currently being negotiated in the Big Data world. "The ability to delete personal information means that you lose the potential for lessons learned," he said. "If you can step away and erase something someone says that is stupid or hurtful, you lose an element of accountability."

Yet if, as Reding maintains, 80% of British citizens are already concerned that data held by companies will be used for purposes other than the reason it was collected, there may be a shift in how much information people are willing to share.

The weakest link is the technology itself. The Target pregnancy case demonstrates that machines can pick up patterns in ways that may have unexpected consequences for individuals. The people who design the systems that collect and analyse the data are now responsible for thinking about data privacy and projecting future outcomes, and may – because they're human – get it wrong.

"These technologies are as neutral as guns," said Adler. "The Big Data guys who want to send you coupons when you're pregnant – because they're nerdy and technologists – probably don't realise that pregnancy is a sensitive issue." And sensitivities shift throughout an individual's lifespan and, more broadly, social norms shift over time.

Fundamentally, privacy means the same thing in an era of Big Data as it always has, but the capacity of machines to capture, store, process, synthesise and analyse details about everyone has forced new boundaries. It is unlikely that people will stop sharing data in exchange for services that are viewed as valuable.

Big Data offers undeniable opportunities, but requires a delicate balance between the right to knowledge and the right of the individual. Privacy norms will demand that new systems of trust be built into technology design.

© 2012 Guardian News and Media Limited or its affiliated companies. All rights reserved.

Thursday, April 19, 2012

Kauffman: Leverage big data to control healthcare costs

Opening Up Big Data is the Big Solution to Curing Health Care Ills, according to Kauffman Report

Media Contacts:
Rose Levy, 646-660-8641, rose@goldinsolutions.com, Goldin Solutions
Barbara Pruitt, 816-932-1288; bpruitt@kauffman.org, Kauffman Foundation

Kauffman Foundation task force offers incremental approaches to unlocking obstacles to efficient health care reform

WASHINGTON (April 19, 2012) – Cost trends in U.S. health care consistently increase at about 2.5 percentage points faster than the general rate of inflation – clearly an unsustainable rate. To address what it called "America's most urgent public policy problem," the Ewing Marion Kauffman Foundation released a report at The Atlantic's fourth annual Health Care Forum in Washington today that focuses on improving the cost-benefit balance in American health care through open access to medical data. The conference will be live streamed at http://events.theatlantic.com, and tweeters can follow the conversation on @Atlantic_Live, with the hashtag #HealthCare2012.

The report, "Valuing Health Care: Improving Productivity and Quality," is based on the recommendations of 31 experts from related fields, whom the Kauffman Foundation convened to reframe thinking around the question, "How can the productivity and value of American health care be increased, in both the short-term and long-term?"


While acknowledging that there's no shortage of reports and recommendations for health care reform, the task force took a unique approach to tackling health care value and productivity challenges.

"Rather than look for a 'one-shot-fix' solution, the task force focused on incremental reforms that cumulatively can both reduce costs and enhance the value of health care delivered to Americans, regardless of whether and how the Affordable Care Act is implemented," said Robert Litan, vice president of research and policy at the Kauffman Foundation and a task force co-organizer. "The underlying thread to the recommendations is leveraging big medical data."


"Using proper safeguards, we need to open the information that is locked in medical offices, hospitals and the files of pharmaceutical and insurance companies," said John Wilbanks, Kauffman senior fellow and an author of the report. "For example, combining larger datasets on drug response with genomic data on patients could steer therapies to the people they are most likely to help. This could substantially reduce the need for trial-and-error medicine, with all its discomforts, high costs and sometimes tragically wrong guesses."

Specifically, the report recommends:

· Unleashing the power of information by breaking down silos and encouraging data sharing between research centers, medical offices, pharmaceutical companies, insurance firms and others; and that a new corps of data entrepreneurs be incentivized to collect and analyze existing medical data to discover and then disseminate new therapies
.

· Funding more translational, cross-cutting research, with larger average grants made available to larger teams, many of them with participants from multiple institutions; and requiring collaboration across research institutions.

· Reforming medical malpractice systems to streamline new drug approvals and remove counter-productive restrictions on health insurance premiums.

· Empowering patients by, among other means, providing unbiased information on treatment options' benefits and drawbacks, and helping them make choices about the relevant lifestyle implications and risk-reward tradeoffs.


Further, the task force contends, health care delivery deserves its own national research program, one focused on comparative efficiency research. More efficiency (with acceptable quality guidelines) leads to profitability, and corrects the easy practice of simply passing costs down the health care stream.

Jeff Jonas - Big Data Q&A for the Data Protection Law and Policy Newsletter

Jeff Jonas, April 18, 2012

I have been given permission to re-publish an interview I did with the Data Protection & Law Policy newsletter. Also to appear in the e-Finance and Payments Law & Policy newsletter.

May be of interest to some of my readers.

[INTERVIEW]

1. Data protection challenge of the future: what is Big Data?

The three V’s - Volume, Velocity, and Variety – are the essential characteristics of “Big Data”. While data protection and privacy laws are still busy catching up with technologies of yesterday, Big Data is growing at a lightning speed on a daily basis. How can companies deal with the data protection challenges brought about by Big Data in order to truly benefit from the opportunities introduced by Big Data? First, one must truly grasp what is Big Data. We interview Jeff Jonas, Chief Scientist at IBM Entity Analytics, to obtain his perspectives and definition of Big Data, and his experience handling Big Data.

2. When did data become big?

Big Data did not become big overnight. What I think happened is data started getting generated faster than organizations could get their hands around it. Then one day you simple wake up and feel like you are drowning in data. On that day, data felt big.

3. Please explain and elaborate on the characteristics of Big Data?

Big Data means different things to different people.

Personally, my favorite definition is: “something magical happens when very large corpuses of data come together.” Some example of this can be seen at Google, for example Google flu trends and Google translate. In my own work, I witnessed this first in 2006. In this particular system, the system started getting higher quality predictions and faster as it ingested more data. This is so counter intuitive. The easiest way to explain this though is to consider the familiar process of putting a puzzle together at home. Why is it do you think the last few pieces are as easy as the first few – even though you have more data in front of you then ever before? Same thing really that is happening in my systems these days. It’s rather exciting to tell you the truth.

To elaborate briefly on the new physics of Big Data, I pinpointed the three phenomena of Big Data physics in my blog entry - Big Data. New Physics – drawing from my personal experience of 14 years of designing and deploying a number of multi-billion row context accumulating systems:

1. Better Prediction. Simultaneously lower false positives and lower false negatives

2. Bad data good. More specifically, natural variability in data including spelling errors, transposition errors, and even professionally fabricated lies – all helpful.

3. More data faster. Less compute effort as the database gets bigger.

Another definition of Big Data is related to the ability for organizations to harness data sets previously believed to be “too large to handle.” Historically, Big Data means too many rows, too much storage and too much cost for organizations who lack the tools and ability to really handle data of such quantity. Today, we are seeing ways to explore and iterate cheaply over Big Data.

4. When did data become big for you? What is your “Big Data” processing experience?

As previously mentioned, for me, Big Data is about the magical things that happen when a critical mass is reached. To be honest, Big Data does not feel big to me unless it is hard to process and make sense of. A few billion rows here and a few billion rows there – such volumes once seemed a lot of data to me. Then helping organizations think about dealing volumes of 100 million or more records a day seemed like a lot. Today, when I think about the volumes at Google and Facebook, I think: “Now that really is Big Data!”

My personal interest and primary focus on Big Data these days is: how to make sense of data in real time, that is fast enough to do something about the transaction while the transaction is still happening. While you swipe that credit card, there is only a few seconds to decide if that is you or maybe someone pretending to be you. If an unauthorized user is inside your network, and data starts getting pumped out, an organization needs sub-second “sense and respond” capabilities. End of day batch processes producing great answers is simply late!

5. What are the technologies currently adopted to process Big Data?

The availability of Big Data technologies seems to be growing by leaps and bounds and on many fronts. We are seeing a large corporate investments resulting in commercial products – at IBM two examples would be IBM InfoSphere Streams for Big Data in motion and IBM InfoSphere Big Insights for pattern discovery over data at rest. There are also many Big Data open source efforts under way for example HADOOP, Cassandra and Lucene. If one were to divide these into types one would find some well suited for streaming analytics and others for batch analytics. Some help organizations harness structured data while others are ideal for unstructured data. One thing is for sure – there are many options, and there will be many more choices to come as Big Data continues to get investment.

6. How can companies benefit from the use of Big Data?

I’d like to think consumers benefit too, just to be clear. To illustrate my point, I find it very helpful when Google responds to my search with “did you mean ______”. To pull this very smart stunt, Google must remember the typographical errors of the world, and that I do believe would qualify as Big Data. Moreover, I think health care is benefiting from Big Data, or let’s hope so. Organizations like financial institutions and insurance companies are benefitting from Big Data also by using these insights to run more efficient operations and mitigate risks.

We, you and I, are responsible in part for generating so much Big Data. These social media platforms we use to speak our mind and stay connected are responsible for massive volumes of data. Companies know this and are paying attention. For example, my friend’s wife complained on Twitter about a specific company’s service. Not long thereafter they reached out to her because they too were listening. They fixed the problem and she was as happy as ever. How did the company benefit? They kept a customer.

7. What is the trend of processing Big Data?

I think a lot of Big Data systems are running as periodic batch processes, for example, once a week or once a month. My suspicion is as these systems begin to generate more and more relevant insight, it will not be long before the users say: “Why did I have to wait until the end of the week to learn that? They already left the web site.”; or, “I already denied their loan when it is now clear I should have granted them that loan.”

8. What are the complications dealing with the privacy implications brought about by Big Data compare to average sized data?

There are lots of privacy complications that come along with Big Data. Consumers, for example, often want to know what data an organization collects and the purpose of the collection. Something that further complicates this: I think many consumers would be surprised to know what is computationally possible with Big Data. For example, where you are going to be next Thursday at 5:35pm or your three best friends, and which two of them are not on Facebook. Big Data is making it harder to have secrets. To illustrate using lines from my blog entry - Using Transparency As A Mask – ‘Unlike two decades ago, humans are now creating huge volumes of extraordinarily useful data as they self-annotate their relationships and yours, their photographs and yours, their thoughts and their thoughts about you … and more. With more data, comes better understanding and prediction. The convergence of data might reveal your “discreet” rendezvous or the fact you are no longer on speaking terms your best friend. No longer secret is your visit to the porn store and the subsequent change in your home’s late night energy profile, another telling story about who you are … again out of the bag, and little you can do about it. Pity … you thought that all of this information was secret.’

9. What are the privacy concerns & threats Big Data might bring about - to companies and to individuals whose data are contained in 'Big Data'?

My number one recommendation to organizations is “Avoid Consumer Surprise.”

That said, my concern is many consumers don’t seem to give a hoot. When is the last time you actually read the privacy statement or terms of use on your favourite social media site? I think in the future we’ll see Big Data being used to make the services offered even more irresistible. Your Internet searches will become custom crafted lenses. As a student of privacy and someone building Privacy by Design (PbD) into my inventions, I think about these things all the time.

10. How are companies currently applying privacy protection principles before/after Big Data has been processed?

I think there are many best practices being adopted. One of my favorites involves letting consumers opt-in instead of opting them in automatically and then requiring them to opt-out. One new thing I would like to see become a new best practice is: a place on the web site, for example my bank, where I can see a list of third parties whom my bank has shared my data with. I think this transparency would be good and certainly would make consumers more aware.

11. What is “Big Data”, according to Jeff Jonas?

Big Data is a pile of data so big - and harnessed so well - that it becomes possible to make substantially better predictions, for example, what web page would be the absolute best web page to place first on your results, just for you.

Big Data: Splunk's Data With Destiny


Rolfe Winkler, The Wall Street Journal, April 18, 2012

WSJ's Rolfe Winkler makes a stop on Mean Street to discuss the upcoming IPO of tech company Splunk. He and Evan Newmark ponder if a tech bubble is looming closer. Photo: Getty Images.

Splunking is the new googling. And it could make investors a tidy sum of money.The new verb tossed around by information-technology pros comes courtesy of Splunk, a startup specializing in data analysis that will open for trading Thursday after staging its initial public offering. The company is one of many capitalizing on the explosion of information, a trend being referred to as "Big Data."

 

Splunk's particular specialty is collecting and processing so-called "machine data." From web sites to cell phones to smart meters to GPS equipment, machines create little bits of information all the time. So much gets created, it is often lost after being recorded on a server somewhere.

When you click through an e-commerce web site, you visit lots of different product pages, put items in a shopping cart, and maybe disappear without buying anything. A company that could follow its customers to determine why they don't complete their order might be able to isearchable—is similar to what Splunk does with machine data. Its software is available free on a trial basis to start. Often someone inside a company will start using it, find it useful and then others will start using it themselves for their own projects. Splunk starts charging as more data gets plugged in. Yet the product is still cheaper than many older software alternatives currently on the market. The formula has worked well so far. Splunk reported $121 million of revenue in the fiscal year that ended in January, up 83% from the prior year. That growth has excited IPO investors.

Originally Splunk planned to price its shares in a range between $8 to $10, but has since bumped up the target to between $11 and $13. Investors lucky to get in on the shares around that price could see them pop significantly.


At first glance, the pricing seems aggressive. The valuation of the company net of cash would be around $1.3 billion at a $13 share price. Splunk is unprofitable.

Yet combine Splunk's growth rate with the appeal of its technology and the firm looks a mouth-watering takeover target for a larger software company like BMC Software, BMC 0.00% International Business Machines IBM -0.38% or Hewlett-Packard HPQ -0.02% .

A valuation of 10 times forward revenue would be in line with previous deals, for storage-software companies 3PAR and Isilon Systems. Assuming Splunk grows at, say, 70% this fiscal year, that would translate to a roughly $20 share price.


To be sure, the company has faced growing pains. Older versions of its software were buggy. And as it has expanded, Splunk has had trouble keeping up with customer demand for support. Yet it has mostly overcome these issues.

Look for the company to cash in nicely as a result.

Wednesday, April 18, 2012

"Coursera" open now

Education for Everyone.

We offer courses from the top universities, for free.

Learn from world-class professors, watch high quality lectures, achieve mastery via interactive exercises, and collaborate with a global community of students.

New Online University, headed by Larry Summers

TechCrunch, April,  2012

The Minerva Project is aiming to rethink the role of colleges and universities, taking into account the ways in which the Web has completely altered the distribution of and access to information.

The Minerva Project aims to offer a liberal arts education that is defined by an “extraordinarily rigorous” learning and admissions process. Not only does Minerva want to attract the same bright young minds that attend Harvard, Yale, and Stanford, it wants a global student body, both at home and abroad.

Just like traditional institutions, Minerva will be a four-year university, with two semesters, and four classes per semester. But, for the first year, Minerva students will live in their home countries, learning the core curriculum, so that by their sophomore year, in spite of language differences, all the students will have the same basics.

Then, from the start of their sophomore year through graduation, students will be encouraged to live in a new country (or at the very least, a new city) every semester. In this way, Minerva wants its education to be informed by experience and by the resources available online: “We’re not going to offer a single foreign language class, but if you’re not trilingual by the end of your four years, you won’t graduate,” Nelson says.

Sunday, April 15, 2012

Sergey Brin: Web freedom facing greatest threat ever

Series: Battle for the internet

Exclusive: Threats range from governments trying to control citizens to the rise of Facebook and Apple-style 'walled gardens'

Ian Katz  guardian.co.uk,  April 15, 2012 
 
The principles of openness and universal access that underpinned the creation of the internet three decades ago are under greater threat than ever, according to Google co-founder Sergey Brin.

In an interview with the Guardian, Brin warned there were "very powerful forces that have lined up against the open internet on all sides and around the world". "I am more worried than I have been in the past," he said. "It's scary."

The threat to the freedom of the internet comes, he claims, from a combination of governments increasingly trying to control access and communication by their citizens, the entertainment industry's attempts to crack down on piracy, and the rise of "restrictive" walled gardens such as Facebook and Apple, which tightly control what software can be released on their platforms.

The 38-year-old billionaire, whose family fled antisemitism in the Soviet Union, was widely regarded as having been the driving force behind Google's partial pullout from China in 2010 over concerns about censorship and cyber-attacks. He said five years ago he did not believe China or any country could effectively restrict the internet for long, but now says he has been proven wrong. "I thought there was no way to put the genie back in the bottle, but now it seems in certain areas the genie has been put back in the bottle," he said.

He said he was most concerned by the efforts of countries such as China, Saudi Arabia and Iran to censor and restrict use of the internet, but warned that the rise of Facebook and Apple, which have their own proprietary platforms and control access to their users, risked stifling innovation and balkanising the web.

"There's a lot to be lost," he said. "For example, all the information in apps – that data is not crawlable by web crawlers. You can't search it."

Brin's criticism of Facebook is likely to be controversial, with the social network approaching an estimated $100bn (£64bn) flotation. Google's upstart rival has seen explosive growth: it has signed up half of Americans with computer access and more than 800 million members worldwide.

Brin said he and co-founder Larry Page would not have been able to create Google if the internet was dominated by Facebook. "You have to play by their rules, which are really restrictive," he said. "The kind of environment that we developed Google in, the reason that we were able to develop a search engine, is the web was so open. Once you get too many rules, that will stifle innovation."

He criticised Facebook for not making it easy for users to switch their data to other services. "Facebook has been sucking down Gmail contacts for many years," he said.

Brin's comments come on the first day of a week-long Guardian investigation of the intensifying battle for control of the internet being fought across the globe between governments, companies, military strategists, activists and hackers.

From the attempts made by Hollywood to push through legislation allowing pirate websites to be shut down, to the British government's plans to monitor social media and web use, the ethos of openness championed by the pioneers of the internet and worldwide web is being challenged on a number of fronts.

In China, which now has more internet users than any other country, the government recently introduced new "real identity" rules in a bid to tame the boisterous microblogging scene. In Russia, there are powerful calls to rein in a blogosphere blamed for fomenting a wave of anti-Vladimir Putin protests. It has been reported that Iran is planning to introduce a sealed "national internet" from this summer.

Ricken Patel, co-founder of Avaaz, the 14 million-strong online activist network which has been providing communication equipment and training to Syrian activists, echoed Brin's warning: "We've seen a massive attack on the freedom of the web. Governments are realising the power of this medium to organise people and they are trying to clamp down across the world, not just in places like China and North Korea; we're seeing bills in the United States, in Italy, all across the world."

Writing in the Guardian on Monday, outspoken Chinese artist and activist Ai Weiwei says the Chinese government's attempts to control the internet will ultimately be doomed to failure. "In the long run," he says, "they must understand it's not possible for them to control the internet unless they shut it off – and they can't live with the consequences of that."

Amid mounting concern over the militarisation of the internet and claims – denied by Beijing – that China has mounted numerous cyber-attacks on US military and corporate targets, he said it would be hugely difficult for any government to defend its online "territory".

"If you compare the internet to the physical world, there really aren't any walls between countries," he said. "If Canada wanted to send tanks into the US there is nothing stopping them and it's the same on the internet. It's hopeless to try to control the internet."

He reserved his harshest words for the entertainment industry, which he said was "shooting itself in the foot, or maybe worse than in the foot" by lobbying for legislation to block sites offering pirate material.

He said the Sopa and Pipa bills championed by the film and music industries would have led to the US using the same technology and approach it criticised China and Iran for using. The entertainment industry failed to appreciate people would continue to download pirated content as long as it was easier to acquire and use than legitimately obtained material, he said.
"I haven't tried it for many years but when you go on a pirate website, you choose what you like; it downloads to the device of your choice and it will just work – and then when you have to jump through all these hoops [to buy legitimate content], the walls created are disincentives for people to buy," he said.

Brin acknowledged that some people were anxious about the amount of their data that was now in the reach of US authorities because it sits on Google's servers. He said the company was periodically forced to hand over data and sometimes prevented by legal restrictions from even notifying users that it had done so.

He said: "We push back a lot; we are able to turn down a lot of these requests. We do everything possible to protect the data. If we could wave a magic wand and not be subject to US law, that would be great. If we could be in some magical jurisdiction that everyone in the world trusted, that would be great … We're doing it as well as can be done."