Showing posts with label Information Economy. Show all posts
Showing posts with label Information Economy. Show all posts

Wednesday, December 12, 2012

Digital Universe is Expanding


Quote from FT below: “Only 0.4 per cent of that data are actually being analysed.”

Paul Taylor, The Financial Times, December 11, 2012 

The “digital universe” continues to expand rapidly and will double every two years through 2020, but only a small fraction of this ‘big data’ is being analysed, which leaves a huge untapped opportunity for companies and other organisations.

According to a report published on Tuesday by IDC, the market research firm, a large ‘big data gap’ is opening up between the rapidly expanding digital universe – a measure of all the digital data created, replicated and consumed in a single year. That data holds potential analytic value that could have major business, health and societal impact.

Only 0.4 per cent of that data are actually being analysed.

The reason, IDC suggests, is that the majority of new data being generated is unstructured – for example video surveillance footage, audio files or social media – meaning we know little about the data until it is characterised and tagged. The industry currently faces a shortage of the talent and technology needed to tag and analyse unstructured data.

IDC’s sixth annual report on the “Digital Universe,” sponsored by EMC, the storage market leader, notes that today the digital universe comprises “the images and videos on mobile phones uploaded to YouTube, digital movies populating the pixels of our high-definition TVs, banking data swiped in an ATM, security footage at airports and major events such as the Olympic Games, subatomic collisions recorded by the Large Hadron Collider at CERN, transponders recording highway tolls, voice calls zipping through digital phone lines, and texting as a widespread means of communications.

“With the rise of big data awareness and analytics technology, the digital universe in 2012 has taken on the feel of a tangible geography,” notes the report, “vast, barely charted place full of promise and danger.”
This untapped value could be found in patterns in social media usage, correlations in scientific data from discrete studies, medical information intersected with sociological data, faces in security footage, and so on.
However, even with a generous estimate, the amount of information in the digital universe that is “tagged” accounts for only about 3 per cent of the digital universe in 2012, and that which is analysed is half a per cent of the digital universe. “Herein is the promise of “big data” technology – the extraction of value from the large untapped pools of data in the digital universe,” says the report.

The reports authors conclude that “our digital universe in 2020 will be bigger than ever, more valuable than ever, and more volatile than ever.”

Not surprisingly, IDC predicts that big data will be a big boon for the IT industry. “Web sites that gather significant data need to find ways to monetise this asset. Data scientists must be absolutely sure that the intersection of disparate data sets yields repeatable results if new businesses are going to emerge and thrive. Further, companies that deliver the most creative and meaningful ways to display the results of big data analytics will be coveted and sought after.”

Among the report’s main findings it suggests:

n From 2005 to 2020, the digital universe will grow by a factor of 300 – more than 5,200 gigabytes for every man, woman, and child in 2020. From now until 2020, the digital universe will about double every two years.

n The investment in managing, containing, studying and storing the bits in the digital universe will only grow by 40 per cent between 2012 and 2020. As a result, the investment per gigabyte during that same period will drop from $2.00 to $0.20.

n Between 2012 and 2020, emerging markets’ share of the expanding digital universe will grow from 36 per cent to 62 per cent.

n A majority of the information in the digital universe, 68 per cent in 2012, is created and consumed by consumers – watching digital TV, interacting with social media, sending camera phone images and videos between devices and around the Internet, and so on. Yet enterprises have liability or responsibility for nearly 80 per cent of the information in the digital universe.

n Only a small fraction of the digital universe has been explored for analytic value. IDC estimates that by 2020, as much as 33 per cent of the digital universe will contain information that might be valuable if analysed, compared to 23 per cent today.

n By 2020, nearly 40 per of the information in the digital universe will be “touched” by cloud computing providers – meaning that a byte will be stored or processed in a cloud somewhere in its journey from originator to disposal.

n The proportion of data in the digital universe that requires protection is growing faster than the digital universe itself, from less than a third in 2010 to more than 40 per in 2020.

n The amount of information individuals create themselves – for example writing documents, taking pictures, downloading music – is far less than the amount of information being created about them in the digital universe.

See results of EMC study at http://www.emc.com/leadership/digital-universe/index.htm

Sunday, December 2, 2012

Digital Privacy in the Big Data Era: Microsoft's Data Protection Keynote


Ms. Smith, Network World, December 2, 2012

There are several Internet security experts who agree with Steve Rambam's claim [1] that "Privacy is dead - get over it." Yet other privacy and security experts such as Bruce Schneier completely disagree. In The Value of Privacy [2] Schneier wrote, "Privacy protects us from abuses by those in power, even if we're doing nothing wrong at the time of surveillance." When it comes to data protection and protecting people's privacy in the digital age, Europe is far more advanced than America. [3]

In fact, the head of France's data protection agency, Isabelle Falque-Pierrotin, did an excellent job summing it up as: "In Europe, we consider privacy a fundamental right. That doesn't mean it is exclusive of other rights, but economic rights are not superior to privacy." The New York Times also reported [4] that she said in the United States, "personal data are seen as raw material for business."

In November, Microsoft's Chief Privacy Officer Brendon Lynch said [5] of the IAPP European Data Protection Congress 2012 [6], "One area of strong consensus was the tremendous potential the digital economy holds for companies on both sides of the pond. Accordingly, it's important to strike the right balance between data protection with business growth through interoperability between privacy regulation in the EU, U.S. and elsewhere."

Many privacy advocates cringe when hearing the word "balance," such as striking a balance between security and privacy. Hopefully people won't come to cringe when they hear the word balance applied to big data security protections and privacy. As Bruce Schneier wrote [2] way back in 2006:

Too many wrongly characterize the debate as "security versus privacy." The real choice is liberty versus control. Tyranny, whether it arises under threat of foreign physical attack or under constant domestic authoritative scrutiny, is still tyranny. Liberty requires security without intrusion, security plus privacy. Widespread police surveillance is the very definition of a police state. And that's why we should champion privacy even when we have nothing to hide.

Whether people realize it or not, big data is not privacy-friendly even when it is supposedly anonymized or contains obfuscated PII (Personally Identifiable Information) data. Researchers have shown that "linkability threats" can re-identity individuals. Since it boils down to the fact that you are not anonymous when it comes to big data [7], Microsoft has developed "Differential Privacy for everyone" [download PDF [8]].
In the IAPP keynote address [download PDF [9]], Lynch made some excellent and thought-provoking privacy points regarding big data. He said:

Data is the fuel that drives all of these powerful technologies, but what can be done with the data today can at times seem enormously helpful or enormously threatening. Consider two scenarios shown here. In the first case, I am using my phone in a grocery store to find out more about the items on the shelves and it is mashing up that with my private data to personalize my experience. So here I downloaded a recipe and customized it for my dietary needs. If it's a trusted system, that's a great experience. On the other hand, consider the US company, Target, which recently generated a lot of press about its pregnancy prediction score. This was based on what people were purchasing in Target stores, they are able to indicate a shopper that appeared to be pregnant. The concern about how Target can figure out such details about customers shopping in its stores, who are not explicitly sharing that information, is the concern. And what does it do with those insights? In this particular case, they sent some mailers to the individual involved - it was a teenage girl and her father was very offended that they were wrongly marketing to her, but it eventually did come out that she was in fact pregnant. Target knew a lot more than her father knew.

Peter Cullen, Microsoft's Chief Privacy Strategist, wrote [10] about "notice and consent" as a means of privacy protection and how data privacy frameworks need to "focus on the 'harms' or 'impacts' of data use, which should not only include physical and financial injury, but also broader concepts such as reputational or social harm."

Yet after showing a video that highlighted data transfers in today's world at the IAPP conference, Lynch said, "How could there possibly be meaningful notice and consent mechanisms in place for every transfer of data that was involved?" He added, "It would seem that advances in technology and the rise of big data can create amazing societal benefits but they can also strain traditional notions of secrecy and the notice and consent approach to privacy protection."

In his keynote, Lynch said:

Some technology and internet companies today take the position that privacy is dead, or at least that privacy is an outdated concept that people need to get over so technology companies can help them reap the benefits of sharing as much information as possible. But we disagree that privacy is not relevant or desirable, in this sensor-driven, social everywhere, big data world that we are heading towards. People today expect strong privacy protections because they are increasingly aware of, and concerned about, the digital trails they leave behind online and indeed there's plenty of evidence that people still care deeply about privacy.

Of course people care about privacy. Europe continues to illustrate this to the world by taking a hard stance when data is used without "informed consent" and when users cannot "opt out." Lynch believes we need to not only protect privacy in regards to big data, but also that people "need an updated notion of privacy and data protection principles, one that shifts from a focus on secrecy to a more nuanced approach, based on reasonable consumer expectations, context and a greater emphasis on how personal information is used."

Big data definitely represents significant threats to personal privacy. Let's hope this "shift" and "updated notion of privacy" won't include the word "balance" that puts individuals on the losing end as it generally has when the government talks of striking a balance between security and privacy.

Tuesday, November 13, 2012

How 'Social Intelligence' Can Guide Decisions


By offering decision makers rich real-time data, social media is giving some companies fresh strategic insight.

Martin Harrysson, Estelle Metayer, and Hugo Sarrazin, McKinsey Quarterly, November 2012

In many companies, marketers have been first movers in social media, tapping into it for insights on how consumers think and behave. As social technologies mature and organizations become convinced of their power, we believe they will take on a broader role: informing competitive strategy. In particular, social media should help companies overcome some limits of old-school intelligence gathering, which typically involves collecting information from a range of public and propriety sources, distilling insights using time-tested analytic methods, and creating reports for internal company “clients” often “siloed” by function or business unit.

Today, many people who have expert knowledge and shape perceptions about markets are freely exchanging data and viewpoints through social platforms. By identifying and engaging these players, employing potent Web-focused analytics to draw strategic meaning from social-media data, and channeling this information to people within the organization who need and want it, companies can develop a “social intelligence” that is forward looking, global in scope, and capable of playing out in real time.

This isn’t to suggest that “social” will entirely displace current methods of intelligence gathering. But it should emerge as a strong complement. As it does, social-intelligence literacy will become a critical asset for C-level executives and board members seeking the best possible basis for their decisions.

In this article, we explore four distinct ways social technologies can augment the intelligence-gathering approaches of companies. As Exhibit 1 makes clear, social media has little effect on some aspects of the intelligence cycle—in particular, the need to identify priorities for exploration and decision making over the next 6 to 12 months, as well as the use of assembled information to make unbiased decisions. But social technologies can play a surprisingly central role in how information is sourced, collected, analyzed, and distributed.

Tuesday, October 23, 2012

Gartner Says Big Data Creates Big Jobs: 4.4 Million IT Jobs Globally to Support Big Data By 2015



Analysts Discuss Key Issues Facing the IT Industry During Gartner Symposium/ITxpo 2012, October 21-25, in Orlando

ORLANDO, Fla.--(BUSINESS WIRE)-- Worldwide IT spending is forecast to surpass $3.7 trillion in 2013, a 3.8 percent increase from 2012 projected spending of $3.6 trillion, but it's the outlook for big data that is creating much excitement, according to Gartner, Inc.

"By 2015, 4.4 million IT jobs globally will be created to support big data, generating 1.9 million IT jobs in the United States," said Peter Sondergaard, senior vice president at Gartner and global head of Research. "In addition, every big data-related role in the U.S. will create employment for three people outside of IT, so over the next four years a total of 6 million jobs in the U.S. will be generated by the information economy."

"But there is a challenge. There is not enough talent in the industry. Our public and private education systems are failing us. Therefore, only one-third of the IT jobs will be filled. Data experts will be a scarce, valuable commodity," Mr. Sondergaard said. "IT leaders will need immediate focus on how their organization develops and attracts the skills required. These jobs will be needed to grow your business. These jobs are the future of the new information economy."

Mr. Sondergaard provided the latest outlook for the IT industry today to an audience of more than 8,000 CIOs and IT leaders at Gartner Symposium/ITxpo, which is taking place here through October 25. He said the IT industry is entering the Nexus of Forces, which includes a confluence and integration of cloud, social collaboration, mobile and information.

"This is a time of accelerating change, where your current IT architecture will be rendered obsolete," Mr. Sondergaard said. "You must lead through this change, selectively destroy low impact systems, and aggressively change your IT cost structure. This is the New World of the Nexus, the next age of computing."

Cloud
The cloud is the carrier for the three other Forces: mobile is personal cloud, social media is only possible via the cloud, and big data is the killer app for the cloud. Cloud will be the permanent fixture, the foundation.

"Cloud is not merely about cost-cutting, the end game is not just cheap on-demand services. In fact, 90 percent of these services are still subscription based, not pay-as-you-go," Mr. Sondergaard said. "We are just at the beginning of realizing the cost benefits of cloud, but organizations moving to the cloud are also attracted by the new capabilities they do not get today. It is bringing new approaches to designing applications, specifically for the cloud, and providing more resilience by architecturing failure as a design concept. Cloud also teaches us about services and service levels, and the contrast between what the business wants for outcomes versus IT's old methods of getting there."

Mobile
In 2016, more than 1.6 billion smart mobile devices will be purchased globally. Two-thirds of the mobile workforce will own a smartphone, and 40 percent of the workforce will be mobile. The challenge for IT leaders is determining what to do with this new channel to their customers and employees.

"Mobile is about computing at the right time, in the moment. It is the point of entry for all applications, delivering personalized, contextual experiences," Mr. Sondergaard said. "It means: marketing gets more time with the customer; employees become more productive; and process flows get dramatically cut."

In less than two years, iPads will be more common in business than Blackberries. Mr. Sondergaard said some CIOs are now placing orders for tens of thousands of iPads at a time. Productivity is the driver. Two years from now, 20 percent of sales organizations will use tablets as the primary mobile platform for their field sales force. As a result, by 2018, 70 percent of mobile workers will use a tablet or a hybrid device that has tablet-like characteristics.

Gartner forecasts that in 2016, half of all non-PC devices will be purchased by employees. By the end of the decade, half of all devices in business will be purchased by employees.

Social Computing
In the next three years, the dominant consumer social networks will the limits of their growth. However, social computing will become even more important. Companies are establishing social media as a discipline. Gartner predicts that in three years, 10 organizations will each spend more than $1 billion on social media.

"Social computing is moving from being just on the outside of the organization to being at the core of business operations," Mr. Sondergaard said. "It is changing the fundamentals of management: how you establish a sense of purpose and motivate people to act. Social computing will move organizations from hierarchical structures and defined teams to communities that can cross any organizational boundary."

Big Data
By tapping a continual stream of information from internal and external sources, businesses today have an endless array of new opportunities for: transforming decision-making; discovering new insights; optimizing the business; and innovating their industries.

Big data creates a new layer in the economy which is all about information, turning information, or data, into revenue. This will accelerate growth in the global economy and create jobs.

"Big data is about looking ahead, beyond what everybody else sees," Mr. Sondergaard said. "You need to understand how to deal with hybrid data, meaning the combination of structured and unstructured data, and how you shine a light on 'dark data.' Dark data is the data being collected, but going unused despite its value. Leading organizations of the future will be distinguished by the quality of their predictive algorithms. This is the CIO challenge, and opportunity."

About Gartner Symposium/ITxpo
Gartner Symposium/ITxpo is the world's most important gathering of CIOs and senior IT executives. This event delivers independent and objective content with the authority and weight of the world's leading IT research and advisory organization, and provides access to the latest solutions from key technology providers. Gartner's annual Symposium/ITxpo events are key components of attendees' annual planning efforts. IT executives rely on Gartner Symposium/ITxpo to gain insight into how their organizations can use IT to address business challenges and improve operational efficiency.

Additional information about Gartner Symposium/ITxpo in Orlando, is available at www.gartner.com/symposium/us. Video replays of keynotes and sessions are available on Gartner Events on Demand at www.gartnerondemand.com. Follow news, photos and video coming from Gartner Symposium/ITxpo on Facebook at http://www.facebook.com/GartnerSymposium, and on Twitter at http://twitter.com/Gartner_inc and using #GartnerSym.

Friday, October 19, 2012

Will Big Data decide the election?


A new book traces the recent history of data mining in political campaigns. A review of The Victory Lab, by Sasha Issenberg.

Chip Lebovitz, Fortune, October 19, 2012

FORTUNE -- There's a powerful vignette in Sasha Issenberg's The Victory Lab in which political consultant Alexander Gage presents his new data targeting system to Mitt Romney's 2002 gubernatorial campaign.

Gage has combined consumer records with political voting history to identify potential Romney supporters among nontraditional Republican voting blocks. Gage sees his work as revolutionary -- a first in politics, and potentially a first anywhere. Yet just as he completes his presentation, Romney's deputy campaign manager Alex Dunn raises his hand and deadpans, "You mean you don't do this in politics."

Dunn's surprise was not out of place in 2001. And while much has changed since then, politics remains an analog art in many ways. The Victory Lab charts the recent history of political data mining designed to identify persuadable voters and swing elections.

Politico has called Issenberg's book "Moneyball for politics." There are obvious parallels between political operatives like Gage and the protagonist of Moneyball, Oakland Athletics general manager Billy Beane. 

Both pioneered data-centric techniques that gave their respective organizations a leg up over the competition.

Yet in his 14 years as GM, Billy Beane has yet to make it to the World Series. By contrast, the work of Gage and other data-mining political consultants has translated directly into electoral success. These operatives may not have singlehandedly won the past two presidential elections, but the side that had the most advanced data-driven mobilization efforts went a perfect 2-0.

The George W. Bush campaign's sophisticated data mining helped the president win reelection in 2004. Four years later, Barack Obama defeated John McCain in part because the Obama campaign outclassed McCain's operatives on the data front.

Issenberg seems to have interviewed everyone who's anyone in the field of political statistics, but his subject matter doesn't necessarily lend itself to prose. Numbers have power, but too many of them can addle rather than inform. After reading the book, you'll probably remember the outcome of a few experiments, but be stuck scouring your brain for the results of the rest. Readers are advised to keep a pen and pencil handy when reading The Victory Lab, so they can jot down key stats.

Yet Issenberg's material is sufficiently gripping that you'll want to keep turning the pages, even if it means deciphering the results of yet another randomized trial. For example, micro-targeting efforts on behalf of Senator Michael Bennett (D-Colo.)'s 2010 Senate reelection campaign likely spurred 25,000 additional Bennett votes. That might not seem like a big number, until you consider that the race was decided by 15,000 votes.

Issenberg avoids sweeping predictions about the future of political data mining. That's probably wise, given that political professionals tend to define the future as Election Day. Campaigns at their simplest are a one-day snapshot of personal preference. The Victory Lab reminds us, however, that every individual choice is the result of a thousand factors, many of them subject to manipulation by political operatives.


Monday, October 15, 2012

McKinsey Anthology: Government Designed for New Times


To explore the approaches that governments around the world are taking to common problems, this anthology convenes political leaders and civil servants, economists and policy experts, generalists and specialists.

McKinsey Anthology, 2012
               
Contents
        Transforming government
·       Tony Blair—Leading transformation in the 21st century
·       François-Daniel Migeon—Interview: Transforming government in France
·       Frank-Jürgen Weise—Behind the German jobs miracle
·       Tim Brown—Quick take: Designing a tech-enabled government
·       Michael Fullan—Transforming schools an entire system at a time
·       Todd Park—Interview: Unleashing government's 'innovation mojo'
·       Diana Farrell—Government designed for new times
        Innovating government services

·       James Fishkin—What the people think when they're really thinking
·       Nandan Nilekani—Interview: For every citizen, an identity
·       Matthew Taylor—Citizens: The untapped resource
·       Susan Zielinski—The new mobility
·       Peter Shergold—A social contract for government
·       Wim Elfrink—The smart-city solution
·       Karan Bhatia—Quick take: Building the world's infrastructure
·       Salman Khan—Teaching for the new millennium
·       Xie Chengxiang—Interview: Home for the urban poor
·       Elena Berkowitz and Blaise Warren—Quick take: How Estonia became E-stonia    

        Building new competencies
·       Douglas Holtz-Eakin—Fiscal management fix: Simple math—and a very big stick
·       Coen Teulings—Why politicans prefer austerity to long-term fiscal reform
·       Lu Mai—The urbanization solution
·       Göran Persson—How to tame a budget crisis
·       Peter Ho—Coping with complexity
·       Mohamed Ibrahim—Better data, better policy making    
        Understanding government in new times

·       Daron Acemoglu—The servant state
·       Parag Khanna—The rise of hybrid governance
·       Neil deGrasse Tyson—Why exploration matters—and why the government should pay  
        for it
·       Ray O. Johnson—Quick take: The research imperative
·       Hernando de Soto—Interview: Building a nation of owners
·       Nicolas Berggruen and Nathan Gardels—A middle way for governance       

Wednesday, October 10, 2012

Harnessing Data as a New Source of Growth: Big Data Analytics and Policies


OECD Headquarters, Paris, France, October 22, 2012

Introduction
The confluence of several key socio-economic and technological trends is resulting in the generation of huge streams of data every day. The major trends include:

·       The increasing migration of social and economic activities on line: Social network site Facebook, for example, now counts over 900 million active participants around the world generating together more than 1 500 status updates every second about their interests and whereabouts. In 2011, e-commerce platform eBay collected data on more than 100 million active users including the 6 million new goods they offered every day.

·       The strong decline in the cost of data collection, storage, transportation, and processing: The average cost of consumer hard disk drives (HDDs) per gigabyte, for example, dropped on average by almost 40% per year between 1998 (USD 56 per gigabyte) and 2012 (USD 0.05 per gigabyte). In 1995, as another example, consumers in France paid USD 75 equivalent per month for a dial-up (56 Kb/s) connection, while in 2011 they paid the equivalence of USD 33 per month for a broadband (51 Mb/s) connection, which was almost 1 000 times faster.

·       The increasing deployment of “smart” ICT applications such as smart grids and smart transportations based on machine-to-machine (M2M) communication: Connecting one million homes to a smart grid may produce as much as 11 gigabytes of data per day. In order to accommodate for hourly readings through smart meters, a network with a minimum capacity of up to 1 Mbit/s dedicated to M2M communication is needed.

·       The continued expansion of mobile communication: In 2011, there were 780 million smart phones worldwide capable of collecting and transmitting geo-location data, which generated more than 600 petabytes (millions of gigabytes) of data every month. It is estimated that the global data traffic generated by mobile communication (including M2M enabled smart devices) will almost double every year to reach 11 exabyte (billions of gigabytes) per month by 2016.

The collection and exploitation of these large data flows through data analytics is leading to a shift towards a data-driven socio-economic model commonly referred to as “big data”. In this model, data is a core asset that provides a huge resource for new industries, processes, services and goods leading to significant competitive advantage. In business, for example, data analytics are increasingly being used in a wide number of operations ranging from optimising the value chain and manufacturing production to more efficiently using labour and improving customer relationships. It is estimated that firms, which adopt data-driven decision-making, for example, have output and productivity that is 5-6% higher than what would be expected given their other investments and information technology usage. These firms also perform better in terms of asset utilisation, return on equity and market value.

To unlock the potential of big data and data analytics (big data analytics), OECD countries need to ensure the development of coherent policies and practices around the collection, transportation, storage and use of data, most prominently in areas related to privacy protection. New data sources, new actors and the increasing ease with which personal data can be collected, linked and processed, seriously challenge the effective implementation of current frameworks for privacy protection. But the potential implications for policy also spill over into many other domains; including among others open access to data, intellectual property rights, competition, skills and employment, infrastructure, and measurement.

Objective
First launched in 2005, Technology Foresight Forums are an annual event organised by the OECD Committee for Information, Computer, and Communications Policy (ICCP) to help identify opportunities and challenges for the Internet Economy posed by technical developments. The 2012 Technology Foresight Forum will focus on the potential of big data analytics as a new source of growth, which could help generate significant economic and social benefits. It will put big data analytics in the context of emerging trends discussed at the last three Foresight Forums, namely mobile communications (2011), smart ICTs (2010), and cloud computing (2009), to highlight the confluence of trends leading towards a data-driven economy (see Figure below).
The confluence of major trends enabling data as a new source of growth

The 2012 Foresight Forum will then discuss the social and economic issues related to big data analytics that may put at risk fundamental values and warrant a review of current policy frameworks, most prominently those aimed at ensuring the protection of privacy, intellectual property, and competition. Other policy areas that will be addressed include health care, science and research, government administration and labour markets.

This year’s Foresight Forum will contribute to OECD’s ongoing horizontal project entitled New Sources of Growth: Intangible Assets (NSG) as well as to the follow-up horizontal project on Knowledge-Based Capital: Seizing the Benefits of New Sources of Growth, which will be launched in 2013. Both projects aim at (i) providing structured evidence of the economic value of intangible assets, including data, as a new source of growth, and (ii) improving understanding of current and emerging challenges for policies related to the increasing relevance of these intangible assets.

Modalities
The Foresight Forum represents a collaborative effort of policy makers, business, civil society, and the Internet technical community and thus provides opportunities to liaise with experts in the fields covered. To increase interactivity with a broader public, the Foresight Forum will be supported by various participative web technologies, including a webcast, and Twitter feeds. Further details will be announced closer to the event.

Agenda
The 2012 Foresight Forum will include two morning and three afternoon sessions: The first (morning) session will introduce data analytics in the context of the key technological and socio-economic trends, namely i) cloud computing; ii) smart ICT applications; and iii) the Internet of Things. The following session will then focus on the socio-economic implications of harnessing data as a new source for growth. The afternoon sessions will take a more in-depth look at specific areas that will include: i) science and research (including public health); ii) marketing and competition; and iii) public administration. Each of these sessions will be taken as a mean to discuss specific potential policy opportunities and challenges ranging from i) privacy and consumer protection; ii) intellectual property rights; iii) open access to data; iv) competition; and v) skills and employment. A concluding session will then highlight the main policy implications to be examined by the OECD in the context of its programme of work for 2013-14.

Thursday, September 20, 2012

Big Data for All


Omer Tene, Concurring Opinions, September 20, 2012

Much has been written over the past couple of years about “big data” (See, for example, here and here and here). In a new article, Big Data for All: Privacy and User Control in the Age of Analytics, which will be published in the Northwestern Journal of Technology and Intellectual Property, Jules Polonetsky and I try to reconcile the inherent tension between big data business models and individual privacy rights. We argue that going forward, organizations should provide individuals with practical, easy to use access to their information, so they can become active participants in the data economy. In addition, organizations should be required to be transparent about the decisional criteria underlying their data processing activities.

The term “big data” refers to advances in data mining and the massive increase in computing power and data storage capacity, which have expanded by orders of magnitude the scope of information available for organizations. Data are now available for analysis in raw form, escaping the confines of structured databases and enhancing researchers’ abilities to identify correlations and conceive of new, unanticipated uses for existing information. In addition, the increasing number of people, devices, and sensors that are now connected by digital networks has revolutionized the ability to generate, communicate, share, and access data.

Data creates enormous value for the world economy, driving innovation, productivity, efficiency and growth. In the article, we flesh out some compelling use cases for big data analysis. Consider, for example, a group of medical researchers who were able to parse out a harmful side effect of a combination of medications, which were used daily by millions of Americans, by analyzing massive amounts of online search queries. Or scientists who analyze mobile phone communications to better understand the needs of people who live in settlements or slums in developing countries.

At the same time, the “data deluge” presents formidable privacy concerns. Protecting privacy become harder as information is multiplied and shared ever more widely among multiple parties around the world. As more information regarding individuals’ health, financials, location, electricity use and online activity percolates, concerns arise about profiling, tracking, discrimination, exclusion, government surveillance and loss of control. From a more technical legal angle, big data challenges some of the most fundamental concepts of privacy law, including the definition of “personally identifiable information”, the role of individual control, and the principles of data minimization and purpose limitation.

In our article, we make the case for providing individuals with usable access to their data. The call for transparency is not new, of course. Rather the emphasis is on access to data in usable format, which can work to create value to individuals. Transparency and access alone have not emerged as potent tools because individuals do not care for, and cannot afford to indulge in transparency and access for their own sake (see one oft-cited counterexample here). The enabler of transparency and access is the ability to use the information and benefit from it in a tangible way. This will be achieved through “featurization” or “app-ification” of privacy. Organizations should build as many dials and levers as needed for individuals to engage with their data.

We expect that “featurization” of big data, harnessing its immense force for not only organizational but also individual benefit, will unleash a wave of innovation and create a market for personal data applications. The technological groundwork has already been completed with mash-ups and real-time APIs making it easier for organizations to combine information from different sources and services into a single user experience. Regardless of lingering questions concerning who – if anyone – “owns” the information, we think that fairness dictates that individuals enjoy beneficial use of the data about them.

Our second proposal would require organizations to disclose the decisional criteria underpinning their data analytics machinery. In a big data world, it is often not the data but rather the inferences drawn from them that give cause for concern. Inaccurate, manipulative or discriminatory conclusions may be drawn from perfectly innocuous, accurate data. Much like in quantum physics, the observer in big data analysis can affect the results of her research by defining the data set, proposing a hypothesis or writing an algorithm. At the end of the day, big data analysis is an interpretative process, in which one’s identity and perspective informs one’s results. Like any interpretative process, it is subject to error, inaccuracy and bias. Louis Brandeis, who together with Samuel Warren “invented” the legal right to privacy in 1890, has also written that “[s]unlight is said to be the best of disinfectants”. We trust if the existence and uses of databases were visible to the public, organizations would be more likely to avoid unethical or socially unacceptable uses of data.

Tuesday, September 11, 2012

Erik Brynjolfsson and Andrew McAfee : Big Data's Management Revolution

Big Data's Management Revolution

by Erik Brynjolfsson and Andrew McAfee  |  10:05 AM September 11, 2012

http://blogs.hbr.org/cs/2012/09/big_datas_management_revolutio.html

Big data has the potential to revolutionize management. Simply put, because of big data, managers can measure, and hence know, radically more about their businesses, and directly translate that knowledge into improved decision making and performance. Of course, companies such as Google and Amazon are already doing this. After all, we expect companies that were born digital to accomplish things that business executives could only dream of a generation ago. But in fact the use of big data has the potential to transform traditional businesses as well.

We've seen big data used in supply chain management to understand why a carmaker's defect rates in the field suddenly increased, in customer service to continually scan and intervene in the health care practices of millions of people, in planning and forecasting to better anticipate online sales on the basis of a data set of product characteristics, and so on.

Here's how two companies, both far from Silicon Valley upstarts, used new flows of information to radically improve performance.

Case #1: Using Big Data to Improve Predictions
Minutes matter in airports. So does accurate information about flight arrival times: If a plane lands before the ground staff is ready for it, the passengers and crew are effectively trapped, and if it shows up later than expected, the staff sits idle, driving up costs. So when a major U.S. airline learned from an internal study that about 10% of the flights into its major hub had at least a 10-minute gap between the estimated time of arrival and the actual arrival time — and 30% had a gap of at least five minutes — it decided to take action.

At the time, the airline was relying on the aviation industry's long-standing practice of using the ETAs provided by pilots. The pilots made these estimates during their final approach to the airport, when they had many other demands on their time and attention. In search of a better solution, the airline turned to PASSUR Aerospace, a provider of decision-support technologies for the aviation industry.

In 2001 PASSUR began offering its own arrival estimates as a service called RightETA. It calculated these times by combining publicly available data about weather, flight schedules, and other factors with proprietary data the company itself collected, including feeds from a network of passive radar stations it had installed near airports to gather data about every plane in the local sky.

PASSUR started with just a few of these installations, but by 2012 it had more than 155. Every 4.6 seconds it collects a wide range of information about every plane that it "sees." This yields a huge and constant flood of digital data. What's more, the company keeps all the data it has gathered over time, so it has an immense body of multidimensional information spanning more than a decade. RightETA essentially works by asking itself "What happened all the previous times a plane approached this airport under these conditions? When did it actually land?"

After switching to RightETA, the airline virtually eliminated gaps between estimated and actual arrival times. PASSUR believes that enabling an airline to know when its planes are going to land and plan accordingly is worth several million dollars a year at each airport. It's a simple formula: Using big data leads to better predictions, and better predictions yield better decisions.

Case #2: Using Big Data to Drive Sales
A couple of years ago, Sears Holdings came to the conclusion that it needed to generate greater value from the huge amounts of customer, product, and promotion data it collected from its Sears, Craftsman, and Lands' End brands. Obviously, it would be valuable to combine and make use of all these data to tailor promotions and other offerings to customers, and to personalize the offers to take advantage of local conditions.

Valuable, but difficult: Sears required about eight weeks to generate personalized promotions, at which point many of them were no longer optimal for the company. It took so long mainly because the data required for these large-scale analyses were both voluminous and highly fragmented — housed in many databases and "data warehouses" maintained by the various brands.

In search of a faster, cheaper way, Sears Holdings turned to the technologies and practices of big data. As one of its first steps, it set up a Hadoop cluster. This is simply a group of inexpensive commodity servers whose activities are coordinated by an emerging software framework called Hadoop (named after a toy elephant in the household of Doug Cutting, one of its developers).

Sears started using the cluster to store incoming data from all its brands and to hold data from existing data warehouses. It then conducted analyses on the cluster directly, avoiding the time-consuming complexities of pulling data from various sources and combining them so that they can be analyzed. This change allowed the company to be much faster and more precise with its promotions.

According to the company's CTO, Phil Shelley, the time needed to generate a comprehensive set of promotions dropped from eight weeks to one, and is still dropping. And these promotions are of higher quality, because they're more timely, more granular, and more personalized. Sears's Hadoop cluster stores and processes several petabytes of data at a fraction of the cost of a comparable standard data warehouse.

These aren't just a few flashy examples. We believe there is a more fundamental transformation of the economy happening. We've become convinced that almost no sphere of business activity will remain untouched by this movement.

Without question, many barriers to success remain. There are too few data scientists to go around. The technologies are new and in some cases exotic. It's too easy to mistake correlation for causation and to find misleading patterns in the data. The cultural challenges are enormous, and, of course, privacy concerns are only going to become more significant. But the underlying trends, both in the technology and in the business payoff, are unmistakable.

The evidence is clear: Data-driven decisions tend to be better decisions. In sector after sector, companies that embrace this fact will pull away from their rivals. We can't say that all the winners will be harnessing big data to transform decision making. But the data tell us that's the surest bet.

This blog post was excerpted from the authors' upcoming article "Big Data: The Management Revolution," which will appear in the October issue of Harvard Business Review.
_____________________

BIG DATA INSIGHT CENTER

·       Use Big Data to Find New Micromarkets

·       Integrate Data Into Products, or Get Left Behind

·       How to Avoid the Big Data "Gotcha's"

·       Big Data, Analytics and the Path from Insights to Value

Here Comes the Data Economy


New companies are creating services using government data on health care, education, and more.

Alexander B. Howard, Slate, September 10, 2012


We're living in the exabyte age, where the actions of billions of humans using the Web and their mobile devices are creating massive amounts of big data to collect, store, analyze, and put to work.

If big data is a strategic resource, as has been suggested, then many national and state governments have public reserves that can be tapped for the public good in this young century's version of the industrial revolution. Given that the United States economy is still coming out of the worst recession and financial shock since the Great Depression, supporting civic and tech entrepreneurs enjoys political support from both sides of the aisle.

Entrepreneurs, big and small, are mashing up data from the rapidly expanding collection of sources and building new businesses on it or improve their existing services, like Zillow or Google Maps or Consumer Reports or Bloomberg Government. In a time when job creation is critical, using public sector information to create jobs isn’t an aim to dismiss lightly, although the terms and conditions under which such activity occurs must be clear to all actors involved, to avoid the creation of new monopolies based upon artificial scarcity.

My publisher, long-time open source and open government advocate Tim O'Reilly, has asked how government can act as a platform to enable people inside and outside government to innovate on top of it. One answer is certainly releasing open data. In that context, open data and application programming interfaces, more commonly known as APIs, increasingly look like fundamental infrastructure for digital government in the 21st century.

There's good reason to think that open data could have an overall effect on the economy akin to open source and small business. Gartner, the IT research analysis firm, recently highlighted how open data creates value in the public and private sector.

You may not realize it, but services you use on a daily basis have been built upon data released by the government. Weather data collected by the National Oceanic and Atmospheric Association has an annual estimated economic value of $10 billion, according to U.S. Chief Information Officer Steven VanRoekel and U.S. Chief Technology Officer Todd Park. NOAA data sets are used by Weather.com, Weather Underground, and the Weather Channel—and the nation's farmers consult these forecasts to manage both their crops and the risks of loss. VanRoekel and Park estimate the annual economic value of the data from the U.S. global positioning system at some $90 billion. From companies like TomTom or Garmin to dashboard GPS systems to smartphones and associated location-based applications, GPS data sets are baked into an expanding number of services and products.

Now, as Park seeks to scale open data across the federal government, we’re on the verge of the next generation of services driven by open data, which will involve everything from energy to health care to consumer finance to transit sectors. The challenge is that the cities and federal agencies that hold vast amounts of data may not always understand the value of the information they hold or how to create or sustain businesses using it. That's where open innovation in the public sector and the dynamism of entrepreneurs will play an important role in making the people's data more useful to the people.

BrightScope is a notable example of what dogged persistence can create. The California startup made a profitable business using government data to help the American people understand the fees associated with their 401(k)s. Last May, BrightScope went further, launching financial adviser pages based on open government data from the Securities and Exchange Commission and the Financial Industry Regulatory Authority, the largest independent securities regulator in the United States. Previously, financial adviser profiles could only be found through exact queries at an obscure URL on the regulators' websites. Now, information that citizens care about—the records of financial advisers in their geographic region—is available where they're looking for it: in search engine results.

Just as labor and regulatory data fuels BrightScope's business, there's an expanding number of startups that are tapping into other data released so-called “smart disclosure” initiatives. Smart disclosure is when a private company or government agency provides a person with periodic access to his or her own data in open formats that enable them to easily put the information to use. Startups like Billshrink.com and Hello Wallet are already using a combination of private sector and public sector data to enhance consumer finance decisions. The success of such consumer finance startups suggests an important lesson: The most successful apps and services will combine government, industry, and user-generated data.

The key open data story to watch in the federal government, however, centers on health care. McKinsey and Associates estimates the annual economic value of big, open liquid health data at about $350 billion annually. The explosion of mHealth apps are just the beginning of the disruption in health care from open health data. The effort to revolutionize the health care industry by making health data as useful as weather data is still in its infancy—but the early results are promising. iTriage, which was acquired by Aetna, is enabling people to make better mobile health care decisions where and when they need to do so. It uses a combination of government and private sector data to evaluable symptoms or conditions and point users to nearby medical care. Another startup, Castlight, is analyzing health care data to empower patients, acting like Kayak.com for those who want more transparency about costs. In May, Castlight completed a $100 million round of financing.

But for these sorts of initiatives to take off, entrepreneurs and regulators will have to work together to get contextual consent right and inform patients about the reuse of their data. Transparency is crucial to building a health data commons and thriving startup ecosytem based upon it.

If that balance can be struck, there's considerable potential for entrepreneurs to create better civic interfaces for many digital services. If open government data have helped build new tools, open data disclosed by private companies could create even more value for citizens. But currently, few businesses release anonymized data in an open, usable format. It will soon be time for the government to step in, convene stakeholders, and answer some key questions: How can we create uniform standards that will allow entrepreneurs and developers to innovate? When should data be licensed? Most of the big data releases we have seen come from finance, with bank records or stock trades. But there are significant opportunities to help both entrepreneurs and empowered consumers in health care, energy, education, and telecommunications, to name just a few.

Just as the glowing blue dot on the maps in our smartphone screens revolutionized how we navigate the world, similar "blue dots" could emerge for health care, finance, energy, and any product or service that is regulated or cataloged by government and industry. First, however, they'll need to open the data.

Also in the Future Tense package on government and open data: why Yelp and the government should share data; what a burger mob tells us about the future of democracy; and how Mexico is using open data to move beyond its authoritarian past.


Friday, August 24, 2012

Don't Build a Database of Ruin


Paul Ohm, Harvard Business Review, August 23, 2012

Many businesses today find themselves locked in an arms race with competitors to see who can convert customer secrets into the most pennies. To try to win, they are building perfect digital dossiers, to use a phrase coined by Daniel Solove, massive data stores containing hundreds, if not thousands or tens of thousands, of facts about every member of our society. 
In my work, I've argued that these databases will grow to connect every individual to at least one closely guarded secret. This might be a secret about a medical condition, family history, or personal preference. It is a secret that, if revealed, would cause more than embarrassment or shame; it would lead to serious, concrete, devastating harm. And these companies are combining their data stores, which will give rise to a single, massive database. I call this the Database of Ruin. Once we have created this database, it is unlikely we will ever be able to tear it apart.

I have become convinced that my earlier, bleak predictions about the Database of Ruin were in fact understated, arriving before it was clear how Big Data would accelerate the problem. Consider the most famous recent example of big data's utility in invading personal privacy: Target's analytics team can determine which shoppers are pregnant, and even predict their delivery dates, by detecting subtle shifts in purchasing habits. This is only one of countless similarly invasive Big Data efforts being pursued. In the absence of intervention, soon companies will know things about us that we do not even know about ourselves. This is the exciting possibility of Big Data, but for privacy, it is a recipe for disaster.

If we stick to our current path, the Database of Ruin will become an inevitable fixture of our future landscape, one that will be littered with lives ruined by the exploitation of data assembled for profit. But we can chart a different course, in various ways. I think our brightest engineers can develop innovative privacy-enhancing technologies which will enable new techniques for data analytics that minimize costs to privacy. I hope that public institutions and industry, through self-regulation, will devise ways to better balance the burdens on privacy and the benefits of Big Data. If nothing else, I anticipate that society will slowly develop new norms for engaging with the massive amount of information collected about us, creating informal rules governing when and how it is appropriate to release, collect, and use data, the way minors have learned to speak and listen carefully on social networks.

But every one of these correctives requires the same thing: time. We need to slow things down, to give our institutions, individuals, and processes the time they need to find new and better solutions. The only way we will buy this time is if companies learn to say, "no" to some of the privacy-invading innovations they're pursuing. Executives should require those who work for them to justify new invasions of privacy against a heavy burden, weighing them against not only the financial upside, but also against the potential costs to individuals, society, and the firm's reputation. Companies should do this not only as matter of good corporate social responsibility, but also because it will likely square with the government's recommendations for protecting privacy, which seem to advise caution and deliberation, under the banner of "context."

Earlier this year, Federal government officials released two privacy reports — the White House's White Paper and the FTC's Final Privacy Report — that together describe a national privacy policy for the foreseeable future. Although the two reports vary on some particulars, they both point to context as a central, important, and fundamental measuring stick we should use to assess decisions that bear on personal privacy.

The FTC report offers three broad recommendations: Privacy by Design, Simplified Choice for Businesses and Consumers, and Greater Transparency. In discussing the second recommendation — a call for simplified and more transparent choice — the FTC suggests a carve out. "Companies do not need to provide choice before collecting and using consumer data for practices that are consistent with the context of the transaction or the company's relationship with the consumer, or are required or specifically authorized by law." Under this standard, it might be "consistent with the context," for a company in a direct business relationship with a customer to use that customer's information to deliver ads for its other services, but it might be inconsistent with the context — thus requiring notice and choice — to sell that information to third-party advertisers, the FTC explains.

Similarly, the White House white paper defines a "Consumer Privacy Bill of Rights," which would protect, among other things, "Respect for Context." "Consumers have a right to expect that companies will collect, use, and disclose personal data in ways that are consistent with the context in which consumers provide the data," the paper explains.

These parallel pronouncements mean that companies that deal with personal information (meaning all companies, really) need to focus much more often than they have on the history of privacy practices in their industries. Although neither report defines in depth what it means by the word "context," to me the message seems to be: do not push the privacy envelope. Companies that use personal information in ways that go well beyond the practices of their competitors risk crossing the line from responsible steward to reckless abuser of consumer privacy.

The lesson is plain: compete vigorously and beat your competitors in every legitimate way, except when it comes to privacy invasion. Too many companies have learned this lesson the hard way, launching invasive new services that have triggered class action lawsuits, Congressional inquiries, and media firestorms. These companies knew that they were treading where others had feared to go. This may have felt like an exciting opportunity. It should have felt instead like perilous risk-taking, because it meant hurtling beyond the contextual borderlands defined by past practice.