Thursday, October 11, 2012

Big Data: The End of Privacy or a New Beginning?


Rubinstein, Ira. "Big Data: The End of Privacy or a New Beginning?" New York University, Information Law Institute (October 5, 2012).

From the abstract“Big data” refers to novel ways in which organizations, including government and businesses, combine diverse digital data sets and then use statistics and other data mining techniques to extract from them both hidden information and surprising correlations. While big data promises significant economic and social benefits, it also raises serious privacy concerns. In particular, big data challenges the Fair Information Practices (FIPs) as embodied in various privacy laws including the EU Data Protection Directive. This past January, the European Commission released a proposal to reform and replace the Directive by adopting a new Regulation. In this paper, I argue that this Regulation relies too heavily on the discredited informed choice model, and therefore fails to fully engage with the coming big data tsunami. My contention is that when this advancing wave arrives, it will so overwhelm the core privacy principles of informed choice and data minimization on which the Directive rests that reform efforts alone will prove inadequate. Rather, an adequate response must combine legal reform with encouragement of new business models premised on consumer empowerment and supported by a personal data ecosystem. This new business model is important for two reasons: first, existing business models have proven time and again that privacy regulation is no match for them. Businesses inevitably collect and use more and more personal data, and while consumers realize many benefits in exchange, there is little doubt that businesses, not consumers, control the market in personal data with their own interests in mind. Second, a new business model, which I describe in this paper, promises to stand processing of personal data on its head by shifting control over both the collection and use of data from firms to individuals. This “control shift” — and this alone — stands a chance of making the FIPs efficacious by giving individuals the capacity to benefit from big data and hence the motivation to learn about and control how their data is collected and used, while also enabling businesses to profit from a new breed of services that are both data-intensive and imbued with privacy values." Read more

Walking the Talk: Philanthropy 'Does' Big Data


Bradford K. Smith, PhilanTopic, October 9, 2012

(Bradford K. Smith is president of the Foundation Center. In his last post, he took a closer look at the China Foundation Center's new Foundation Transparency Index.

With the modestly labeled "Reporting Commitment," fifteen of America's largest foundations are transforming the practice of philanthropy. From today on, information about their grants will be made available on a near-real-time basis, as entirely open data and coded to a common geographical standard, making it easy to see the communities, regions, and countries that benefit from those grants. The initiative's simple name should not deceive: this is big. The participants -- Annenberg, Carnegie, Gates, Getty, Hewlett, Packard, MacArthur, Mott, Robert Wood Johnson, and six others -- provide nearly 12 percent of the $46 billion in grants made by American foundations each year. To see the Reporting Commitment in action, take a quick look at Glasspockets, the transparency Web site of the Foundation Center, then read on.


What makes the Reporting Commitment so transformative? Let's break it down.

A Bold Idea -- Real-Time Reporting
The fifteen participating foundations have committed to electronically report their current grants data to the Foundation Center on at least a quarterly basis. As pragmatic as this may sound, it's a dramatic departure from the norm for the field. All the 76,000 private foundations in America file 990-PF tax returns in which they provide information on their grants. They have up to a year after the close of their fiscal year to file these returns, the Foundation Center eventually gets them from the IRS as image files and converts them into a more usable format, cleans and codes the data, and insures public access through databases and research reports. In a world where value is being created exponentially by analyzing enormous real-time data sets generated through search logs, consumer purchases, and Facebook "likes," philanthropy remains an industry with $640 billion in assets that relies on two-year old data to understand its own grant trends.

The Foundation Center has convinced more than seven hundred foundations to electronically submit their grants information through its eGrant Reporting Program, covering more than 20 percent of total foundation giving. Although this provides the field with current-year grants data, most participating foundations report on an annual basis. The Reporting Commitment takes this effort one important step forward by having participating foundations report at least quarterly -- with some reporting weekly, even daily.

A Radical Idea -- Open Data
By and large, foundations tend to think of open data and transparency as something they should fund rather than do. There are lots of reasons for this, including the private nature of foundations, the cultural legacy of keeping a low profile and "letting our good works speak for themselves," and sensitivity surrounding some of the issues addressed by foundation grants. Notwithstanding, the ability of foundations to not call attention to themselves is being steadily eroded by the ease of finding, displaying, and circulating information in a densely networked, digital age. Meanwhile, sectors and institutions with which foundations increasingly collaborate, such as the World Bank and foreign aid donors, are barreling ahead with initiatives like the Open Aid Partnership and Publish What You Fund.

The fifteen Reporting Commitment foundations have chosen to get ahead of the curve by taking the radical step of making their grants data entirely open. Under the agreement forged among them, they will either submit their data in machine-readable format or have the Foundation Center convert it so that it can be "harvested" by computers and used by developers to create apps, dashboards, visualizations, and things we haven't yet imagined. To make it easier, Glasspockets features a query builder that allows users to construct their own search and then "grab" the resulting data via an API.

A Strategic Idea -- GeoCoding
Some five years ago, when the Foundation Center started visualizing foundation grants data on interactive online maps, the most common reaction was, "You only show the location of the grantee organization, not the geographic focus of the grant." There was a reason for this: the vast majority of foundations, even those that electronically submit their data to the Foundation Center, do not include any coding for geographic area served. And even when there was a clue embedded in the grant description, there was no single standard that foundations used to describe the world; commonly used phrases such as "Deep South," "Middle East," and "developing countries" do not have agreed-upon definitions. That's why so many mapping visualizations (including our own) consign grants with insufficient or no geographic coding to big bubbles floating around in the ocean.

The Reporting Commitment foundations want to be able to compare their grants data with other participating foundations' data, from the community level all the way up to the continental level, to better identify gaps and areas of overlap and be more strategic about their giving. Thus they have agreed to use the GeoTree developed by the Foundation Center as an open geographic standard for use by philanthropy and the social sector. Geographic coding, or geocoding, as it is commonly known, requires a degree of specificity and decision making (i.e., how to handle grants that benefit multiple locations) that is something of a new discipline for most foundations. An interactive mapping tool on Glasspockets allows users to filter and search more than 3,800 grants by city/town, state/province, country, continent, or keyword. As participating foundations geocode more and more of their grants, the volume of data visualized on this map will expand.

A Mission-Critical Idea -- Transparency
When the Foundation Center was created in 1956 as a response to McCarthy-era hearings on philanthropy, transparency meant collecting printed reports from foundations and organizing them in file cabinets for public inspection. Today, it increasingly means open data. For an organization that has built a successful business model that relies on revenue from subscription databases to sustain an enormous volume of free information and services provided to more than nine million users, this may seem like risky business -- and it is. But the future of the Foundation Center requires disrupting its role as a data publisher. In the end, it is the Foundation Center's ability to analyze and combine multiple streams of information and analysis that adds value to data. And it is technology and networks that will allow the center to deliver knowledge into the hands of organizations and individuals who can leverage it to change the world.

Thanks to the vision, leadership, and hard work of the fifteen Reporting Commitment foundations, philanthropy has taken a crucial public step. Other foundations wishing to join the commitment can get started by contacting the Foundation Center. Later this year and again in 2013, the Foundation Center plans to release new and exciting forms of open data. While philanthropy may have been slow to get there, it is finally entering the era of Big Data.
-- Brad Smith



Wednesday, October 10, 2012

Big Data, Bigger Outcomes


Healthcare is embracing the big data movement, hoping to revolutionize HIM by distilling vast collections of data for specific analysis

Lorraine Fernandes, Michele O'Conner and Victoria Weaver, Journal of the American Health Information Management Association, October 2012

One only needs to open a recent conference brochure, read an electronic newsletter, or preview marketing materials to appreciate that “Big Data” is getting a lot of buzz in healthcare—as well as many other sectors of the global economy. Big Data tries to make sense out of information overload, and provides new insights from the growing volumes and sources of data with the goal of answering business, operational, and clinical questions in near-real time. As technology grows, the various types of data available for research grow with it. Big Data solutions aim to harness large and complex collections of digital data and extract focused knowledge and insights from it. In healthcare, experts say Big Data empowers caregivers, scientists, and management to make better decisions that have the potential to save lives, improve efficiencies, and decrease costs. Big Data also has the potential to revolutionize the way health information management (HIM) professionals collect, store, and transmit data.

“Today’s episode-oriented discrete data does not allow us to be as prescriptive as we need to be in delivering better healthcare and empowering consumers,” says Lisa Khorey, vice president of enterprise systems and data management, information technology at the University of Pittsburgh Medical Center. “Medicine can get closer to the action when it is prescriptive, predictive, and precise. Big Data allows organizations to focus on wellness and standardize care processes.”

Big Data Basics
Big Data can be defined by reviewing its basic characteristics, sometimes referred to as the 3 Vs: volume, velocity, and variety.

·       Volume refers to the rapid rate at which data is growing. In 2020 it is estimated there will be 44 times more data than in 2009—35 zettabytes compared to 800,000 petabytes. Big Data techniques and software work to manage large data blocks and make sense of the information.

·       Velocity represents the increasing frequency with which data is delivered. Data such as social media, monitoring and sensing devices, and embedded chips— now in every imaginable device from refrigerators and airplanes to bodily implants—all add to the growing mounds of available data.

·       Variety signifies the many forms in which data exists. In healthcare this includes unstructured data in text format, scanned documents, streams of data from monitoring devices, email or text messages, and audio and video from images and procedures that add to the wide variety of existing structured healthcare data.

The intrigue of Big Data technologies in many industries, including healthcare, is its promise to transform how an industry operates. Scott Schumacher, PhD, IBM chief scientist and distinguished engineer, says these technologies can allow physicians to have predictive analytics that can lead to both long-term and immediate care decisions.

“Technologies aimed at the first V, volume, support the analysis of the large quantity of data required for meaningful statistics and finer grained personalization,” Schumacher says. “The second V, velocity, delivers the transformational promise of Big Data through predictive analytics tied to real-time measurements.
“The third V, variety, leverages natural language processing, semantic normalization using standard ontologies, and image and video extraction to bring more and varied evidence into analytic systems.”

How Big Data Helps Healthcare
Big Data has tremendous potential to add value in all healthcare settings. Big Data solutions can help organizations personalize care, engage patients, reduce variability and costs, and improve quality. Once Big Data is managed and integrated, organizations can apply analytics to better understand the clinical and operational states of their business based on historical and current trends, and predict what might occur in the future with a trusted level of reliability.

Personalization, whether based on genomic data, standard test data, or a combination of the two, requires the integration and analysis of much larger volumes of data than is used today, Khorey says.

“Big Data provides a rich context to shape many areas of healthcare, especially genomics where massive amounts of data are required and costs are rapidly decreasing,” she says.

While these technologies center on vast collections of data, they can also be used for select and specific analysis. For example, Big Data can be used to define patient populations at a level of granularity previously unobtainable, according to Dr. Richard Tayrien, DO, FACOL, chief health information officer for the Hospital Corporation of America. By referencing a patient to a cohort of several million similar patients, aligned by hundreds of clinical features and modeled through numerous therapeutic pathways, Big Data tools can be used to find outcomes that are predicted with a high degree of sensitivity and specificity, Tayrien says.
“Big Data solutions can result in personalized medicine that makes a dramatic difference by redirecting the care of a patient toward the most favorable outcome before predictably sustaining an adverse clinical event,” he says.

Big Data solutions will benefit healthcare providers, payers, research, and government organizations. The following is an overview of what Big Data delivers for each of these sections of the healthcare industry.

Providers Get Patient-Specific Best Practices
Healthcare providers have massive amounts of unstructured data in the form of images, scanned documents, and encounter or progress notes. Big Data solutions enable providers to analyze unstructured data in its native state, integrate it with structured data, and address priorities based on their findings. Priorities may include care pattern identification that aids in process modifications; predictive identification of risk factors to avoid never or sentinel events and untoward outcomes; and comparisons of images, procedures, and surgeries to improve education, research, and care.

Kristen Wilson-Jones, vice president of data and online services for Sutter Health, describes Big Data as a means for provider organizations to apply “mass personalization” principles to healthcare in ways similar to those used in consumer product design and manufacturing.

“Big Data will allow traditional claims and procedure data to be integrated with data created outside of healthcare to break down artificial barriers between healthcare settings,” Wilson-Jones says. “For example, data from grocery store purchases, social media, and personal preferences can be integrated to better understand what impacts individual and population health.”

These new insights can improve health at many levels, Wilson-Jones feels. With Big Data, best practices are more readily identified, variability decreases, and costs and quality are enhanced by providers, delivering a truly personalized patient experience.

Payers Leverage Data Pool
Payers have massive amounts of claims data they would like to harness to provide insights that improve wellness, patient compliance, fraud detection, and enable early warning to negative patient trends. Whether they are private payers or the government, payers increasingly use incentive programs to reward better outcomes while controlling costs. Many also want to utilize social media as a wellness and patient intervention tool that drives lifestyle changes, improves care, and reduces costs. Big Data solutions enable payers to integrate high volumes of different varieties and sources of data to enable these diverse initiatives.

Research Enabled with Unprecedented Reach
Research that requires the integration of large amounts of data has historically been underserved due to computational limitations. With Big Data solutions, researchers can contextually integrate and correlate large amounts of information automatically to gain faster insights.

For example, the State University of New York (SUNY) at Buffalo has deployed a Big Data solution to better understand the complex causes of multiple sclerosis. The system combines and analyzes variables such as diet, exercise, living, and working conditions, as well as clinical and genetic data. This approach used to take days of computing time, but now takes minutes due to the advanced computing power of today’s systems.

“Big Data allows us to take our research to a new level,” says Dr. Murali Ramanathan, PhD, lead researcher at SUNY Buffalo. “We can now rapidly analyze larger data sets including thousands of genetic variations, many environmental factors, and the interaction between them to gain valuable new insights that weren’t possible before.”

Benefit of Using Vast Government Data Stores
Government organizations may be the biggest beneficiary of Big Data solutions. Organizations already have vast stores of data sitting in data warehouse silos. With Big Data solutions these data silos can be quickly integrated to provide valuable insights such as detection of fraud and abuse patterns, identification of best practices for safer and more efficient care delivery, and better epidemiology surveillance.

“Proceeding with the implementation of Big Data healthcare solutions requires organizations to make a cultural commitment to use data to improve quality and reduce waste,” Wilson-Jones comments. “Information must be recognized as the strategic enterprise asset it is, and must be mastered and governed to break down the large number of silos and barriers in today’s healthcare systems.”

Ensuring Success for Big Data Solutions
Big Data solutions can provide significant benefits, but to ensure their successful implementation healthcare organizations need to take the following four steps:

1. Establish data governance, define data objectives
Before organizations implement Big Data solutions, stakeholders should convene an executive council made up of senior leadership to develop an information governance model that clearly defines Big Data objectives and expected outcomes, as well as drives Big Data initiatives.

“Data must be managed and treated as a strategic enterprise asset, and data governance or active management of the data should be vital, especially in light of Big Data,” Wilson-Jones says.


An effective Big Data governance program should include the basic tenets of people, process, and policies. Specific people that should be included are data stewards, who can assist with the interpretation and use of data, and a data governance council that provides representation for key stakeholders across the organization. Special consideration must be given to the new automated processes, inferences, metrics, and monitoring tools provided by Big Data solutions. Policies and procedures will also be required that govern the use of data, define the required actions and quality control processes, and optimize, secure, and leverage information as an enterprise asset by aligning the objectives of multiple functions.

2. Identify data and information requirements
Once an organization has established an information governance model, its next step is to identify where all of the required data resides, what information should be gleaned from data, and how data will be leveraged to help prevent adverse situations, improve care, and keep patients healthy. Most structured healthcare data, estimated to make up 20 percent of all data in a healthcare facility, resides in automated systems such as the hospital information system, the radiology information system, laboratory systems, etc. The remaining 80 percent of healthcare data consists of rich unstructured data that historically has only been leveraged using labor-intensive processes or, more commonly, has not been leveraged at all.

Big Data solutions provide healthcare organizations with the ability to access and analyze unstructured data to assist them in making more informed decisions and reducing errors and missed opportunities. However, unstructured data introduces new challenges for data stewards, specifically verifying that new information is extracted correctly (i.e., proper handling of negations such as “… tests indicate lack of evidence of …”) and that individual patient records are accurate. To properly identify and remediate errors, organizations will need to develop and deploy new data mining tools.

Organizations need to understand what data they will use today, and any potential data that they may want to access in the future. This can include data from mobile or remote devices, implanted devices, text messages and e-mails between patients and providers, and data from third parties or health information exchanges. Organizations will also need to establish a data acquisition roadmap based on business and analysis priorities.

3. Normalize, integrate, and organize Big Data solutions
After all data sources have been identified, a plan needs to be developed for how data will be normalized, integrated with, and organized into the Big Data solution. The plan should address technology requirements as well as business objectives, and must ensure that data are accurate and complete. Big Data solutions present even greater challenges than traditional data and analytic solutions as the volume of data is multiplied many times. The quality of many data sources accessed may have never been evaluated before.

For individually-focused analytics, most Big Data solutions require a complete view of patient and provider data. The ability to recognize relationships between patients and providers, households, payers, and organizations may also be helpful but difficult to achieve given the number of data sources. Any Big Data solution should support the systems, data, and information needs that organizations have today, but also must be configurable and flexible enough to adapt and meet future requirements.

4. Protect security and privacy of Big Data
Data privacy and security must also be a key component of any Big Data solution. All systems, data flows, and information lifecycles must be accounted for and the privacy of personally identifiable information protected. Organizations need to consider what types of information they expect to generate, and whether it will be individually identified or population-based. Data that are used for population-based clinical research to detect diseases or disease patterns usually masks or removes the identities of individuals before the database is populated with the clinical information. But due diligence should be taken to check if the information is de-identified before using the data.

Since healthcare Big Data solutions may use data from many different sources and be predictive and inferential in nature, there may be uncertainty within an organization about how to apply privacy and security mandates like the HIPAA requirements, the Fair Credit Reporting Act, and the Federal Trade Commission’s Fair Information Practice Principles (FIPPs). The best way to address privacy concerns or requirements is for Big Data solutions to support FIPPs. FIPPs are industry-agnostic, basic information privacy principles that can guide the thorny discussions that may be required when analytic projects cross industries, data sources, and data types.

“FIPPs are a roadmap for good data stewardship and the foundation for regulations or policies understood and practiced around the globe,” says Deven McGraw, JD, director of the Health Privacy Project, Center for Democracy and Technology, and member of the Office of the National Coordinator for Health IT’s Health IT Policy Committee. “Since many organizations will deploy healthcare Big Data solutions that use data from outside their walls, they must be able to assure consumers that they have put the appropriate privacy practices in place and that only authorized personnel can access data.”

Big Data’s HIM Opportunities
The move to Big Data solutions provides HIM professionals with significant opportunities for advancement. Those professionals who have an understanding of Big Data and know how to apply HIM principles and data management skills to Big Data implementations will have the most growth opportunities.

Big Data offers HIM professionals the chance to play a strategic role in crafting the next level of healthcare information management, and act as key stakeholders in advancing the strategic use of Big Data across the healthcare ecosystem.

As the industry transforms, it becomes essential for HIM professionals to move beyond the principles of record maintenance and documentation and develop an understanding for data transport, mapping processes, and other Big Data characteristics. Continuing education can help to expand individual knowledge and expertise in health informatics, data management, clinical vocabularies, and data standards—all important aspects of Big Data solution planning. For example, being well-versed in key concepts such as the Systematized Nomenclature of Medicine (SNOMED) classification system and the Logical Observation Identifiers Names and Codes (LOINC) can empower an HIM professional to champion the use of data across systems and facilitate interoperability.

From a broad perspective, HIM professionals should ensure that industry leadership understands the value that HIM brings to Big Data. Not only is it important for HIM professionals to get involved in Big Data planning, but they must come prepared to work with the organizational team and address data and information on a whole new level.

References
Office of Science and Technology Policy, Executive Office of the President of the United States of America. “Obama administration unveils ‘Big Data’ initiative: announces $200 million in new R&D investments.” March 29, 2012. http://www.whitehouse.gov/sites/default/files/microsites/ostp/big_data_press_release_final_2.pdf.
US Department of Health and Human Services National Institutes of Health. “1000 Genomes Project data available on Amazon Cloud.” March 29, 2012. http://www.nih.gov/news/health/mar2012/nhgri-29.htm.
Federal Trade Commission. “Fair Information Practice Principles.” http://www.ftc.gov/reports/privacy3/fairinfo.shtm.
Lorraine Fernandes (lfernand@us.ibm.com) is global healthcare industry ambassador and Michele O’Connor (moconno@us.ibm.com) is global MDM sales at IBM. Victoria Weaver (victoria.weaver@hcahealthcare.com) is assistant vice president, clinical data management at HCA.

Article citation:
Fernandes, Lorraine; O’Connor, Michele; Weaver, Victoria. "Big Data, Bigger Outcomes." Journal of AHIMA 83, no.10 (October 2012): 38-43.      


Brookings: How to Maintain a Competitive Internet


This paper is being released in conjunction with a Center for Technology Innovation event on October 10, 2012.

EXECUTIVE SUMMARY
The Internet has become an essential vehicle for communications, electronic commerce, and entrepreneurship. A McKinsey report found that the Internet provided 21 percent of the GDP growth over the past five years in 13 different countries. As enumerated by the Boston Consulting Group, the Internet currently generates 4.1 percent of Gross Domestic Product; in some countries, the percentage is double that.  By 2016, analysts estimate that the digital economy will comprise $4.2 trillion among G-20 nations, up from $2.3 trillion in 2010.  

The Web offers several features that drive its usefulness for consumers and businesses:  interconnectivity, openness, scalability, and efficiency. Interconnectivity is important because the Internet links users across the globe.  Americans can order goods from shops in Europe or Asia, and vice versa. The openness and growth possibilities allow entrepreneurs to scale up quickly. And since it offers these benefits in a ubiquitous manner, it is a remarkably efficient vehicle for communications and service delivery.

To protect these virtues, a number of academic experts and business leaders have concluded that the government should be cautious about applying competition law to the Internet market.  They argue we should have a “hands-off” competition policy given the rapidly changing nature of digital technology, the complexity of networked industries, the slow pace of government decision-making, the lack of substantive knowledge on the part of regulators, and the globalization of service delivery. 

In this paper, we argue that robust competition policy, including the application of law and enforcement, are vital to ensure the continuing benefits of Internet communications and commerce. Competition is good for consumers, and we need to protect against threats to open competition in Internet markets in order to maintain its beneficial features.  It is important to have antitrust enforcement and fair, transparent, and non-discriminatory market behavior to gain the full benefits of the Internet. We need public policies that promote consumer choice and encourage innovation without stifling competition.
Download » (PDF)

Harnessing Data as a New Source of Growth: Big Data Analytics and Policies


OECD Headquarters, Paris, France, October 22, 2012

Introduction
The confluence of several key socio-economic and technological trends is resulting in the generation of huge streams of data every day. The major trends include:

·       The increasing migration of social and economic activities on line: Social network site Facebook, for example, now counts over 900 million active participants around the world generating together more than 1 500 status updates every second about their interests and whereabouts. In 2011, e-commerce platform eBay collected data on more than 100 million active users including the 6 million new goods they offered every day.

·       The strong decline in the cost of data collection, storage, transportation, and processing: The average cost of consumer hard disk drives (HDDs) per gigabyte, for example, dropped on average by almost 40% per year between 1998 (USD 56 per gigabyte) and 2012 (USD 0.05 per gigabyte). In 1995, as another example, consumers in France paid USD 75 equivalent per month for a dial-up (56 Kb/s) connection, while in 2011 they paid the equivalence of USD 33 per month for a broadband (51 Mb/s) connection, which was almost 1 000 times faster.

·       The increasing deployment of “smart” ICT applications such as smart grids and smart transportations based on machine-to-machine (M2M) communication: Connecting one million homes to a smart grid may produce as much as 11 gigabytes of data per day. In order to accommodate for hourly readings through smart meters, a network with a minimum capacity of up to 1 Mbit/s dedicated to M2M communication is needed.

·       The continued expansion of mobile communication: In 2011, there were 780 million smart phones worldwide capable of collecting and transmitting geo-location data, which generated more than 600 petabytes (millions of gigabytes) of data every month. It is estimated that the global data traffic generated by mobile communication (including M2M enabled smart devices) will almost double every year to reach 11 exabyte (billions of gigabytes) per month by 2016.

The collection and exploitation of these large data flows through data analytics is leading to a shift towards a data-driven socio-economic model commonly referred to as “big data”. In this model, data is a core asset that provides a huge resource for new industries, processes, services and goods leading to significant competitive advantage. In business, for example, data analytics are increasingly being used in a wide number of operations ranging from optimising the value chain and manufacturing production to more efficiently using labour and improving customer relationships. It is estimated that firms, which adopt data-driven decision-making, for example, have output and productivity that is 5-6% higher than what would be expected given their other investments and information technology usage. These firms also perform better in terms of asset utilisation, return on equity and market value.

To unlock the potential of big data and data analytics (big data analytics), OECD countries need to ensure the development of coherent policies and practices around the collection, transportation, storage and use of data, most prominently in areas related to privacy protection. New data sources, new actors and the increasing ease with which personal data can be collected, linked and processed, seriously challenge the effective implementation of current frameworks for privacy protection. But the potential implications for policy also spill over into many other domains; including among others open access to data, intellectual property rights, competition, skills and employment, infrastructure, and measurement.

Objective
First launched in 2005, Technology Foresight Forums are an annual event organised by the OECD Committee for Information, Computer, and Communications Policy (ICCP) to help identify opportunities and challenges for the Internet Economy posed by technical developments. The 2012 Technology Foresight Forum will focus on the potential of big data analytics as a new source of growth, which could help generate significant economic and social benefits. It will put big data analytics in the context of emerging trends discussed at the last three Foresight Forums, namely mobile communications (2011), smart ICTs (2010), and cloud computing (2009), to highlight the confluence of trends leading towards a data-driven economy (see Figure below).
The confluence of major trends enabling data as a new source of growth

The 2012 Foresight Forum will then discuss the social and economic issues related to big data analytics that may put at risk fundamental values and warrant a review of current policy frameworks, most prominently those aimed at ensuring the protection of privacy, intellectual property, and competition. Other policy areas that will be addressed include health care, science and research, government administration and labour markets.

This year’s Foresight Forum will contribute to OECD’s ongoing horizontal project entitled New Sources of Growth: Intangible Assets (NSG) as well as to the follow-up horizontal project on Knowledge-Based Capital: Seizing the Benefits of New Sources of Growth, which will be launched in 2013. Both projects aim at (i) providing structured evidence of the economic value of intangible assets, including data, as a new source of growth, and (ii) improving understanding of current and emerging challenges for policies related to the increasing relevance of these intangible assets.

Modalities
The Foresight Forum represents a collaborative effort of policy makers, business, civil society, and the Internet technical community and thus provides opportunities to liaise with experts in the fields covered. To increase interactivity with a broader public, the Foresight Forum will be supported by various participative web technologies, including a webcast, and Twitter feeds. Further details will be announced closer to the event.

Agenda
The 2012 Foresight Forum will include two morning and three afternoon sessions: The first (morning) session will introduce data analytics in the context of the key technological and socio-economic trends, namely i) cloud computing; ii) smart ICT applications; and iii) the Internet of Things. The following session will then focus on the socio-economic implications of harnessing data as a new source for growth. The afternoon sessions will take a more in-depth look at specific areas that will include: i) science and research (including public health); ii) marketing and competition; and iii) public administration. Each of these sessions will be taken as a mean to discuss specific potential policy opportunities and challenges ranging from i) privacy and consumer protection; ii) intellectual property rights; iii) open access to data; iv) competition; and v) skills and employment. A concluding session will then highlight the main policy implications to be examined by the OECD in the context of its programme of work for 2013-14.

Monday, October 8, 2012

The Benefits of Open Data - Evidence from Economic Research


Guo Xu, Open Economics, October 3, 2012

This contribution is by Guo Xu (OKFN Economics and LSE) and the first part of the blog series “Mainstreaming Open Economics”.

Looking back to the Open Knowledge Festival 2012 in September, there’s an impression that openness is everywhere: There are working groups on Open Science and Open Linguistics, topic streams on Gender and Diversity in Openness, and events like Open Prom and Open Sauna: Open Knowledge and Open Data, it seems, is omnipresent.

Looking beyond the Open Knowledge community, however, the situation is very different: In Economics, for example, not many know what “open data”, “open access” or “Open Economics” exactly mean. Indeed, not many even care. A common reaction is: “Yes, it sounds interesting and important, but does it really matter? And why should I care about it?”

In this post, I would like to give some hard evidence on the positive role of opening up information has had in economics, and sketch ideas for how to involve economists – professional or in training – to mainstream ideas of openness. The blog post is divided into three parts: The first part looks at economic research on open data. The second part looks at the impact of open data on economic research. The third part discusses challenges and ways forward.

The real world impacts of open information
Making information accessible to the public can improve public service delivery. In countries where corruption is pervasive, services and funds often do not reach the frontline provider. And even if services do reach the people, the quality of services provided is often shockingly poor: Survey evidence from Bangladesh, Ecuador, India, Peru and Uganda found absence rates as high as 20% and 35% for school teachers and health workers. In many cases, the staff is poorly trained.

Releasing data on service delivery in this case can help reduce corruption and improve public services. In Uganda, researchers provided information to parents by publishing funding data for a random subset of schools in local newspapers. In consequence, corruption decreased significantly, while schooling outcomes improved substantially. Similar evidence in health delivery and redistributive policies suggest that providing information can help the public to discipline public service providers, improving the quality of services.

Information can also expose corrupt politicians: The Federal Government of Brazil, for example, began to select and audit municipalities at random, releasing audit reports to the media. Researchers found that the audit outcomes had a significant impact on the reelection probability of politicians: Those exposed for corruption were punished at the ballots, and the impact was most pronounced in areas where the dissemination of information was favoured by local radio.

A story from fishermen in South India provides another example of how information can improve market efficiency: Studying the adoption of mobile phones in Kerala, researchers have found convincing evidence that access to information through mobile phones helped fishermen sell their catch at the market where the price was highest (and fish most demanded): Instead of sailing to a port and simply hoping for a good price, fishermen were empowered by technology to make informed decisions on how to trade.

Finally, the benefits of transparency are not only restricted to reducing corruption and lowering the cost of information: A comparative study finds that transparency – measured by accuracy and frequency of macroeconomic information released to the public – leads to lower borrowing costs in sovereign bond markets. Open data pays off in many ways – in many different contexts.

These are just a few selective examples on how cutting-edge economic research has identified the benefits of openness in a diverse range of situations. The cases I presented are not based on correlations, but carefully established causal relationships, leaving – at least within the context studied – little doubt that information matters – big time. Perhaps most importantly, these cases have also shown that open data must be understood in a broad sense: These interventions do not take advantage of linked data, do not use CSVs that are shared through Facebook or Twitter – often, these interventions are simple solutions that ultimately help improving the everyday lives of the people.

Big Data: A Short History


How we arrived at a term to describe the potential and peril of today's data deluge.

Uri Friedman, Foreign Policy, November 2012

Humans have been whining about being bombarded with too much information since the advent of clay tablets. The complaint in Ecclesiastes that "of making many books there is no end" resonated in the Renaissance, when the invention of the printing press flooded Western Europe with what an alarmed Erasmus called "swarms of new books." But the digital revolution -- with its ever-growing horde of sensors, digital devices, corporate databases, and social media sites -- has been a game-changer, with 90 percent of the data in the world today created in the last two years alone. In response, everyone from marketers to policymakers has begun embracing a loosely defined term for today's massive data sets and the challenges they present: Big Data. While today's information deluge has enabled governments to improve security and public services, it has also sowed fears that Big Data is just another euphemism for Big Brother.

1887-1890
American statistician Herman Hollerith invents an electric machine that reads holes punched into paper cards to tabulate 1890 census data, revolutionizing the concept of a national head count, which had originated with the Babylonians in 3800 B.C. The device, which enables the United States to complete its census in one year instead of eight, spreads globally as the age of modern data processing begins.


1935-1937
President Franklin D. Roosevelt's Social Security Act launches the U.S. government on its most ambitious data-gathering project ever, as IBM wins a government contract to keep employment records on 26 million working Americans and 3 million employers. "Imagine the vast army of clerks which will be necessary to keep these records," Republican presidential candidate Alf Landon scoffs. "Another army of field investigators will be necessary to check up on the people whose records are not clear."


1943
At Bletchley Park, a British facility dedicated to breaking Nazi codes during World War II, engineers develop a series of groundbreaking mass data-processing machines, culminating in the first programmable electronic computer. The device, named "Colossus," searches for patterns in intercepted messages by reading paper tape at 5,000 characters per second -- reducing a process that had previously taken weeks to a matter of hours. Deciphered information on German troop formations later helps the Allies during their D-Day invasion.


1961
The U.S. National Security Agency (NSA), a nine-year-old intelligence agency with more than 12,000 cryptologists, confronts information overload during the espionage-saturated Cold War, as it begins collecting and processing signals intelligence automatically with computers while struggling to digitize a backlog of records stored on analog magnetic tape in warehouses. (In July 1961 alone, the agency receives 17,000 reels of tape.)


1965-1966
The U.S. government secretly studies a plan to transfer all government records -- including 742 million tax returns and 175 million sets of fingerprints -- to magnetic computer tape at a single national data center, though the plan is later scrapped amid public concern about bringing "Orwell's '1984' at least as close as 1970," as one report puts it. The outcry inspires the 1974 Privacy Act, which places limits on federal agencies' sharing of personal information.


1989
British computer scientist Tim Berners-Lee proposes leveraging the Internet, pioneered by the U.S. government in the 1960s, to share information globally through a "hypertext" system called the World Wide Web. "The information contained would grow past a critical threshold," he writes, "so that the usefulness [of] the scheme would in turn encourage its increased use."


August 1996
"We are developing a supercomputer that will do more calculating in a second than a person with a hand-held calculator can do in 30,000 years." --U.S. President Bill Clinton


1997
NASA researchers Michael Cox and David Ellsworth use the term "big data" for the first time to describe a familiar challenge in the 1990s: supercomputers generating massive amounts of information -- in Cox and Ellsworth's case, simulations of airflow around aircraft -- that cannot be processed and visualized. "[D]ata sets are generally quite large, taxing the capacities of main memory, local disk, and even remote disk," they write. "We call this the problem of big data."


2002
After the 9/11 attacks, the U.S. government, which has already dabbled in mining large volumes of data to thwart terrorism, escalates these efforts. Former national security advisor John Poindexter leads a Defense Department effort to fuse existing government data sets into a "grand database" that sifts through communications, criminal, educational, financial, medical, and travel records to identify suspicious individuals. Congress shutters the program a year later due to civil liberties concerns, though components of the initiative are simply shifted to other agencies.


2004
The 9/11 Commission calls for unifying counterterrorism agencies "in a network-based information sharing system" that is quickly inundated with data. By 2010, the NSA's 30,000 employees will be intercepting and storing 1.7 billion emails, phone calls, and other communications daily. Meanwhile, with retailers amassing information on customers' shopping and personal habits, Wal-Mart boasts a cache of 460 terabytes -- more than double the amount of data on the Internet at the time.


2007-2008
As social networks proliferate, technology bloggers and professionals breathe new life into the "big data" concept. "This is a world where massive amounts of data and applied mathematics replace every other tool that might be brought to bear," Wired's Chris Anderson writes in "The End of Theory." Government agencies, some of the United States' top computer scientists report, "should be deeply involved in the development and deployment of big-data computing, since it will be of direct benefit to many of their missions."


January 2009
The Indian government establishes the Unique Identification Authority of India to fingerprint, photograph, and take an iris scan of all 1.2 billion people in the country and assign each person a 12-digit ID number, funneling the data into the world's largest biometric database. Officials say it will improve the delivery of government services and reduce corruption, but critics worry about the government profiling individuals and sharing intimate details about their personal lives.


May 2009
U.S. President Barack Obama's administration launches data.gov as part of its Open Government Initiative. The website's more than 445,000 data sets go on to fuel websites and smartphone apps that track everything from flights to product recalls to location-specific unemployment, inspiring governments from Kenya to Britain to launch similar initiatives.


July 2009
Reacting to the global financial crisis, U.N. Secretary-General Ban Ki-moon pledges to create an alert system that captures "real-time data on the impact of the economic crisis on the poorest nations." The U.N. Global Pulse program has conducted research on how to predict everything from spiraling prices to disease outbreaks by analyzing data from sources such as mobile phones and social networks.


August 2010
"There were 5 exabytes of information created by the entire world between the dawn of civilization and 2003. Now that same amount is created every two days." --Google CEO Eric Schmidt


February 2011
Scanning 200 million pages of information, or 4 terabytes of disk storage, in a matter of seconds, IBM's Watson computer system defeats two human challengers in the quiz show Jeopardy!. The New York Times later dubs this moment a "triumph of Big Data computing."


March 2012
The Obama administration announces a $200 million Big Data Research and Development Initiative in response to a U.S. government report calling for every federal agency to have a "'big data' strategy." The National Institutes of Health puts a data set of the Human Genome Project in Amazon's computer cloud, while the Defense Department pledges to develop "autonomous" defense systems that can "learn from experience." CIA Director David Petraeus, marveling that the "'digital dust' to which we have access is being delivered by the equivalent of dump trucks," discusses a post-Arab Spring agency effort to collect and analyze global social media feeds through cloud computing.


July 2012
U.S. Secretary of State Hillary Clinton announces a public-private partnership called "Data 2X" to collect statistics on women and girls' economic, political, and social status around the world. "Data not only measures progress -- it inspires it," she explains. "Once you start measuring problems, people are more inclined to take action to fix them because nobody wants to end up at the bottom of a list of rankings." Let the Big Data race begin.

Saturday, October 6, 2012

Predicting the Tuture through Online Data Mining


Santiago Zabala, Al Jazeera, October 5, 2012

It is often said philosophers are either late when it comes to comment upon new technological innovations or in advance, that is, so early that they actually seem to predict them. When they are late, it's usually because they prefer to carefully examine the new innovations in order to achieve insightful analysis, and when they foresee such discoveries, it arises from an ethical concern over the direction the world is taking.

In other words, their insight is not expressed by envisioning the day a software company manages to predict the future by scanning information from the internet and therefore framing our freedom, but rather by working through the existential consequences these innovations might have upon our life. 

This is probably why among Martin Heidegger's greatest concerns when it came to technological innovations was the formation of existential conditions where, as he said, the "lack of emergency is the only emergency". In this condition human beings would be completely "uprooted" from the earth, that is, "framed" ("Ge-stell") by a technological power they are no longer able to control.

As it turns out, the software company Recorded Future (which has recently been praised by Wired , the MIT Technology Review and other media outlets, after the CIA and Google invested millions in their services) seems to be offering its clients something similar: a world where emergencies, that is, future events, can be calculated in advance. But how does this start-up actually function, and why are the German philosopher’s concerns relevant to its services? 

Recorded Future
Recorded Future is based in Gothenburg and has offices in London, Boston, Arlington and New York. A team of 20 computer scientists, statisticians and experts in linguistics "calculate" the future. While Yahoo, Google and Bing use links to connect and rank different web pages, Recorded Future goes further by scouring (in real time) thousands of available information sources such as blogs, websites and Twitter comments in order to find "invisible links", that is, relationships among actions, people and institutions that refer to related events in the future. 

Even though this might not seem particularly relevant, considering that we can also predict next week's weather by searching through different weather stations, if we look at the amount of data this company is capable of analysing and relating in just a few hours, it becomes clear it can obtain better data than public internet users have access to.

The information we need to predict whether tomorrow it will rain is limited by the number of weather sites available, but sites that might refer to upcoming anti-American demonstrations in the Middle East are infinitely more numerous given the political, economic and military aspects of these sorts of events.

After mining from the web all the related people ("Bashar al-Assad"), places ("Syria") and activities ("military interventions") that refer to a possible demonstration, Recorded Future uses algorithms to predict when and where a demonstration will occur. 

An example of a predicted demonstration is available in a video on the company website which illustrates how its powerful engines monitor these protests not only in the Middle East, but also in South America and North Africa. The fact that Google and the CIA have already invested millions in this company is an indication that it will be used to conserve certain interest against others as the example above indicates.
The different fee levels for the customers of Recorded Future are probably related to the quality and quantity of information they wish to purchase, making this, and similar companies, at the service of the wealthiest and most powerful. 

Lack of emergencies
From a philosophical point of view, the most interesting feature of all this is not that these demonstrations can be predicted, but rather how technology has finally uprooted and dislodged man from the world, that is, has given human existence to a power beyond human control. The secured, comfortable and calculated environment that allowed the creation of society has now become so functional and rationalised that we cannot help but become victims by existing in it.

This existential dilemma does not arise from the fact that it’s finally possible to organise all the things the web already knows about the future, which could certainly become useful to prevent diseases or famine, but rather that a private company now means to know everything, that is, all human projects. 

We have entered an age where only those framed within the approved interests of Recorded Future clients will be able to live freely, that is, without being predicted. But how free is an existence that is completely revealed to the modern "lack of emergencies"?

As Heidegger explained , emergencies do not arise when something doesn't function correctly, but rather when "everything functions … and propels everything more and more toward further functioning". It's within this logic that as soon as something critical to the interests of those who can afford it fails to function, Recorded Future will alert its customers, who will then take the appropriate measures to conserve the previous condition.

In sum, Heidegger's concerns over a world lacking "emergencies" more than 50 years ago was meant to point out how technologies such as that employed by Recorded Future (and similar companies ) aim to avoid the future, that is, to change the world.

Santiago Zabala is ICREA Research Professor of Philosophy at the University of Barcelona. His books include The Hermeneutic Nature of Analytic Philosophy (2008), The Remains of Being (2009) and most recently, Hermeneutic Communism (2011, co-authored with G Vattimo), all published by Columbia University Press. 

Wednesday, October 3, 2012

TechAmerica Big Data Commission Report:


The Big Data Commission released its findings October 3rd to help answer these questions and many more. The report, “Demystifying Big Data: A Practical Guide to Transforming the Business of Government,” provides the government’s senior policy and decision makers with a comprehensive roadmap to using Big Data to better serve the American people.

More info:

Data in the world is doubling every 18 months. Across government everyone is talking about the concept of Big Data, and how this new technology will transform the way Washington does business. Looking past the excitement, many questions remained unanswered- until now. The TechAmerica Foundation’s Big Data Commission has worked diligently to put together the most comprehensive report on Big Data of its kind. What is Big Data, really? How is it defined? What capabilities are required to succeed? How do you use Big Data to make intelligent decisions? How will agencies effectively govern and secure huge volumes of information, while protecting privacy and civil liberties? And perhaps most importantly, what value will it really deliver to the US Government and the citizenry we serve?

The Big Data Commission released its findings October 3rd to help answer these questions and many more. The report, “Demystifying Big Data: A Practical Guide to Transforming the Business of Government,” provides the government’s senior policy and decision makers with a comprehensive roadmap to using Big Data to better serve the American people.

For more information about the Big Data Commission or other related activities, please contact Chris Wilson, chris.wilson@techamericafoundation.org or (202) 682-4451.

Commission Leadership
                          
The Big Data Commission is chaired by Steven A. Mills, Senior Vice President and Group Executive, Software & Systems, IBM Corporation and Steve Lucas, Global Executive Vice President and General Manager, SAP Database & Technology. Serving as vice chairs of the commission are Teresa Carlson, Vice president Global Public Sector, Amazon Web Services and Bill Perlowitz, Chief Technology Officer, Science, Technology and Engineering Group, Wyle.


Serving as the Academic Co-Chairs will be Dr. Michael Rappa, founding director of the Institute for Advanced Analytics at North Carolina State University and principal architect of its Master of Science in Analytics degree; and Dr. Leo Irakliotis, Dean and National Director for the College of Information Technology at Western Governors University (WGU).

Commission Members
The commission membership is comprised of leading experts on big data and representing both industry and academia, along with a government advisory board with the objective of providing guidance on how Government Agencies should be leveraging Big Data to address their most critical business imperatives, and how Big Data can drive U.S. innovation and competitiveness. A full list of the commission membership can be found here.

Case Studies
·       AM Biotech
·       IRS Compliance Data Warehouse
·       NARA ERA
·       NASA Human Spaceflight Imagery
·       NOAA NWS
·       TerraEchos
·       University of Ontario
·       Vesta Wind Energy