Showing posts with label Information sharing. Show all posts
Showing posts with label Information sharing. Show all posts

Thursday, January 17, 2013

Open Access: Aaron Swartz's illusion over research


John Gapper, The Financial Times, January 16, 2013

Deleting private companies from the equation might allow savings but could reduce efficiency
The death of the internet activist Aaron Swartz at the age of 26 has rightly evoked tributes to his creativity and selflessness. Swartz, who faced jail for illegally downloading millions of academic papers from an electronic library, committed suicide last week.

Five years ago, Swartz signed a “guerrilla open access manifesto” in which he complained of “the world’s entire scientific and cultural heritage” being “digitised and locked up by a handful of private corporations” such as Reed Elsevier. He advised computer hackers to “take information, wherever it is stored, make our copies and share them with the world”.

In 2010, he disguised his identity and exploited the electronic network of the Massachusetts Institute of Technology to download most of the database of Jstor, a non-profit group that digitises academic journals and articles. He did not share or sell the material – he later handed it back – but prosecutors took the manifesto seriously and charged him with fraud.

Mr Swartz worked on projects from the news aggregator Reddit to the Creative Commons open copyright licence, and was widely liked and admired. But, in his analysis of academic research and publishing, he suffered from an illusion.

Free access to academic research – the system Mr Swartz advocated – could bring public benefits. It would enable anyone to read, analyse and build upon privately and publicly funded research. However, someone would still need to pay for it and the costs to universities such as MIT and Oxford would rise, not fall.

Critics of the current system, under which research libraries pay up to $50,000 annually to use online databases, tend to blame profiteering by companies such as Reed Elsevier and Springer for this cost. George Monbiot, the activist and Guardian writer, describes it as “pure rentier capitalism”, arguing that people should “throw off these parasitic overlords and liberate the research that belongs to us”.

Allied to this is the belief that publishing costs have fallen heavily in the shift from print to digital. Elsevier, the scientific publishing arm of Reed Elsevier, made profits of £352m on revenue of £978m in the first half of 2012 – an operating margin of 36 per cent. Remove the capitalists and distribute research through public utilities, and surely swaths of cost would disappear?

Well, perhaps. Elsevier could certainly do with a bit more competition. Its fee structure is opaque and it publishes journals in which academics vie to be published. It has what Warren Buffett calls a moat – it is a 130-year-old business with 20 per cent of the market that is hard to attack.

It did not, however, steal this advantage. It acquired it from the 1960s and 1970s onwards as research universities saved money by outsourcing their costly and subscale publishing presses. Elsevier employs 7,000 editors, manages a network of some 500,000 peer reviewers (whom it does not pay), publishes 300,000 new articles a year and runs a 100-terabyte database.

Printing is only a small part of the cost of academic publishing. The bulk lies in the labour-intensive business of editing and reviewing submissions (rejecting two-thirds of them) and managing data. These costs are similar for open access publishers such as the Public Library of Science (Plos) in San Francisco, a competitor to Elsevier.

An independent study by the Research Information Network in the UK found that the shift from print to digital may save £1bn globally – worth having but only 12 per cent of total costs. Removing private companies from the equation might allow further savings, but it might equally reduce efficiency.

In any case, there will still be a hefty bill. About 90 per cent of the industry operates on subscription – the model Swartz so hated. The other 10 per cent is now open access, under which researchers (or research funders) have to pay journals between $1,000 and $5,000 an article to cover publishing costs. Anyone can then read it free.

Open access is appealing and is supported both by research funds, such as the US National Institutes of Health and the UK Wellcome Trust, and by the UK government. The trust believes it makes no sense to invest £700m each year on research without paying an extra £10m to make it widely available.

Research is largely read by other academics at the moment, most of whom have access through libraries. But there could be big benefits to broadening reach – Plos One, the science journal, is a trove of fascinating material.

That said, open access mostly transfers the bill. The Research Information Network estimated that, if the market moves to 90 per cent open access, total costs would fall by £560m but universities would pay more. The UK would save £128m in library subscriptions but contribute £213m in fees because its universities publish a lot of research.

Open access also has its pitfalls. In the 1970s the credit rating industry turned from investors subscribing to ratings to bond issuers paying. That established open access but also gave agencies a motive to please issuers with good ratings, culminating in the triple A rating of flimsy mortgage-backed securities.

Open access journals have a similar incentive to widen access and dilute quality. It is worth noting that Plos One publishes 24,000 pieces of research every year – it accepts any submission that meets the hurdle of “valid science” – while the most prestigious journals (including other Plos titles) publish 200.

If Swartz’s sad death shifts the balance further toward open access, that will be a worthy legacy. But someone will always pay.

john.gapper@ft.com



Tuesday, October 23, 2012

Building A Culture Around Big Data


Deanna Glick, AOL Government, October 16, 2012

A report released today by the Partnership for Public Service aims to educate federal managers on how agencies can do just that. The report, From Data to Decisions II: Building an Analytics Culture, examines how to best use data – not anecdotes – to base decisions.

Building on an original report released last November that examined how several federal agencies use data, the new report identifies strategies for how to develop and grow an analytics culture within agencies and incorporate it into how federal workers perform the mission. It profiles seven agencies using analytics to achieve better results and the strategies used in a budget-cutting climate.

Both reports were joint efforts between the partnership and the IBM Center for The Business of Government.

"By sharing compelling stories of how agencies are developing, growing and sustaining their analytics and performance-management approaches, we hope to shed light on key steps and processes that are transferable to other agencies," the report states.

To complete the report, the organizations studies how agencies are using analytics; how they got started; what conditions helped to grow their approaches; what challenges arose and why; and what success looks like.

"We found many parallels in approach across agencies and programs," according to the report. "Driven by budget realities and the push for more data-driven actions, agency managers were examining their programs in a disciplined, comprehensive way to determine how they conduct their business."

The report features details of analytics efforts at agencies within the departments of Homeland Security, Health and Human Services, Interior, Defense and Treasury.

A common successful first step in creating a culture around analytics, researchers found, was agencies tying specific activities directly to what they are intended to achieve and linking them to goals. Focusing on these details help agencies employ a data-driven approach to managing programs, identify critical information to gauge progress and results, and ensure that only those activities that are key or essential to meeting desired results are performed.

To improve airport security, for example, a federal security director with the Transportation Security Administration worked with a team to break down the job of a transportation security officer at checkpoint and baggage areas. After analyzing and brainstorming around specific tasks related to the job, his team identified more than 1,300 knowledge areas, values and skills for a transportation security officer. Based on this analysis, they identified vulnerabilities in security screening and uncovered weaknesses in training, procedures or technology. They then pinpointed what could be improved through training and better application of procedures or policy and where technology could support improved performance.

"By instituting these types of systematic processes, agencies start building analytic cultures so they can look critically at what they do and thoroughly understand how their activities can lead to better results," the report states. "The reward for their meticulous appraisal is the enhanced ability to serve the American public cost-effectively and efficiently."

For more news and insights on innovations at work in government, please sign up for the AOL Gov newsletter. For the quickest updates, like us on Facebook.

Thursday, October 18, 2012

Data from Health Care Reviews Could Power "Yelp for Health Care" Startups


Data-driven decision engines will need patient experience to complete the feedback loop.

Alex Howard, O'Reilly Radar, October 17, 2012

Given where my work and health has taken me this year, I’ve been thinking much more about the relationship of the Internet and health data to accountability and patient-driven health care.

When I was looking for a place in Maine to go for care this summer, I went online to look at my options. I consulted hospital data from the government at HospitalCompare.HHS.gov and patient feedback data on Yelp, and then made a decision based upon proximity and those ratings. If I had been closer to where I live in Washington D.C., I would also have consulted friends, peers or neighbors for their recommendations of local medical establishments.

My brush with needing to find health care when I was far from home reminded me of the prism that collective intelligence can now provide for the treatment choices we make, if we have access to the Internet.

Patients today are sharing more of their health data and experiences online voluntarily, which in turn means that the Internet is shaping health care. There’s a growing phenomenon of “e-patients” and caregivers going online to find communities and information about illness and disability.

Aided by search engines and social media, newly empowered patients are discussing health conditions with others suffering from disease and sickness — and they’re taking that peer-to-peer health care knowledge into their doctors’ offices with them, frequently on mobile devices. E-patients are sharing their health data of their own volition because they have a serious health condition, want to get healthy, and are willing.

From the perspective of practicing physicians and hospitals, the trend of patients contributing to and consulting on online forums adds the potential for errors, fraud, or misunderstanding. And yet, I don’t think there’s any going back from a networked future of peer-to-peer health care, anymore than we can turn back the dial on networked politics or disaster response.

What’s needed in all three of these areas is better data that informs better data-driven decisions. Some of that data will come from industry, some from government, and some from citizens.

This fall, the Obama administration proposed a system for patients to report medical mistakes. The system would create a new “consumer reporting system for patient safety” that would enable patients to tell the federal government about unsafe practices or errors. This kind of review data, if validated by government, could be baked into the next generation of consumer “choice engines,” adding another layer for people, like me, searching for care online.

There are precedents for the collection and publishing of consumer data, including the Consumer Product Safety Commission’s public complaint database at SaferProducts.gov and the Consumer Financial Protection Bureau’s complaint database. Each met with initial resistance by industry but have successfully gone online without massive abuse or misuse, at least to date.

It will be interesting to see how medical associations, hospitals and doctors react. Given that such data could amount to government collecting data relevant to thousands of “Yelps for health care,” there’s both potential and reason for caution. Health care is a bit different than product safety or consumer finance, particularly with respect to how a patient experiences or understands his or her treatment or outcomes for a given injury or illness. For those that support or oppose this approach, there is an opportunity for public comment on proposed data collection at the Federal Register.

The power of performance data
Combining patients review data with government-collected performance data could be quite powerful in helping to drive better decisions and adding more transparency to health care.
In the United Kingdom, officials are keen to find the right balance between open data, transparency and prosperity.

“David Cameron, the Prime Minister, has made open data a top priority because of the evidence that this public asset can transform outcomes and effectiveness, as well as accountability,” said Tim Kelsey, in an interview this year. He used to head up the United Kingdom’s transparency and open data efforts and now works at its National Health Service.

“There is a good evidence base to support this,” said Kelsey. “Probably the most famous example is how, in cardiac surgery, surgeons on both sides of the Atlantic have reduced the number of patient deaths through comparative analysis of their outcomes.”

More data collected by patients, advocates, governments and industry could help to shed light on the performance of more physicians and clinics engaged in other expensive and lifesaving surgeries and associated outcomes.

Should that be extrapolated across the medical industry, it’s a safe bet that some medical practices or physicians will use whatever tools or legislative influence they have to fight or discredit websites, services or data that puts them in a poor light. This might parallel the reception that BrightScope’s profiles of financial advisors have received in industry.

When I talked recently with Dr. Atul Gawande about health data and care givers, he said more transparency in these areas is crucial:

“As long as we are not willing to open up data to let people see what the results are, we will never actually learn. The experience of what happens in fields where the data is open is that it’s the practitioners themselves that use it.”

In that context, health data will be the backbone of the disruption in health care ahead. Part of that change will necessarily have to come from health care entrepreneurs and watchdogs connecting code to research. In the future, a move to open science and perhaps establish a health data commons could accelerate that change.

The ability of caregivers and patients alike to make better data-driven decisions is limited by access to data. To make a difference, that data will also need to be meaningful to both the patient and the clinician, said Dr. Gawande. He continued:

“[Health data] needs to be able to connect the abstract world of data to the physical world of what really happens, which means it has to be timely data. A six-month turnaround on data is not great. Part of what has made Wal-Mart powerful, for example, is they took retail operations from checking their inventory once a month to checking it once a week and then once a day and then in real-time, knowing exactly what’s on the shelves and what’s not. That equivalent is what we’ll have to arrive at if we’re to make our systems work. Timeliness, I think, is one of the under-recognized but fundamentally powerful aspects because we sometimes over prioritize the comprehensiveness of data and then it’s a year old, which doesn’t make it all that useful. Having data that tells you something that happened this week, that’s transformative.”

Health data, in other words, will need to be open, interoperable, timely, higher quality, baked into the services that people use, and put at the fingertips of caregivers, as US CTO Todd Park explains in the video below:

There is more that needs to be done than simply putting “how to live better” information online or into an app. To borrow a phrase from Robert Kirkpatrick, for data to change health care, we’ll need to apply the wisdom of the crowds, the power of algorithms and the intuition of experts to find meaning in health data and help patients and caregivers alike make better decisions.

That isn’t to say that health data, once published, can’t be removed or filtered. Witness the furor over the removal of a malpractice database from the Internet last year, along with its restoration.

But as more data about doctors, services, drugs, hospitals and insurance companies goes online, the ability of those institutions to control public perception of the institutions will shift, just as it has with government and media. Given
flaws in devices or poor outcomes, patients deserve such access, accountability and insight.


Enabling better health-data-driven decisions to happen across the world will be far from easy. It is, however, a future worth building toward.

Monday, October 15, 2012

McKinsey Anthology: Government Designed for New Times


To explore the approaches that governments around the world are taking to common problems, this anthology convenes political leaders and civil servants, economists and policy experts, generalists and specialists.

McKinsey Anthology, 2012
               
Contents
        Transforming government
·       Tony Blair—Leading transformation in the 21st century
·       François-Daniel Migeon—Interview: Transforming government in France
·       Frank-Jürgen Weise—Behind the German jobs miracle
·       Tim Brown—Quick take: Designing a tech-enabled government
·       Michael Fullan—Transforming schools an entire system at a time
·       Todd Park—Interview: Unleashing government's 'innovation mojo'
·       Diana Farrell—Government designed for new times
        Innovating government services

·       James Fishkin—What the people think when they're really thinking
·       Nandan Nilekani—Interview: For every citizen, an identity
·       Matthew Taylor—Citizens: The untapped resource
·       Susan Zielinski—The new mobility
·       Peter Shergold—A social contract for government
·       Wim Elfrink—The smart-city solution
·       Karan Bhatia—Quick take: Building the world's infrastructure
·       Salman Khan—Teaching for the new millennium
·       Xie Chengxiang—Interview: Home for the urban poor
·       Elena Berkowitz and Blaise Warren—Quick take: How Estonia became E-stonia    

        Building new competencies
·       Douglas Holtz-Eakin—Fiscal management fix: Simple math—and a very big stick
·       Coen Teulings—Why politicans prefer austerity to long-term fiscal reform
·       Lu Mai—The urbanization solution
·       Göran Persson—How to tame a budget crisis
·       Peter Ho—Coping with complexity
·       Mohamed Ibrahim—Better data, better policy making    
        Understanding government in new times

·       Daron Acemoglu—The servant state
·       Parag Khanna—The rise of hybrid governance
·       Neil deGrasse Tyson—Why exploration matters—and why the government should pay  
        for it
·       Ray O. Johnson—Quick take: The research imperative
·       Hernando de Soto—Interview: Building a nation of owners
·       Nicolas Berggruen and Nathan Gardels—A middle way for governance       

Thursday, October 11, 2012

Walking the Talk: Philanthropy 'Does' Big Data


Bradford K. Smith, PhilanTopic, October 9, 2012

(Bradford K. Smith is president of the Foundation Center. In his last post, he took a closer look at the China Foundation Center's new Foundation Transparency Index.

With the modestly labeled "Reporting Commitment," fifteen of America's largest foundations are transforming the practice of philanthropy. From today on, information about their grants will be made available on a near-real-time basis, as entirely open data and coded to a common geographical standard, making it easy to see the communities, regions, and countries that benefit from those grants. The initiative's simple name should not deceive: this is big. The participants -- Annenberg, Carnegie, Gates, Getty, Hewlett, Packard, MacArthur, Mott, Robert Wood Johnson, and six others -- provide nearly 12 percent of the $46 billion in grants made by American foundations each year. To see the Reporting Commitment in action, take a quick look at Glasspockets, the transparency Web site of the Foundation Center, then read on.


What makes the Reporting Commitment so transformative? Let's break it down.

A Bold Idea -- Real-Time Reporting
The fifteen participating foundations have committed to electronically report their current grants data to the Foundation Center on at least a quarterly basis. As pragmatic as this may sound, it's a dramatic departure from the norm for the field. All the 76,000 private foundations in America file 990-PF tax returns in which they provide information on their grants. They have up to a year after the close of their fiscal year to file these returns, the Foundation Center eventually gets them from the IRS as image files and converts them into a more usable format, cleans and codes the data, and insures public access through databases and research reports. In a world where value is being created exponentially by analyzing enormous real-time data sets generated through search logs, consumer purchases, and Facebook "likes," philanthropy remains an industry with $640 billion in assets that relies on two-year old data to understand its own grant trends.

The Foundation Center has convinced more than seven hundred foundations to electronically submit their grants information through its eGrant Reporting Program, covering more than 20 percent of total foundation giving. Although this provides the field with current-year grants data, most participating foundations report on an annual basis. The Reporting Commitment takes this effort one important step forward by having participating foundations report at least quarterly -- with some reporting weekly, even daily.

A Radical Idea -- Open Data
By and large, foundations tend to think of open data and transparency as something they should fund rather than do. There are lots of reasons for this, including the private nature of foundations, the cultural legacy of keeping a low profile and "letting our good works speak for themselves," and sensitivity surrounding some of the issues addressed by foundation grants. Notwithstanding, the ability of foundations to not call attention to themselves is being steadily eroded by the ease of finding, displaying, and circulating information in a densely networked, digital age. Meanwhile, sectors and institutions with which foundations increasingly collaborate, such as the World Bank and foreign aid donors, are barreling ahead with initiatives like the Open Aid Partnership and Publish What You Fund.

The fifteen Reporting Commitment foundations have chosen to get ahead of the curve by taking the radical step of making their grants data entirely open. Under the agreement forged among them, they will either submit their data in machine-readable format or have the Foundation Center convert it so that it can be "harvested" by computers and used by developers to create apps, dashboards, visualizations, and things we haven't yet imagined. To make it easier, Glasspockets features a query builder that allows users to construct their own search and then "grab" the resulting data via an API.

A Strategic Idea -- GeoCoding
Some five years ago, when the Foundation Center started visualizing foundation grants data on interactive online maps, the most common reaction was, "You only show the location of the grantee organization, not the geographic focus of the grant." There was a reason for this: the vast majority of foundations, even those that electronically submit their data to the Foundation Center, do not include any coding for geographic area served. And even when there was a clue embedded in the grant description, there was no single standard that foundations used to describe the world; commonly used phrases such as "Deep South," "Middle East," and "developing countries" do not have agreed-upon definitions. That's why so many mapping visualizations (including our own) consign grants with insufficient or no geographic coding to big bubbles floating around in the ocean.

The Reporting Commitment foundations want to be able to compare their grants data with other participating foundations' data, from the community level all the way up to the continental level, to better identify gaps and areas of overlap and be more strategic about their giving. Thus they have agreed to use the GeoTree developed by the Foundation Center as an open geographic standard for use by philanthropy and the social sector. Geographic coding, or geocoding, as it is commonly known, requires a degree of specificity and decision making (i.e., how to handle grants that benefit multiple locations) that is something of a new discipline for most foundations. An interactive mapping tool on Glasspockets allows users to filter and search more than 3,800 grants by city/town, state/province, country, continent, or keyword. As participating foundations geocode more and more of their grants, the volume of data visualized on this map will expand.

A Mission-Critical Idea -- Transparency
When the Foundation Center was created in 1956 as a response to McCarthy-era hearings on philanthropy, transparency meant collecting printed reports from foundations and organizing them in file cabinets for public inspection. Today, it increasingly means open data. For an organization that has built a successful business model that relies on revenue from subscription databases to sustain an enormous volume of free information and services provided to more than nine million users, this may seem like risky business -- and it is. But the future of the Foundation Center requires disrupting its role as a data publisher. In the end, it is the Foundation Center's ability to analyze and combine multiple streams of information and analysis that adds value to data. And it is technology and networks that will allow the center to deliver knowledge into the hands of organizations and individuals who can leverage it to change the world.

Thanks to the vision, leadership, and hard work of the fifteen Reporting Commitment foundations, philanthropy has taken a crucial public step. Other foundations wishing to join the commitment can get started by contacting the Foundation Center. Later this year and again in 2013, the Foundation Center plans to release new and exciting forms of open data. While philanthropy may have been slow to get there, it is finally entering the era of Big Data.
-- Brad Smith



Friday, August 24, 2012

Don't Build a Database of Ruin


Paul Ohm, Harvard Business Review, August 23, 2012

Many businesses today find themselves locked in an arms race with competitors to see who can convert customer secrets into the most pennies. To try to win, they are building perfect digital dossiers, to use a phrase coined by Daniel Solove, massive data stores containing hundreds, if not thousands or tens of thousands, of facts about every member of our society. 
In my work, I've argued that these databases will grow to connect every individual to at least one closely guarded secret. This might be a secret about a medical condition, family history, or personal preference. It is a secret that, if revealed, would cause more than embarrassment or shame; it would lead to serious, concrete, devastating harm. And these companies are combining their data stores, which will give rise to a single, massive database. I call this the Database of Ruin. Once we have created this database, it is unlikely we will ever be able to tear it apart.

I have become convinced that my earlier, bleak predictions about the Database of Ruin were in fact understated, arriving before it was clear how Big Data would accelerate the problem. Consider the most famous recent example of big data's utility in invading personal privacy: Target's analytics team can determine which shoppers are pregnant, and even predict their delivery dates, by detecting subtle shifts in purchasing habits. This is only one of countless similarly invasive Big Data efforts being pursued. In the absence of intervention, soon companies will know things about us that we do not even know about ourselves. This is the exciting possibility of Big Data, but for privacy, it is a recipe for disaster.

If we stick to our current path, the Database of Ruin will become an inevitable fixture of our future landscape, one that will be littered with lives ruined by the exploitation of data assembled for profit. But we can chart a different course, in various ways. I think our brightest engineers can develop innovative privacy-enhancing technologies which will enable new techniques for data analytics that minimize costs to privacy. I hope that public institutions and industry, through self-regulation, will devise ways to better balance the burdens on privacy and the benefits of Big Data. If nothing else, I anticipate that society will slowly develop new norms for engaging with the massive amount of information collected about us, creating informal rules governing when and how it is appropriate to release, collect, and use data, the way minors have learned to speak and listen carefully on social networks.

But every one of these correctives requires the same thing: time. We need to slow things down, to give our institutions, individuals, and processes the time they need to find new and better solutions. The only way we will buy this time is if companies learn to say, "no" to some of the privacy-invading innovations they're pursuing. Executives should require those who work for them to justify new invasions of privacy against a heavy burden, weighing them against not only the financial upside, but also against the potential costs to individuals, society, and the firm's reputation. Companies should do this not only as matter of good corporate social responsibility, but also because it will likely square with the government's recommendations for protecting privacy, which seem to advise caution and deliberation, under the banner of "context."

Earlier this year, Federal government officials released two privacy reports — the White House's White Paper and the FTC's Final Privacy Report — that together describe a national privacy policy for the foreseeable future. Although the two reports vary on some particulars, they both point to context as a central, important, and fundamental measuring stick we should use to assess decisions that bear on personal privacy.

The FTC report offers three broad recommendations: Privacy by Design, Simplified Choice for Businesses and Consumers, and Greater Transparency. In discussing the second recommendation — a call for simplified and more transparent choice — the FTC suggests a carve out. "Companies do not need to provide choice before collecting and using consumer data for practices that are consistent with the context of the transaction or the company's relationship with the consumer, or are required or specifically authorized by law." Under this standard, it might be "consistent with the context," for a company in a direct business relationship with a customer to use that customer's information to deliver ads for its other services, but it might be inconsistent with the context — thus requiring notice and choice — to sell that information to third-party advertisers, the FTC explains.

Similarly, the White House white paper defines a "Consumer Privacy Bill of Rights," which would protect, among other things, "Respect for Context." "Consumers have a right to expect that companies will collect, use, and disclose personal data in ways that are consistent with the context in which consumers provide the data," the paper explains.

These parallel pronouncements mean that companies that deal with personal information (meaning all companies, really) need to focus much more often than they have on the history of privacy practices in their industries. Although neither report defines in depth what it means by the word "context," to me the message seems to be: do not push the privacy envelope. Companies that use personal information in ways that go well beyond the practices of their competitors risk crossing the line from responsible steward to reckless abuser of consumer privacy.

The lesson is plain: compete vigorously and beat your competitors in every legitimate way, except when it comes to privacy invasion. Too many companies have learned this lesson the hard way, launching invasive new services that have triggered class action lawsuits, Congressional inquiries, and media firestorms. These companies knew that they were treading where others had feared to go. This may have felt like an exciting opportunity. It should have felt instead like perilous risk-taking, because it meant hurtling beyond the contextual borderlands defined by past practice.

At the Intersection of Big Data and Healthcare: What 7.2 Million Medical Records Can Tell Us

Kenneth Hines, Computing Computer Consortium Blog, August 23, 2012

We’ve featured lots of stories about Big Data over the last several months, but here’s a fascinating new one that illustrates the value of Big Data analytics in addressing important national priorities. Researchers at SENSEable City Lab – a new research initiative of the Massachusetts Institute of Technology — together with colleagues at GE Healthymagination have analyzed data from over 7 million electronic medical records, illustrating in a powerful visual the (sometimes surprising) relationships between medical conditions on the basis of the frequency of co-occurrences. They’re calling this extensive disease network the “Health  InfoScape.”

When you have heartburn, do you also feel nauseous? Or if you’re experiencing insomnia, do you tend to put on a few pounds, or more? By combing through 7.2 million of our electronic medical records, we have created a disease network to help illustrate relationships between various conditions and how common those connections are…

We often have a tendency to think of illness as an isolated event, but our first analysis details the numerous (sometimes unexpected) associations that exist around any given condition. This gives us new insight as to how closely connected some seemingly un-related health conditions might be. Such results force us to re-examine conventional categories of disease classification, as the boundaries between traditional disease categories are thoroughly blurred.

Our initial results are a mix of the expected and the unexpected — simultaneously challenging and reaffirming our preconceptions of health pattern, within individuals and across the U.S.
Click here to “take a look by condition or condition category and gender to uncover interesting associations.”

And there’s yet more opportunity for researchers here:

Now that we have a succinct picture of the human health network in the country, we will continue our investigation by delving deeper into how the environments around us factor into these results.

The Health InfoScape constitutes a perfect example of the role of Big Data science and engineering into the future.

Sunday, August 12, 2012

The Debate Over 'Re-Identification' Of Health Information: What Do We Risk?


Daniel Barth-Jones, Health Affairs, August 10, 2012

Dateline: May 18, 1996 – The collapse and attack. Massachusetts Governor William Weld wasn’t feeling well under his commencement cap and gown. He was about to receive an honorary doctorate from Bentley College and give their keynote graduation address. But, unbeknownst to him, he would instead make a critical contribution to the privacy of our health information. As he stepped forward to the podium, it wasn’t what Weld said that now protects your health privacy, but rather what he did: He teetered and collapsed unconscious before a shocked audience.


Weld recovered quickly and the incident might have passed quietly but for an MIT graduate student. Latanya Sweeney’s studies had brought to her attention hospital data released to researchers by the Massachusetts Group Insurance Commission (GIC) for the purpose of improving healthcare and controlling costs. Federal Trade Commission Senior Privacy Adviser Paul Ohm provides a gripping account of Sweeney’s now famous re-identification of Weld’s hospitalization data using voter list information in his 2010 paper “Broken Promises of Privacy.”

It would be difficult to overstate the influence of the Weld voter list attack on health privacy policy in the United States – it had a direct impact on the development of the de-identification provisions in the HIPAA Privacy rule. However, careful examination of the demographics in Cambridge, MA at the time of the re-identification attempt indicates that Weld was most likely re-identifiable only because he was a public figure who experienced a highly publicized hospitalization rather than there being any actual certainty about the accuracy of his attempted re-identification using the Cambridge voter data.

The Cambridge population was nearly 100,000 and the voter list contained only 54,000 of these residents, so the voter linkage could not provide sufficient evidence to allege any definitive re-identification. Because the logic underlying re-identification depends critically on being able to demonstrate that a person within a health data set is the only person in the larger population who has a set of combined “quasi-identifier” characteristics that could potentially re-identify them, re-identification attempts face a strong challenge in being able to create a complete and accurate population register. Furthermore, the same methodological flaws that undermined the certainty of the Weld re-identification continue to create far-reaching systemic challenges for all re-identification attempts – a fact which must be understood by public policy-makers seeking to realistically assess current privacy risks posed by HIPAA de-identified data. (The full details of these technical issues for re-identification risk assessment are available in a more lengthy review.)

With the benefit of hindsight, it is apparent that the Weld/Cambridge re-identification has served as an important illustration of privacy risks that were not adequately controlled prior to the 2003 HIPAA Privacy Rule. Still, a broader policy debate continues to rage between some voices, like Ohm, alleging that computer scientists can re-identify individuals hidden in anonymized data with “astonishing ease,” and others who view de-identified data as an essential foundation for a host of envisioned advances under healthcare reform.

Nowhere is this tension more evident within the health policy arena than in the recent proposal by the Office of the National Coordinator for Health Information Technology (ONC) for standards, services, and policies enabling secure health information exchange over the Internet to support the Nationwide Health Information Network (NwHIN). Motivated by concern that perceived re-identification risks could “undermine trust”, ONC proposes that de-identified health information could not be used or disclosed for any commercial purpose, a policy which would be certain to unleash a Pandora’s box of unintended consequences. Yet ONC also broadcasts their skepticism regarding purported re-identification risks by noting that they have been “somewhat exaggerated”.

Because a vast array of healthcare improvements and medical research critically depend on de-identified health information, the essential public policy challenge then is to accurately assess the current state of privacy protections for de-identified data, and properly balance both risks and benefits to maximum effect.

Re-Identification Risks Today Under the HIPAA Privacy Rule

HHS appropriately responded to the concerns raised by the Weld/Cambridge voter list privacy attack and, through the HIPAA Privacy Rules, acted to help prevent re-identification attempts.

In 2007, testifying before the Ad Hoc Workgroup on Secondary Uses of Health Data of the National Committee on Vital and Health Statistics, Dr. Latanya Sweeney reported that 0.04 percent (4 in 10,000) of the individuals in the U.S. population within data sets de-identified using the “Safe Harbor” method could be identified on the basis of their year of birth, gender and three-digit ZIP code. To provide some perspective, this risk falls slightly above the lifetime odds of being struck by lightning (one in 10,000).

Further boosting our confidence that re-identification is not a trivial task under today’s protections, a 2010 study estimated re-identification risks under the HIPAA Safe Harbor rule on a state-by-state basis using voter registration data. The percentage of a state’s population estimated to be vulnerable (i.e., not definitively re-identified, but potentially re-identifiable) ranged from 0.01 percent to 0.25 percent.

Another likely source of ONC’s skepticism about re-identification risks comes from ONC’s own 2011 study examining an attack on HIPAA de-identified data under realistic conditions, testing whether HIPAA Safe Harbor de-identified data could be combined with external data to re-identify patients. The study was performed under practical and plausible conditions and verified the re-identifications against direct identifiers—a crucial step often missing from this sort of study. The team began with a set of about 15,000 de-identified patient records. The experiment showed a match for only two of the fifteen thousand individuals (a re-identification rate of 0.013 percent), and even when maximally strong assumptions were made about the possible knowledge of the hypothetical intruder, the re-identification risk (under the questionable assumption that re-identification would even be attempted) was likely to be less than 0.22 percent.

Re-identification risks under the HIPAA Privacy Rule have been reduced to the point that most people wouldn’t (and shouldn’t) lose any sleep over the issue.

What’s At Stake For The Future Of Health Care?

Balancing privacy protection and scientific accuracy. Considerable costs come with incorrectly evaluating the true risks of re-identification under current HIPAA protections. It is essential to understand that de-identification comes at a cost to the scientific accuracy and quality of the healthcare decisions that will be made based on research using de-identified data. Balancing disclosure risks and statistical accuracy is crucial because some popular de-identification methods, such as “k-anonymity methods,” can unnecessarily, and often undetectably, degrade the accuracy of de-identified data for multivariate statistical analyses. This problem is well understood by statisticians and computer scientists, but not well-appreciated in the public policy arena. Poorly conducted de-identification and the overuse of de-identification methods in cases where they do not produce real privacy protections can quickly lead to “bad science” and damaging policy decisions.

Even worse, if we abandon the use of de-identified data because we falsely believe that de-identification cannot provide valuable privacy protections, we will lose the rich benefits that come from analysis of de-identified health data. Jane Yakowitz, a University of Arizona Law School Professor, wrote extensively on this topic in her paper, “Tragedy of the Data Commons,” and addresses the societal costs in information flow and knowledge growth that would follow the abandonment of a realistic assessment of the risks of re-identification.

The reality is that, while one can point to very few, if any, cases of persons who have been harmed by attacks with verified re-identifications, virtually every member of our society has routinely benefited from the use of de-identified health information. De-identified health data is the workhorse that supports numerous healthcare improvements and a wide variety of medical research activities. But just as we cannot identify the specific people who have had their lives saved by speed limit laws, we may fail to realize that we owe our lives to the ongoing research and health system improvements achieved with de-identified data. Hopefully, advancements will continue to accrue in generations to come, but unfounded fears of re-identification could derail this progress.

In my own career as an HIV epidemiologist, I have heightened concerns not only for the very important personal privacy of individuals, but also for the serious tragedies that would occur if fears about de-identification led to a failure to detect and control the next emerging infectious disease that begins to spread globally. If we abandon the use of de-identified data simply because of unwarranted fears regarding privacy risks under today’s HIPAA protections, the consequences of such misguided public policy could be truly disastrous. Privacy advocates and policymakers alike must better understand that, rather than posing new privacy risks, using de-identified data under HIPAA results in vast (thousands-fold) improvements in our individual privacy protection and also sustains a rich public good in research and healthcare improvements.

This critical role that de-identified health information plays in improving healthcare is becoming increasingly more widely recognized, but properly balancing the competing goals of protecting patient privacy while also preserving the accuracy of research requires policy makers to realistically assess both sides of this coin. De-identification policy must achieve an ethical equipoise between potential privacy harms and the very real benefits that result from the advancement of science and healthcare improvements which are accomplished with de-identified data. Properly implemented de-identification complying with the HIPAA de-identification provisions goes a long way toward promoting such a reasonable balance, but I would suggest that there is still room for further improvements in this regard.

Where should we go from here? Because re-identification attacks could still put rare but very real people—with names, faces, and personal lives—at risk of potential privacy harms, we should actively prohibit re-identification, and require those with access to de-identified data to guard and use it appropriately.

HHS Office of Civil Rights (OCR) regulators have promised to provide new guidance in the near future for the de-identification of health data in response to a Congressional mandate to do so. HHS OCR regulators should consider whether it is appropriate for de-identified data to fall entirely outside of the purview of the Privacy Rule, or whether, like the so called “Limited Data Sets” (LDSs), which have been stripped of 16 types of direct identifiers, de-identified data should be subject to certain terms in required Data Use Agreements (DUAs) or subject to direct HHS mandates for use conditions. Effective parallels to the LDS DUA can be carefully constructed to provide assurances which help to further limit re-identification concerns, but which also impose little unnecessary burden on appropriate uses of de-identified data.

Several recommended best practices for the use of de-identified data that should be considered by regulators as possible mandatory de-identified data use conditions include:
.

1. Prohibiting of the re-identification, or attempted re-identification, of individuals and their relatives, family or household members. We should establish civil and criminal penalties for unauthorized re-identification of de-identified data (and for limited data sets). A carefully designed prohibition on re-identification attempts could still allow re-identification research approved by Institutional Review Boards (IRBs) to be conducted, but would ban re-identification attempts conducted without essential human subjects research protections.

2. Requiring parties who wish to link new data elements (which might increase re-identification risks) with data de-identified under the Statistical De-identification provision of the Privacy Rule to confirm that the data remains de-identified.

3. Specifying that HIPAA de-identification status would expire if, at any time, the data contains data elements specified within an evolving Safe Harbor list. The Safe Harbor list should be periodically updated by HHS to include any new “quasi-identifiers” for which population registries of sufficient completeness and accuracy might be reasonably constructed.

4. Formally specifying that for statistically de-identified data, anticipated data recipients must always comply with specified time limits, data use restrictions, qualifications or conditions set forth in the statistical de-identification determination associated with the data.

5. Requiring those holding and using de-identified data to implement and maintain appropriate data security and privacy policies, procedures and associated physical, technical and administrative safeguards as needed to assure that this data is: (a) accessed only by personnel or parties who have agreed to abide by the foregoing conditions, and (b) will remain de-identified in accordance with HIPAA de-identification provisions.

6. Requiring those transferring de-identified data to third parties to enter into data use agreements which would oblige those receiving the data to also hold to the conditions list here, thus maintaining an important “chain-of-trust” data stewardship principal accompanying de-identified data throughout its uses.

Data use requirements of the sort suggested above would impose only modest impositions on the use of de-identified data and would help to provide recourse for actions against data intruders and parties who have not properly managed those very small re-identification risks that might still be associated with de-identified data.

Conclusion

William Weld’s 1997 “re-identification” had an important impact on improving healthcare privacy because it led to regulations that help to importantly protect patients from re-identification risks. But the Weld saga does not reflect the privacy risks that exist under the HIPAA Privacy rules today. We should not let today’s de minimus re-identification risks cause us to abandon our use of de-identified to protect privacy, save lives and continue to improve our healthcare system.

Hopefully, HHS regulators issuing impending de-identification guidance and considering the role of de-identified data for the NwHIN will correctly recognize that substantive protections for de-identification have already been importantly achieved and will carefully balance the substantial societal benefits that result from our ability to conduct analyses, innovate, and improve our healthcare systems using de-identified health data.

Monday, April 2, 2012

Social Networking Heads to the Office

New business applications allow far-flung workers to collaborate in a way they often can't with email

Shayndi Raice, The Wall Street Journal, April 2, 2012

Facebook Inc. has changed the way people socialize on the Web. Now the same concept is changing the way people interact at work.


New social-networking applications aimed at the workplace borrow from Facebook in that they enable workers to set up profiles, form groups and "follow" each other's status updates. But the purpose isn't just social connection; rather, it is to increase productivity by making it easier for employees to identify who does what within an organization and to share their knowledge.

"It's very difficult to collaborate on email and figure out who the experts are," says Rob Koplowitz, an analyst at Forrester Research Inc. "A lot of organizations are finding that working more socially…magnifies the value of information."

These new business tools are based on the idea that email isn't always the best method of communication for widely dispersed workers who need to collaborate on projects or stay abreast of other initiatives within their organizations.

"Email is a 50-year-old technology" and an inefficient way to get work done, says Tony Zingale, chief executive of Palo Alto, Calif.-based Jive Software Inc., which went public in December and whose software allows organizations to maintain their own social networks, among other things. "Collaborative communications are more effective," he says.

Among the other companies selling social-networking applications for businesses are Yammer Inc., Tibco Software Inc. and Salesforce.com Inc., which offers a service called Chatter.

Rob Zell, a leadership-development specialist with Dallas-based 7-Eleven Inc., says the company has about 2,000 employees using Yammer. The convenience-store operator says it deployed the application in May 2011 to help field consultants, who work with local franchise owners, share their knowledge and learn best practices from one another.

Someone might post a picture of a display that worked particularly well in one franchise location, so that others can see it and try the same approach in their areas.

"I use the analogy of the virtual water cooler," Mr. Zell says. "People talk about what's going on in an informal way and have some formal documentation to keep track of best practices."

The companies selling these applications recommend that users create virtual groups so employees can follow a specific topic, branch of the company or type of job.

"Groups are the key," says David Sacks, chief executive of San Francisco-based Yammer, which has four million users. Information-technology professionals can create the groups, but employees should be permitted to create them, as well, so that people can follow topics they are interested in without having to sort through a flood of information, Mr. Sacks says.

In January 2011, Tibco Software launched a service called tibbr that not only enables users to follow what colleagues are doing, it also allows them to follow data, such as the status of an order or invoice.

The service automatically surfaces data so that those interested in keeping track of, for example, orders are notified every time one is filled, rather than having to go look for that information.

"We can bring that information and post it on your proverbial wall," says Ram Menon, the president of social computing for Palo Alto, Calif.-based Tibco, which says it has 70 customers and close to a million users.

Brian Frezza, co-founder and chief executive of Silicon Valley-based biotech start-up Emerald Therapeutics, says that before he started using software from San Francisco-based Asana that allows people associated with tasks to see what others are doing and comment on it, he spent most of his time in 30-minute meetings with his eight employees. There was no easy way for people to know what others had accomplished or had placed as their top priorities, he says.

Mr. Frezza says he now keeps his Asana window open all day, so he can follow what his workers are doing and how they are progressing on various projects. "From the management side, it's staying on track," he says. "It puts things into coherent threads."

Making It Fun

Getting employees to rely more on these social-networking tools for communication and less on email can be a challenge, according to company executives. They say the key to pushing adoption is making sure that senior managers are using the tools. Employees will use a social-networking service if they know their chief executive is on there, listening to their feedback and input, they say.

Making the tools fun to use also helps.

Melissa Madura-Altmann, a vice president of communications at Prudential Real Estate Investors, a unit of Prudential Financial Inc., says she tried to get employees excited about using Jive's software by assigning points to people when they posted on the site.

Employees reacted very competitively, she says, trying to rack up points even though there was no prize involved. "It's fun to see that kind of reaction," she says.

Fostering Unity

Wayne Shurts, chief information officer at Supervalu Inc., a grocery giant with stores in 48 states, says the idea to install Yammer for the company's 15,000 associates came from Chief Executive Craig Herkert, who returned from a Microsoft Corp. CEO conference in Seattle in the fall of 2010 inspired by the idea of using social-media tools to unify employees at his various companies.

Some of Supervalu's workers were already making use of a free version of Yammer's product, so "it was sort of like the top met the bottom in this beautiful way," says Mr. Shurts.

Supervalu, which owns brands such as Albertson's in the U.S. Northwest and Shaw's in New England, was created through a series of acquisitions, so the company was looking for a way to foster better communication and a sense of shared culture.

"The real benefit for us was to break down those walls and start to act and operate as Supervalu," says Mr. Shurts.

Ms. Raice is a staff reporter in The Wall Street Journal's San Francisco bureau. She can be reached at shayndi.raice@wsj.com.