Showing posts with label Government. Show all posts
Showing posts with label Government. Show all posts
Friday, November 30, 2012
Dashboards and Big Data: Before Fruit Ninja, Cybernetics
Will Wiles, The New York Times, November 29, 2012
THIS spring, Fraser Nelson, editor of The Spectator, chastised David Cameron, Britain’s prime minister, for spending too much time on his iPad. “One of his senior advisers says the P.M. spends ‘a crazy, scary amount of time playing Fruit Ninja,’ ” Mr. Nelson wrote in The Daily Telegraph. Such is Mr. Cameron’s devotion that he ordered the creation of an iPad app that would allow him to monitor the British economy.
According to the Web site of the Cabinet Office, the application is now complete. Called the No. 10 Dashboard, after the prime minister’s residence, it gives “the prime minister, other ministers, and senior Whitehall officials an at-a-glance overview of everything that’s happening in government and elsewhere” — stock prices, housing and jobs data, information on the performance of government departments, as well as the “political context”: polls, commentary and indications of the national mood from sources like Twitter. The prime minister liked having “quotable facts about what was going on,” said Alice Newton, one of the app’s developers.
The site included an example of the kind of data the app offers — a graph of Britain’s stuttering gross domestic product — something you’d expect Mr. Cameron to be able to see whenever he closes his eyes, let alone on his iPad.
The No. 10 Dashboard could be the White House Dashboard; Mr. Cameron plans to show the app to President Obama at the Group of 8 summit meeting.
The whole story is bathed in the white heat of 21st-century digital technology. “Ours is the first generation of digital natives coming through to work in government,” Ms. Newton said. But there’s nothing new about it. If Mr. Obama follows in Mr. Cameron’s footsteps, he should know that he’s also on the trail of another leader — the ill-fated socialist president of Chile, Salvador Allende.
What’s the connection between men so widely separated by ideology, geography and time? After assuming power in 1970, Allende’s administration was faced with economic paralysis, inequality and civil strife. The usual socialist prescription of nationalization of industry was applied, but Allende also looked to an unexpected quarter: the relatively new and niche science of cybernetics. Cybernetics is the study of control and communication in large, complex systems, be they organisms, machines or organizations. It spans management theory, information technology, psychology, biology and sociology. The Chilean government approached a British cybernetician named Stafford Beer and asked him to build a cybernetic hub for the management of the country’s economy — something that had never been attempted before, or since.
A recent history, “Cybernetic Revolutionaries,” by Eden Medina, gives the subsequent project, Cybersyn, the serious attention it deserves. The system, once up and running, would have channeled data from Chile’s nationalized industries into an operations room in Santiago, where Allende’s ministers would have made informed decisions in chairs with control panels built into the armrests. Photos of this room still retain a veneer of giddy futurity. White surfaces, pared-down interfaces, rounded corners — a dash of Apple in the mix.
Ms. Medina contends that rather than a tool of technocratic centralization, Cybersyn was democratic in design and intent. Workers would have access to the heart of government and could inform decision makers. The aim was transparency and economic and social homeostasis, a natural equilibrium. Mr. Beer called the system the “Liberty Machine.” But Allende’s regime, overthrown by a C.I.A.-backed military coup in 1973, didn’t last long enough to see Cybersyn operate.
The No. 10 Dashboard taps into the same desire to master available information — a desire that has only grown as the amount of information in circulation has increased. Where Cybersyn needed dedicated national infrastructure and rooms full of equipment, the app runs on a hand-held device. And yet, the dashboard is actually less sophisticated.
It is not truly cybernetic because it lacks a mechanism to translate all that data into action. It can display information; it cannot consult and control. Less a driver’s dashboard, it is more a window out of which a passenger can observe the national scenery speeding past. Government might be given the illusion of hand-held, one-stop manageability, but no actual managing is going on.
The app could thus be an apt metaphor for politicians reduced to spectators by the surges and shocks of the globalized world. Mr. Cameron should remember that there’s at least one other instance of government-by-app: the team that worked on restructuring Greece’s debt used iPads too, equipped with an app purpose-built for the job.
However that turns out, we can at least say this: in terms of distractions, these apps are marginally more useful than Fruit Ninja.
Will Wiles is the author of the novel “Care of Wooden Floors.”
Friday, October 19, 2012
Will Big Data decide the election?
A new book traces the recent history of data mining in political campaigns. A review of The Victory Lab, by Sasha Issenberg.
Chip Lebovitz, Fortune, October 19, 2012
FORTUNE -- There's a powerful vignette in Sasha Issenberg's The Victory Lab in which political consultant Alexander Gage presents his new data targeting system to Mitt Romney's 2002 gubernatorial campaign.
Gage has combined consumer records with political voting history to identify potential Romney supporters among nontraditional Republican voting blocks. Gage sees his work as revolutionary -- a first in politics, and potentially a first anywhere. Yet just as he completes his presentation, Romney's deputy campaign manager Alex Dunn raises his hand and deadpans, "You mean you don't do this in politics."
Dunn's surprise was not out of place in 2001. And while much has changed since then, politics remains an analog art in many ways. The Victory Lab charts the recent history of political data mining designed to identify persuadable voters and swing elections.
Politico has called Issenberg's book "Moneyball for politics." There are obvious parallels between political operatives like Gage and the protagonist of Moneyball, Oakland Athletics general manager Billy Beane.
Both pioneered data-centric techniques that gave their respective organizations a leg up over the competition.
Yet in his 14 years as GM, Billy Beane has yet to make it to the World Series. By contrast, the work of Gage and other data-mining political consultants has translated directly into electoral success. These operatives may not have singlehandedly won the past two presidential elections, but the side that had the most advanced data-driven mobilization efforts went a perfect 2-0.
The George W. Bush campaign's sophisticated data mining helped the president win reelection in 2004. Four years later, Barack Obama defeated John McCain in part because the Obama campaign outclassed McCain's operatives on the data front.
Issenberg seems to have interviewed everyone who's anyone in the field of political statistics, but his subject matter doesn't necessarily lend itself to prose. Numbers have power, but too many of them can addle rather than inform. After reading the book, you'll probably remember the outcome of a few experiments, but be stuck scouring your brain for the results of the rest. Readers are advised to keep a pen and pencil handy when reading The Victory Lab, so they can jot down key stats.
Yet Issenberg's material is sufficiently gripping that you'll want to keep turning the pages, even if it means deciphering the results of yet another randomized trial. For example, micro-targeting efforts on behalf of Senator Michael Bennett (D-Colo.)'s 2010 Senate reelection campaign likely spurred 25,000 additional Bennett votes. That might not seem like a big number, until you consider that the race was decided by 15,000 votes.
Issenberg avoids sweeping predictions about the future of political data mining. That's probably wise, given that political professionals tend to define the future as Election Day. Campaigns at their simplest are a one-day snapshot of personal preference. The Victory Lab reminds us, however, that every individual choice is the result of a thousand factors, many of them subject to manipulation by political operatives.
Monday, October 15, 2012
McKinsey Anthology: Government Designed for New Times
To explore the approaches that governments around the world are taking to common problems, this anthology convenes political leaders and civil servants, economists and policy experts, generalists and specialists.
McKinsey Anthology, 2012
Contents
Transforming government
· Tony Blair—Leading transformation in the 21st century
· François-Daniel Migeon—Interview: Transforming government in France
· Frank-Jürgen Weise—Behind the German jobs miracle
· Tim Brown—Quick take: Designing a tech-enabled government
· Michael Fullan—Transforming schools an entire system at a time
· Todd Park—Interview: Unleashing government's 'innovation mojo'
· Diana Farrell—Government designed for new times
Innovating government services
· James Fishkin—What the people think when they're really thinking
· Nandan Nilekani—Interview: For every citizen, an identity
· Matthew Taylor—Citizens: The untapped resource
· Susan Zielinski—The new mobility
· Peter Shergold—A social contract for government
· Wim Elfrink—The smart-city solution
· Karan Bhatia—Quick take: Building the world's infrastructure
· Salman Khan—Teaching for the new millennium
· Xie Chengxiang—Interview: Home for the urban poor
· Elena Berkowitz and Blaise Warren—Quick take: How Estonia became E-stonia
Building new competencies
· Douglas Holtz-Eakin—Fiscal management fix: Simple math—and a very big stick
· Coen Teulings—Why politicans prefer austerity to long-term fiscal reform
· Lu Mai—The urbanization solution
· Göran Persson—How to tame a budget crisis
· Peter Ho—Coping with complexity
· Mohamed Ibrahim—Better data, better policy making
Understanding government in new times
· Daron Acemoglu—The servant state
· Parag Khanna—The rise of hybrid governance
· Neil deGrasse Tyson—Why exploration matters—and why the government should pay
for it
· Ray O. Johnson—Quick take: The research imperative
· Hernando de Soto—Interview: Building a nation of owners
· Nicolas Berggruen and Nathan Gardels—A middle way for governance
Wednesday, October 3, 2012
TechAmerica Big Data Commission Report:
The Big Data Commission released its findings October 3rd to help answer these questions and many more. The report, “Demystifying Big Data: A Practical Guide to Transforming the Business of Government,” provides the government’s senior policy and decision makers with a comprehensive roadmap to using Big Data to better serve the American people.
More info:
Data in the world is doubling every 18 months. Across government everyone is talking about the concept of Big Data, and how this new technology will transform the way Washington does business. Looking past the excitement, many questions remained unanswered- until now. The TechAmerica Foundation’s Big Data Commission has worked diligently to put together the most comprehensive report on Big Data of its kind. What is Big Data, really? How is it defined? What capabilities are required to succeed? How do you use Big Data to make intelligent decisions? How will agencies effectively govern and secure huge volumes of information, while protecting privacy and civil liberties? And perhaps most importantly, what value will it really deliver to the US Government and the citizenry we serve?
The Big Data Commission released its findings October 3rd to help answer these questions and many more. The report, “Demystifying Big Data: A Practical Guide to Transforming the Business of Government,” provides the government’s senior policy and decision makers with a comprehensive roadmap to using Big Data to better serve the American people.
For more information about the Big Data Commission or other related activities, please contact Chris Wilson, chris.wilson@techamericafoundation.org or (202) 682-4451.
Commission Leadership
The Big Data Commission is chaired by Steven A. Mills, Senior Vice President and Group Executive, Software & Systems, IBM Corporation and Steve Lucas, Global Executive Vice President and General Manager, SAP Database & Technology. Serving as vice chairs of the commission are Teresa Carlson, Vice president Global Public Sector, Amazon Web Services and Bill Perlowitz, Chief Technology Officer, Science, Technology and Engineering Group, Wyle.
Serving as the Academic Co-Chairs will be Dr. Michael Rappa, founding director of the Institute for Advanced Analytics at North Carolina State University and principal architect of its Master of Science in Analytics degree; and Dr. Leo Irakliotis, Dean and National Director for the College of Information Technology at Western Governors University (WGU).
Commission Members
The commission membership is comprised of leading experts on big data and representing both industry and academia, along with a government advisory board with the objective of providing guidance on how Government Agencies should be leveraging Big Data to address their most critical business imperatives, and how Big Data can drive U.S. innovation and competitiveness. A full list of the commission membership can be found here.
Case Studies
· AM Biotech
· IRS Compliance Data Warehouse
· NARA ERA
· NASA Human Spaceflight Imagery
· NOAA NWS
· TerraEchos
· University of Ontario
· Vesta Wind Energy
Thursday, September 27, 2012
New Talk by Clay Shirky -
Clay Shirky’s Ted Talk “How the Internet will (one day) transform government”
The open-source world has learned to deal with a flood of new, oftentimes divergent, ideas using hosting services like GitHub -- so why can’t governments? In this rousing talk Clay Shirky shows how democracies can take a lesson from the Internet, to be not just transparent but also to draw on the knowledge of all their citizens.
Clay Shirky argues that the history of the modern world could be rendered as the history of ways of arguing, where changes in media change what sort of arguments are possible -- with deep social and political implications
Saturday, August 18, 2012
Data and Analytics Key to Health Reform, but Challenges Stand in Way
Kate Ackerman, iHealthBeat, August 15, 2012
NATIONAL HARBOR, Md. -- At the eHealth Initiative's National Forum on Data and Analytics in Healthcare last week, stakeholders discussed the importance of data and analytics in implementing health reform, as well as the challenges associated with it.
Jennifer Covich Bordenick, CEO of the eHealth Initiative, said, "Our survey, CIOs and members kept telling us that they are concerned about analytics. They don't feel they have the tools necessary to meet the demands of accountable care and meaningful use." She noted that "93% of the CIOs believe it is very important, but 72% don't feel their organizations have what they need to meet the analytical needs."
By convening experts in health data and analytics, the forum aimed to highlight organizations that are leading the way and facilitate conversations around the need for improvement, she said.
Federal Government Touts Data and Analytics To Support Health Reform
Niall Brennan -- director of the Policy and Data Analysis Group at CMS -- told attendees that health data analytics is "absolutely central" to everything that has to do with health reform.
Brennan -- who stepped in for U.S. Chief Technology Officer Todd Park to give the afternoon's keynote speech -- highlighted the federal government's efforts "to make the data more helpful," while not compromising individuals' privacy.
He said that in the past CMS was "overly conservative" in terms of data release. However, in the last few years -- in large part because of the Affordable Care Act -- the agency has made great strides in liberating health data, he said.
"It might not look like it on the outside, but we are literally constantly pushing the envelope," Brennan said.
He cited the Blue Button Initiative and HealthData.gov as examples of the government's efforts to make privacy-protected health information available to the public.
He also noted that Section 10332 of the ACA authorizes the release of Medicare fee-for-service data to qualified entities if they agree to combine the CMS data with claims data from other sources to compile performance reports.
While Brennan touted the federal government's release of more health care data, he acknowledged, "You can have all the data in the world, but if you don't have the right analytics ... it's just a bunch of useless numbers."
Brennan said the federal government -- from CMS to HHS to the White House -- is "very, very committed" to data and analytics.
Survey Finds Industry Still Has Far To Go
At the forum, Jason Goldwater -- vice president of programs and research at eHI -- offered a sneak peek into the results of a survey eHI conducted with the College of Healthcare Information and Management Executives to get a picture of the types of data and analytics being used in health care.
The survey, which was conducted in July, focused on four areas:
· Types of data used;
· Types of analytic functions used;
· Types of functions needed; and
· Challenges to the use of data and analytics.
When asked what data their organizations actively exchange:
· 76.6% of respondents said lab results;
· 74.5% said demographics;
· 70.2% said discharge summaries;
· 46.8% said allergy information;
· 36.2% said continuity of care documents; and
· 36.2% said problem lists.
The survey found that most health care organizations are focusing their resources on retrospective analysis, with 58.3% citing that as the area in which they direct the majority of their analytical resources. According to the survey, 16.7% of respondents said their organizations direct the majority of their analytical resources toward real-time decision support, 13.9% said optimization and efficiency and 2.8% cited predictive analytics.
When asked what type of analytical functions their organizations primarily use:
· 87.5% of respondents said ad-hoc queries;
· 61.1% said data mining;
· 56.9% said data warehousing;
· 34.7% said exploratory data analysis;
· 30.6% said on-line analytical processing; and
· 23.6% said predictive modeling.
Goldwater said that despite health care organizations' interest in data and analytics, respondents cited several challenges, including:
· Lack of standardized data across systems;
· Lack of a system infrastructure to support analytics;
· Cost of analytical software;
· Concerns about privacy and security of the data; and
· Limited utility of the results to the organizations.
Goldwater said that eHI, CHIME and McKesson will host a webinar on Aug. 30 to discuss the survey results in more detail and that eHI will release an issue brief in the fall.
Speakers Highlight Challenges Associated With Health Data and Analytics
Several speakers offered real-life examples of the challenges cited by respondents to the eHI/CHIME survey.
Jason Williams -- vice president of business analytics at RelayHealth -- said that health care cost and quality reforms require inter-stakeholder transparency and that there needs to be "more emphasis" in that area. He also cited the demand for business and technology analysts and the need to find a balanced approach to privacy as areas for improvement.
Micky Tripathi -- president and CEO of the Massachusetts eHealth Collaborative and chair of eHI's Board of Directors -- said that variations in EHR systems can be problematic for analytics.
Brendan Mullen -- senior director of PINNACLE, the American College of Cardiology's outpatient registry -- noted that its system integration tool has to look at 41 locations in the NextGen EHR system to determine if a physician provided patient education on heart failure.
Tripathi said that as vendors are working to address the issues, new measures are coming down the pipeline.
For the Healthcare Information and Management Systems Society Conference in February, MAeHC compared the results of its certified Quality Data Center with the Office of the National Coordinator for Health IT-sponsored, open-source popHealth tool to evaluate meaningful use quality measures. The tools used the same exact data, but they did not produce the same results for any of the 44 measures, Tripathi said.
He explained that further investigation found several reasons for the discrepancies, including the definition of the continuity of care document, coding and mapping, and the interpretation of certain measures, such as age.
Tripathi said, "We have a ton of work to do," adding that the industry needs to keep pushing along.
Covich Bordenick told iHealthBeat, "It was clear by the end of that day that this was just the start of a conversation that is going to take years to explore," adding, "eHealth Initiative is going to help unpack this issue."
MORE ON THE WEB
· National Forum on Data and Analytics in Healthcare
· HealthData.gov
· "Data Analysis and the Future of Health Care" (Tibken, Wall Street Journal, 4/16).
Sunday, August 12, 2012
The Debate Over 'Re-Identification' Of Health Information: What Do We Risk?
Daniel Barth-Jones, Health Affairs, August 10, 2012
Dateline: May 18, 1996 – The collapse and attack. Massachusetts Governor William Weld wasn’t feeling well under his commencement cap and gown. He was about to receive an honorary doctorate from Bentley College and give their keynote graduation address. But, unbeknownst to him, he would instead make a critical contribution to the privacy of our health information. As he stepped forward to the podium, it wasn’t what Weld said that now protects your health privacy, but rather what he did: He teetered and collapsed unconscious before a shocked audience.
Weld recovered quickly and the incident might have passed quietly but for an MIT graduate student. Latanya Sweeney’s studies had brought to her attention hospital data released to researchers by the Massachusetts Group Insurance Commission (GIC) for the purpose of improving healthcare and controlling costs. Federal Trade Commission Senior Privacy Adviser Paul Ohm provides a gripping account of Sweeney’s now famous re-identification of Weld’s hospitalization data using voter list information in his 2010 paper “Broken Promises of Privacy.”
It would be difficult to overstate the influence of the Weld voter list attack on health privacy policy in the United States – it had a direct impact on the development of the de-identification provisions in the HIPAA Privacy rule. However, careful examination of the demographics in Cambridge, MA at the time of the re-identification attempt indicates that Weld was most likely re-identifiable only because he was a public figure who experienced a highly publicized hospitalization rather than there being any actual certainty about the accuracy of his attempted re-identification using the Cambridge voter data.
The Cambridge population was nearly 100,000 and the voter list contained only 54,000 of these residents, so the voter linkage could not provide sufficient evidence to allege any definitive re-identification. Because the logic underlying re-identification depends critically on being able to demonstrate that a person within a health data set is the only person in the larger population who has a set of combined “quasi-identifier” characteristics that could potentially re-identify them, re-identification attempts face a strong challenge in being able to create a complete and accurate population register. Furthermore, the same methodological flaws that undermined the certainty of the Weld re-identification continue to create far-reaching systemic challenges for all re-identification attempts – a fact which must be understood by public policy-makers seeking to realistically assess current privacy risks posed by HIPAA de-identified data. (The full details of these technical issues for re-identification risk assessment are available in a more lengthy review.)
With the benefit of hindsight, it is apparent that the Weld/Cambridge re-identification has served as an important illustration of privacy risks that were not adequately controlled prior to the 2003 HIPAA Privacy Rule. Still, a broader policy debate continues to rage between some voices, like Ohm, alleging that computer scientists can re-identify individuals hidden in anonymized data with “astonishing ease,” and others who view de-identified data as an essential foundation for a host of envisioned advances under healthcare reform.
Nowhere is this tension more evident within the health policy arena than in the recent proposal by the Office of the National Coordinator for Health Information Technology (ONC) for standards, services, and policies enabling secure health information exchange over the Internet to support the Nationwide Health Information Network (NwHIN). Motivated by concern that perceived re-identification risks could “undermine trust”, ONC proposes that de-identified health information could not be used or disclosed for any commercial purpose, a policy which would be certain to unleash a Pandora’s box of unintended consequences. Yet ONC also broadcasts their skepticism regarding purported re-identification risks by noting that they have been “somewhat exaggerated”.
Because a vast array of healthcare improvements and medical research critically depend on de-identified health information, the essential public policy challenge then is to accurately assess the current state of privacy protections for de-identified data, and properly balance both risks and benefits to maximum effect.
Re-Identification Risks Today Under the HIPAA Privacy Rule
HHS appropriately responded to the concerns raised by the Weld/Cambridge voter list privacy attack and, through the HIPAA Privacy Rules, acted to help prevent re-identification attempts.
In 2007, testifying before the Ad Hoc Workgroup on Secondary Uses of Health Data of the National Committee on Vital and Health Statistics, Dr. Latanya Sweeney reported that 0.04 percent (4 in 10,000) of the individuals in the U.S. population within data sets de-identified using the “Safe Harbor” method could be identified on the basis of their year of birth, gender and three-digit ZIP code. To provide some perspective, this risk falls slightly above the lifetime odds of being struck by lightning (one in 10,000).
Further boosting our confidence that re-identification is not a trivial task under today’s protections, a 2010 study estimated re-identification risks under the HIPAA Safe Harbor rule on a state-by-state basis using voter registration data. The percentage of a state’s population estimated to be vulnerable (i.e., not definitively re-identified, but potentially re-identifiable) ranged from 0.01 percent to 0.25 percent.
Another likely source of ONC’s skepticism about re-identification risks comes from ONC’s own 2011 study examining an attack on HIPAA de-identified data under realistic conditions, testing whether HIPAA Safe Harbor de-identified data could be combined with external data to re-identify patients. The study was performed under practical and plausible conditions and verified the re-identifications against direct identifiers—a crucial step often missing from this sort of study. The team began with a set of about 15,000 de-identified patient records. The experiment showed a match for only two of the fifteen thousand individuals (a re-identification rate of 0.013 percent), and even when maximally strong assumptions were made about the possible knowledge of the hypothetical intruder, the re-identification risk (under the questionable assumption that re-identification would even be attempted) was likely to be less than 0.22 percent.
Re-identification risks under the HIPAA Privacy Rule have been reduced to the point that most people wouldn’t (and shouldn’t) lose any sleep over the issue.
What’s At Stake For The Future Of Health Care?
Balancing privacy protection and scientific accuracy. Considerable costs come with incorrectly evaluating the true risks of re-identification under current HIPAA protections. It is essential to understand that de-identification comes at a cost to the scientific accuracy and quality of the healthcare decisions that will be made based on research using de-identified data. Balancing disclosure risks and statistical accuracy is crucial because some popular de-identification methods, such as “k-anonymity methods,” can unnecessarily, and often undetectably, degrade the accuracy of de-identified data for multivariate statistical analyses. This problem is well understood by statisticians and computer scientists, but not well-appreciated in the public policy arena. Poorly conducted de-identification and the overuse of de-identification methods in cases where they do not produce real privacy protections can quickly lead to “bad science” and damaging policy decisions.
Even worse, if we abandon the use of de-identified data because we falsely believe that de-identification cannot provide valuable privacy protections, we will lose the rich benefits that come from analysis of de-identified health data. Jane Yakowitz, a University of Arizona Law School Professor, wrote extensively on this topic in her paper, “Tragedy of the Data Commons,” and addresses the societal costs in information flow and knowledge growth that would follow the abandonment of a realistic assessment of the risks of re-identification.
The reality is that, while one can point to very few, if any, cases of persons who have been harmed by attacks with verified re-identifications, virtually every member of our society has routinely benefited from the use of de-identified health information. De-identified health data is the workhorse that supports numerous healthcare improvements and a wide variety of medical research activities. But just as we cannot identify the specific people who have had their lives saved by speed limit laws, we may fail to realize that we owe our lives to the ongoing research and health system improvements achieved with de-identified data. Hopefully, advancements will continue to accrue in generations to come, but unfounded fears of re-identification could derail this progress.
In my own career as an HIV epidemiologist, I have heightened concerns not only for the very important personal privacy of individuals, but also for the serious tragedies that would occur if fears about de-identification led to a failure to detect and control the next emerging infectious disease that begins to spread globally. If we abandon the use of de-identified data simply because of unwarranted fears regarding privacy risks under today’s HIPAA protections, the consequences of such misguided public policy could be truly disastrous. Privacy advocates and policymakers alike must better understand that, rather than posing new privacy risks, using de-identified data under HIPAA results in vast (thousands-fold) improvements in our individual privacy protection and also sustains a rich public good in research and healthcare improvements.
This critical role that de-identified health information plays in improving healthcare is becoming increasingly more widely recognized, but properly balancing the competing goals of protecting patient privacy while also preserving the accuracy of research requires policy makers to realistically assess both sides of this coin. De-identification policy must achieve an ethical equipoise between potential privacy harms and the very real benefits that result from the advancement of science and healthcare improvements which are accomplished with de-identified data. Properly implemented de-identification complying with the HIPAA de-identification provisions goes a long way toward promoting such a reasonable balance, but I would suggest that there is still room for further improvements in this regard.
Where should we go from here? Because re-identification attacks could still put rare but very real people—with names, faces, and personal lives—at risk of potential privacy harms, we should actively prohibit re-identification, and require those with access to de-identified data to guard and use it appropriately.
HHS Office of Civil Rights (OCR) regulators have promised to provide new guidance in the near future for the de-identification of health data in response to a Congressional mandate to do so. HHS OCR regulators should consider whether it is appropriate for de-identified data to fall entirely outside of the purview of the Privacy Rule, or whether, like the so called “Limited Data Sets” (LDSs), which have been stripped of 16 types of direct identifiers, de-identified data should be subject to certain terms in required Data Use Agreements (DUAs) or subject to direct HHS mandates for use conditions. Effective parallels to the LDS DUA can be carefully constructed to provide assurances which help to further limit re-identification concerns, but which also impose little unnecessary burden on appropriate uses of de-identified data.
Several recommended best practices for the use of de-identified data that should be considered by regulators as possible mandatory de-identified data use conditions include:
.
1. Prohibiting of the re-identification, or attempted re-identification, of individuals and their relatives, family or household members. We should establish civil and criminal penalties for unauthorized re-identification of de-identified data (and for limited data sets). A carefully designed prohibition on re-identification attempts could still allow re-identification research approved by Institutional Review Boards (IRBs) to be conducted, but would ban re-identification attempts conducted without essential human subjects research protections.
2. Requiring parties who wish to link new data elements (which might increase re-identification risks) with data de-identified under the Statistical De-identification provision of the Privacy Rule to confirm that the data remains de-identified.
3. Specifying that HIPAA de-identification status would expire if, at any time, the data contains data elements specified within an evolving Safe Harbor list. The Safe Harbor list should be periodically updated by HHS to include any new “quasi-identifiers” for which population registries of sufficient completeness and accuracy might be reasonably constructed.
4. Formally specifying that for statistically de-identified data, anticipated data recipients must always comply with specified time limits, data use restrictions, qualifications or conditions set forth in the statistical de-identification determination associated with the data.
5. Requiring those holding and using de-identified data to implement and maintain appropriate data security and privacy policies, procedures and associated physical, technical and administrative safeguards as needed to assure that this data is: (a) accessed only by personnel or parties who have agreed to abide by the foregoing conditions, and (b) will remain de-identified in accordance with HIPAA de-identification provisions.
6. Requiring those transferring de-identified data to third parties to enter into data use agreements which would oblige those receiving the data to also hold to the conditions list here, thus maintaining an important “chain-of-trust” data stewardship principal accompanying de-identified data throughout its uses.
Data use requirements of the sort suggested above would impose only modest impositions on the use of de-identified data and would help to provide recourse for actions against data intruders and parties who have not properly managed those very small re-identification risks that might still be associated with de-identified data.
Conclusion
William Weld’s 1997 “re-identification” had an important impact on improving healthcare privacy because it led to regulations that help to importantly protect patients from re-identification risks. But the Weld saga does not reflect the privacy risks that exist under the HIPAA Privacy rules today. We should not let today’s de minimus re-identification risks cause us to abandon our use of de-identified to protect privacy, save lives and continue to improve our healthcare system.
Hopefully, HHS regulators issuing impending de-identification guidance and considering the role of de-identified data for the NwHIN will correctly recognize that substantive protections for de-identification have already been importantly achieved and will carefully balance the substantial societal benefits that result from our ability to conduct analyses, innovate, and improve our healthcare systems using de-identified health data.
Monday, June 18, 2012
'Big Data' disguises digital doubts (USA today)
Dan Vergano, USA TODAY, June 16, 2012
Buzzwords don't come any bigger than "Big Data," which promises to reveal the secrets hidden within big blocks of data held by companies, governments and musty old archives.
But maybe Big Data has an Achilles' heel, some experts warn, despite its Big promises.
"The initiative we are launching today promises to transform our ability to use Big Data for scientific discovery, environmental and biomedical research, education and national security," said presidential science adviser John Holdren, announcing a $200 million effort in March by six federal agencies to uncork the power of Big Data.
Holdren compared Big Data's advent to the invention of supercomputers and the Internet. But what is Big Data really? Starting from scientists struggling to analyze massive amounts of genetic data, as they did in the Human Genome Project a decade ago, or astronomical data, such as the survey of more than 930,000 galaxies undertaken by the Sloan Digital Sky Survey, Big Data has blossomed into a constellation of computer science approaches to handling, visualizing and blending together "big" sets of data.
For example, police forces from Honolulu to New York have looked at combinations of crime tips submitted via Facebook, Twitter and text messages to identify "hotspots" for muggings and other felonies. Amazon famously tracks masses of book purchases to suggest new buys to like-minded readers. The Defense Department hopes to weave together information from a new generation of battlefield sensors at speeds 100 times faster than today using Big Data techniques.
Such efforts have blossomed in the Facebook era, where poking through troves of customer data is seen as the key to unlocking sales. In medicine, a two-day "Health Datapalooza" held this month in the nation's capital drew together federal officials, former Senate majority leader Bill Frist, R-Tenn., and Wired Magazine executive editor Thomas Goetz, to talk health data. If your gene map can be compared instantly to the genomes of millions of other folks in coming decades, for example, the hope is that medicine finely tuned to your medical needs will result.
A Sciencejournal study last year introduced the notion of "culturomics," using Big Data — Google's millions of searchable digitized books in its case — to reveal, "linguistic and cultural phenomena that were reflected in the English language between 1800 and 2000."
Data. Data. Data. So much of it is out there, tracked from the moment you look up a dentist on a website, take a trip through a highway tollbooth to sit in the chair, pay your bill at the reception desk and post your toothache experience afterward on Facebook. Can an ad for a toothbrush be far from your in-box?
"The only problem is that a lot of the Big Data isn't really data," says anthropologist Robert Albro of American University in Washington D.C., who studies how culture affects public policy. "It's a mash-up of all kinds of numbers that started out as data, but they don't necessarily mean anything once they have been removed from where they started out." In the social sciences, he says, researchers have learned over the last century that half the battle in any study is carefully explaining your data's origins. "Once you leave that behind, there is a risk you'll be wrong, and a risk that the decisions you make based on being wrong will affect people in negative ways."
In anthropology, one historical example of a problem comes from turn-of-the-century attempts to pigeonhole people in far-off nations into tribes or countries, using data in categories now understood as far too simplistic. That effort contributed to European countries inventing imaginary borders or people in nations such as Rwanda. Public health researchers in the 20th century pigeonholed poor people into "defective" categories, based on bogus data, during the "eugenics" movement aimed at breeding better human beings that led to the involuntary sterilization of perhaps 60,000 people nationwide by the 1960s.
More recently, University of Wisconsin-Madison, communications scholars have warned that Google's search recommendations (the list of suggested searches that pop up when you start typing a word using the popular search engine) actually bend people's perception. Looking at nanotechnology, for example, the study showed that top search suggestions over a few years turned away from business to health concerns. The search recommendations were actually steering more people to look into less-reliable nanotechnology health-issue websites, they found. "Google is shaping the reality we experience in the suggestions it makes, pointing us away from the most accurate information and towards the most popular," study lead author Dietram Scheufele told USA TODAY in 2010.
Still, what's so wrong about using Big Data to find crime hotspots or books you might like? "Nothing. There is obviously immense promise there, as long as the data is kept to uses for which its limits are understood," Albro suggests. However, he worries that since so much of the data out there start out as "market research" — likes or dislikes when it comes to buying things — that removing data from advertising-focused troves and translating it into health care or planning for new roads or "culture" will essentially turn everyone into consumers, rather than citizens, in the minds of planners.
"We can't even agree on what 'culture' is, and now we're going to have 'culturomics.' Isn't that a little ambitious?" Albro asks. "Now we have claims that tweets predicted the 'Arab Spring,' which turns out to be questionable, or can detect 'sentiment' or 'mood,' which are even fuzzier or lazier words. We need to be a little cautious here." (In their defense, the "culturomics" study authors do urge caution on folks using their approach.)
But Albro worries that the hype over Big Data is warming up for Big Disappointment down the road. Much of the criticism of Big Data heard now focuses on privacy issues: who is using your data, or whether it will really help sell stuff. "Those are useful discussions, but we really need to talk a little more deeply about data," Albro says. "It's more than 'Garbage In, Garbage Out,' it's about how we shape the digital world."
Buzzwords don't come any bigger than "Big Data," which promises to reveal the secrets hidden within big blocks of data held by companies, governments and musty old archives.
But maybe Big Data has an Achilles' heel, some experts warn, despite its Big promises.
"The initiative we are launching today promises to transform our ability to use Big Data for scientific discovery, environmental and biomedical research, education and national security," said presidential science adviser John Holdren, announcing a $200 million effort in March by six federal agencies to uncork the power of Big Data.
Holdren compared Big Data's advent to the invention of supercomputers and the Internet. But what is Big Data really? Starting from scientists struggling to analyze massive amounts of genetic data, as they did in the Human Genome Project a decade ago, or astronomical data, such as the survey of more than 930,000 galaxies undertaken by the Sloan Digital Sky Survey, Big Data has blossomed into a constellation of computer science approaches to handling, visualizing and blending together "big" sets of data.
For example, police forces from Honolulu to New York have looked at combinations of crime tips submitted via Facebook, Twitter and text messages to identify "hotspots" for muggings and other felonies. Amazon famously tracks masses of book purchases to suggest new buys to like-minded readers. The Defense Department hopes to weave together information from a new generation of battlefield sensors at speeds 100 times faster than today using Big Data techniques.
Such efforts have blossomed in the Facebook era, where poking through troves of customer data is seen as the key to unlocking sales. In medicine, a two-day "Health Datapalooza" held this month in the nation's capital drew together federal officials, former Senate majority leader Bill Frist, R-Tenn., and Wired Magazine executive editor Thomas Goetz, to talk health data. If your gene map can be compared instantly to the genomes of millions of other folks in coming decades, for example, the hope is that medicine finely tuned to your medical needs will result.
A Sciencejournal study last year introduced the notion of "culturomics," using Big Data — Google's millions of searchable digitized books in its case — to reveal, "linguistic and cultural phenomena that were reflected in the English language between 1800 and 2000."
Data. Data. Data. So much of it is out there, tracked from the moment you look up a dentist on a website, take a trip through a highway tollbooth to sit in the chair, pay your bill at the reception desk and post your toothache experience afterward on Facebook. Can an ad for a toothbrush be far from your in-box?
"The only problem is that a lot of the Big Data isn't really data," says anthropologist Robert Albro of American University in Washington D.C., who studies how culture affects public policy. "It's a mash-up of all kinds of numbers that started out as data, but they don't necessarily mean anything once they have been removed from where they started out." In the social sciences, he says, researchers have learned over the last century that half the battle in any study is carefully explaining your data's origins. "Once you leave that behind, there is a risk you'll be wrong, and a risk that the decisions you make based on being wrong will affect people in negative ways."
In anthropology, one historical example of a problem comes from turn-of-the-century attempts to pigeonhole people in far-off nations into tribes or countries, using data in categories now understood as far too simplistic. That effort contributed to European countries inventing imaginary borders or people in nations such as Rwanda. Public health researchers in the 20th century pigeonholed poor people into "defective" categories, based on bogus data, during the "eugenics" movement aimed at breeding better human beings that led to the involuntary sterilization of perhaps 60,000 people nationwide by the 1960s.
More recently, University of Wisconsin-Madison, communications scholars have warned that Google's search recommendations (the list of suggested searches that pop up when you start typing a word using the popular search engine) actually bend people's perception. Looking at nanotechnology, for example, the study showed that top search suggestions over a few years turned away from business to health concerns. The search recommendations were actually steering more people to look into less-reliable nanotechnology health-issue websites, they found. "Google is shaping the reality we experience in the suggestions it makes, pointing us away from the most accurate information and towards the most popular," study lead author Dietram Scheufele told USA TODAY in 2010.
Still, what's so wrong about using Big Data to find crime hotspots or books you might like? "Nothing. There is obviously immense promise there, as long as the data is kept to uses for which its limits are understood," Albro suggests. However, he worries that since so much of the data out there start out as "market research" — likes or dislikes when it comes to buying things — that removing data from advertising-focused troves and translating it into health care or planning for new roads or "culture" will essentially turn everyone into consumers, rather than citizens, in the minds of planners.
"We can't even agree on what 'culture' is, and now we're going to have 'culturomics.' Isn't that a little ambitious?" Albro asks. "Now we have claims that tweets predicted the 'Arab Spring,' which turns out to be questionable, or can detect 'sentiment' or 'mood,' which are even fuzzier or lazier words. We need to be a little cautious here." (In their defense, the "culturomics" study authors do urge caution on folks using their approach.)
But Albro worries that the hype over Big Data is warming up for Big Disappointment down the road. Much of the criticism of Big Data heard now focuses on privacy issues: who is using your data, or whether it will really help sell stuff. "Those are useful discussions, but we really need to talk a little more deeply about data," Albro says. "It's more than 'Garbage In, Garbage Out,' it's about how we shape the digital world."
Wednesday, May 23, 2012
New Report: Leveraging Data Analytics in Federal Organizations
Helena Sims and Steven Sossei, Corporate Partner Advisory Group, May 2012
“Data analytics is a powerful tool that can help government agencies reduce fraud, waste and abuse. The commercial sector has used data analytics for years to improve decision making, achieve better financial outcomes and improve customer service. The use of data analytics is growing at a rapid rate. The International Data Corporation, a provider of market intelligence in the information technology field, estimates that the business analytics market for software, hardware and consulting services is expected to grow at an 8 percent rate worldwide, reaching nearly $33 billion in 2012.
AGA set out to determine how the federal government is using data analytics and what it is doing with the resulting information…. we learned that some federal agencies have embraced data analytics and have demonstrated the benefits of integrating analytics tools into their operations. As a result, the federal government is in a position to build on the analytic advances it has already made. Some organizations are poised to share their capabilities with other federal organizations and possibly with other levels of government that implement federally funded programs. However, there is no clear plan to leverage the government’s investment in data analytics."
Subscribe to:
Posts (Atom)