Showing posts with label Transparency. Show all posts
Showing posts with label Transparency. Show all posts
Wednesday, January 16, 2013
Vivek Kundra: Release Data, Even If It's Imperfect
Vivek Kundra, former CIO of the federal government, says organizations will be more innovative if their data is rapidly released and shared, even if it’s imperfect.
Real-time data sharing promotes competitiveness and accountability, Kundra told CIO Journal Editor Michael Hickins at the Wall Street Journal CIO Network conference in San Diego. Soon after joining the federal government in 2009, Kundra created a Web-based tool that allows the public to track the progress of federal IT projects. “The default is that people will go after things that may not be 100% accurate,” Kundra said. “My view is it’s much, much better for the government to put out data that’s not 100% accurate, then to hold it in a secretive and opaque way.”
Kundra, now an emerging markets chief at Salesforce.com , said the release of government data is helping the private sector create a new wave of innovative apps, like applications that will help patients choose better hospitals. Those apps are built atop anonymized Medicare information. Kundra says he had a simple litmus test for assessing the risk of releasing data, which may include imperfections. “Assessing the risk, I asked a very simple question: were we using this data to make public policy decisions or investments. And inevitably, the answer would be yes,” Kundra said. “Then the question was, we’re using it, but why can’t let the American people see it? My view was let’s put it out there because we have to be able to trust that there’s somebody much smarter than us…it becomes a feedback loop that makes the data quality improve, and it makes the processes improve internally.”
Thursday, October 11, 2012
Walking the Talk: Philanthropy 'Does' Big Data
Bradford K. Smith, PhilanTopic, October 9, 2012
(Bradford K. Smith is president of the Foundation Center. In his last post, he took a closer look at the China Foundation Center's new Foundation Transparency Index.
With the modestly labeled "Reporting Commitment," fifteen of America's largest foundations are transforming the practice of philanthropy. From today on, information about their grants will be made available on a near-real-time basis, as entirely open data and coded to a common geographical standard, making it easy to see the communities, regions, and countries that benefit from those grants. The initiative's simple name should not deceive: this is big. The participants -- Annenberg, Carnegie, Gates, Getty, Hewlett, Packard, MacArthur, Mott, Robert Wood Johnson, and six others -- provide nearly 12 percent of the $46 billion in grants made by American foundations each year. To see the Reporting Commitment in action, take a quick look at Glasspockets, the transparency Web site of the Foundation Center, then read on.
What makes the Reporting Commitment so transformative? Let's break it down.
A Bold Idea -- Real-Time Reporting
The fifteen participating foundations have committed to electronically report their current grants data to the Foundation Center on at least a quarterly basis. As pragmatic as this may sound, it's a dramatic departure from the norm for the field. All the 76,000 private foundations in America file 990-PF tax returns in which they provide information on their grants. They have up to a year after the close of their fiscal year to file these returns, the Foundation Center eventually gets them from the IRS as image files and converts them into a more usable format, cleans and codes the data, and insures public access through databases and research reports. In a world where value is being created exponentially by analyzing enormous real-time data sets generated through search logs, consumer purchases, and Facebook "likes," philanthropy remains an industry with $640 billion in assets that relies on two-year old data to understand its own grant trends.
The Foundation Center has convinced more than seven hundred foundations to electronically submit their grants information through its eGrant Reporting Program, covering more than 20 percent of total foundation giving. Although this provides the field with current-year grants data, most participating foundations report on an annual basis. The Reporting Commitment takes this effort one important step forward by having participating foundations report at least quarterly -- with some reporting weekly, even daily.
A Radical Idea -- Open Data
By and large, foundations tend to think of open data and transparency as something they should fund rather than do. There are lots of reasons for this, including the private nature of foundations, the cultural legacy of keeping a low profile and "letting our good works speak for themselves," and sensitivity surrounding some of the issues addressed by foundation grants. Notwithstanding, the ability of foundations to not call attention to themselves is being steadily eroded by the ease of finding, displaying, and circulating information in a densely networked, digital age. Meanwhile, sectors and institutions with which foundations increasingly collaborate, such as the World Bank and foreign aid donors, are barreling ahead with initiatives like the Open Aid Partnership and Publish What You Fund.
The fifteen Reporting Commitment foundations have chosen to get ahead of the curve by taking the radical step of making their grants data entirely open. Under the agreement forged among them, they will either submit their data in machine-readable format or have the Foundation Center convert it so that it can be "harvested" by computers and used by developers to create apps, dashboards, visualizations, and things we haven't yet imagined. To make it easier, Glasspockets features a query builder that allows users to construct their own search and then "grab" the resulting data via an API.
A Strategic Idea -- GeoCoding
Some five years ago, when the Foundation Center started visualizing foundation grants data on interactive online maps, the most common reaction was, "You only show the location of the grantee organization, not the geographic focus of the grant." There was a reason for this: the vast majority of foundations, even those that electronically submit their data to the Foundation Center, do not include any coding for geographic area served. And even when there was a clue embedded in the grant description, there was no single standard that foundations used to describe the world; commonly used phrases such as "Deep South," "Middle East," and "developing countries" do not have agreed-upon definitions. That's why so many mapping visualizations (including our own) consign grants with insufficient or no geographic coding to big bubbles floating around in the ocean.
The Reporting Commitment foundations want to be able to compare their grants data with other participating foundations' data, from the community level all the way up to the continental level, to better identify gaps and areas of overlap and be more strategic about their giving. Thus they have agreed to use the GeoTree developed by the Foundation Center as an open geographic standard for use by philanthropy and the social sector. Geographic coding, or geocoding, as it is commonly known, requires a degree of specificity and decision making (i.e., how to handle grants that benefit multiple locations) that is something of a new discipline for most foundations. An interactive mapping tool on Glasspockets allows users to filter and search more than 3,800 grants by city/town, state/province, country, continent, or keyword. As participating foundations geocode more and more of their grants, the volume of data visualized on this map will expand.
A Mission-Critical Idea -- Transparency
When the Foundation Center was created in 1956 as a response to McCarthy-era hearings on philanthropy, transparency meant collecting printed reports from foundations and organizing them in file cabinets for public inspection. Today, it increasingly means open data. For an organization that has built a successful business model that relies on revenue from subscription databases to sustain an enormous volume of free information and services provided to more than nine million users, this may seem like risky business -- and it is. But the future of the Foundation Center requires disrupting its role as a data publisher. In the end, it is the Foundation Center's ability to analyze and combine multiple streams of information and analysis that adds value to data. And it is technology and networks that will allow the center to deliver knowledge into the hands of organizations and individuals who can leverage it to change the world.
Thanks to the vision, leadership, and hard work of the fifteen Reporting Commitment foundations, philanthropy has taken a crucial public step. Other foundations wishing to join the commitment can get started by contacting the Foundation Center. Later this year and again in 2013, the Foundation Center plans to release new and exciting forms of open data. While philanthropy may have been slow to get there, it is finally entering the era of Big Data.
-- Brad Smith
Monday, October 8, 2012
The Benefits of Open Data - Evidence from Economic Research
Guo Xu, Open Economics, October 3, 2012
This contribution is by Guo Xu (OKFN Economics and LSE) and the first part of the blog series “Mainstreaming Open Economics”.
Looking back to the Open Knowledge Festival 2012 in September, there’s an impression that openness is everywhere: There are working groups on Open Science and Open Linguistics, topic streams on Gender and Diversity in Openness, and events like Open Prom and Open Sauna: Open Knowledge and Open Data, it seems, is omnipresent.
Looking beyond the Open Knowledge community, however, the situation is very different: In Economics, for example, not many know what “open data”, “open access” or “Open Economics” exactly mean. Indeed, not many even care. A common reaction is: “Yes, it sounds interesting and important, but does it really matter? And why should I care about it?”
In this post, I would like to give some hard evidence on the positive role of opening up information has had in economics, and sketch ideas for how to involve economists – professional or in training – to mainstream ideas of openness. The blog post is divided into three parts: The first part looks at economic research on open data. The second part looks at the impact of open data on economic research. The third part discusses challenges and ways forward.
The real world impacts of open information
Making information accessible to the public can improve public service delivery. In countries where corruption is pervasive, services and funds often do not reach the frontline provider. And even if services do reach the people, the quality of services provided is often shockingly poor: Survey evidence from Bangladesh, Ecuador, India, Peru and Uganda found absence rates as high as 20% and 35% for school teachers and health workers. In many cases, the staff is poorly trained.
Releasing data on service delivery in this case can help reduce corruption and improve public services. In Uganda, researchers provided information to parents by publishing funding data for a random subset of schools in local newspapers. In consequence, corruption decreased significantly, while schooling outcomes improved substantially. Similar evidence in health delivery and redistributive policies suggest that providing information can help the public to discipline public service providers, improving the quality of services.
Information can also expose corrupt politicians: The Federal Government of Brazil, for example, began to select and audit municipalities at random, releasing audit reports to the media. Researchers found that the audit outcomes had a significant impact on the reelection probability of politicians: Those exposed for corruption were punished at the ballots, and the impact was most pronounced in areas where the dissemination of information was favoured by local radio.
A story from fishermen in South India provides another example of how information can improve market efficiency: Studying the adoption of mobile phones in Kerala, researchers have found convincing evidence that access to information through mobile phones helped fishermen sell their catch at the market where the price was highest (and fish most demanded): Instead of sailing to a port and simply hoping for a good price, fishermen were empowered by technology to make informed decisions on how to trade.
Finally, the benefits of transparency are not only restricted to reducing corruption and lowering the cost of information: A comparative study finds that transparency – measured by accuracy and frequency of macroeconomic information released to the public – leads to lower borrowing costs in sovereign bond markets. Open data pays off in many ways – in many different contexts.
These are just a few selective examples on how cutting-edge economic research has identified the benefits of openness in a diverse range of situations. The cases I presented are not based on correlations, but carefully established causal relationships, leaving – at least within the context studied – little doubt that information matters – big time. Perhaps most importantly, these cases have also shown that open data must be understood in a broad sense: These interventions do not take advantage of linked data, do not use CSVs that are shared through Facebook or Twitter – often, these interventions are simple solutions that ultimately help improving the everyday lives of the people.
Big Data: A Short History
How we arrived at a term to describe the potential and peril of today's data deluge.
Uri Friedman, Foreign Policy, November 2012
Humans have been whining about being bombarded with too much information since the advent of clay tablets. The complaint in Ecclesiastes that "of making many books there is no end" resonated in the Renaissance, when the invention of the printing press flooded Western Europe with what an alarmed Erasmus called "swarms of new books." But the digital revolution -- with its ever-growing horde of sensors, digital devices, corporate databases, and social media sites -- has been a game-changer, with 90 percent of the data in the world today created in the last two years alone. In response, everyone from marketers to policymakers has begun embracing a loosely defined term for today's massive data sets and the challenges they present: Big Data. While today's information deluge has enabled governments to improve security and public services, it has also sowed fears that Big Data is just another euphemism for Big Brother.
1887-1890
American statistician Herman Hollerith invents an electric machine that reads holes punched into paper cards to tabulate 1890 census data, revolutionizing the concept of a national head count, which had originated with the Babylonians in 3800 B.C. The device, which enables the United States to complete its census in one year instead of eight, spreads globally as the age of modern data processing begins.
1935-1937
President Franklin D. Roosevelt's Social Security Act launches the U.S. government on its most ambitious data-gathering project ever, as IBM wins a government contract to keep employment records on 26 million working Americans and 3 million employers. "Imagine the vast army of clerks which will be necessary to keep these records," Republican presidential candidate Alf Landon scoffs. "Another army of field investigators will be necessary to check up on the people whose records are not clear."
1943
At Bletchley Park, a British facility dedicated to breaking Nazi codes during World War II, engineers develop a series of groundbreaking mass data-processing machines, culminating in the first programmable electronic computer. The device, named "Colossus," searches for patterns in intercepted messages by reading paper tape at 5,000 characters per second -- reducing a process that had previously taken weeks to a matter of hours. Deciphered information on German troop formations later helps the Allies during their D-Day invasion.
1961
The U.S. National Security Agency (NSA), a nine-year-old intelligence agency with more than 12,000 cryptologists, confronts information overload during the espionage-saturated Cold War, as it begins collecting and processing signals intelligence automatically with computers while struggling to digitize a backlog of records stored on analog magnetic tape in warehouses. (In July 1961 alone, the agency receives 17,000 reels of tape.)
1965-1966
The U.S. government secretly studies a plan to transfer all government records -- including 742 million tax returns and 175 million sets of fingerprints -- to magnetic computer tape at a single national data center, though the plan is later scrapped amid public concern about bringing "Orwell's '1984' at least as close as 1970," as one report puts it. The outcry inspires the 1974 Privacy Act, which places limits on federal agencies' sharing of personal information.
1989
British computer scientist Tim Berners-Lee proposes leveraging the Internet, pioneered by the U.S. government in the 1960s, to share information globally through a "hypertext" system called the World Wide Web. "The information contained would grow past a critical threshold," he writes, "so that the usefulness [of] the scheme would in turn encourage its increased use."
August 1996
"We are developing a supercomputer that will do more calculating in a second than a person with a hand-held calculator can do in 30,000 years." --U.S. President Bill Clinton
1997
NASA researchers Michael Cox and David Ellsworth use the term "big data" for the first time to describe a familiar challenge in the 1990s: supercomputers generating massive amounts of information -- in Cox and Ellsworth's case, simulations of airflow around aircraft -- that cannot be processed and visualized. "[D]ata sets are generally quite large, taxing the capacities of main memory, local disk, and even remote disk," they write. "We call this the problem of big data."
2002
After the 9/11 attacks, the U.S. government, which has already dabbled in mining large volumes of data to thwart terrorism, escalates these efforts. Former national security advisor John Poindexter leads a Defense Department effort to fuse existing government data sets into a "grand database" that sifts through communications, criminal, educational, financial, medical, and travel records to identify suspicious individuals. Congress shutters the program a year later due to civil liberties concerns, though components of the initiative are simply shifted to other agencies.
2004
The 9/11 Commission calls for unifying counterterrorism agencies "in a network-based information sharing system" that is quickly inundated with data. By 2010, the NSA's 30,000 employees will be intercepting and storing 1.7 billion emails, phone calls, and other communications daily. Meanwhile, with retailers amassing information on customers' shopping and personal habits, Wal-Mart boasts a cache of 460 terabytes -- more than double the amount of data on the Internet at the time.
2007-2008
As social networks proliferate, technology bloggers and professionals breathe new life into the "big data" concept. "This is a world where massive amounts of data and applied mathematics replace every other tool that might be brought to bear," Wired's Chris Anderson writes in "The End of Theory." Government agencies, some of the United States' top computer scientists report, "should be deeply involved in the development and deployment of big-data computing, since it will be of direct benefit to many of their missions."
January 2009
The Indian government establishes the Unique Identification Authority of India to fingerprint, photograph, and take an iris scan of all 1.2 billion people in the country and assign each person a 12-digit ID number, funneling the data into the world's largest biometric database. Officials say it will improve the delivery of government services and reduce corruption, but critics worry about the government profiling individuals and sharing intimate details about their personal lives.
May 2009
U.S. President Barack Obama's administration launches data.gov as part of its Open Government Initiative. The website's more than 445,000 data sets go on to fuel websites and smartphone apps that track everything from flights to product recalls to location-specific unemployment, inspiring governments from Kenya to Britain to launch similar initiatives.
July 2009
Reacting to the global financial crisis, U.N. Secretary-General Ban Ki-moon pledges to create an alert system that captures "real-time data on the impact of the economic crisis on the poorest nations." The U.N. Global Pulse program has conducted research on how to predict everything from spiraling prices to disease outbreaks by analyzing data from sources such as mobile phones and social networks.
August 2010
"There were 5 exabytes of information created by the entire world between the dawn of civilization and 2003. Now that same amount is created every two days." --Google CEO Eric Schmidt
February 2011
Scanning 200 million pages of information, or 4 terabytes of disk storage, in a matter of seconds, IBM's Watson computer system defeats two human challengers in the quiz show Jeopardy!. The New York Times later dubs this moment a "triumph of Big Data computing."
March 2012
The Obama administration announces a $200 million Big Data Research and Development Initiative in response to a U.S. government report calling for every federal agency to have a "'big data' strategy." The National Institutes of Health puts a data set of the Human Genome Project in Amazon's computer cloud, while the Defense Department pledges to develop "autonomous" defense systems that can "learn from experience." CIA Director David Petraeus, marveling that the "'digital dust' to which we have access is being delivered by the equivalent of dump trucks," discusses a post-Arab Spring agency effort to collect and analyze global social media feeds through cloud computing.
July 2012
U.S. Secretary of State Hillary Clinton announces a public-private partnership called "Data 2X" to collect statistics on women and girls' economic, political, and social status around the world. "Data not only measures progress -- it inspires it," she explains. "Once you start measuring problems, people are more inclined to take action to fix them because nobody wants to end up at the bottom of a list of rankings." Let the Big Data race begin.
Saturday, October 6, 2012
Predicting the Tuture through Online Data Mining
Santiago Zabala, Al Jazeera, October 5, 2012
It is often said philosophers are either late when it comes to comment upon new technological innovations or in advance, that is, so early that they actually seem to predict them. When they are late, it's usually because they prefer to carefully examine the new innovations in order to achieve insightful analysis, and when they foresee such discoveries, it arises from an ethical concern over the direction the world is taking.
In other words, their insight is not expressed by envisioning the day a software company manages to predict the future by scanning information from the internet and therefore framing our freedom, but rather by working through the existential consequences these innovations might have upon our life.
This is probably why among Martin Heidegger's greatest concerns when it came to technological innovations was the formation of existential conditions where, as he said, the "lack of emergency is the only emergency". In this condition human beings would be completely "uprooted" from the earth, that is, "framed" ("Ge-stell") by a technological power they are no longer able to control.
As it turns out, the software company Recorded Future (which has recently been praised by Wired , the MIT Technology Review and other media outlets, after the CIA and Google invested millions in their services) seems to be offering its clients something similar: a world where emergencies, that is, future events, can be calculated in advance. But how does this start-up actually function, and why are the German philosopher’s concerns relevant to its services?
Recorded Future
Recorded Future is based in Gothenburg and has offices in London, Boston, Arlington and New York. A team of 20 computer scientists, statisticians and experts in linguistics "calculate" the future. While Yahoo, Google and Bing use links to connect and rank different web pages, Recorded Future goes further by scouring (in real time) thousands of available information sources such as blogs, websites and Twitter comments in order to find "invisible links", that is, relationships among actions, people and institutions that refer to related events in the future.
Even though this might not seem particularly relevant, considering that we can also predict next week's weather by searching through different weather stations, if we look at the amount of data this company is capable of analysing and relating in just a few hours, it becomes clear it can obtain better data than public internet users have access to.
The information we need to predict whether tomorrow it will rain is limited by the number of weather sites available, but sites that might refer to upcoming anti-American demonstrations in the Middle East are infinitely more numerous given the political, economic and military aspects of these sorts of events.
After mining from the web all the related people ("Bashar al-Assad"), places ("Syria") and activities ("military interventions") that refer to a possible demonstration, Recorded Future uses algorithms to predict when and where a demonstration will occur.
An example of a predicted demonstration is available in a video on the company website which illustrates how its powerful engines monitor these protests not only in the Middle East, but also in South America and North Africa. The fact that Google and the CIA have already invested millions in this company is an indication that it will be used to conserve certain interest against others as the example above indicates.
The different fee levels for the customers of Recorded Future are probably related to the quality and quantity of information they wish to purchase, making this, and similar companies, at the service of the wealthiest and most powerful.
Lack of emergencies
From a philosophical point of view, the most interesting feature of all this is not that these demonstrations can be predicted, but rather how technology has finally uprooted and dislodged man from the world, that is, has given human existence to a power beyond human control. The secured, comfortable and calculated environment that allowed the creation of society has now become so functional and rationalised that we cannot help but become victims by existing in it.
This existential dilemma does not arise from the fact that it’s finally possible to organise all the things the web already knows about the future, which could certainly become useful to prevent diseases or famine, but rather that a private company now means to know everything, that is, all human projects.
We have entered an age where only those framed within the approved interests of Recorded Future clients will be able to live freely, that is, without being predicted. But how free is an existence that is completely revealed to the modern "lack of emergencies"?
As Heidegger explained , emergencies do not arise when something doesn't function correctly, but rather when "everything functions … and propels everything more and more toward further functioning". It's within this logic that as soon as something critical to the interests of those who can afford it fails to function, Recorded Future will alert its customers, who will then take the appropriate measures to conserve the previous condition.
In sum, Heidegger's concerns over a world lacking "emergencies" more than 50 years ago was meant to point out how technologies such as that employed by Recorded Future (and similar companies ) aim to avoid the future, that is, to change the world.
Santiago Zabala is ICREA Research Professor of Philosophy at the University of Barcelona. His books include The Hermeneutic Nature of Analytic Philosophy (2008), The Remains of Being (2009) and most recently, Hermeneutic Communism (2011, co-authored with G Vattimo), all published by Columbia University Press.
Thursday, September 27, 2012
New Talk by Clay Shirky -
Clay Shirky’s Ted Talk “How the Internet will (one day) transform government”
The open-source world has learned to deal with a flood of new, oftentimes divergent, ideas using hosting services like GitHub -- so why can’t governments? In this rousing talk Clay Shirky shows how democracies can take a lesson from the Internet, to be not just transparent but also to draw on the knowledge of all their citizens.
Clay Shirky argues that the history of the modern world could be rendered as the history of ways of arguing, where changes in media change what sort of arguments are possible -- with deep social and political implications
Subscribe to:
Posts (Atom)