Showing posts with label Open Data. Show all posts
Showing posts with label Open Data. Show all posts

Friday, January 18, 2013

Facebook's Other Big Disruption


Quentin Hardy, The New York Times, January 17, 2013

Facebook just made a potentially game-changing announcement. It got less fanfare than Tuesday’s announcement that it is going into the social search business, but this other announcement may have bigger long-term implications for the technology industry.

Put simply, some of the world’s biggest computing systems just got a little cheaper, and a lot easier to configure. As a consequence, the companies that supply the hardware to these systems may have to scramble to remain as profitable. The reason is a Facebook-led open source project.

In 2011 Facebook began the Open Compute Project, an effort among technology companies to use open-source computer hardware. Tech companies similarly shared intellectual property with Linux software, which lowered costs and spurred innovation. Facebook’s project has attracted many significant participants, including Goldman Sachs, Arista Networks, Rackspace, Hewlett-Packard and Dell.

At a user summit on Wednesday Intel, another key member of the Open Compute Project, announced it would release to the group a silicon-based optical system that enables the data and computing elements in a rack of computer servers to communicate at 100 gigabits a second. That is significantly faster than conventional wire-based methods, and uses about half the power.

More important, it means that elements of memory and processing that now must be fixed closely together can be separated within a rack, and used as needed for different kinds of tasks. There is a lot of waste in data centers today simply because, when there is an upgrade in servers, lots of other associated data-processing hardware has to be changed, too.

There were other announcements, like a computer motherboard called Grouphug that allows different manufacturers’ chips to be interchanged without altering other parts of the machine. Before, they were custom made. Put together, such innovations potentially lower the cost and complexity of running big and small data centers to an extent that works for a lot of companies.

“Who wouldn’t want a cheaper, more efficient server?” said Frank Frankovsky, vice president of hardware design at Facebook, and the chairman of Open Compute. “The problem we’re solving is much larger than Facebook’s own challenges. There is a massive amount of data in the world that people expect to have processed quickly.”

To be sure, it’s in Facebook’s interest to attack expensive hardware. The company makes money from a service that requires hundreds of thousands of computer servers distributed in big centers around the world. Google and Amazon.com, which are not members of the project, maintain proprietary systems which they apparently felt gave them a competitive edge.

For Facebook, the difference seems to be more in the software. To the extent hardware costs drop, that’s great for them. Mr. Frankovsky argued that, while “this puts challenges on the incumbents” in hardware, “it also helps them. They have a finite number of engineering resources, and this way they hear from a community about whether there is an interest for a product.” Intel may hope to benefit from its open-source release, since it could see an overall rise in demand for its chips with the move toward cheaper computing.

The real test is whether Facebook can increase the number of potential buyers for Open Compute equipment. “The question is, can they extend this beyond a few Web businesses like Facebook and Rackspace, or a few financial exercises at Goldman, and bring this to industries like oil or aerospace?” said Matt Eastwood, an analyst with IDC, a technology research firm. “That will take it from 20 or 30 companies to hundreds of companies.”

The issue isn’t so much a technical one, he argues, as it is one of getting corporate information technology professionals interested in radical design changes. Mr. Frankovsky is aware of the problem. Recently he and his colleagues led a seminar in Texas for BP, Shell and other oil giants on how they could use Open Compute hardware in their data centers.

This will not change things dramatically this year, and possibly even next, but over the long haul it could remake a lot of businesses. Linux, remember, was around for several years as a minor player, but eventually undid Sun Microsystems and others.

Thursday, January 17, 2013

Open Access: Aaron Swartz's illusion over research


John Gapper, The Financial Times, January 16, 2013

Deleting private companies from the equation might allow savings but could reduce efficiency
The death of the internet activist Aaron Swartz at the age of 26 has rightly evoked tributes to his creativity and selflessness. Swartz, who faced jail for illegally downloading millions of academic papers from an electronic library, committed suicide last week.

Five years ago, Swartz signed a “guerrilla open access manifesto” in which he complained of “the world’s entire scientific and cultural heritage” being “digitised and locked up by a handful of private corporations” such as Reed Elsevier. He advised computer hackers to “take information, wherever it is stored, make our copies and share them with the world”.

In 2010, he disguised his identity and exploited the electronic network of the Massachusetts Institute of Technology to download most of the database of Jstor, a non-profit group that digitises academic journals and articles. He did not share or sell the material – he later handed it back – but prosecutors took the manifesto seriously and charged him with fraud.

Mr Swartz worked on projects from the news aggregator Reddit to the Creative Commons open copyright licence, and was widely liked and admired. But, in his analysis of academic research and publishing, he suffered from an illusion.

Free access to academic research – the system Mr Swartz advocated – could bring public benefits. It would enable anyone to read, analyse and build upon privately and publicly funded research. However, someone would still need to pay for it and the costs to universities such as MIT and Oxford would rise, not fall.

Critics of the current system, under which research libraries pay up to $50,000 annually to use online databases, tend to blame profiteering by companies such as Reed Elsevier and Springer for this cost. George Monbiot, the activist and Guardian writer, describes it as “pure rentier capitalism”, arguing that people should “throw off these parasitic overlords and liberate the research that belongs to us”.

Allied to this is the belief that publishing costs have fallen heavily in the shift from print to digital. Elsevier, the scientific publishing arm of Reed Elsevier, made profits of £352m on revenue of £978m in the first half of 2012 – an operating margin of 36 per cent. Remove the capitalists and distribute research through public utilities, and surely swaths of cost would disappear?

Well, perhaps. Elsevier could certainly do with a bit more competition. Its fee structure is opaque and it publishes journals in which academics vie to be published. It has what Warren Buffett calls a moat – it is a 130-year-old business with 20 per cent of the market that is hard to attack.

It did not, however, steal this advantage. It acquired it from the 1960s and 1970s onwards as research universities saved money by outsourcing their costly and subscale publishing presses. Elsevier employs 7,000 editors, manages a network of some 500,000 peer reviewers (whom it does not pay), publishes 300,000 new articles a year and runs a 100-terabyte database.

Printing is only a small part of the cost of academic publishing. The bulk lies in the labour-intensive business of editing and reviewing submissions (rejecting two-thirds of them) and managing data. These costs are similar for open access publishers such as the Public Library of Science (Plos) in San Francisco, a competitor to Elsevier.

An independent study by the Research Information Network in the UK found that the shift from print to digital may save £1bn globally – worth having but only 12 per cent of total costs. Removing private companies from the equation might allow further savings, but it might equally reduce efficiency.

In any case, there will still be a hefty bill. About 90 per cent of the industry operates on subscription – the model Swartz so hated. The other 10 per cent is now open access, under which researchers (or research funders) have to pay journals between $1,000 and $5,000 an article to cover publishing costs. Anyone can then read it free.

Open access is appealing and is supported both by research funds, such as the US National Institutes of Health and the UK Wellcome Trust, and by the UK government. The trust believes it makes no sense to invest £700m each year on research without paying an extra £10m to make it widely available.

Research is largely read by other academics at the moment, most of whom have access through libraries. But there could be big benefits to broadening reach – Plos One, the science journal, is a trove of fascinating material.

That said, open access mostly transfers the bill. The Research Information Network estimated that, if the market moves to 90 per cent open access, total costs would fall by £560m but universities would pay more. The UK would save £128m in library subscriptions but contribute £213m in fees because its universities publish a lot of research.

Open access also has its pitfalls. In the 1970s the credit rating industry turned from investors subscribing to ratings to bond issuers paying. That established open access but also gave agencies a motive to please issuers with good ratings, culminating in the triple A rating of flimsy mortgage-backed securities.

Open access journals have a similar incentive to widen access and dilute quality. It is worth noting that Plos One publishes 24,000 pieces of research every year – it accepts any submission that meets the hurdle of “valid science” – while the most prestigious journals (including other Plos titles) publish 200.

If Swartz’s sad death shifts the balance further toward open access, that will be a worthy legacy. But someone will always pay.

john.gapper@ft.com



Wednesday, January 16, 2013

Vivek Kundra: Release Data, Even If It's Imperfect



Vivek Kundra, former CIO of the federal government, says organizations will be more innovative if their data is rapidly released and shared, even if it’s imperfect.

Real-time data sharing promotes competitiveness and accountability, Kundra told CIO Journal Editor Michael Hickins at the Wall Street Journal CIO Network conference in San Diego. Soon after joining the federal government in 2009, Kundra created a Web-based tool that allows the public to track the progress of federal IT projects. “The default is that people will go after things that may not be 100% accurate,” Kundra said. “My view is it’s much, much better for the government to put out data that’s not 100% accurate, then to hold it in a secretive and opaque way.”

Kundra, now an emerging markets chief at Salesforce.com , said the release of government data is helping the private sector create a new wave of innovative apps, like applications that will help patients choose better hospitals. Those apps are built atop anonymized Medicare information. Kundra says he had a simple litmus test for assessing the risk of releasing data, which may include imperfections. “Assessing the risk, I asked a very simple question: were we using this data to make public policy decisions or investments. And inevitably, the answer would be yes,” Kundra said. “Then the question was, we’re using it, but why can’t let the American people see it? My view was let’s put it out there because we have to be able to trust that there’s somebody much smarter than us…it becomes a feedback loop that makes the data quality improve, and it makes the processes improve internally.”

Friday, December 21, 2012

The Best Open Data Releases of 2012


Picture (Device Independent Bitmap)Emily Badger, The Atlantic, December 19, 2012

Last year, Cities named ten of its favorite metro datasets of 2011 from cities across North America, illustrating the breadth of what we might learn (regarding mosquito traps! misplaced vehicles! energy consumption!) in the still relatively young field of urban open data. For this year's installment, we're going one step further. Sure, raw data is great. But useful tools, maps and data visualizations built with said data are even better.

Below, you'll find our picks for 2012's best open data releases from municipal vaults, with an emphasis on tools that can be used by anyone, not just developers and data geeks. If we missed your favorite, please add it in the comments.

1. Crime in Philadelphia. Philadelphia snuck onto our 2012 list just under the wire, publishing last week a big data set on all major crimes in the city dating back to January 1, 2006. The data is now updated daily, covering incidents of homicide, rape, robbery, assault and theft. Philadelphia now joins Chicago, which released 10 years of crime data last year. Baltimore has a similar dataset. For Philadelphians more interested in live trends than historic ones, the city is now also mapping recent crimes (see the above graphic). A smart bonus feature: when you click on an individual incident, the map gives you an opportunity to "submit a tip" to the Philadelphia Police Department.

2. Bikeshare rides in Boston. Boston’s Hubway bikeshare system published a massive file of historic trip data earlier this year, then invited riders and developers to turn the information into something useful with a data visualization challenge. This map comes from one of the winners, Ari Ofsevit, showing the average speeds across different routes between bikeshare stations.

Capital Bikeshare in Washington also publishes trip data. Nice Ride in Minneapolis did so earlier this year as well, although that release ran into privacy complications when it turned out the anonymoized data wasn’t so anonymous after all.

3. Public transit in Atlanta. Earlier this year, we wrote about the handful of large metros in the U.S. that were still not opening up their GTFS files of public transit data to anyone other than Google. Atlanta was one of the notable holdouts. In October, however, the Metropolitan Atlanta Rapid Transit Authority finally published this feed, making it possible for developers anywhere – and not just Google Maps – to produce apps, maps, trip planners or other tools with the city’s real-time transit data.

4. Blight in New Orleans. Code for America helped the city build a web tool this year mapping and tracking blighted properties all over town to help neighbors and community groups keep track of the status of abandoned and code-violating properties. With Blight Status, residents can for the first time follow blighted properties through the process of inspections, legal hearings, judgment and resolution. As a result, the city is able to convey that it’s actually working on the problem, while residents are given some reassurance of that progress.

5. Pedestrian injuries in San Francisco. San Francisco’Department of Public Health publishes a slew of data and maps (see their great Sustainable Communities Indicators website tracking everything from air quality to food access). One project in particular has had a significant impact. San Francisco’s High-Injury Corridors map tracks data on pedestrian injuries across the city. But instead of mapping individual collisions, the tool weights pedestrian deaths more heavily than other injuries and highlights injury-prone corridors instead of intersections. “The story that this map tells is that 5 percent of the city streets account for 55 percent of the injuries weighted for severity,” says Rajiv Bhatia, director of environmental health for the city’s Department of Public Health. “That was a transformative map to both the pedestrian safety stakeholders, to the police department and the transportation agency.” This is the map he’s talking about:
Picture (Device Independent Bitmap)
San Francisco High-Injury Corridors map.

“The transportation agency realized that if you overlay this map with a map of where the city of San Francisco has made their traffic calming and traffic safety investments,” Bhatia adds, “you’ll see almost no overlap.”
6. Green roofs in Chicago. The city has identified 359 vegetated roofs across town, with a total surface area of more than 5 million square feet. The city’s data portal now publishes data complete with location, dimensions and satellite imagery for all of them, as well as a map of their locations. Perhaps the most well-known green roof in town? This one, above Chicago’s City Hall:

Chicago City Hall, via the city's green roofs map.

7. Rat sightings in New York City. No, New York doesn't keep a special call-in line or data log just for rat spottings across the city. This data comes instead from the broader database of 311 service requests to city hall over the past three years. The good news? "This information is automatically updated daily." And you can use the New York data portal to map the results. But we're holding out for someone else to do this with a bunch of tiny rat icons.
8. Tsunami sirens in Honolulu. Honolulu published data on the location of dozens of warning sirens around the island, and Code for America helped to build an app on top of the data allowing local citizens to "adopt" a siren in the same way that other communities have created adopt-a-hydrant programs. In this case, instead of volunteering to dig out hydrants during snow storms, Honolulu residents can take responsibility for listening to siren tests and reporting problems ahead of any Tsunami. As this map shows, about half of the sirens (those in green) have already been adopted:
Picture (Device Independent Bitmap)
Honolulu Adopt-a-Siren.

9. Dangerous dogs in Austin. "Declared Dangerous Dogs" in Austin are court-ordered to be restrained at all times and are required to wear large tags identifying them as such. "They have attacked in the past," warns the city. "If they attack again the court could order them put to sleep." Want to know where they live? The city has now mapped them, complete with useful dog descriptions. Watch out, for instance, on Daleview Drive for "Nibbles," a female red-and-tan Golden Retriever/Chow mix.

10. Fixed speed cameras in Baltimore. Baltimore has dozens of these things around town waiting to snap photos of aggressive drivers. The penalty? A $40 mailed citation for going more than 12 miles per hour over the speed limit. But at least the city is up-front about where these cameras are located. This map and dataset on Baltimore's open-data portal identifies the intersection, coordinates and even driving direction (southbound, eastbound, etc.) for all of them. We can imagine such data might come in handy in any number of navigation apps.

Friday, December 7, 2012

Panjiva Uses Government Data to Build a Global Search Engine for Commerce


Successful startups look to solve a problem first, then look for the datasets they need.

Alex Howard, O'Reilly Strata, December 6, 2012


“If you go back to how we got started,” mused Josh Green, “government data really is at the heart of that story.” Green, who co-founded Panjiva with Jim Psota in 2006, was demonstrating the newest version of Panjiva.com to me over the web, thinking back to the startup’s origins in Cambridge, Mass.

At first blush, the search engine for products, suppliers and shipping services didn’t have a clear connection to the open data movement I’d been chronicling over the past several years. His account of the back story of the startup is a case study that aspiring civic entrepreneurs, Congress and the White House should take to heart.

“I think there are a lot of entrepreneurs who start with datasets,” said Green, “but it’s hard to start with datasets and build business. You’re better off starting with a problem that needs to be solved and then going hunting for the data that will solve it. That’s the experience I had.”

The problem that the founders of Panjiva wanted to help address was one that many other entrepreneurs face: how do you connect with companies in far away places? Green came to the realization that a better solution was needed in the same way that many people who come up with an innovative idea do: he had a frustrating experience and wanted to scratch his own itch. When he was working at an electronics company earlier in his career, his boss asked him to find a supplier they could do business with in China.

“I thought I could do that, but I was stunned by the lack of reliable information,” said Green. “At that moment, I realized we were talking about a problem that should be solvable. At a time when people are interested in doing business globally, there should be reliable sources of information. So, let’s build that.”
Today, Panjiva has created a higher tech way to find overseas suppliers. The way they built it, however, deserves more attention.

Government data as a platform
By 2009, the startup had an initial product they could bring to market and launched a search engine that used government data as a platform for international trade. An importer could type in “patio furniture” and
determine who shipped it and who their customers were. The company chose a freemium model, where search is available for free but relationships between suppliers are only available to subscribers. The mapping of relationships between buyers and suppliers is where Panjiva delivered added value on top of public data.

That added value is crucial, given that competitors can also request and use the dataset. “Companies have been packaging and reselling this data in one way or another for years, ” said Green. “If you looked at this data, people are going to find value. It’s typically folks in the shipping industry, who want to know what’s going into ports or moving on different shipping lines. For us, the central purpose of the data was something different and required more work.”

That work paid off. In 2010, Panjiva built a search engine for global commerce that worked. Today, they have more than 100,000 users in 190 countries using its free service and some 3,700 companies subscribing to the paid version, including 42 Fortune 500 companies.

Notably, the Department of Homeland Security (DHS) itself is also a paying subscriber. Green declined to disclose the terms of relationships with all of Panjiva’s partners or data suppliers, some of which include nonprofits. Some users “do a revenue share, some are paying for data, others are providing data because they think there’s public good for that data being on the platform,” he said.

Panjiva competes with ImportGenius, Zepol, AliBaba and PIERS. Green credits PIERS for extracting similar value from customs datasets.

The turning point
When they started, the first approach that Green and his co-founder decided to take was to build a “Yelp for global trade” that would be based on feedback from people who work with companies. Unfortunately for the young startup, they couldn’t get off the starting blocks in generating reviews, much less reach critical mass.

They also encountered a new problem: even if they were able to get ratings of exporters and suppliers, how would they ensure the reviews came from people who had actually done business with the entities being rated? In retrospect, that focus was a bit silly, said Green, because they couldn’t get engagement, but talking about how to solve it led them to an unexpected answer: government data.

That direction came from a meeting where a staffer for a trade promotion organization told them it was straightforward to get shipping data on what’s coming into the country from the United States Customs Agency, which is now part of the Department of Homeland Security.

“It was a turning point for the company,” said Green. “We realized there was a dataset available to the public for a fee. They make available data about shipments that enter the U.S. While not all data is made available to the public and there are a bunch of limitations, the data that is made available is amazing. There’s about 10 million shipping records every year, typically including who is sending goods, who is receiving goods, what’s inside, and how much is inside a container.”

While useful, these government datasets do come with inherent limitations, cautioned Green. For one, they only contain data about shipments coming into the United States, not what’s going into Europe or Asia. For another, the data made available to the public only covers shipments made by boat, which is about about half of the shipments that come into the United States.

“It’s unfortunate that government cannot make available data on other modes of transport,” observed Green, with a hint of frustration in his voice. “That leaves out truck, rail, and air. Congress actually attempted to clarify that the regulations that govern this data weren’t just about boats but applied to air. Thus far, DHS hasn’t acted.”

Given the lens that has been focused on trade deficits between other countries and the United States in recent decades, there’s also a political angle to the market intelligence Panjiva provides that Congress and taxpayers may find of interest. For instance, Panjiva data showed global trade growth slowing in the first part of 2012.

“What we’ve organized, by its nature, gives us insight on companies around the world that serve the U.S. market,” said Green, “We’re helping people find overseas suppliers. Why not help find suppliers here at home? It turns out there’s a similar story on export data that’s supposed to be made to the public as well. DHS has a hard time with that as well. We can’t get the data.”

Data availability is also affected by the actions of the companies themselves, which have the ability to petition the government to hide shipments that are coming to them. “In about a third of the cases, you cannot see who is sending and receiving the goods,” said Green. “Government can see, but what’s released to the public has information pulled from it.”

This government data comes at a cost
Accessing this public data comes at a cost of some $100 per day, which is the service fee DHS charges for providing a daily CD-ROM. Each disc includes one day’s worth of shipments, which is generally around 30,000 shipping records. Panjiva started requesting data on July 1, 2007, and now has a little over five years of records.

“This data, on a record-by-record basis, is interesting,” said Green. “If you can organize, it’s phenomenal. If you can associate with companies, can say this company has experience with these supplies and this company has experience with these customers, it’s very useful in deciding if a company is a good fit. You can see by customers if they’re reasonably high quality.”

Making those CD-ROMs into a useful, searchable resource, however, was far from a simple matter of just inserting them into an optical drive and moving their contents into a structured database.

“Jim and a team of engineers went to work organizing the datasets initially,” said Green. “They were very hard to work with — absurdly messy. Think about the number of ways you can misname a Chinese factory. It was really problematic. You need to build company profiles, correct for misspellings and variations on names. We spent years getting that right.” Eventually, Panjiva was able to automate the process of ingesting the data from the CD-ROMs, building an algorithm to take the data and clean it up.

Making data a strategic asset
Panjiva’s initial foray, which created a search engine for customs data, didn’t meet with strong demand out of the gate. As they refined the product, it generated what Green described as a “nice business.” The startup was profitable, in other words, but its leadership aspired to build something bigger.

The direction they took was driven by user feedback. When Panjiva also asked its users about how they were making buying decisions, they saw a pattern emerge that looked like a bigger opportunity.

“Users started with Panjiva then went to search for additional information on B2B sites or on Google,” said Green. “We heard this process and it sounded a lot like the experience consumers had searching for flights before search engines or Kayak.com — except that instead of airline sites, people are going to B2B sites. The difference is it’s not just every airline. It’s like every flight has its own website.”

The founders now have raised just under $10 million from Battery Ventures and Harrison Metal, and invested it in technology and data acquisition. They’ve now grown their engineering team to 10 people, out of a total of 50 or so current employees. The engineering team is focused on improving search and enriching Panjiva’s data with other sources, beyond government data.

This October, the startup relaunched Panjiva.com with another layer: data supplied by the companies themselves.

“We have a database of six million companies spread around the world and contact information on four million companies,” said Green. “We have product photos for 34 million products. There was a lot of investment required to do that, but none of this would be possible if we hadn’t had a backbone of data that came from the U.S. government.”


Since Panjiva added global search, Green said that traffic to the search engine has gone up 50%.
The data sources that Panjiva integrated were also driven by customer interest. As the founders shared their product with potential subscribers, they kept hearing the same thing: 1) “that’s awesome” and 2) “I’d like more data.”

“We loved the first one and hated the second,” said Green. “In retrospect, we should have loved both. The second one was a roadmap for us to build them a really great differentiated product.”

When they asked users exactly which kinds of data would make the service more useful, a map to the future of the company emerged.

The first was operational data. “Customs data is a perfect example,” said Green. “It gives you a sense of what companies have done and their track record.”

The second was financial data. “Sure, a company has experience, but are they financially healthy?” asked Green. “Some of that you can infer, but there’s other things you can use. We’ve partnered with Dun & Bradstreet and Experian to pull that data into our platform.”

The third was positive and negative data about a company. “That includes getting certified as financially responsible,” said Green. “We’ve partnered with nonprofits and added that data, showing you information about companies doing wrong, including a blacklist of illicit global trade.”

The key insight that anyone interested in building a business on top of government data should take away here is to go beyond.

What happens if government data becomes open?
Green thinks that Panjiva is well-positioned to be both competitive and profitable, even if DHS decided to start publishing customs data online. “We don’t worry that much about data becoming more accessible,” he said, “even if government data becomes free. It’s not the $36,500 per year to buy the data — it’s the engineering talent to clear it up. That’s a massive problem, and it wouldn’t be as simple as getting the data.”
Panjiva is betting that the investments they’ve made in technology, talent and — crucially — combining so many different data sources have created a differentiated product that solves a problem for its customers.

“We’re not trying to build out a data business where we’re reselling government data,” said Green. “We’re trying to build a platform where serious buyers and sellers can connect. We’re now going to the world’s most important buyers. We have two revenue streams: selling premium access to data and selling access to suppliers who want it. The starting point for customers is $99 per month, going up to $10,000 per month for unlimited access for an unlimited number of users, then services that we sell on the top.”

The experience that Panjiva has had with government data and building a business using it has left Green with a strong perspective on what works — and what doesn’t.

“We don’t think there are infinite numbers of possibilities in terms of ways to build sustainable value with public data,” he said. “One is to take datasets that are commoditizeable and add value. Another is to feed the creation of more data. Another is to build a service. Another is to create network effects, where the data is the honey that attracts the bees.”

Most important, Green suggested, is to use public data to solve a problem that’s both hard and important. For Panjiva, that means making global trade more efficient and more transparent.

“There is a future where information is consolidated and accessible to people making key decisions, from a buying or regulatory standpoint,” he said. “Once that happens — and we’re close — there’s potentially a place where there’s a race to the top instead of the bottom, in terms of supply chain records. That will make a difference when you’re under scrutiny. Right now, the fragmentation of data is the ally of bad behavior. Our hope is to change that reality.”

Thursday, October 11, 2012

Walking the Talk: Philanthropy 'Does' Big Data


Bradford K. Smith, PhilanTopic, October 9, 2012

(Bradford K. Smith is president of the Foundation Center. In his last post, he took a closer look at the China Foundation Center's new Foundation Transparency Index.

With the modestly labeled "Reporting Commitment," fifteen of America's largest foundations are transforming the practice of philanthropy. From today on, information about their grants will be made available on a near-real-time basis, as entirely open data and coded to a common geographical standard, making it easy to see the communities, regions, and countries that benefit from those grants. The initiative's simple name should not deceive: this is big. The participants -- Annenberg, Carnegie, Gates, Getty, Hewlett, Packard, MacArthur, Mott, Robert Wood Johnson, and six others -- provide nearly 12 percent of the $46 billion in grants made by American foundations each year. To see the Reporting Commitment in action, take a quick look at Glasspockets, the transparency Web site of the Foundation Center, then read on.


What makes the Reporting Commitment so transformative? Let's break it down.

A Bold Idea -- Real-Time Reporting
The fifteen participating foundations have committed to electronically report their current grants data to the Foundation Center on at least a quarterly basis. As pragmatic as this may sound, it's a dramatic departure from the norm for the field. All the 76,000 private foundations in America file 990-PF tax returns in which they provide information on their grants. They have up to a year after the close of their fiscal year to file these returns, the Foundation Center eventually gets them from the IRS as image files and converts them into a more usable format, cleans and codes the data, and insures public access through databases and research reports. In a world where value is being created exponentially by analyzing enormous real-time data sets generated through search logs, consumer purchases, and Facebook "likes," philanthropy remains an industry with $640 billion in assets that relies on two-year old data to understand its own grant trends.

The Foundation Center has convinced more than seven hundred foundations to electronically submit their grants information through its eGrant Reporting Program, covering more than 20 percent of total foundation giving. Although this provides the field with current-year grants data, most participating foundations report on an annual basis. The Reporting Commitment takes this effort one important step forward by having participating foundations report at least quarterly -- with some reporting weekly, even daily.

A Radical Idea -- Open Data
By and large, foundations tend to think of open data and transparency as something they should fund rather than do. There are lots of reasons for this, including the private nature of foundations, the cultural legacy of keeping a low profile and "letting our good works speak for themselves," and sensitivity surrounding some of the issues addressed by foundation grants. Notwithstanding, the ability of foundations to not call attention to themselves is being steadily eroded by the ease of finding, displaying, and circulating information in a densely networked, digital age. Meanwhile, sectors and institutions with which foundations increasingly collaborate, such as the World Bank and foreign aid donors, are barreling ahead with initiatives like the Open Aid Partnership and Publish What You Fund.

The fifteen Reporting Commitment foundations have chosen to get ahead of the curve by taking the radical step of making their grants data entirely open. Under the agreement forged among them, they will either submit their data in machine-readable format or have the Foundation Center convert it so that it can be "harvested" by computers and used by developers to create apps, dashboards, visualizations, and things we haven't yet imagined. To make it easier, Glasspockets features a query builder that allows users to construct their own search and then "grab" the resulting data via an API.

A Strategic Idea -- GeoCoding
Some five years ago, when the Foundation Center started visualizing foundation grants data on interactive online maps, the most common reaction was, "You only show the location of the grantee organization, not the geographic focus of the grant." There was a reason for this: the vast majority of foundations, even those that electronically submit their data to the Foundation Center, do not include any coding for geographic area served. And even when there was a clue embedded in the grant description, there was no single standard that foundations used to describe the world; commonly used phrases such as "Deep South," "Middle East," and "developing countries" do not have agreed-upon definitions. That's why so many mapping visualizations (including our own) consign grants with insufficient or no geographic coding to big bubbles floating around in the ocean.

The Reporting Commitment foundations want to be able to compare their grants data with other participating foundations' data, from the community level all the way up to the continental level, to better identify gaps and areas of overlap and be more strategic about their giving. Thus they have agreed to use the GeoTree developed by the Foundation Center as an open geographic standard for use by philanthropy and the social sector. Geographic coding, or geocoding, as it is commonly known, requires a degree of specificity and decision making (i.e., how to handle grants that benefit multiple locations) that is something of a new discipline for most foundations. An interactive mapping tool on Glasspockets allows users to filter and search more than 3,800 grants by city/town, state/province, country, continent, or keyword. As participating foundations geocode more and more of their grants, the volume of data visualized on this map will expand.

A Mission-Critical Idea -- Transparency
When the Foundation Center was created in 1956 as a response to McCarthy-era hearings on philanthropy, transparency meant collecting printed reports from foundations and organizing them in file cabinets for public inspection. Today, it increasingly means open data. For an organization that has built a successful business model that relies on revenue from subscription databases to sustain an enormous volume of free information and services provided to more than nine million users, this may seem like risky business -- and it is. But the future of the Foundation Center requires disrupting its role as a data publisher. In the end, it is the Foundation Center's ability to analyze and combine multiple streams of information and analysis that adds value to data. And it is technology and networks that will allow the center to deliver knowledge into the hands of organizations and individuals who can leverage it to change the world.

Thanks to the vision, leadership, and hard work of the fifteen Reporting Commitment foundations, philanthropy has taken a crucial public step. Other foundations wishing to join the commitment can get started by contacting the Foundation Center. Later this year and again in 2013, the Foundation Center plans to release new and exciting forms of open data. While philanthropy may have been slow to get there, it is finally entering the era of Big Data.
-- Brad Smith



Monday, October 8, 2012

The Benefits of Open Data - Evidence from Economic Research


Guo Xu, Open Economics, October 3, 2012

This contribution is by Guo Xu (OKFN Economics and LSE) and the first part of the blog series “Mainstreaming Open Economics”.

Looking back to the Open Knowledge Festival 2012 in September, there’s an impression that openness is everywhere: There are working groups on Open Science and Open Linguistics, topic streams on Gender and Diversity in Openness, and events like Open Prom and Open Sauna: Open Knowledge and Open Data, it seems, is omnipresent.

Looking beyond the Open Knowledge community, however, the situation is very different: In Economics, for example, not many know what “open data”, “open access” or “Open Economics” exactly mean. Indeed, not many even care. A common reaction is: “Yes, it sounds interesting and important, but does it really matter? And why should I care about it?”

In this post, I would like to give some hard evidence on the positive role of opening up information has had in economics, and sketch ideas for how to involve economists – professional or in training – to mainstream ideas of openness. The blog post is divided into three parts: The first part looks at economic research on open data. The second part looks at the impact of open data on economic research. The third part discusses challenges and ways forward.

The real world impacts of open information
Making information accessible to the public can improve public service delivery. In countries where corruption is pervasive, services and funds often do not reach the frontline provider. And even if services do reach the people, the quality of services provided is often shockingly poor: Survey evidence from Bangladesh, Ecuador, India, Peru and Uganda found absence rates as high as 20% and 35% for school teachers and health workers. In many cases, the staff is poorly trained.

Releasing data on service delivery in this case can help reduce corruption and improve public services. In Uganda, researchers provided information to parents by publishing funding data for a random subset of schools in local newspapers. In consequence, corruption decreased significantly, while schooling outcomes improved substantially. Similar evidence in health delivery and redistributive policies suggest that providing information can help the public to discipline public service providers, improving the quality of services.

Information can also expose corrupt politicians: The Federal Government of Brazil, for example, began to select and audit municipalities at random, releasing audit reports to the media. Researchers found that the audit outcomes had a significant impact on the reelection probability of politicians: Those exposed for corruption were punished at the ballots, and the impact was most pronounced in areas where the dissemination of information was favoured by local radio.

A story from fishermen in South India provides another example of how information can improve market efficiency: Studying the adoption of mobile phones in Kerala, researchers have found convincing evidence that access to information through mobile phones helped fishermen sell their catch at the market where the price was highest (and fish most demanded): Instead of sailing to a port and simply hoping for a good price, fishermen were empowered by technology to make informed decisions on how to trade.

Finally, the benefits of transparency are not only restricted to reducing corruption and lowering the cost of information: A comparative study finds that transparency – measured by accuracy and frequency of macroeconomic information released to the public – leads to lower borrowing costs in sovereign bond markets. Open data pays off in many ways – in many different contexts.

These are just a few selective examples on how cutting-edge economic research has identified the benefits of openness in a diverse range of situations. The cases I presented are not based on correlations, but carefully established causal relationships, leaving – at least within the context studied – little doubt that information matters – big time. Perhaps most importantly, these cases have also shown that open data must be understood in a broad sense: These interventions do not take advantage of linked data, do not use CSVs that are shared through Facebook or Twitter – often, these interventions are simple solutions that ultimately help improving the everyday lives of the people.

Big Data: A Short History


How we arrived at a term to describe the potential and peril of today's data deluge.

Uri Friedman, Foreign Policy, November 2012

Humans have been whining about being bombarded with too much information since the advent of clay tablets. The complaint in Ecclesiastes that "of making many books there is no end" resonated in the Renaissance, when the invention of the printing press flooded Western Europe with what an alarmed Erasmus called "swarms of new books." But the digital revolution -- with its ever-growing horde of sensors, digital devices, corporate databases, and social media sites -- has been a game-changer, with 90 percent of the data in the world today created in the last two years alone. In response, everyone from marketers to policymakers has begun embracing a loosely defined term for today's massive data sets and the challenges they present: Big Data. While today's information deluge has enabled governments to improve security and public services, it has also sowed fears that Big Data is just another euphemism for Big Brother.

1887-1890
American statistician Herman Hollerith invents an electric machine that reads holes punched into paper cards to tabulate 1890 census data, revolutionizing the concept of a national head count, which had originated with the Babylonians in 3800 B.C. The device, which enables the United States to complete its census in one year instead of eight, spreads globally as the age of modern data processing begins.


1935-1937
President Franklin D. Roosevelt's Social Security Act launches the U.S. government on its most ambitious data-gathering project ever, as IBM wins a government contract to keep employment records on 26 million working Americans and 3 million employers. "Imagine the vast army of clerks which will be necessary to keep these records," Republican presidential candidate Alf Landon scoffs. "Another army of field investigators will be necessary to check up on the people whose records are not clear."


1943
At Bletchley Park, a British facility dedicated to breaking Nazi codes during World War II, engineers develop a series of groundbreaking mass data-processing machines, culminating in the first programmable electronic computer. The device, named "Colossus," searches for patterns in intercepted messages by reading paper tape at 5,000 characters per second -- reducing a process that had previously taken weeks to a matter of hours. Deciphered information on German troop formations later helps the Allies during their D-Day invasion.


1961
The U.S. National Security Agency (NSA), a nine-year-old intelligence agency with more than 12,000 cryptologists, confronts information overload during the espionage-saturated Cold War, as it begins collecting and processing signals intelligence automatically with computers while struggling to digitize a backlog of records stored on analog magnetic tape in warehouses. (In July 1961 alone, the agency receives 17,000 reels of tape.)


1965-1966
The U.S. government secretly studies a plan to transfer all government records -- including 742 million tax returns and 175 million sets of fingerprints -- to magnetic computer tape at a single national data center, though the plan is later scrapped amid public concern about bringing "Orwell's '1984' at least as close as 1970," as one report puts it. The outcry inspires the 1974 Privacy Act, which places limits on federal agencies' sharing of personal information.


1989
British computer scientist Tim Berners-Lee proposes leveraging the Internet, pioneered by the U.S. government in the 1960s, to share information globally through a "hypertext" system called the World Wide Web. "The information contained would grow past a critical threshold," he writes, "so that the usefulness [of] the scheme would in turn encourage its increased use."


August 1996
"We are developing a supercomputer that will do more calculating in a second than a person with a hand-held calculator can do in 30,000 years." --U.S. President Bill Clinton


1997
NASA researchers Michael Cox and David Ellsworth use the term "big data" for the first time to describe a familiar challenge in the 1990s: supercomputers generating massive amounts of information -- in Cox and Ellsworth's case, simulations of airflow around aircraft -- that cannot be processed and visualized. "[D]ata sets are generally quite large, taxing the capacities of main memory, local disk, and even remote disk," they write. "We call this the problem of big data."


2002
After the 9/11 attacks, the U.S. government, which has already dabbled in mining large volumes of data to thwart terrorism, escalates these efforts. Former national security advisor John Poindexter leads a Defense Department effort to fuse existing government data sets into a "grand database" that sifts through communications, criminal, educational, financial, medical, and travel records to identify suspicious individuals. Congress shutters the program a year later due to civil liberties concerns, though components of the initiative are simply shifted to other agencies.


2004
The 9/11 Commission calls for unifying counterterrorism agencies "in a network-based information sharing system" that is quickly inundated with data. By 2010, the NSA's 30,000 employees will be intercepting and storing 1.7 billion emails, phone calls, and other communications daily. Meanwhile, with retailers amassing information on customers' shopping and personal habits, Wal-Mart boasts a cache of 460 terabytes -- more than double the amount of data on the Internet at the time.


2007-2008
As social networks proliferate, technology bloggers and professionals breathe new life into the "big data" concept. "This is a world where massive amounts of data and applied mathematics replace every other tool that might be brought to bear," Wired's Chris Anderson writes in "The End of Theory." Government agencies, some of the United States' top computer scientists report, "should be deeply involved in the development and deployment of big-data computing, since it will be of direct benefit to many of their missions."


January 2009
The Indian government establishes the Unique Identification Authority of India to fingerprint, photograph, and take an iris scan of all 1.2 billion people in the country and assign each person a 12-digit ID number, funneling the data into the world's largest biometric database. Officials say it will improve the delivery of government services and reduce corruption, but critics worry about the government profiling individuals and sharing intimate details about their personal lives.


May 2009
U.S. President Barack Obama's administration launches data.gov as part of its Open Government Initiative. The website's more than 445,000 data sets go on to fuel websites and smartphone apps that track everything from flights to product recalls to location-specific unemployment, inspiring governments from Kenya to Britain to launch similar initiatives.


July 2009
Reacting to the global financial crisis, U.N. Secretary-General Ban Ki-moon pledges to create an alert system that captures "real-time data on the impact of the economic crisis on the poorest nations." The U.N. Global Pulse program has conducted research on how to predict everything from spiraling prices to disease outbreaks by analyzing data from sources such as mobile phones and social networks.


August 2010
"There were 5 exabytes of information created by the entire world between the dawn of civilization and 2003. Now that same amount is created every two days." --Google CEO Eric Schmidt


February 2011
Scanning 200 million pages of information, or 4 terabytes of disk storage, in a matter of seconds, IBM's Watson computer system defeats two human challengers in the quiz show Jeopardy!. The New York Times later dubs this moment a "triumph of Big Data computing."


March 2012
The Obama administration announces a $200 million Big Data Research and Development Initiative in response to a U.S. government report calling for every federal agency to have a "'big data' strategy." The National Institutes of Health puts a data set of the Human Genome Project in Amazon's computer cloud, while the Defense Department pledges to develop "autonomous" defense systems that can "learn from experience." CIA Director David Petraeus, marveling that the "'digital dust' to which we have access is being delivered by the equivalent of dump trucks," discusses a post-Arab Spring agency effort to collect and analyze global social media feeds through cloud computing.


July 2012
U.S. Secretary of State Hillary Clinton announces a public-private partnership called "Data 2X" to collect statistics on women and girls' economic, political, and social status around the world. "Data not only measures progress -- it inspires it," she explains. "Once you start measuring problems, people are more inclined to take action to fix them because nobody wants to end up at the bottom of a list of rankings." Let the Big Data race begin.

Thursday, September 27, 2012

New Talk by Clay Shirky -


Clay Shirky’s Ted Talk “How the Internet will (one day) transform government”

The open-source world has learned to deal with a flood of new, oftentimes divergent, ideas using hosting services like GitHub -- so why can’t governments? In this rousing talk Clay Shirky shows how democracies can take a lesson from the Internet, to be not just transparent but also to draw on the knowledge of all their citizens.

Clay Shirky argues that the history of the modern world could be rendered as the history of ways of arguing, where changes in media change what sort of arguments are possible -- with deep social and political implications

Tuesday, September 11, 2012

Here Comes the Data Economy


New companies are creating services using government data on health care, education, and more.

Alexander B. Howard, Slate, September 10, 2012


We're living in the exabyte age, where the actions of billions of humans using the Web and their mobile devices are creating massive amounts of big data to collect, store, analyze, and put to work.

If big data is a strategic resource, as has been suggested, then many national and state governments have public reserves that can be tapped for the public good in this young century's version of the industrial revolution. Given that the United States economy is still coming out of the worst recession and financial shock since the Great Depression, supporting civic and tech entrepreneurs enjoys political support from both sides of the aisle.

Entrepreneurs, big and small, are mashing up data from the rapidly expanding collection of sources and building new businesses on it or improve their existing services, like Zillow or Google Maps or Consumer Reports or Bloomberg Government. In a time when job creation is critical, using public sector information to create jobs isn’t an aim to dismiss lightly, although the terms and conditions under which such activity occurs must be clear to all actors involved, to avoid the creation of new monopolies based upon artificial scarcity.

My publisher, long-time open source and open government advocate Tim O'Reilly, has asked how government can act as a platform to enable people inside and outside government to innovate on top of it. One answer is certainly releasing open data. In that context, open data and application programming interfaces, more commonly known as APIs, increasingly look like fundamental infrastructure for digital government in the 21st century.

There's good reason to think that open data could have an overall effect on the economy akin to open source and small business. Gartner, the IT research analysis firm, recently highlighted how open data creates value in the public and private sector.

You may not realize it, but services you use on a daily basis have been built upon data released by the government. Weather data collected by the National Oceanic and Atmospheric Association has an annual estimated economic value of $10 billion, according to U.S. Chief Information Officer Steven VanRoekel and U.S. Chief Technology Officer Todd Park. NOAA data sets are used by Weather.com, Weather Underground, and the Weather Channel—and the nation's farmers consult these forecasts to manage both their crops and the risks of loss. VanRoekel and Park estimate the annual economic value of the data from the U.S. global positioning system at some $90 billion. From companies like TomTom or Garmin to dashboard GPS systems to smartphones and associated location-based applications, GPS data sets are baked into an expanding number of services and products.

Now, as Park seeks to scale open data across the federal government, we’re on the verge of the next generation of services driven by open data, which will involve everything from energy to health care to consumer finance to transit sectors. The challenge is that the cities and federal agencies that hold vast amounts of data may not always understand the value of the information they hold or how to create or sustain businesses using it. That's where open innovation in the public sector and the dynamism of entrepreneurs will play an important role in making the people's data more useful to the people.

BrightScope is a notable example of what dogged persistence can create. The California startup made a profitable business using government data to help the American people understand the fees associated with their 401(k)s. Last May, BrightScope went further, launching financial adviser pages based on open government data from the Securities and Exchange Commission and the Financial Industry Regulatory Authority, the largest independent securities regulator in the United States. Previously, financial adviser profiles could only be found through exact queries at an obscure URL on the regulators' websites. Now, information that citizens care about—the records of financial advisers in their geographic region—is available where they're looking for it: in search engine results.

Just as labor and regulatory data fuels BrightScope's business, there's an expanding number of startups that are tapping into other data released so-called “smart disclosure” initiatives. Smart disclosure is when a private company or government agency provides a person with periodic access to his or her own data in open formats that enable them to easily put the information to use. Startups like Billshrink.com and Hello Wallet are already using a combination of private sector and public sector data to enhance consumer finance decisions. The success of such consumer finance startups suggests an important lesson: The most successful apps and services will combine government, industry, and user-generated data.

The key open data story to watch in the federal government, however, centers on health care. McKinsey and Associates estimates the annual economic value of big, open liquid health data at about $350 billion annually. The explosion of mHealth apps are just the beginning of the disruption in health care from open health data. The effort to revolutionize the health care industry by making health data as useful as weather data is still in its infancy—but the early results are promising. iTriage, which was acquired by Aetna, is enabling people to make better mobile health care decisions where and when they need to do so. It uses a combination of government and private sector data to evaluable symptoms or conditions and point users to nearby medical care. Another startup, Castlight, is analyzing health care data to empower patients, acting like Kayak.com for those who want more transparency about costs. In May, Castlight completed a $100 million round of financing.

But for these sorts of initiatives to take off, entrepreneurs and regulators will have to work together to get contextual consent right and inform patients about the reuse of their data. Transparency is crucial to building a health data commons and thriving startup ecosytem based upon it.

If that balance can be struck, there's considerable potential for entrepreneurs to create better civic interfaces for many digital services. If open government data have helped build new tools, open data disclosed by private companies could create even more value for citizens. But currently, few businesses release anonymized data in an open, usable format. It will soon be time for the government to step in, convene stakeholders, and answer some key questions: How can we create uniform standards that will allow entrepreneurs and developers to innovate? When should data be licensed? Most of the big data releases we have seen come from finance, with bank records or stock trades. But there are significant opportunities to help both entrepreneurs and empowered consumers in health care, energy, education, and telecommunications, to name just a few.

Just as the glowing blue dot on the maps in our smartphone screens revolutionized how we navigate the world, similar "blue dots" could emerge for health care, finance, energy, and any product or service that is regulated or cataloged by government and industry. First, however, they'll need to open the data.

Also in the Future Tense package on government and open data: why Yelp and the government should share data; what a burger mob tells us about the future of democracy; and how Mexico is using open data to move beyond its authoritarian past.