Showing posts with label security. Show all posts
Showing posts with label security. Show all posts

Friday, December 7, 2012

Panjiva Uses Government Data to Build a Global Search Engine for Commerce


Successful startups look to solve a problem first, then look for the datasets they need.

Alex Howard, O'Reilly Strata, December 6, 2012


“If you go back to how we got started,” mused Josh Green, “government data really is at the heart of that story.” Green, who co-founded Panjiva with Jim Psota in 2006, was demonstrating the newest version of Panjiva.com to me over the web, thinking back to the startup’s origins in Cambridge, Mass.

At first blush, the search engine for products, suppliers and shipping services didn’t have a clear connection to the open data movement I’d been chronicling over the past several years. His account of the back story of the startup is a case study that aspiring civic entrepreneurs, Congress and the White House should take to heart.

“I think there are a lot of entrepreneurs who start with datasets,” said Green, “but it’s hard to start with datasets and build business. You’re better off starting with a problem that needs to be solved and then going hunting for the data that will solve it. That’s the experience I had.”

The problem that the founders of Panjiva wanted to help address was one that many other entrepreneurs face: how do you connect with companies in far away places? Green came to the realization that a better solution was needed in the same way that many people who come up with an innovative idea do: he had a frustrating experience and wanted to scratch his own itch. When he was working at an electronics company earlier in his career, his boss asked him to find a supplier they could do business with in China.

“I thought I could do that, but I was stunned by the lack of reliable information,” said Green. “At that moment, I realized we were talking about a problem that should be solvable. At a time when people are interested in doing business globally, there should be reliable sources of information. So, let’s build that.”
Today, Panjiva has created a higher tech way to find overseas suppliers. The way they built it, however, deserves more attention.

Government data as a platform
By 2009, the startup had an initial product they could bring to market and launched a search engine that used government data as a platform for international trade. An importer could type in “patio furniture” and
determine who shipped it and who their customers were. The company chose a freemium model, where search is available for free but relationships between suppliers are only available to subscribers. The mapping of relationships between buyers and suppliers is where Panjiva delivered added value on top of public data.

That added value is crucial, given that competitors can also request and use the dataset. “Companies have been packaging and reselling this data in one way or another for years, ” said Green. “If you looked at this data, people are going to find value. It’s typically folks in the shipping industry, who want to know what’s going into ports or moving on different shipping lines. For us, the central purpose of the data was something different and required more work.”

That work paid off. In 2010, Panjiva built a search engine for global commerce that worked. Today, they have more than 100,000 users in 190 countries using its free service and some 3,700 companies subscribing to the paid version, including 42 Fortune 500 companies.

Notably, the Department of Homeland Security (DHS) itself is also a paying subscriber. Green declined to disclose the terms of relationships with all of Panjiva’s partners or data suppliers, some of which include nonprofits. Some users “do a revenue share, some are paying for data, others are providing data because they think there’s public good for that data being on the platform,” he said.

Panjiva competes with ImportGenius, Zepol, AliBaba and PIERS. Green credits PIERS for extracting similar value from customs datasets.

The turning point
When they started, the first approach that Green and his co-founder decided to take was to build a “Yelp for global trade” that would be based on feedback from people who work with companies. Unfortunately for the young startup, they couldn’t get off the starting blocks in generating reviews, much less reach critical mass.

They also encountered a new problem: even if they were able to get ratings of exporters and suppliers, how would they ensure the reviews came from people who had actually done business with the entities being rated? In retrospect, that focus was a bit silly, said Green, because they couldn’t get engagement, but talking about how to solve it led them to an unexpected answer: government data.

That direction came from a meeting where a staffer for a trade promotion organization told them it was straightforward to get shipping data on what’s coming into the country from the United States Customs Agency, which is now part of the Department of Homeland Security.

“It was a turning point for the company,” said Green. “We realized there was a dataset available to the public for a fee. They make available data about shipments that enter the U.S. While not all data is made available to the public and there are a bunch of limitations, the data that is made available is amazing. There’s about 10 million shipping records every year, typically including who is sending goods, who is receiving goods, what’s inside, and how much is inside a container.”

While useful, these government datasets do come with inherent limitations, cautioned Green. For one, they only contain data about shipments coming into the United States, not what’s going into Europe or Asia. For another, the data made available to the public only covers shipments made by boat, which is about about half of the shipments that come into the United States.

“It’s unfortunate that government cannot make available data on other modes of transport,” observed Green, with a hint of frustration in his voice. “That leaves out truck, rail, and air. Congress actually attempted to clarify that the regulations that govern this data weren’t just about boats but applied to air. Thus far, DHS hasn’t acted.”

Given the lens that has been focused on trade deficits between other countries and the United States in recent decades, there’s also a political angle to the market intelligence Panjiva provides that Congress and taxpayers may find of interest. For instance, Panjiva data showed global trade growth slowing in the first part of 2012.

“What we’ve organized, by its nature, gives us insight on companies around the world that serve the U.S. market,” said Green, “We’re helping people find overseas suppliers. Why not help find suppliers here at home? It turns out there’s a similar story on export data that’s supposed to be made to the public as well. DHS has a hard time with that as well. We can’t get the data.”

Data availability is also affected by the actions of the companies themselves, which have the ability to petition the government to hide shipments that are coming to them. “In about a third of the cases, you cannot see who is sending and receiving the goods,” said Green. “Government can see, but what’s released to the public has information pulled from it.”

This government data comes at a cost
Accessing this public data comes at a cost of some $100 per day, which is the service fee DHS charges for providing a daily CD-ROM. Each disc includes one day’s worth of shipments, which is generally around 30,000 shipping records. Panjiva started requesting data on July 1, 2007, and now has a little over five years of records.

“This data, on a record-by-record basis, is interesting,” said Green. “If you can organize, it’s phenomenal. If you can associate with companies, can say this company has experience with these supplies and this company has experience with these customers, it’s very useful in deciding if a company is a good fit. You can see by customers if they’re reasonably high quality.”

Making those CD-ROMs into a useful, searchable resource, however, was far from a simple matter of just inserting them into an optical drive and moving their contents into a structured database.

“Jim and a team of engineers went to work organizing the datasets initially,” said Green. “They were very hard to work with — absurdly messy. Think about the number of ways you can misname a Chinese factory. It was really problematic. You need to build company profiles, correct for misspellings and variations on names. We spent years getting that right.” Eventually, Panjiva was able to automate the process of ingesting the data from the CD-ROMs, building an algorithm to take the data and clean it up.

Making data a strategic asset
Panjiva’s initial foray, which created a search engine for customs data, didn’t meet with strong demand out of the gate. As they refined the product, it generated what Green described as a “nice business.” The startup was profitable, in other words, but its leadership aspired to build something bigger.

The direction they took was driven by user feedback. When Panjiva also asked its users about how they were making buying decisions, they saw a pattern emerge that looked like a bigger opportunity.

“Users started with Panjiva then went to search for additional information on B2B sites or on Google,” said Green. “We heard this process and it sounded a lot like the experience consumers had searching for flights before search engines or Kayak.com — except that instead of airline sites, people are going to B2B sites. The difference is it’s not just every airline. It’s like every flight has its own website.”

The founders now have raised just under $10 million from Battery Ventures and Harrison Metal, and invested it in technology and data acquisition. They’ve now grown their engineering team to 10 people, out of a total of 50 or so current employees. The engineering team is focused on improving search and enriching Panjiva’s data with other sources, beyond government data.

This October, the startup relaunched Panjiva.com with another layer: data supplied by the companies themselves.

“We have a database of six million companies spread around the world and contact information on four million companies,” said Green. “We have product photos for 34 million products. There was a lot of investment required to do that, but none of this would be possible if we hadn’t had a backbone of data that came from the U.S. government.”


Since Panjiva added global search, Green said that traffic to the search engine has gone up 50%.
The data sources that Panjiva integrated were also driven by customer interest. As the founders shared their product with potential subscribers, they kept hearing the same thing: 1) “that’s awesome” and 2) “I’d like more data.”

“We loved the first one and hated the second,” said Green. “In retrospect, we should have loved both. The second one was a roadmap for us to build them a really great differentiated product.”

When they asked users exactly which kinds of data would make the service more useful, a map to the future of the company emerged.

The first was operational data. “Customs data is a perfect example,” said Green. “It gives you a sense of what companies have done and their track record.”

The second was financial data. “Sure, a company has experience, but are they financially healthy?” asked Green. “Some of that you can infer, but there’s other things you can use. We’ve partnered with Dun & Bradstreet and Experian to pull that data into our platform.”

The third was positive and negative data about a company. “That includes getting certified as financially responsible,” said Green. “We’ve partnered with nonprofits and added that data, showing you information about companies doing wrong, including a blacklist of illicit global trade.”

The key insight that anyone interested in building a business on top of government data should take away here is to go beyond.

What happens if government data becomes open?
Green thinks that Panjiva is well-positioned to be both competitive and profitable, even if DHS decided to start publishing customs data online. “We don’t worry that much about data becoming more accessible,” he said, “even if government data becomes free. It’s not the $36,500 per year to buy the data — it’s the engineering talent to clear it up. That’s a massive problem, and it wouldn’t be as simple as getting the data.”
Panjiva is betting that the investments they’ve made in technology, talent and — crucially — combining so many different data sources have created a differentiated product that solves a problem for its customers.

“We’re not trying to build out a data business where we’re reselling government data,” said Green. “We’re trying to build a platform where serious buyers and sellers can connect. We’re now going to the world’s most important buyers. We have two revenue streams: selling premium access to data and selling access to suppliers who want it. The starting point for customers is $99 per month, going up to $10,000 per month for unlimited access for an unlimited number of users, then services that we sell on the top.”

The experience that Panjiva has had with government data and building a business using it has left Green with a strong perspective on what works — and what doesn’t.

“We don’t think there are infinite numbers of possibilities in terms of ways to build sustainable value with public data,” he said. “One is to take datasets that are commoditizeable and add value. Another is to feed the creation of more data. Another is to build a service. Another is to create network effects, where the data is the honey that attracts the bees.”

Most important, Green suggested, is to use public data to solve a problem that’s both hard and important. For Panjiva, that means making global trade more efficient and more transparent.

“There is a future where information is consolidated and accessible to people making key decisions, from a buying or regulatory standpoint,” he said. “Once that happens — and we’re close — there’s potentially a place where there’s a race to the top instead of the bottom, in terms of supply chain records. That will make a difference when you’re under scrutiny. Right now, the fragmentation of data is the ally of bad behavior. Our hope is to change that reality.”

Sunday, December 2, 2012

Digital Privacy in the Big Data Era: Microsoft's Data Protection Keynote


Ms. Smith, Network World, December 2, 2012

There are several Internet security experts who agree with Steve Rambam's claim [1] that "Privacy is dead - get over it." Yet other privacy and security experts such as Bruce Schneier completely disagree. In The Value of Privacy [2] Schneier wrote, "Privacy protects us from abuses by those in power, even if we're doing nothing wrong at the time of surveillance." When it comes to data protection and protecting people's privacy in the digital age, Europe is far more advanced than America. [3]

In fact, the head of France's data protection agency, Isabelle Falque-Pierrotin, did an excellent job summing it up as: "In Europe, we consider privacy a fundamental right. That doesn't mean it is exclusive of other rights, but economic rights are not superior to privacy." The New York Times also reported [4] that she said in the United States, "personal data are seen as raw material for business."

In November, Microsoft's Chief Privacy Officer Brendon Lynch said [5] of the IAPP European Data Protection Congress 2012 [6], "One area of strong consensus was the tremendous potential the digital economy holds for companies on both sides of the pond. Accordingly, it's important to strike the right balance between data protection with business growth through interoperability between privacy regulation in the EU, U.S. and elsewhere."

Many privacy advocates cringe when hearing the word "balance," such as striking a balance between security and privacy. Hopefully people won't come to cringe when they hear the word balance applied to big data security protections and privacy. As Bruce Schneier wrote [2] way back in 2006:

Too many wrongly characterize the debate as "security versus privacy." The real choice is liberty versus control. Tyranny, whether it arises under threat of foreign physical attack or under constant domestic authoritative scrutiny, is still tyranny. Liberty requires security without intrusion, security plus privacy. Widespread police surveillance is the very definition of a police state. And that's why we should champion privacy even when we have nothing to hide.

Whether people realize it or not, big data is not privacy-friendly even when it is supposedly anonymized or contains obfuscated PII (Personally Identifiable Information) data. Researchers have shown that "linkability threats" can re-identity individuals. Since it boils down to the fact that you are not anonymous when it comes to big data [7], Microsoft has developed "Differential Privacy for everyone" [download PDF [8]].
In the IAPP keynote address [download PDF [9]], Lynch made some excellent and thought-provoking privacy points regarding big data. He said:

Data is the fuel that drives all of these powerful technologies, but what can be done with the data today can at times seem enormously helpful or enormously threatening. Consider two scenarios shown here. In the first case, I am using my phone in a grocery store to find out more about the items on the shelves and it is mashing up that with my private data to personalize my experience. So here I downloaded a recipe and customized it for my dietary needs. If it's a trusted system, that's a great experience. On the other hand, consider the US company, Target, which recently generated a lot of press about its pregnancy prediction score. This was based on what people were purchasing in Target stores, they are able to indicate a shopper that appeared to be pregnant. The concern about how Target can figure out such details about customers shopping in its stores, who are not explicitly sharing that information, is the concern. And what does it do with those insights? In this particular case, they sent some mailers to the individual involved - it was a teenage girl and her father was very offended that they were wrongly marketing to her, but it eventually did come out that she was in fact pregnant. Target knew a lot more than her father knew.

Peter Cullen, Microsoft's Chief Privacy Strategist, wrote [10] about "notice and consent" as a means of privacy protection and how data privacy frameworks need to "focus on the 'harms' or 'impacts' of data use, which should not only include physical and financial injury, but also broader concepts such as reputational or social harm."

Yet after showing a video that highlighted data transfers in today's world at the IAPP conference, Lynch said, "How could there possibly be meaningful notice and consent mechanisms in place for every transfer of data that was involved?" He added, "It would seem that advances in technology and the rise of big data can create amazing societal benefits but they can also strain traditional notions of secrecy and the notice and consent approach to privacy protection."

In his keynote, Lynch said:

Some technology and internet companies today take the position that privacy is dead, or at least that privacy is an outdated concept that people need to get over so technology companies can help them reap the benefits of sharing as much information as possible. But we disagree that privacy is not relevant or desirable, in this sensor-driven, social everywhere, big data world that we are heading towards. People today expect strong privacy protections because they are increasingly aware of, and concerned about, the digital trails they leave behind online and indeed there's plenty of evidence that people still care deeply about privacy.

Of course people care about privacy. Europe continues to illustrate this to the world by taking a hard stance when data is used without "informed consent" and when users cannot "opt out." Lynch believes we need to not only protect privacy in regards to big data, but also that people "need an updated notion of privacy and data protection principles, one that shifts from a focus on secrecy to a more nuanced approach, based on reasonable consumer expectations, context and a greater emphasis on how personal information is used."

Big data definitely represents significant threats to personal privacy. Let's hope this "shift" and "updated notion of privacy" won't include the word "balance" that puts individuals on the losing end as it generally has when the government talks of striking a balance between security and privacy.

Monday, November 19, 2012

WSJ - CEO Council on Big Data


Big Data: Opportunities and Risks Co-Chairs
Tim Armstrong, Chairman and CEO, AOL Inc.
Dominic Barton, Global Managing Director, McKinsey & Co.
R. Marcelo Claure, Chairman, President and CEO, Brightstar Corp.

Subject Expert
David J. Rothkopf, President and CEO, Garten Rothkopf

Big Data: The Top Four Recommendations
1. Big Data Is Opportunity
Industry already recognizes that the advent of Big Data is a potential engine of significant economic growth. Policy makers must address issues such as security and privacy but not restrict this opportunity. Consumers, government and companies all have a role to play in defining policy. Countries need to recognize that the treatment of data is ultimately an issue of national competitiveness.


2. Create Rules of the Road
The government should define which activities concerning data are legal or illegal. For activities that are legal, consumers should then define what is public or private in their own cases through opting in or out in terms of what they choose to reveal about themselves.


3. Enact Cybersecurity Standards
The private sector, government, the military and consumers should jointly develop detailed standards and incidence-reporting practices for cybersecurity, building on existing industry best practices.


4. Create a Global Data Treaty
The U.S. should play a leadership role in coordinating with international bodies to pass a treaty to establish clear, universal standards on data privacy and ownership.


It's a phenomenon commonly referred to as Big Data, and it has generated widespread debate over a host of issues. What should companies do with such data? How can it be used profitably? How should it be treated? Should there be laws governing its use? Internationally recognized safeguards for consumer privacy? And how do U.S. companies stay competitive as more countries learn how to process and interpret such data?

The Wall Street Journal's John Bussey moderated a task-force discussion about just such issues. Here are edited excerpts of their presentations to the CEO Council:

Opportunity First
JOHN BUSSEY: Our group ended up with essentially two principles and two action items. Dominic, if you could take our first, please?

DOMINIC BARTON: Our group felt very strongly that this is a huge opportunity, and we shouldn't focus on how to protect or regulate Big Data before we recognize how important it is. Companies that use Big Data effectively get about a 6% productivity improvement versus others.

We're at the early stages. It's in every sector. And we all felt it can be the next wave of productivity growth if we use this data effectively.

Depending on how you protect or use data, it can actually lead to a country's competitive advantage. We've got small city-states like Singapore and Abu Dhabi that are allowing, for example, consumer medical information to be publicly available to get innovation.
So let's not lose sight of how important it can be for productivity improvement as we think about the protection.

MR. BUSSEY: Marcelo, our second principle.
MARCELO CLAURE: First of all, we had a fascinating group. We had a tremendous amount of interaction. And we quickly realized that all of us live in a hyper-connected world. By that I mean we're passing a huge amount of data every day through a connected car, a connected cellphone, a connected tablet, a connected home.

Big Data: The Top Four Recommendations
1. Big Data Is Opportunity
Industry already recognizes that the advent of Big Data is a potential engine of significant economic growth. Policy makers must address issues such as security and privacy but not restrict this opportunity. Consumers, government and companies all have a role to play in defining policy. Countries need to recognize that the treatment of data is ultimately an issue of national competitiveness.


2. Create Rules of the Road
The government should define which activities concerning data are legal or illegal. For activities that are legal, consumers should then define what is public or private in their own cases through opting in or out in terms of what they choose to reveal about themselves.


3. Enact Cybersecurity Standards
The private sector, government, the military and consumers should jointly develop detailed standards and incidence-reporting practices for cybersecurity, building on existing industry best practices.


4. Create a Global Data Treaty
The U.S. should play a leadership role in coordinating with international bodies to pass a treaty to establish clear, universal standards on data privacy and ownership.


Companies around the world increasingly collect and process vast amounts of customer data, particularly data tracking consumer behavior on the Web.

 We see it as a tremendous opportunity, but what is the government role?

Some members of our group argued that government shouldn't be allowed to regulate or do anything related to Big Data. But the rest of the group agreed that we've got to define the basics of what the government is going to do.

And where we came out at the end was, government should have the role of defining what is legal and what is illegal in the use of Big Data. And that's where the government should stop. We think it should be left up to the consumer to define what he or she wants to share and doesn't want to share, what is public and what is private.

Limiting government control to what's legal and illegal will allow the consumer to make more choices. As a consumer, you should be able to determine how much of your Facebook profile you want to share, or whether you want to share information about where you shopped last.

Security and Treaty
MR. BUSSEY: And Tim has our two action items.

TIM ARMSTRONG: The first one is probably more important than it seems on the surface, especially for private companies. I think we started this as kind of a government conversation, but mostly it's private companies that actually own the infrastructure that makes the company or the country work.

So cybersecurity is something that really needs to be dealt with. And specifically, having standards around this is important in two different directions.

One is standards of what happens when something gets attacked. How do you report it, and what is the infrastructure to help companies with that?

The second piece which is important and which came up in our discussion is when your data lives outside the country, or when you're going to do things with data in other countries. There are groups like CFIUS [the U.S. government's Committee on Foreign Investment in the U.S., which reviews foreign investment deemed to have an effect on national security] which are inside the government, which, if you haven't dealt with them, you will, over these type of data assets. Cybersecurity is really important.

The last action item is coming up with a global data treaty. The U.S. is probably the largest economy involved in Big Data right now, and we think it's important for the U.S. to take a leading role in defining what some of the Big Data policies and standards should be in such a treaty. We would hope that we, the U.S., could define a basic treaty to start off, and then other countries could either adopt or augment the treaty over time.