Wednesday, December 12, 2012
Social Media a Healthcare Data Gold Mine
April Dembosky, The Financial Times, December 11, 2012
Bill Schmarzo envisions a future where holiday photos posted on Facebook become a gauge of a person’s weight loss or gain over time.
The chief technology officer for EMC’s consultant services acknowledges that privacy advocates are unlikely to allow his fantasy to become reality, but the technology that can measure minute body changes in photographs and feed it into someone’s electronic health record already exists.
“The scary thing is if that data could be used to deny care and insurability,” Mr Schmarzo says.
Healthcare companies are loathe to tread into such sensitive territory, but they are keenly aware of the gold mine of health data stored in people’s social media accounts.
“Studies have shown that people are more willing to share more private medical information in social media than they’re willing to share with their medical providers,” says Martin Kohn, chief medical scientist at IBM.
Pharmaceutical companies already analyse social media sites to track reports of side-effects of their drugs. That data can help correct formulas more quickly than waiting for the results of years-long clinical trials; it can also be used to set prices and test marketing slogans.
Public health officials are also interested in social media data as a source of information on disease outbreaks. Several start-ups are working on algorithms that study Facebook, Twitter, and blog posts to track early signs of infectious disease outbreaks, as people generally complain to their friends long before public health agencies can collect doctors’ reports and issue official warnings.
One outbreak of the Norovirus stomach bug at a student journalism conference in Canada was live-tweeted earlier this year, with posts such as “Motion that nobody else on this bus puke” and “36 hours without leaving my hotel room” signalling its coming and going before any traditional surveillance system had noted it.
The trouble with social media data are that they are fragmented and incomplete, warns James Kaufman, manager of public health research at IBM’s Silicon Valley lab, and there are limits to the types of insights that can be drawn from such patient-reported data.
“Colds, flus, sure people report that,” he says, “but no one’s going to report Aids or haemorrhoids.”
Making Dollars and Sense of the Open Data Economy
Is the push to free up government data resulting in economic activity and startup creation?
Alex Howard, O'Reilly Radar, December 11, 2012
Over the past several years, I’ve been writing about how government data is moving into the marketplaces, underpinning ideas, products and services. Open government data and application programming interfaces to distribute it, more commonly known as APIs, increasingly look like fundamental public infrastructure for digital government in the 21st century.
What I’m looking for now is more examples of startups and businesses that have been created using open data or that would not be able to continue operations without it. If big data is a strategic resource, it’s important to understand how and where organizations are using it for public good, civic utility and economic benefit.
Sometimes government data has been proactively released, like the federal government’s work to revolutionize the health care industry by making health data as useful as weather data or New York City’s approach to becoming a data platform.
In other cases, startups like Panjiva or BrightScope have liberated government data through Freedom of Information Act requests and automated means. By doing so, they’ve helped the American people and global customers understand the supply chain, the fees associated with 401(k) plans and the history of financial advisors.
I’ve hypothesized that open data will have an overall effect on the economy akin to that of open source and small business. Gartner’s research has posited that open data creates value in the public and private sector. If government acts as a platform to enable people inside and outside government to innovate on top of it, what are the outcomes?
Over the past four years, the world has heard a rising chorus for raw data from voices like the creator of the World Wide Web, Tim Berners-Lee, and the chief technology officer of the United States, Todd Park. Park, in particular, has been working to scale open data across the federal government as the nation’s “entrepreneur in residence.”
McKinsey and Associates estimated the annual economic value of big, open liquid health data at some $350 billion annually. While that number is eye-opening, which companies and startups stand to change health care using open health data?
Some examples are clear, from mobile apps like iTriage (now owned by Aetna) to Castlight, but they aren’t sufficient to understand what’s happening out there.
Other promising startups are in the consumer finance space, where so-called “smart disclosure” initiatives are enabling people to put their personal data to use. Startups like Billshrink.com and Hello Wallet now are enabling people to make smarter financial decisions.
I know there are more stories out there, and in sectors beyond health care and consumer finance — including transit, energy, education and media. Over the next several months, I’ll be identifying and profiling more civic startups, such as those from the first class in the Code for America accelerator, like Captricity, to specialized search engines, like Zillow, Panjiva and DataMarket.
In the course of that work, I hope to answer some big questions. What are the sustainable business models that successful civic startups are using, whether they use legislative data or other reuse of public sector information? What are the real costs associated with opening up government data to make it usable, both for government and entrepreneurs? And how does it balance against what datasets, at the federal, state or local levels, are the most valuable? Are they open and usable? If so, who’s using them and to what effect? If not, why not?
At the end of this particular project, in February, we’ll publish a report on what I’ve found. In it, I hope to be able to share some answers to several core questions on the topic. Where I need your help is in identifying new startups that are using or consuming government data or in highlighting how existing companies use it in their operations, good or services. Who is doing the most interesting work — and where? If you have research and evidence to share on the questions I posed above, feel free to ring in on that count as well.
Please weigh in through the comments or drop me a line at alex@oreilly.com or at @digiphile on Twitter.
Monday, December 10, 2012
FT Special on Big Data
Big data is the term used to describe the huge volumes of data generated by traditional business activities and from new sources such as social media. Typical big data includes information from store point-of-sale terminals, bank ATMs, Facebook posts, and YouTube videos, writes Paul Taylor.
Paul Taylor, The Financial Times, December 9, 2012
Companies use sophisticated software to analyse this data looking for hidden patterns, trends or other insights that they can use to better tailor their products and services to customers, anticipate demand or improve performance.
Haven’t companies been doing this for years?
Companies and other organisations, including governments, have for many years been collecting, and slowly sifting through so-called structured data. These are data, like a sales ledger, that are already well-organised and therefore relatively easy to process.
But more recently, there has been an explosion in the amount of ‘unstructured’ data like video clips or facebook posts. Their lack of an identifiable structure makes them much more difficult to analyse, but these information promise to provide the most valuable insights for a company. For example, Facebook posts could tell brands what consumers think of their products.
Is big data all that its cracked up to be?
That depends. Like most hot new technology trends, big data is going through a ‘hype cycle’. Initially people expect far too much, then disillusionment sets in, but eventually, the technology begins deliver real benefits as it matures.
Big data has been heavily overhyped and as a result, some companies have invested in the technology without a clear idea of how to use it or how it fits into their business strategies.
On the other hand, some companies, particularly early adopters in retail and financial services, are already reporting significant results from big data projects.
Who will make money out of big data?
Some companies have already used big data to disrupt existing industry models, but the main beneficiaries are likely to be the manufacturers of the computer systems and software that do the heavy-lifting analytical work.
These include established vendors like IBM, Oracle and SAP, together with a raft of startups including Alpine Data Labs, Splunk, and SumoLogic.
Other likely winners are companies that collect and supply the data – there is already a multibillion-dollar data broker industry.
Should I be worried about the privacy implications of Big Data?
Much of the most valuable big data is information about people and the digital trail we all leave when we go online.
By making connections between disparate snippets of information, big data can reveal far more about us than we ever intended. Inevitably that means the collection and use of big data is front and centre of the debate over privacy and the use of personal data.
In some senses however, the privacy genie is already out of the bottle. As a report on big data commissioned by the Obama administration last year noted: “the impact of Big Data has the potential to be as profound as the development of the Internet itself,” and that impact is likely to be felt by all of us.
Friday, December 7, 2012
Panjiva Uses Government Data to Build a Global Search Engine for Commerce
Successful startups look to solve a problem first, then look for the datasets they need.
Alex Howard, O'Reilly Strata, December 6, 2012
“If you go back to how we got started,” mused Josh Green, “government data really is at the heart of that story.” Green, who co-founded Panjiva with Jim Psota in 2006, was demonstrating the newest version of Panjiva.com to me over the web, thinking back to the startup’s origins in Cambridge, Mass.
At first blush, the search engine for products, suppliers and shipping services didn’t have a clear connection to the open data movement I’d been chronicling over the past several years. His account of the back story of the startup is a case study that aspiring civic entrepreneurs, Congress and the White House should take to heart.
“I think there are a lot of entrepreneurs who start with datasets,” said Green, “but it’s hard to start with datasets and build business. You’re better off starting with a problem that needs to be solved and then going hunting for the data that will solve it. That’s the experience I had.”
The problem that the founders of Panjiva wanted to help address was one that many other entrepreneurs face: how do you connect with companies in far away places? Green came to the realization that a better solution was needed in the same way that many people who come up with an innovative idea do: he had a frustrating experience and wanted to scratch his own itch. When he was working at an electronics company earlier in his career, his boss asked him to find a supplier they could do business with in China.
“I thought I could do that, but I was stunned by the lack of reliable information,” said Green. “At that moment, I realized we were talking about a problem that should be solvable. At a time when people are interested in doing business globally, there should be reliable sources of information. So, let’s build that.”
Today, Panjiva has created a higher tech way to find overseas suppliers. The way they built it, however, deserves more attention.
Government data as a platform
By 2009, the startup had an initial product they could bring to market and launched a search engine that used government data as a platform for international trade. An importer could type in “patio furniture” and
determine who shipped it and who their customers were. The company chose a freemium model, where search is available for free but relationships between suppliers are only available to subscribers. The mapping of relationships between buyers and suppliers is where Panjiva delivered added value on top of public data.
That added value is crucial, given that competitors can also request and use the dataset. “Companies have been packaging and reselling this data in one way or another for years, ” said Green. “If you looked at this data, people are going to find value. It’s typically folks in the shipping industry, who want to know what’s going into ports or moving on different shipping lines. For us, the central purpose of the data was something different and required more work.”
That work paid off. In 2010, Panjiva built a search engine for global commerce that worked. Today, they have more than 100,000 users in 190 countries using its free service and some 3,700 companies subscribing to the paid version, including 42 Fortune 500 companies.
Notably, the Department of Homeland Security (DHS) itself is also a paying subscriber. Green declined to disclose the terms of relationships with all of Panjiva’s partners or data suppliers, some of which include nonprofits. Some users “do a revenue share, some are paying for data, others are providing data because they think there’s public good for that data being on the platform,” he said.
Panjiva competes with ImportGenius, Zepol, AliBaba and PIERS. Green credits PIERS for extracting similar value from customs datasets.
The turning point
When they started, the first approach that Green and his co-founder decided to take was to build a “Yelp for global trade” that would be based on feedback from people who work with companies. Unfortunately for the young startup, they couldn’t get off the starting blocks in generating reviews, much less reach critical mass.
They also encountered a new problem: even if they were able to get ratings of exporters and suppliers, how would they ensure the reviews came from people who had actually done business with the entities being rated? In retrospect, that focus was a bit silly, said Green, because they couldn’t get engagement, but talking about how to solve it led them to an unexpected answer: government data.
That direction came from a meeting where a staffer for a trade promotion organization told them it was straightforward to get shipping data on what’s coming into the country from the United States Customs Agency, which is now part of the Department of Homeland Security.
“It was a turning point for the company,” said Green. “We realized there was a dataset available to the public for a fee. They make available data about shipments that enter the U.S. While not all data is made available to the public and there are a bunch of limitations, the data that is made available is amazing. There’s about 10 million shipping records every year, typically including who is sending goods, who is receiving goods, what’s inside, and how much is inside a container.”
While useful, these government datasets do come with inherent limitations, cautioned Green. For one, they only contain data about shipments coming into the United States, not what’s going into Europe or Asia. For another, the data made available to the public only covers shipments made by boat, which is about about half of the shipments that come into the United States.
“It’s unfortunate that government cannot make available data on other modes of transport,” observed Green, with a hint of frustration in his voice. “That leaves out truck, rail, and air. Congress actually attempted to clarify that the regulations that govern this data weren’t just about boats but applied to air. Thus far, DHS hasn’t acted.”
Given the lens that has been focused on trade deficits between other countries and the United States in recent decades, there’s also a political angle to the market intelligence Panjiva provides that Congress and taxpayers may find of interest. For instance, Panjiva data showed global trade growth slowing in the first part of 2012.
“What we’ve organized, by its nature, gives us insight on companies around the world that serve the U.S. market,” said Green, “We’re helping people find overseas suppliers. Why not help find suppliers here at home? It turns out there’s a similar story on export data that’s supposed to be made to the public as well. DHS has a hard time with that as well. We can’t get the data.”
Data availability is also affected by the actions of the companies themselves, which have the ability to petition the government to hide shipments that are coming to them. “In about a third of the cases, you cannot see who is sending and receiving the goods,” said Green. “Government can see, but what’s released to the public has information pulled from it.”
This government data comes at a cost
Accessing this public data comes at a cost of some $100 per day, which is the service fee DHS charges for providing a daily CD-ROM. Each disc includes one day’s worth of shipments, which is generally around 30,000 shipping records. Panjiva started requesting data on July 1, 2007, and now has a little over five years of records.
“This data, on a record-by-record basis, is interesting,” said Green. “If you can organize, it’s phenomenal. If you can associate with companies, can say this company has experience with these supplies and this company has experience with these customers, it’s very useful in deciding if a company is a good fit. You can see by customers if they’re reasonably high quality.”
Making those CD-ROMs into a useful, searchable resource, however, was far from a simple matter of just inserting them into an optical drive and moving their contents into a structured database.
“Jim and a team of engineers went to work organizing the datasets initially,” said Green. “They were very hard to work with — absurdly messy. Think about the number of ways you can misname a Chinese factory. It was really problematic. You need to build company profiles, correct for misspellings and variations on names. We spent years getting that right.” Eventually, Panjiva was able to automate the process of ingesting the data from the CD-ROMs, building an algorithm to take the data and clean it up.
Making data a strategic asset
Panjiva’s initial foray, which created a search engine for customs data, didn’t meet with strong demand out of the gate. As they refined the product, it generated what Green described as a “nice business.” The startup was profitable, in other words, but its leadership aspired to build something bigger.
The direction they took was driven by user feedback. When Panjiva also asked its users about how they were making buying decisions, they saw a pattern emerge that looked like a bigger opportunity.
“Users started with Panjiva then went to search for additional information on B2B sites or on Google,” said Green. “We heard this process and it sounded a lot like the experience consumers had searching for flights before search engines or Kayak.com — except that instead of airline sites, people are going to B2B sites. The difference is it’s not just every airline. It’s like every flight has its own website.”
The founders now have raised just under $10 million from Battery Ventures and Harrison Metal, and invested it in technology and data acquisition. They’ve now grown their engineering team to 10 people, out of a total of 50 or so current employees. The engineering team is focused on improving search and enriching Panjiva’s data with other sources, beyond government data.
This October, the startup relaunched Panjiva.com with another layer: data supplied by the companies themselves.
“We have a database of six million companies spread around the world and contact information on four million companies,” said Green. “We have product photos for 34 million products. There was a lot of investment required to do that, but none of this would be possible if we hadn’t had a backbone of data that came from the U.S. government.”
The data sources that Panjiva integrated were also driven by customer interest. As the founders shared their product with potential subscribers, they kept hearing the same thing: 1) “that’s awesome” and 2) “I’d like more data.”
“We loved the first one and hated the second,” said Green. “In retrospect, we should have loved both. The second one was a roadmap for us to build them a really great differentiated product.”
When they asked users exactly which kinds of data would make the service more useful, a map to the future of the company emerged.
The first was operational data. “Customs data is a perfect example,” said Green. “It gives you a sense of what companies have done and their track record.”
The second was financial data. “Sure, a company has experience, but are they financially healthy?” asked Green. “Some of that you can infer, but there’s other things you can use. We’ve partnered with Dun & Bradstreet and Experian to pull that data into our platform.”
The third was positive and negative data about a company. “That includes getting certified as financially responsible,” said Green. “We’ve partnered with nonprofits and added that data, showing you information about companies doing wrong, including a blacklist of illicit global trade.”
The key insight that anyone interested in building a business on top of government data should take away here is to go beyond.
What happens if government data becomes open?
Green thinks that Panjiva is well-positioned to be both competitive and profitable, even if DHS decided to start publishing customs data online. “We don’t worry that much about data becoming more accessible,” he said, “even if government data becomes free. It’s not the $36,500 per year to buy the data — it’s the engineering talent to clear it up. That’s a massive problem, and it wouldn’t be as simple as getting the data.”
Panjiva is betting that the investments they’ve made in technology, talent and — crucially — combining so many different data sources have created a differentiated product that solves a problem for its customers.
“We’re not trying to build out a data business where we’re reselling government data,” said Green. “We’re trying to build a platform where serious buyers and sellers can connect. We’re now going to the world’s most important buyers. We have two revenue streams: selling premium access to data and selling access to suppliers who want it. The starting point for customers is $99 per month, going up to $10,000 per month for unlimited access for an unlimited number of users, then services that we sell on the top.”
The experience that Panjiva has had with government data and building a business using it has left Green with a strong perspective on what works — and what doesn’t.
“We don’t think there are infinite numbers of possibilities in terms of ways to build sustainable value with public data,” he said. “One is to take datasets that are commoditizeable and add value. Another is to feed the creation of more data. Another is to build a service. Another is to create network effects, where the data is the honey that attracts the bees.”
Most important, Green suggested, is to use public data to solve a problem that’s both hard and important. For Panjiva, that means making global trade more efficient and more transparent.
“There is a future where information is consolidated and accessible to people making key decisions, from a buying or regulatory standpoint,” he said. “Once that happens — and we’re close — there’s potentially a place where there’s a race to the top instead of the bottom, in terms of supply chain records. That will make a difference when you’re under scrutiny. Right now, the fragmentation of data is the ally of bad behavior. Our hope is to change that reality.”
Thursday, December 6, 2012
Ethan Roeder (Big Data Czar for Obama) in NYT: I Am Not Big Brother
Ethan Roeder, The New York Times, December 6, 2012
I’VE grown accustomed to reading inaccurate accounts of my day job. I’m in political data.
If I’m not spying on private citizens through the security cam in the parking garage, I’m probably sifting through their garbage for discarded pages from their diaries or deploying billions of spambots to crack into their e-mail. Reading what others muse about my profession is the opposite of my middle-school experience: people with only superficial information about me make a bunch of assumptions to fill in what’s missing and decide that I’m an all-knowing super-genius.
Sadly for me, this is a bunch of malarkey. You may chafe at how much the online world knows about you, but campaigns don’t know anything more about your online behavior than any retailer, news outlet or savvy blogger.
There are two categories of online data: information users provide explicitly, and stuff they communicate implicitly through their behavior. The explicit data includes e-mails and comments that users share directly. The implicit data comes from “click tracking,” which tells a campaign what buttons are getting pressed and how often. Combined, these two categories of data allow a campaign to put together an online experience that will resonate with as many people as possible, but also to customize the experience so that you are more likely to encounter content that’s relevant to you.
At times it might seem like sorcery to the recipient of a targeted e-mail, but it’s just a product of two simple factors: remembering who you are and remembering what you like.
In the offline world, which is my personal area of expertise, campaigns don’t know much more today than they did 44 years ago. In 1968, George Romney, then a candidate for the Republican presidential nomination, made headlines for using a newfangled “secret weapon” — a voter file. As The Times explained that winter, it is “an electronic data bank” that “contains the only really accurate, up-to-date roster of enrolled New Hampshire Republicans any candidate here has ever compiled, plus pertinent information about all of them.” Imagine how freaked out those New Hampshire Republicans must have been.
Virtually all of the offline data that people like me traffic in is boring, basic and publicly available. Want to know the year of birth for everyone who is registered to vote in Ohio? Just Google “Ohio voter file download.” There you go. I was born in 1976. Now we’re even.
How do we predict whether people are going to vote or not? We look at the voter file. It tells us how often a person votes, although not for whom. Not all strategists agree about how to interpret this information, but the source of the data is no secret.
What’s really new in politics today is not the data itself but how campaigns make sense of it. Cheaper and more plentiful computing power allows campaigns to process far more information than ever before to look for patterns, trends and correlations.
The science of modeling is a modern-day application of a practice that has been around for nearly 200 years: polling. Pollsters ask voters whom they support for president and how strongly. Campaigns then take demographic information about these voters into account in order to make assumptions about the entire population of a given state. The mechanics are exactly the same for public polls and internal campaign analyses. The difference is that the campaigns use statistical techniques to apply these assumptions to individual records in the voter file rather than stopping short and simply assuming that entire sections of the electorate will behave identically.
Contemporary data practice also frees campaigns from having to make assumptions about voters in the first place. In 2011 and 2012, the Obama campaign, with the help of more than two million volunteers, had more than 24 million conversations with voters. Online tools gave Obama supporters resources to help them play a crucial role in their neighborhoods, and a series of “share your story” pages on the campaign Web site provided a venue for voters to communicate directly with the campaign in long form.
All of this feedback doesn’t neatly boil down to a “yes” or “no” in a database — and why should it? Numerous avenues of listening, combined with the digital capacity to hold on to qualitative feedback, make campaigns aware of the differences among voters’ motivations, attitudes, protestations — not just their demographics and voting history. In a nation of over 200 million eligible voters, technology is allowing campaigns to finally see through the fog of the crowd and engage voters one by one.
In other words, there is no giant blue computer sitting on the 101st floor of a sleek skyscraper, surrounded by bubbling tubes of illuminated liquid, spitting out the manifest destiny of America’s voters. Campaigns are moving away from the meaningless labels of pollsters and newsweeklies — “Nascar dads” and “waitress moms” — and moving toward treating each voter as a separate person.
In 2012 you didn’t just have to be an African-American from Akron or a suburban married female age 45 to 54. More and more, the information age allows people to be complicated, contradictory and unique. New technologies and an abundance of data may rattle the senses, but they are also bringing a fresh appreciation of the value of the individual to American politics.
Ethan Roeder was the data director of Obama for America.
Wednesday, December 5, 2012
Why Data is the Key to Better Medicine - And Maybe a Cure for Cancer
Derrick Harris, GigaOm, November 28, 2012
The health care industry might have embraced the big data movement with open arms, but embracing it with open data probably would be more effective. Hospital organizations, researchers and the tech companies serving them have lots of great ideas — and have achieved some great results, too — but, ultimately, efforts to use big data to transform the industry will only be as good as the data these stakeholders have to work with. Right now, that isn’t always everything they need.
Have access, will innovate
Wired published an interesting profile Tuesday morning that exemplifies what’s possible when smart people have access to good data. The piece showcases the work of a man named Fred Trotter who has accessed reams of buried Medicare data via a Freedom of Information Act request and is uncovering some potentially valuable information. Already, the article explains, he has built a “Doctor Social Graph” by analyzing some “60 million relationships between doctors, and how often they refer patients to one another.” His next mission is to build a doctor rating system based on data he’s uncovered about credentials, nursing home inspections and other relevant info.
Elsewhere, companies such as Palo Alto, Calif., startup Apixio are trying to make hospitals more efficient by using semantic analysis to connect the dots between patient charts, electronic medical records, billing data and whatever other sources of information that hospitals generate. (We covered Apixio in early 2011, although the company has significantly expanded its services since then.) In health care, everyone seems to have their own way of doing things, as Apixio natural-language-processing scientist Vishnu Vyas told me recently, so “the variety of the data becomes as important as the volume of the data.”
Linda Drumright, GM of the Clinical Trial Optimization Solutions group at IMS Health, agreed. She explained that her company is able to do its job because it has access to mountains of data from pharmacies, insurance claims, medical records, partners and other sources. All told, it houses 17 petabytes of data spread across 5,000 databases. Her division’s clients, which generally include pharmaceutical and biotech companies running patient trials, need all this data in order to ensure their trials will actually be successful.
One recent customer wasn’t able to recruit test subjects fast enough, she noted, and IMS helped it comb through its criteria about who to or not to include in the trial only to find “that the patient population they were looking for didn’t exist.” As IMS went back and began eliminating criteria and iterating design, it realized that trial never should have begun in the first place.
There are a million ways to think about how to use this data, Drumright said, and as more customers begin to fully understand what they can do with it, her goal is to “make this information accessible in a way where it’s easy at the point where it’s needed, and consumable where it’s needed.”
The key to curing cancer might be more data
But whatever Trotter, Apixio, IMS and others accomplish will have been made possible because they have access to some valuable datasets, albeit not always with great ease. Many individuals who’d like to improve the health care system — if not our health, generally — aren’t so lucky. Take, for example, the world’s genetic researchers. It’s very possible the data they need to discover the medical Holy Grail of a cure for cancer is locked in gene sequence data that only very few people will ever see.
According to University of California, Santa Cruz researcher David Haussler, the limited access that many geneticists and computer scientists like himself have to valuable genetic data is “a crime.”
“We are on the brink of a real new understanding of cancer by being able to sequence cancer genomes,” he told me during a recent interview, but big data will be the key to unlocking it.
There are 1.6 million cases of cancer in the United States every year, Haussler explained, and most of the information from those tumors is being ignored. This is partially because of privacy restrictions about who can access personal medical data and for what purposes, and partially because there isn’t yet a concerted effort to collect the necessary genetic samples. As genome sequencing gets faster and cheaper, he says researchers need access to healthy and cancerous samples from the same person — and as many of these samples as possible — in order to analyze the “astounding” number of molecular changes that occur in every type and variation of cancer.
“We can’t completely understand what we’ll find, but we know we the only way we’ll pull out signal from the noise is to [analyze all these genes],” Haussler said.
Haussler understands the need for privacy regulations, but thinks there’s an opportunity to at least ease some current restrictions on how researchers access data. Even when there are relatively large (if not ideal) datasets available such as with the Cancer Genome Atlas project, researchers must apply to the National Institutes of Health for access, and the data must always remain behind an organizational firewall. Every cancer patient in the country could agree to having their data available to researchers, he said, but as long as that data isn’t accessible over the internet it’s only of limited utility.
He — along with others in the field — thinks cloud computing could be the solution because it gives genetic researchers a central location where they can access and perform computations on the data. Haussler and his team that house the Cancer Genome Atlas and a couple other projects currently have more than 400 terabytes of data and expect to have around 5 petabytes of data eventually. Downloading that is infeasible save for access to high-speed research networks, so “we need a place where people can experiment with these big data problems,” Haussler said.
In the meantime, Haussler and his peers will keep on collecting and accessing genome data however they can. And they’ll keep building software packages and algorithms that analyze that data better and faster than ever before. However, he lamented, “If we had the big data out there in an unrestricted setting, then all the best minds in the world would already be crunching on it.”
Policy Debates in the Internet Age
Jon Peha, Reuters, December 5, 2012
Technology is changing how power struggles are waged between the White House and Congress. For the last few years, negotiations between Democratic and Republican leaders have too often led to stalemate. The battle over how to avert the “fiscal cliff” is the latest example.
Since President Barack Obama’s reelection, he has begun to shift strategies — taking his case directly to the American people as a way to pressure Congress. After all, members of Congress ignore their president without penalty, but ignoring the opinions of their constituents can cost them their jobs.
Presidents Ronald Reagan and Bill Clinton both effectively used television to address the nation when facing off against a House of Representatives controlled by the opposing party. While TV will remain important, going directly to the American people continue to morph in the era of the Internet. Political messages can be customized and narrowly targeted.
Much of the political broadcasting of the past may ultimately be replaced by political narrowcasting. We saw this already during the 2012 presidential campaign — candidates began with broad appeals to the nation and ended up focusing on a relatively few undecided and therefore persuadable groups living in key swing states like Ohio and Virginia.
We may even see specific groups of American citizens playing the role of jury, as they are bombarded with carefully tailored appeals from both sides, while the rest of us remain apart from all the sound and fury.
On the Internet, messages can be customized based on the recipient’s congressional district. This could allow Obama, who never has to run for reelection again, to directly challenge every opposing member of Congress by name during the 2014 elections — especially those who won narrowly in 2012.
Political messages can be fitted to the interests and political leanings of each recipient. For example, as the president makes the case for his energy policy, he could emphasize to some the jobs this policy could create, to others the environmental benefits – depending on their interests.
Obama’s targeted messaging probably begins with the treasure-trove of information collected for his re-election campaign. In addition to email addresses , the re-election campaign staff also knows where supporters live, and even which issues interest them most. In supplying information to voters, the re-election team regularly asked about their interests, in addition to basics like their home zip code and perhaps a Facebook account. The team could also have maintained a record of which position papers a voter views, and the campaign events she signs up for.
When a customized email now asks one of these supporters to contact his representative in Congress about tax reform, for example, the request can include the telephone number to call and specifics on what that representative has said and done about tax policy — perhaps affecting that representative’s standing in the polls in the process.
The request could also urge supporters to use social media such as Facebook and Twitter to influence friends and neighbors.
The Republican Party and many other organizations can, of course, use these same techniques. But the Obama re-election campaign reportedly used sophisticated methods of analysis, and it looks like the GOP probably now lags behind the president’s re-election team in information and organization.
Though email is powerful and free, it is most useful for reaching those who want to be reached. To bring targeted messaging to a broader audience, including those swing voters who will likely decide the 2014 congressional elections, political leaders and other organizations could turn to the advertisements that people see as they browse the Web.
There is a common misconception that these online ads are like billboards, which look the same to everyone driving by. In reality, however, many online ads are selected specifically for the person viewing them, not far from the advertising posters in the sci-fi movie, “Minority Report,” which speak directly to the characters walking by.
Consider, when you access a website, it knows from your IP (Internet protocol) address roughly where you are located, and can show you an ad praising or condemning the member of Congress representing your congressional district. When choosing which ad you should see, this site can also use a variety of techniques to know which websites you frequent, which Google searches you have tried this month, or the value of homes in your neighborhood — all to get a better idea of where you stand on the big issue of the day.
If targeted messaging becomes an important force in policy debates, the impact will depend on which groups are being targeted. As both sides seek to reach beyond their base to swing voters in contested congressional districts, this may even give our leaders a newfound incentive to work together toward shared solutions.
While many activists like to hear the kind of tough rhetoric that leaves little room for compromise, a politician usually gets more support from swing voters with pragmatic proposals that actually can get things done.
Tuesday, December 4, 2012
Zachary Lemnios joins IBM
Steve Lohr, New York Times Bits Blog, December 3, 2012
Last Friday afternoon, a half-hour after he said his goodbyes at the Pentagon, where he was assistant secretary of defense for research and engineering, Zachary Lemnios was discussing why his new job made sense as the next step in his career. “Exactly the right thing to do at the right time,” he said.
Mr. Lemnios joined I.B.M. Research on Monday, as vice president for research strategy. Mr. Lemnios, 58, is leaving his post in the Obama administration, after nearly four years (political appointees typically last two to four years). Before that, he was chief technology officer at M.I.T.‘s Lincoln Laboratory, a federally financed research center for advanced technology with national security applications, and previously a senior official at the Pentagon’s futuristic research arm, the Defense Advanced Research Projects Agency.
His own career charts a path from being a chip guy who would later champion and fund artificial intelligence research. Mr. Lemnios holds four patents on semiconductors that use gallium arsenide, an alternative to silicon.
Of course, he noted, the years of progress in microprocessor design and sensors are essential to the recent advances in artificial intelligence. “The hardware substrate is certainly part of it,” Mr. Lemnios said. “But it is the software for reasoning and learning that really pushes this forward.”
In late March, for example, Mr. Lemnios announced $60 million in new Pentagon-supported Big Data research projects, as part of a $200 million administration initiative in the fast-growing field of trying to use smart technology to make sense of the explosion in data from the Web, sensors and streaming into traditional databases.
Today’s technological limitations, he said, are no longer the data collection tools but the technology for making sense of it. “The real challenge is to handle, understand and use big sources of data,” Mr. Lemnios said.
At Darpa, he promoted initiatives for “cognitive systems,” an approach to artificial intelligence that embraces a learning model of computing rather than more purely statistical methods. I.B.M.’s labs have been at the forefront of “cognitive computing” research in recent years, including projects financed by Darpa.
Mr. Lemnios had praise for I.B.M.’s best-known research project, the Watson question-answering computer. “It moves toward what we think of as understanding information, which is the start of another revolution,” he said.
But Mr. Lemnios said he had also been impressed with the company product and services offerings, which include large contributions from I.B.M.’s research labs, for traffic management, energy conservation and crime prevention. They are part of the company’s Smarter Planet projects. “There is a lot of deep technological meat on the bones of Smarter Planet,” he said.
Having smart people is a crucial ingredient in a successful research operation, but so is research strategy and management, Mr. Lemnios noted. On strategy, he said, “balance across a portfolio is important” — that is, balancing work on current products, research bets that may pay off in five to 10 years, and further out exploratory research.
In good times and bad, he observed, I.B.M. has had the managerial patience to continue financing long-range, exploratory research. That is a model, Mr. Lemnios suggests, that the federal government would do well to follow.
Given the need for budgetary belt-tightening in government, Mr. Lemnios said, “There is great pressure to take funding for exploratory research to pay today’s bills. That can prove to be a shortsighted mistake.”
Social Media Report 2012 (Nielsen)
Social Media Use Exploded in 2012, Led by Pinterest
Stephanie Mlot, PC Magazine, December 3, 2012
It's hard for some to remember life before the Internet, let alone social media, and now it appears that these sites are coming of age.
Consumers spend more time on social networks than any other site — about 20 percent via a PC, and 30 percent using a mobile device, Nielsen's Social Media Report revealed.
Still not impressed? The report also tips a 37 percent increase in the total time spent on social media in the U.S., reaching 121 billion minutes in July, compared to 88 billion in summer 2011.
The ultimate social network, Facebook, remains the most-visited in the U.S., and earned the title of most popular Web brand in the U.S. this year. It reached 152.2 million PC visitors, 78.4 million app users, and 74.3 million mobile Web surfers. That dwarfs social sites in all categories (see graphic below), beating No. 2 Blogger by more than 93.7 million PC users. In the mobile app race, Facebook, Twitter, Foursquare, Google+, and Pinterest carry the top five spots.
Top networks like Facebook and Twitter have profound staying power in a world where social media options are expanding every day, including breakout star Pinterest, which boasted what Nielsen reported is the largest year-over-year increase — 1,698 percent. For more, check out PCMag's Pinterest Board.
Overall Internet connections are on the rise, as well, drawing more than 80 percent increases in mobile Web and app usage in the last year; PCs dropped 4 percent. Meanwhile, tablets, handheld music players like the iPod touch, game consoles, Internet-enabled TVs, and e-readers are slowly boosting connectivity.
With so many mobile options, it appears nearly a third of people ages 18 to 24 write Facebook comments, send tweets, and perhaps even blog from the comfort of their bathroom. Those ages 25 to 34 are more likely to use social networking in the office.
For all of the celebrity accounts and political usage accrued by social media, it seems that people still focus on connecting with friends and family members. More than 60 percent of people turn on Facebook to keep up with someone they know in real life, while 9 percent initiate LinkedIn contact because of a person's physical attractiveness, Nielsen reported.
Report at http://blog.nielsen.com/nielsenwire/social/2012/
Monday, December 3, 2012
Lady Gaga: The Value is in the Details
Ravi Mattu, The Financial Times, November 30, 2012
When Lady Gaga, the singer and social media star – with over 31m Twitter followers, more than anyone else, and over 51m Facebook likes – finished watching a screening of The Social Network, she called her manager Troy Carter. "She said she'd like to build a social network for her fans" – she calls them her little monsters – "and build a community where they could congregate and have conversations," he says. "So I called some of my friends in the Valley."
One of those friends was Joe Lonsdale, co-founder of the Palo Alto-based data management company Palintir. "He said, 'Send me all the data you have.' So, we sent him everything and he said it was the worst data he had ever seen in his life." The problem wasn't the amount of data – they had lots of it, from Ticketmaster, Lady
So, with the help of Mr Lonsdale and backing from Google Ventures among others, Mr Carter created Backplane, a new social media platform on which sits Littlemonsters.com. The site is designed to cater to the "hardcore 1m" Lady Gaga fans because their behaviour is more valuable than trying to decipher what happens to a mass audience.
"Our bet is on the future of micronetworks," he says. "Facebook wasn't wired to build a relationship between fans and artists. It's more about communicating with family and friends and old girlfriends or your classmates; 51m likes doesn't mean we're going to sell 51m albums or concert tickets."
It is not just about deeper insight. It is also about getting rid of the middleman. It is a "misconception when people talk about a direct relationship between artists and their fans or brands and consumers through social media. The reality is that these platforms own the relationship. So as much as you can talk directly to a customer or a fan, you still have this intermediary . . . that controls the data. And at any given time, if they turn it off or they change an algorithm, like Facebook did with its newsfeed algorithm last year, it changes the way you're able to communicate with that fan or customer."
Listening to the sharp-suited 40-year-old at the FT Innovate conference calmly skewer the shortcomings of a technology company he claims is a friend – he is at pains to say that his efforts are complementary rather than competitive, and spoke to Facebook and others before launching – it is hard not to see him as a disrupter taking on the social media elite.
It is one reason he is here. His use of social media, particularly with Lady Gaga, has made him a much sought-after voice among the corporates drowning in the sea of big data.
Given his increasingly high profile, it is sometimes difficult to imagine the entrepreneur's childhood in inner-city Philadelphia. His single-parent mother worked at a hospital for 30 years, cleaning surgical instruments, while raising him and his brothers. She often worked long shifts, starting at 5:30 in the morning, so "the streets were raising us at the same time".
For an African-American boy growing up in that world, there were not a lot of options. "You got a couple of choices: drug dealers were the role models – you didn't have doctors and hedge fund managers that looked like you," he says. Or music. "At that time, hip hop culture was exploding . . . and coming from the family I came from, drugs was not an option."
While that may sound like a scene from The Wire, Mr Carter says being an outsider has been key to his success. "Being born in the adolescent years of hip hop helped us learn about flux. And when you're in an industry that is constantly growing, changing, maturing . . . you get a chance to try different things out and a chance to fail."
In high school, he got to know fellow Philadelphians DJ Jazzy Jeff and the Fresh Prince – the actor Will Smith – and became an assistant carrying the hip hop duo's records from gig to gig, before setting out on his own as a music promoter. It was during this time that he met rapper and producer Sean Combs, now known as P Diddy, who gave him a job at his Bad Boy Records.
Mr Carter says this is where he learnt about the record business and P Diddy's example taught him "that you can be a young black entrepreneur with no college degree or any sort of experience and people will give you a shot in this business".
He later set up a boutique talent management company, which he sold to the Sanctuary Group, then part of Universal Music. He quickly discovered that being in a big company was not for him. "Instead of me being able to be creative with the artists, I was sitting in finance meetings a couple of times a week. It killed my spirit as an entrepreneur."
But he also understood the value of getting the organisational culture right. When he launched Atom Factory, hiring other outsiders was essential. "My COO didn't come from the music industry, my vice-president of creative was actually a schoolteacher," he says. "It was important we had people who came from an outside perspective, who didn't come from selling CDs."
As well as Lady Gaga, the group represents John Legend and Bollywood star Priyanka Chopra as she launches a music career outside India.
Alongside Atom Factory sit AF Square, an angel investment fund with stakes in a number of mostly tech start-ups, including news app Summly, the taxi-hailing app Uber and music streaming site Spotify, and A/Idea, an ideas lab. Mr Carter has announced plans to launch a drink called Pop Water.
He declines to disclose profit figures but says the group has grown 60-70 per cent year on year for the past four years and he is sole shareholder.
Mr Carter is still best known as the manager of Lady Gaga, partly because of how he used technology to circumvent mainstream radio when she struggled to get her music played on it.
Once again, via Littlemonsters.com, she is a beta test with a view to understanding how Backplane could be employed by other companies to build communities. He is working with a shoe brand on tapping into "sneaker culture", for example.
"Right now, we're planting the seeds of an oak tree. What we are planting today, we may not see the full benefit for five to 10 years," he says, pointing to the fact that many of the core tenets of the music business, such as digital rights arrangements, could change.
Still, the data being collected is already informing commercial decisions. For example, until Littlemonsters.com went live six months ago, Lady Gaga had never toured or been promoted in South America, a region that is not a big music market in terms of album sales and downloads. But "once we launched the website, we were able to get a lot of info about fans and specifically the numbers of them in South America", leading to the decision to add a number of dates there to the singer's Born This Way Ball tour.
But however excited one gets about data, Mr Carter offers a word of advice for anyone thinking about using it to tinker with the creative side of the business: don't. "I stay away from the arts . . . writing songs, being creative – those are downloads from god. You can't do data analytics on art."
Sunday, December 2, 2012
Digital Privacy in the Big Data Era: Microsoft's Data Protection Keynote
Ms. Smith, Network World, December 2, 2012
There are several Internet security experts who agree with Steve Rambam's claim [1] that "Privacy is dead - get over it." Yet other privacy and security experts such as Bruce Schneier completely disagree. In The Value of Privacy [2] Schneier wrote, "Privacy protects us from abuses by those in power, even if we're doing nothing wrong at the time of surveillance." When it comes to data protection and protecting people's privacy in the digital age, Europe is far more advanced than America. [3]
In fact, the head of France's data protection agency, Isabelle Falque-Pierrotin, did an excellent job summing it up as: "In Europe, we consider privacy a fundamental right. That doesn't mean it is exclusive of other rights, but economic rights are not superior to privacy." The New York Times also reported [4] that she said in the United States, "personal data are seen as raw material for business."
In November, Microsoft's Chief Privacy Officer Brendon Lynch said [5] of the IAPP European Data Protection Congress 2012 [6], "One area of strong consensus was the tremendous potential the digital economy holds for companies on both sides of the pond. Accordingly, it's important to strike the right balance between data protection with business growth through interoperability between privacy regulation in the EU, U.S. and elsewhere."
Many privacy advocates cringe when hearing the word "balance," such as striking a balance between security and privacy. Hopefully people won't come to cringe when they hear the word balance applied to big data security protections and privacy. As Bruce Schneier wrote [2] way back in 2006:
Too many wrongly characterize the debate as "security versus privacy." The real choice is liberty versus control. Tyranny, whether it arises under threat of foreign physical attack or under constant domestic authoritative scrutiny, is still tyranny. Liberty requires security without intrusion, security plus privacy. Widespread police surveillance is the very definition of a police state. And that's why we should champion privacy even when we have nothing to hide.
Whether people realize it or not, big data is not privacy-friendly even when it is supposedly anonymized or contains obfuscated PII (Personally Identifiable Information) data. Researchers have shown that "linkability threats" can re-identity individuals. Since it boils down to the fact that you are not anonymous when it comes to big data [7], Microsoft has developed "Differential Privacy for everyone" [download PDF [8]].
In the IAPP keynote address [download PDF [9]], Lynch made some excellent and thought-provoking privacy points regarding big data. He said:
Data is the fuel that drives all of these powerful technologies, but what can be done with the data today can at times seem enormously helpful or enormously threatening. Consider two scenarios shown here. In the first case, I am using my phone in a grocery store to find out more about the items on the shelves and it is mashing up that with my private data to personalize my experience. So here I downloaded a recipe and customized it for my dietary needs. If it's a trusted system, that's a great experience. On the other hand, consider the US company, Target, which recently generated a lot of press about its pregnancy prediction score. This was based on what people were purchasing in Target stores, they are able to indicate a shopper that appeared to be pregnant. The concern about how Target can figure out such details about customers shopping in its stores, who are not explicitly sharing that information, is the concern. And what does it do with those insights? In this particular case, they sent some mailers to the individual involved - it was a teenage girl and her father was very offended that they were wrongly marketing to her, but it eventually did come out that she was in fact pregnant. Target knew a lot more than her father knew.
Peter Cullen, Microsoft's Chief Privacy Strategist, wrote [10] about "notice and consent" as a means of privacy protection and how data privacy frameworks need to "focus on the 'harms' or 'impacts' of data use, which should not only include physical and financial injury, but also broader concepts such as reputational or social harm."
Yet after showing a video that highlighted data transfers in today's world at the IAPP conference, Lynch said, "How could there possibly be meaningful notice and consent mechanisms in place for every transfer of data that was involved?" He added, "It would seem that advances in technology and the rise of big data can create amazing societal benefits but they can also strain traditional notions of secrecy and the notice and consent approach to privacy protection."
In his keynote, Lynch said:
Some technology and internet companies today take the position that privacy is dead, or at least that privacy is an outdated concept that people need to get over so technology companies can help them reap the benefits of sharing as much information as possible. But we disagree that privacy is not relevant or desirable, in this sensor-driven, social everywhere, big data world that we are heading towards. People today expect strong privacy protections because they are increasingly aware of, and concerned about, the digital trails they leave behind online and indeed there's plenty of evidence that people still care deeply about privacy.
Of course people care about privacy. Europe continues to illustrate this to the world by taking a hard stance when data is used without "informed consent" and when users cannot "opt out." Lynch believes we need to not only protect privacy in regards to big data, but also that people "need an updated notion of privacy and data protection principles, one that shifts from a focus on secrecy to a more nuanced approach, based on reasonable consumer expectations, context and a greater emphasis on how personal information is used."
Big data definitely represents significant threats to personal privacy. Let's hope this "shift" and "updated notion of privacy" won't include the word "balance" that puts individuals on the losing end as it generally has when the government talks of striking a balance between security and privacy.
Subscribe to:
Posts (Atom)