Showing posts with label Start-Up. Show all posts
Showing posts with label Start-Up. Show all posts
Friday, January 18, 2013
A Cure for Cancer? This ‘Big Data’ Startup Says It Can Deliver
Christina Farr, The Washington Post, January 17
‘Big data’ is one of the most over-used buzzwords in the startup vernacular, and founders rarely have the goods to back it up. So you’ll understand that I was intrigued — but highly skeptical — when an email with the subject line “using data to cure cancer” popped into my inbox.
But Ayasdi, a startup that closed $10 million in venture funding Wednesday, doesn’t just talk the talk. Stanford researchers have been baking the complex algorithms behind Ayasdi (its quirky name means “to seek,” in Cherokee) for over a decade, with the goal of unlocking the hidden value in human genetic data.
In 2008, the founders, Gurjeet Singh, Dr. Gunnar Carlsson, and Harlan Sexton, decided to commercialize the technology.
With the government stepping up its funding for science, they were able to pull in $3.5 million in grants from DARPA, the department of defense agency responsible for building new technology for the military, and the National Science Foundation. The result? A synthesis of machine learning technology and topological data analysis (TDA) that has impressed a score of Silicon Valley investors.
Rather than typing in search-style queries, the tools allow for automated discovery of information. As Dr. Carlsson explained in an interview, “The idea is to answer questions that you didn’t know to ask.”
This year, the 30-person team of engineers will expand its marketing and sales efforts with funding from Khosla Ventures, Floodgate, Data Collective’s Matt Ocko, serial entrepreneur, Steve Blank, and more.
Storied investor Vinod Khosla, who rocked the medical world with the statement that 80 percent of doctors would be replaced by machines, said Ayasdi’s “machine powered intelligence” has the potential to unearth “previously unattainable insights that will help solve some of our most pressing global, social, and economic issues.”
Eric Schadt, Director of the Institute for Genomics and Multiscale Biology, has a team of researchers using the technology to identify the genetic predispositions of many diseases, including cancer, which they hope will help them “glean new insights that will lead to breakthrough drug therapies.”
Related: In the burgeoning field of genomics, entrepreneurs aim to deliver more personalized treatments for life-threatening diseases.
Ayasdi is working with the nation’s top hospitals and medical researchers to uncover more targeted treatments for disease. Singh, the company’s CEO, told me that hospitals and big pharmas are routinely pulling data from public sources — medical researchers are required to publish their data – and combine it with private data to yield new insights.
The data isn’t anything new — it’s the technology that has evolved. “We have automated the discovery of knowledge from data,” said Singh in a phone interview. “We were able to discover a new type of breast cancer without asking questions.”
Singh was referring to a recent breakthrough where Ayasdi mapped 14 variants of breast cancer. Using data collected during a 15 year period, and studied by thousands of scientists, the algorithms discovered a sub-group of patients that have a higher chance of survival based on their genetic profile.
If a patient falls into this sub-group, it is unlikely that they will require chemotherapy.
In another recent partnership, with Mount Sinai Medical Center, Ayasdi was used to point to targeted treatment options for E. Coli sufferers. E. Coli affects more than 265,000 people in the U.S. every year, and millions around the world. It is known in the medical community for developing resistance to many drugs, and doctors are never 100 percent sure if a treatment will work or not.
Mount Sinai is using Ayasdi to analyze the entire E. Coli genome sequence, which includes more than 1 million DNA variants. This will further our understanding of why some types of E. Coli develop resistance to antibiotics and how we can combat the spread of the bacteria.
Singh, a former researcher at Stanford, told me that the company has secured 20 customers in the oil and gas, government, pharmaceutical, and healthcare sectors. Big name customers include Merck, the Food and Drug Administration, and the U.S. Department of Agriculture.
Copyright 2013, VentureBeat
Wednesday, December 12, 2012
Making Dollars and Sense of the Open Data Economy
Is the push to free up government data resulting in economic activity and startup creation?
Alex Howard, O'Reilly Radar, December 11, 2012
Over the past several years, I’ve been writing about how government data is moving into the marketplaces, underpinning ideas, products and services. Open government data and application programming interfaces to distribute it, more commonly known as APIs, increasingly look like fundamental public infrastructure for digital government in the 21st century.
What I’m looking for now is more examples of startups and businesses that have been created using open data or that would not be able to continue operations without it. If big data is a strategic resource, it’s important to understand how and where organizations are using it for public good, civic utility and economic benefit.
Sometimes government data has been proactively released, like the federal government’s work to revolutionize the health care industry by making health data as useful as weather data or New York City’s approach to becoming a data platform.
In other cases, startups like Panjiva or BrightScope have liberated government data through Freedom of Information Act requests and automated means. By doing so, they’ve helped the American people and global customers understand the supply chain, the fees associated with 401(k) plans and the history of financial advisors.
I’ve hypothesized that open data will have an overall effect on the economy akin to that of open source and small business. Gartner’s research has posited that open data creates value in the public and private sector. If government acts as a platform to enable people inside and outside government to innovate on top of it, what are the outcomes?
Over the past four years, the world has heard a rising chorus for raw data from voices like the creator of the World Wide Web, Tim Berners-Lee, and the chief technology officer of the United States, Todd Park. Park, in particular, has been working to scale open data across the federal government as the nation’s “entrepreneur in residence.”
McKinsey and Associates estimated the annual economic value of big, open liquid health data at some $350 billion annually. While that number is eye-opening, which companies and startups stand to change health care using open health data?
Some examples are clear, from mobile apps like iTriage (now owned by Aetna) to Castlight, but they aren’t sufficient to understand what’s happening out there.
Other promising startups are in the consumer finance space, where so-called “smart disclosure” initiatives are enabling people to put their personal data to use. Startups like Billshrink.com and Hello Wallet now are enabling people to make smarter financial decisions.
I know there are more stories out there, and in sectors beyond health care and consumer finance — including transit, energy, education and media. Over the next several months, I’ll be identifying and profiling more civic startups, such as those from the first class in the Code for America accelerator, like Captricity, to specialized search engines, like Zillow, Panjiva and DataMarket.
In the course of that work, I hope to answer some big questions. What are the sustainable business models that successful civic startups are using, whether they use legislative data or other reuse of public sector information? What are the real costs associated with opening up government data to make it usable, both for government and entrepreneurs? And how does it balance against what datasets, at the federal, state or local levels, are the most valuable? Are they open and usable? If so, who’s using them and to what effect? If not, why not?
At the end of this particular project, in February, we’ll publish a report on what I’ve found. In it, I hope to be able to share some answers to several core questions on the topic. Where I need your help is in identifying new startups that are using or consuming government data or in highlighting how existing companies use it in their operations, good or services. Who is doing the most interesting work — and where? If you have research and evidence to share on the questions I posed above, feel free to ring in on that count as well.
Please weigh in through the comments or drop me a line at alex@oreilly.com or at @digiphile on Twitter.
Monday, December 3, 2012
Lady Gaga: The Value is in the Details
Ravi Mattu, The Financial Times, November 30, 2012
When Lady Gaga, the singer and social media star – with over 31m Twitter followers, more than anyone else, and over 51m Facebook likes – finished watching a screening of The Social Network, she called her manager Troy Carter. "She said she'd like to build a social network for her fans" – she calls them her little monsters – "and build a community where they could congregate and have conversations," he says. "So I called some of my friends in the Valley."
One of those friends was Joe Lonsdale, co-founder of the Palo Alto-based data management company Palintir. "He said, 'Send me all the data you have.' So, we sent him everything and he said it was the worst data he had ever seen in his life." The problem wasn't the amount of data – they had lots of it, from Ticketmaster, Lady
So, with the help of Mr Lonsdale and backing from Google Ventures among others, Mr Carter created Backplane, a new social media platform on which sits Littlemonsters.com. The site is designed to cater to the "hardcore 1m" Lady Gaga fans because their behaviour is more valuable than trying to decipher what happens to a mass audience.
"Our bet is on the future of micronetworks," he says. "Facebook wasn't wired to build a relationship between fans and artists. It's more about communicating with family and friends and old girlfriends or your classmates; 51m likes doesn't mean we're going to sell 51m albums or concert tickets."
It is not just about deeper insight. It is also about getting rid of the middleman. It is a "misconception when people talk about a direct relationship between artists and their fans or brands and consumers through social media. The reality is that these platforms own the relationship. So as much as you can talk directly to a customer or a fan, you still have this intermediary . . . that controls the data. And at any given time, if they turn it off or they change an algorithm, like Facebook did with its newsfeed algorithm last year, it changes the way you're able to communicate with that fan or customer."
Listening to the sharp-suited 40-year-old at the FT Innovate conference calmly skewer the shortcomings of a technology company he claims is a friend – he is at pains to say that his efforts are complementary rather than competitive, and spoke to Facebook and others before launching – it is hard not to see him as a disrupter taking on the social media elite.
It is one reason he is here. His use of social media, particularly with Lady Gaga, has made him a much sought-after voice among the corporates drowning in the sea of big data.
Given his increasingly high profile, it is sometimes difficult to imagine the entrepreneur's childhood in inner-city Philadelphia. His single-parent mother worked at a hospital for 30 years, cleaning surgical instruments, while raising him and his brothers. She often worked long shifts, starting at 5:30 in the morning, so "the streets were raising us at the same time".
For an African-American boy growing up in that world, there were not a lot of options. "You got a couple of choices: drug dealers were the role models – you didn't have doctors and hedge fund managers that looked like you," he says. Or music. "At that time, hip hop culture was exploding . . . and coming from the family I came from, drugs was not an option."
While that may sound like a scene from The Wire, Mr Carter says being an outsider has been key to his success. "Being born in the adolescent years of hip hop helped us learn about flux. And when you're in an industry that is constantly growing, changing, maturing . . . you get a chance to try different things out and a chance to fail."
In high school, he got to know fellow Philadelphians DJ Jazzy Jeff and the Fresh Prince – the actor Will Smith – and became an assistant carrying the hip hop duo's records from gig to gig, before setting out on his own as a music promoter. It was during this time that he met rapper and producer Sean Combs, now known as P Diddy, who gave him a job at his Bad Boy Records.
Mr Carter says this is where he learnt about the record business and P Diddy's example taught him "that you can be a young black entrepreneur with no college degree or any sort of experience and people will give you a shot in this business".
He later set up a boutique talent management company, which he sold to the Sanctuary Group, then part of Universal Music. He quickly discovered that being in a big company was not for him. "Instead of me being able to be creative with the artists, I was sitting in finance meetings a couple of times a week. It killed my spirit as an entrepreneur."
But he also understood the value of getting the organisational culture right. When he launched Atom Factory, hiring other outsiders was essential. "My COO didn't come from the music industry, my vice-president of creative was actually a schoolteacher," he says. "It was important we had people who came from an outside perspective, who didn't come from selling CDs."
As well as Lady Gaga, the group represents John Legend and Bollywood star Priyanka Chopra as she launches a music career outside India.
Alongside Atom Factory sit AF Square, an angel investment fund with stakes in a number of mostly tech start-ups, including news app Summly, the taxi-hailing app Uber and music streaming site Spotify, and A/Idea, an ideas lab. Mr Carter has announced plans to launch a drink called Pop Water.
He declines to disclose profit figures but says the group has grown 60-70 per cent year on year for the past four years and he is sole shareholder.
Mr Carter is still best known as the manager of Lady Gaga, partly because of how he used technology to circumvent mainstream radio when she struggled to get her music played on it.
Once again, via Littlemonsters.com, she is a beta test with a view to understanding how Backplane could be employed by other companies to build communities. He is working with a shoe brand on tapping into "sneaker culture", for example.
"Right now, we're planting the seeds of an oak tree. What we are planting today, we may not see the full benefit for five to 10 years," he says, pointing to the fact that many of the core tenets of the music business, such as digital rights arrangements, could change.
Still, the data being collected is already informing commercial decisions. For example, until Littlemonsters.com went live six months ago, Lady Gaga had never toured or been promoted in South America, a region that is not a big music market in terms of album sales and downloads. But "once we launched the website, we were able to get a lot of info about fans and specifically the numbers of them in South America", leading to the decision to add a number of dates there to the singer's Born This Way Ball tour.
But however excited one gets about data, Mr Carter offers a word of advice for anyone thinking about using it to tinker with the creative side of the business: don't. "I stay away from the arts . . . writing songs, being creative – those are downloads from god. You can't do data analytics on art."
Tuesday, November 20, 2012
Who Are the Doctors Most Trusted by Doctors? Big Data Can Tell You.
Ki Mae Heussner, GigaOm, November 16, 2012
ZocDoc, Healthgrades, Vitals, Yelp and other sites can tell you what patients think of their doctors. But finding out in any aggregate way what doctors think of their peers has been much harder, if not near impossible, for patients — up until now.
By accessing information in government databases through FOIA (Freedom of Information Act) requests, healthcare innovators are now able to share connections between doctors that are based on millions of physician referrals — a valuable indicator of who doctors hold in esteem.
Last month, Fred Trotter, a self-identified “hacktivist,” revealed that he had obtained a dataset of Medicare physician referrals through a FOIA request and was making the initial data available to those who supported a Medstartr crowdfunding campaign meant to build out his “DocGraph” and make it freely available. This week, he announced that he not only blew past his $15,000 funding goal, but was launching a second campaign to integrate his current data with an additional dataset.
HealthTap, a Palo Alto-based startup that connects patients with an online network of 17,000 doctors, also this week launched a new feature based partly on Trotter’s data. Called “DOConnect,” it combines Trotter’s Medicare data with physician data from its own site and other sources to give patients a new window into their doctors’ networks.
“This isn’t just friendships and business connections. This is who doctors trust,” said HealthTap co-founder and CEO Ron Gutman. “If you could know who your doctor’s doctor is, if you knew who they would choose, this lets you see that for the first time.”
The new tool, which reflects 25 million doctor referral connections, enables patients to see how many doctors are linked to a particular doctor, as well as their locations. As patients search for new physicians and specialists, being able to see who their current doctors are linked with could help them decide who to visit.
It also gives doctors an opportunity to build online networks that reflect their offline networks, Gutman said. In a post about his “DocGraph” project, Trotter said that his data wasn’t strictly a “referral” data set because, in some cases, doctors might be linked through a patient they both happened to see at the same time, not through an active referral. But Gutman emphasized that HealthTap’s DOConnect considered more than Medicare referrals in mapping connections between doctors.
In releasing the dataset, Trotter said his main goal was to create doctor-rating algorithms that “patients find useful and doctors find fair.” But he also hoped that academics, health policy wonks, entrepreneurs and others would use it to bring more transparency to health care overall.
Todd Park, the U.S. Chief Technology Officer, has frequently talked up the value of “setting data free” and has backed hackathons, “datapaloozas” and other open data initiatives to highlight the need for innovators to use government data for the public good — this is a great example of that vision and, hopefully, points to more similar projects in the future.
“Our goal is to empower the patient, make the system transparent and accountable, and release this d
Tuesday, September 11, 2012
Here Comes the Data Economy
New companies are creating services using government data on health care, education, and more.
Alexander B. Howard, Slate, September 10, 2012
We're living in the exabyte age, where the actions of billions of humans using the Web and their mobile devices are creating massive amounts of big data to collect, store, analyze, and put to work.
If big data is a strategic resource, as has been suggested, then many national and state governments have public reserves that can be tapped for the public good in this young century's version of the industrial revolution. Given that the United States economy is still coming out of the worst recession and financial shock since the Great Depression, supporting civic and tech entrepreneurs enjoys political support from both sides of the aisle.
Entrepreneurs, big and small, are mashing up data from the rapidly expanding collection of sources and building new businesses on it or improve their existing services, like Zillow or Google Maps or Consumer Reports or Bloomberg Government. In a time when job creation is critical, using public sector information to create jobs isn’t an aim to dismiss lightly, although the terms and conditions under which such activity occurs must be clear to all actors involved, to avoid the creation of new monopolies based upon artificial scarcity.
My publisher, long-time open source and open government advocate Tim O'Reilly, has asked how government can act as a platform to enable people inside and outside government to innovate on top of it. One answer is certainly releasing open data. In that context, open data and application programming interfaces, more commonly known as APIs, increasingly look like fundamental infrastructure for digital government in the 21st century.
There's good reason to think that open data could have an overall effect on the economy akin to open source and small business. Gartner, the IT research analysis firm, recently highlighted how open data creates value in the public and private sector.
You may not realize it, but services you use on a daily basis have been built upon data released by the government. Weather data collected by the National Oceanic and Atmospheric Association has an annual estimated economic value of $10 billion, according to U.S. Chief Information Officer Steven VanRoekel and U.S. Chief Technology Officer Todd Park. NOAA data sets are used by Weather.com, Weather Underground, and the Weather Channel—and the nation's farmers consult these forecasts to manage both their crops and the risks of loss. VanRoekel and Park estimate the annual economic value of the data from the U.S. global positioning system at some $90 billion. From companies like TomTom or Garmin to dashboard GPS systems to smartphones and associated location-based applications, GPS data sets are baked into an expanding number of services and products.
Now, as Park seeks to scale open data across the federal government, we’re on the verge of the next generation of services driven by open data, which will involve everything from energy to health care to consumer finance to transit sectors. The challenge is that the cities and federal agencies that hold vast amounts of data may not always understand the value of the information they hold or how to create or sustain businesses using it. That's where open innovation in the public sector and the dynamism of entrepreneurs will play an important role in making the people's data more useful to the people.
BrightScope is a notable example of what dogged persistence can create. The California startup made a profitable business using government data to help the American people understand the fees associated with their 401(k)s. Last May, BrightScope went further, launching financial adviser pages based on open government data from the Securities and Exchange Commission and the Financial Industry Regulatory Authority, the largest independent securities regulator in the United States. Previously, financial adviser profiles could only be found through exact queries at an obscure URL on the regulators' websites. Now, information that citizens care about—the records of financial advisers in their geographic region—is available where they're looking for it: in search engine results.
Just as labor and regulatory data fuels BrightScope's business, there's an expanding number of startups that are tapping into other data released so-called “smart disclosure” initiatives. Smart disclosure is when a private company or government agency provides a person with periodic access to his or her own data in open formats that enable them to easily put the information to use. Startups like Billshrink.com and Hello Wallet are already using a combination of private sector and public sector data to enhance consumer finance decisions. The success of such consumer finance startups suggests an important lesson: The most successful apps and services will combine government, industry, and user-generated data.
The key open data story to watch in the federal government, however, centers on health care. McKinsey and Associates estimates the annual economic value of big, open liquid health data at about $350 billion annually. The explosion of mHealth apps are just the beginning of the disruption in health care from open health data. The effort to revolutionize the health care industry by making health data as useful as weather data is still in its infancy—but the early results are promising. iTriage, which was acquired by Aetna, is enabling people to make better mobile health care decisions where and when they need to do so. It uses a combination of government and private sector data to evaluable symptoms or conditions and point users to nearby medical care. Another startup, Castlight, is analyzing health care data to empower patients, acting like Kayak.com for those who want more transparency about costs. In May, Castlight completed a $100 million round of financing.
But for these sorts of initiatives to take off, entrepreneurs and regulators will have to work together to get contextual consent right and inform patients about the reuse of their data. Transparency is crucial to building a health data commons and thriving startup ecosytem based upon it.
If that balance can be struck, there's considerable potential for entrepreneurs to create better civic interfaces for many digital services. If open government data have helped build new tools, open data disclosed by private companies could create even more value for citizens. But currently, few businesses release anonymized data in an open, usable format. It will soon be time for the government to step in, convene stakeholders, and answer some key questions: How can we create uniform standards that will allow entrepreneurs and developers to innovate? When should data be licensed? Most of the big data releases we have seen come from finance, with bank records or stock trades. But there are significant opportunities to help both entrepreneurs and empowered consumers in health care, energy, education, and telecommunications, to name just a few.
Just as the glowing blue dot on the maps in our smartphone screens revolutionized how we navigate the world, similar "blue dots" could emerge for health care, finance, energy, and any product or service that is regulated or cataloged by government and industry. First, however, they'll need to open the data.
Also in the Future Tense package on government and open data: why Yelp and the government should share data; what a burger mob tells us about the future of democracy; and how Mexico is using open data to move beyond its authoritarian past.
Thursday, April 26, 2012
Big Data's Big Problem: Little Talent
Ben Rooney,The Wall Street Journal, April 26, 2012
It seems that the markets are as much in love with "Big Data"—the ability to acquire, process and sort vast quantities of data in real time—as the technology industry.
The first Big Data initial public offering hit the market last week to roaring approval. Splunk Inc., which helps businesses organize and make sense of all the information they gather, soared 109% on its first day of trading. Big Data, big price.
And this week, in cities in the U.S. and the U.K., Big Data Week events are being held to proselytize the unbelievers.
Big Data refers to the idea that an enterprise can mine all the data it collects right across its operations to unlock golden nuggets of business intelligence. And whereas companies in the past have had to rely on sampling, Big Data, or so the promise goes, means you can use your entire corpus of digitized corporate knowledge. It is, by all accounts, the next big thing.
However, according to a report published last year by McKinsey, there is a problem. "A significant constraint on realizing value from Big Data will be a shortage of talent, particularly of people with deep expertise in statistics and machine learning, and the managers and analysts who know how to operate companies by using insights from Big Data," the report said. "We project a need for 1.5 million additional managers and analysts in the United States who can ask the right questions and consume the results of the analysis of Big Data effectively." What the industry needs is a new type of person: the data scientist.
According to Pat Gelsinger, president and chief operating officer of EMC Corp., the giant U.S. data company, this isn't an unprecedented problem. "IBM started a generation of Cobol programmers," he said, referring to one of the first dominant programming languages.
"Thirty years ago we didn't have computer-science departments; now every quality school on the planet has a CS department. Now nobody has a data-science department; in 30 years every school on the planet will have one."
Hilary Mason, chief scientist for the URL shortening service bit.ly, says a data scientist must have three key skills. "They can take a data set and model it mathematically and understand the math required to build those models; they can actually do that, which means they have the engineering skills…and finally they are someone who can find insights and tell stories from their data. That means asking the right questions, and that is usually the hardest piece."
It is this ability to turn data into information into action that presents the most challenges. It requires a deep understanding of the business to know the questions to ask. The problem that a lot of companies face is that they don't know what they don't know, as former U.S. Defense Secretary Donald Rumsfeld would say. The job of the data scientist isn't simply to uncover lost nuggets, but discover new ones and more importantly, turn them into actions. Providing ever-larger screeds of information doesn't help anyone.
One of the earliest tests for biggish data was applying it to the battlefield. The Pentagon ran a number of field exercises of its Force XXI—a device that allows commanders to track forces on the battlefield—around the turn of the century. The hope was that giving generals "exquisite situational awareness" (i.e. knowing everything about everyone on the battlefield) would turn the art of warfare into a science. What they found was that just giving bad generals more information didn't make them good generals; they were still bad generals, just better informed.
At conference in London this week on the subject, the data scientist was called, only half-jokingly, "a caped superhero."
So where can companies find these superheros? Not from universities, it seems. Nigel Shadbolt, who doubles up as the professor of artificial intelligence at the University of Southampton as well as co-director (along with Tim Berners-Lee) of the U.K.'s Open Data Institute, said the courses don't yet exist. "Bits of it do exist in various departments around the country, and also in businesses, but as an integrated discipline it is only just starting to emerge."
Nor can they be found in recruitment agencies. Rob Grimsey, a director of IT recruitment agency Harvey Nash, said they had limited experience in recruiting data scientists—"which might be a statement in itself about how common these kind of roles are," he added.
One of the problems with Big Data is the fact that it has to deal with real data from the real world, which tends to be messy and difficult to represent. Conventional relational databases are excellent at handling stuff that comes in discreet packets, such as your social security number or a stock price. They are less useful when it comes to, say, the content of a phone call, a video, or an email. Out in the real world, most data is unstructured. Handling this sort of real, messy, scrappy data, isn't so simple.
"People have been doing data mining for years, but that was on the premise that the data was quite well behaved and lived in big relational databases," said Mr. Shadbolt. "How do you deal with data sets that might be very ragged, unreliable, with missing data?"
In the meantime, companies will have to be largely self-taught, said Nick Halstead, CEO of DataSift, one of the U.K. start-ups actually doing Big Data. When recruiting, he said that the ability to ask questions about the data is the key, not mathematical prowess. "You have to be confident at the math, but one of our top people used to be an architect".
But Fernando Lucini, chief architect for Autonomy Corp., a U.K. software maker recently acquired by Hewlett-Packard Co., is much more optimistic. Mr. Lucini said the industry is fretting unnecessarily and should have more confidence in its own abilities. Most of these problems can be tackled through algorithms, he said, which coincidentally is the promise of Autonomy. "The problem can be solved by better tools. The tools need to help you understand the data. They can do the heavy lifting for you so that anyone in a business can use them and ask the questions they need to answer."
Sunday, April 22, 2012
Big Data age puts privacy in question as information becomes currency
Exploiting Big Data's opportunities will need a delicate balance between the right to knowledge and the right of the individual
When Social Calendar users give personal details about themselves or their friends, the data ends up in Walmart's hands. Photograph: Marc F Henning/Alamy
This month, the US chain Walmart bought the startup Social Calendar, one of the most popular calendar apps on Facebook, which lets users record special events, birthdays and anniversaries. More than 15 million registered users have posted over 110m personal notifications, and users receive email reminders totalling over 10m a month.
Of course, when a Social Calendar user listed a friend's birthday or details of a holiday to Malaga, she or he probably had no idea the information would end up in the hands of a US supermarket. But now it will be cross-referenced with Walmart's own data, plus any other databases that are available, to generate a compelling profile of individual Social Calendar users and their non-Social Calendar-using friends.
The second decade of the 21st century is epitomised by Big Data. From the status updates, friendship connections and preferences generated by Facebook and Twitter to search strings on Google, locations on mobile phones and purchasing history on store cards, this is data that's too big to compute easily, yet is so rich that it is being used by institutions in the public and private sectors to identify what people want before they are even aware they want it.
The most important thing for data holders in the Big Data age is the kind of information they have access to. Facebook's projected $100bn value is based on the data it offers people who want to exploit its social graph. Its holdings include more than 800m records about who's in a user's social circle, relationship information, likes, dislikes, public and private messages and even physiological characteristics.
Google's recent privacy policy change has integrated the various accounts an individual maintains, creating a single profile that includes intentions from its search engine and the connections identified from its social network Google+; preferences and interests from mail, documents or YouTube; and location from its maps and mobile phone operating system.
Aggregated, this data can prove powerful. "Given enough data, intelligence and power, corporations and government can connect dots in ways that only previously existed in science fiction," said Alexander Howard, government 2.0 correspondent at the technology publisher O'Reilly Media.
In a trend that is remarkably similar to the plotline of Philip K Dick's Minority Report, Big Data is being used to predict social unrest or criminal intent. For example, Pax, an experimental system developed by the documentary maker and historian Brian Lapping, predicts the conditions for uprisings using aggregated search terms in different regions of the world. The analysed intelligence is then sold to governments, which can act accordingly.
The systems used to parse, synthesise, assimilate and make sense of the information are starting to make sophisticated connections and learn patterns. Big Data proponents view this as an opportunity to observe behaviours in real time, draw real-time conclusions and affect real-time change. Yet their conclusions can trip into areas that require human sensibilities to truly understand their implications.
In one recent high-profile example, a Minneapolis man discovered his teenage daughter was pregnant because coupons for baby food and clothing were arriving at his address from the US superstore Target. The girl, who had not registered her pregnancy with the chain, had been identified by a system that looked for pregnancy patterns in her purchase behaviour. "Data can say quite a lot," said Howard. "Though one has to be very careful to verify quality and balance it with human expertise and intuition."
In an infamous case in 2006, anonymised search terms released into the public domain by AOL were quickly de-anonymised, identifying individual searchers. And last month, police in New York used a photo from Facebook in combination with their own photo files and facial recognition software to arrest a man for attempted murder.
"People give out their data often without thinking about it," said the European commission vice-president Viviane Reding. "They have no idea that it will be sold to third parties." So users continue to populate databases such as Social Calendar with increasingly valuable personal information that, as commercial property, can be transferred to a new company with a different privacy ethos.
Privacy is not about control over personal data, according to the web theorist Danah Boyd, but the control individuals think they have. "People seek privacy so that they can make themselves vulnerable in order to gain something: personal support, knowledge, friendship," she said at the WWW conference in 2010. Increasingly, people are gaining services that deliver value, relevance and connection – as Google and Facebook do – in exchange for their personal information.
Expectations of privacy are being renegotiated. "When I grew up in Greensborough, Alabama, the population was 1,200," said Jim Adler, chief privacy officer and general manager of data systems at the information commerce firm Intelius. "If you cut school, everyone knew it by dinner. The expectation of privacy was low.
"Now, the expectation of privacy that we've had before Big Data, and our parents had, has been pulled away."
To some degree, this is happening because web users and web developers may not share a universal sense of what is and what is not private. As Boyd put it, privacy is contextual. An individual may be willing to share what they had for breakfast on Twitter, divulge where they are via FourSquare or record every keystroke made on their computers since 1998, but they wouldn't want information about their health or their children's whereabouts made public.
It becomes even more complicated when the users of software systems and architectures are a global population but the privacy expectations have been put in place by primarily US services. "Our expectations of privacy in the US versus Europe are very different," said Adler. "We are currently negotiating which is more important: the rights of the individual or the rights of knowledge."
In the EU, Reding has campaigned for the "right to be forgotten", already part of the 1995 data protection directive, which establishes by law that private data is the property of the individual and must be deleted from a system on request at any time. "More and more people feel uncomfortable about being traced everywhere, about a brave new world," she said. Information held by public bodies, however, remains exempt.
Reding's motivation is primarily to maintain a business ecosystem friendly for foreign investment. "This isn't about the reputation of the individual," she explained. "It's about the reputation of the companies. Data is their currency.
"What we're aiming for is privacy by design," she said. Companies should initiate a hallmark system that informs users that the privacy policy adheres to the guidelines. This, she argued, would ensure that people continue to share their data.
Sceptics like Adler argue that the right to be forgotten is flawed because it ignores how social boundaries are currently being negotiated in the Big Data world. "The ability to delete personal information means that you lose the potential for lessons learned," he said. "If you can step away and erase something someone says that is stupid or hurtful, you lose an element of accountability."
Yet if, as Reding maintains, 80% of British citizens are already concerned that data held by companies will be used for purposes other than the reason it was collected, there may be a shift in how much information people are willing to share.
The weakest link is the technology itself. The Target pregnancy case demonstrates that machines can pick up patterns in ways that may have unexpected consequences for individuals. The people who design the systems that collect and analyse the data are now responsible for thinking about data privacy and projecting future outcomes, and may – because they're human – get it wrong.
"These technologies are as neutral as guns," said Adler. "The Big Data guys who want to send you coupons when you're pregnant – because they're nerdy and technologists – probably don't realise that pregnancy is a sensitive issue." And sensitivities shift throughout an individual's lifespan and, more broadly, social norms shift over time.
Fundamentally, privacy means the same thing in an era of Big Data as it always has, but the capacity of machines to capture, store, process, synthesise and analyse details about everyone has forced new boundaries. It is unlikely that people will stop sharing data in exchange for services that are viewed as valuable.
Big Data offers undeniable opportunities, but requires a delicate balance between the right to knowledge and the right of the individual. Privacy norms will demand that new systems of trust be built into technology design.
© 2012 Guardian News and Media Limited or its affiliated companies. All rights reserved.
Thursday, April 19, 2012
Big Data: Splunk's Data With Destiny
Rolfe Winkler, The Wall Street Journal, April 18, 2012
WSJ's Rolfe Winkler makes a stop on Mean Street to discuss the upcoming IPO of tech company Splunk. He and Evan Newmark ponder if a tech bubble is looming closer. Photo: Getty Images.
Splunking is the new googling. And it could make investors a tidy sum of money.The new verb tossed around by information-technology pros comes courtesy of Splunk, a startup specializing in data analysis that will open for trading Thursday after staging its initial public offering. The company is one of many capitalizing on the explosion of information, a trend being referred to as "Big Data."
Splunk's particular specialty is collecting and processing so-called "machine data." From web sites to cell phones to smart meters to GPS equipment, machines create little bits of information all the time. So much gets created, it is often lost after being recorded on a server somewhere.
When you click through an e-commerce web site, you visit lots of different product pages, put items in a shopping cart, and maybe disappear without buying anything. A company that could follow its customers to determine why they don't complete their order might be able to isearchable—is similar to what Splunk does with machine data. Its software is available free on a trial basis to start. Often someone inside a company will start using it, find it useful and then others will start using it themselves for their own projects. Splunk starts charging as more data gets plugged in. Yet the product is still cheaper than many older software alternatives currently on the market. The formula has worked well so far. Splunk reported $121 million of revenue in the fiscal year that ended in January, up 83% from the prior year. That growth has excited IPO investors.
Originally Splunk planned to price its shares in a range between $8 to $10, but has since bumped up the target to between $11 and $13. Investors lucky to get in on the shares around that price could see them pop significantly.
At first glance, the pricing seems aggressive. The valuation of the company net of cash would be around $1.3 billion at a $13 share price. Splunk is unprofitable.
Yet combine Splunk's growth rate with the appeal of its technology and the firm looks a mouth-watering takeover target for a larger software company like BMC Software, BMC 0.00% International Business Machines IBM -0.38% or Hewlett-Packard HPQ -0.02% .
A valuation of 10 times forward revenue would be in line with previous deals, for storage-software companies 3PAR and Isilon Systems. Assuming Splunk grows at, say, 70% this fiscal year, that would translate to a roughly $20 share price.
To be sure, the company has faced growing pains. Older versions of its software were buggy. And as it has expanded, Splunk has had trouble keeping up with customer demand for support. Yet it has mostly overcome these issues.
Look for the company to cash in nicely as a result.
Originally Splunk planned to price its shares in a range between $8 to $10, but has since bumped up the target to between $11 and $13. Investors lucky to get in on the shares around that price could see them pop significantly.
At first glance, the pricing seems aggressive. The valuation of the company net of cash would be around $1.3 billion at a $13 share price. Splunk is unprofitable.
Yet combine Splunk's growth rate with the appeal of its technology and the firm looks a mouth-watering takeover target for a larger software company like BMC Software, BMC 0.00% International Business Machines IBM -0.38% or Hewlett-Packard HPQ -0.02% .
A valuation of 10 times forward revenue would be in line with previous deals, for storage-software companies 3PAR and Isilon Systems. Assuming Splunk grows at, say, 70% this fiscal year, that would translate to a roughly $20 share price.
To be sure, the company has faced growing pains. Older versions of its software were buggy. And as it has expanded, Splunk has had trouble keeping up with customer demand for support. Yet it has mostly overcome these issues.
Look for the company to cash in nicely as a result.
Monday, March 19, 2012
Big Data for the Rest of Us, In One Startup
Quentin Hardy, March 19, 2012, The New York Times
The best insight into the current state of the Big Data business may be a five person start-up in Palo Alto, Calif. While the industry mostly gathers lots of data, or stores it in new cloud-based databases, the startup, ClearStory Data, is building the biggest possible base of data consumers.
ClearStory, which was formed last summer in Palo Alto, Calif., is making software for ordinary business professionals. They will be able to blend their own corporate data with the large amounts of publicly-available data, in search of new statistical insights. There just aren’t enough statisticians, the company figures, to address the demand corporations will have for Big Data.
“We’re about making data consumable,” says Sharmila Shahani-Mulligan, ClearStory’s founder and chief executive. “The world is talking about the size of Big Data sources, but at the end of the day it will be about the ease of consumption.” Indeed, the McKinsey Global Institute has projected that the U.S. needs up to 190,000 people with “analytical expertise” than it has.
The idea of combining proprietary corporate data with public data is hardly new; fast food companies have long looked at census data to figure out where to put outlets, and what to feed people there. The public data sources for ClearStory, will largely be the new category of “Data Marts,” like Datasift, Factual, the Windows Azure Data Marketplace from Microsoft, and Infochimps. These data marts collectively have perhaps millions of data types, from Twitter feeds to mean temperatures and store locations.
“There aren’t enough experts at companies to handle all these data sources,” Ms. Shahani-Mulligan says. “Brick and mortar stores in the Midwest can’t get the kind of data scientists LinkedIn has. We can give them that.” The company offers a range of algorithms to wrangle the data. Longer term, ClearStory hopes to get companies to start sharing more of their own internal data too, what Ms. Shahani-Mulligan calls “the gold locked up in 30 years of relational databases.”
Drawing off all of the data sources could be powerful, if ClearStory can manage to maintain an even level of quality and reliability among its data sources. That already bedevils many data marts. If ClearStory offers sloppy data to customers, however, it will hurt ClearStory.
Another challenge will be designing an intuitive Web-based user experience, so the average marketer or communications professional isn’t flummoxed by statistics. Good “U.E.” people are among the toughest people to find in tech.
ClearStory is a good trend barometer in other ways. The product, which is expected to be ready for public consumption late this summer, will be offered free to casual users. The idea is that enough customers far away from the corporate information technology department will come to depend on ClearStory, and eventually I.T. will want to incorporate a more sophisticated paid version. It is a plan that worked for Yammer, a corporate messaging tool, and Dropbox, with offers storage in the cloud, as well as the Linux operating system.
Ms. Shahani-Mulligan, a marketing veteran of Netscape, Opsware, and Cloudera, is backed by Google Ventures, Andreessen Horowitz, Khosla Ventures, among others. That makes it one of the more “hot backer”-compliant ventures: Google Ventures is specializing in companies with a statistical bent, Andreessen Horowitz, is assembling an ecology of companies in cloud computing, and Vinod Khosla is a renowned tech visionary.
Even with backers like that, ClearStory is making itself visible a lot earlier than most startups risking it will hit the radar of potential competitors even before it has a product. In another sign of the times, and the boom (or, possibly, the bubble) in Big Data, ClearStory needs the publicity in order to hire in an increasingly competitive talent market, however.
“There is a war for talent,” says Ben Horowitz, ClearStory’s backer at Andreessen Horowitz. “Getting the first five people is easy, but the next 20 is hard, when you are selling on your vision and momentum.” Besides, he says, “big companies don’t react to potential threats. They only react to real threats. There are just too many threats out there.”
The best insight into the current state of the Big Data business may be a five person start-up in Palo Alto, Calif. While the industry mostly gathers lots of data, or stores it in new cloud-based databases, the startup, ClearStory Data, is building the biggest possible base of data consumers.
ClearStory, which was formed last summer in Palo Alto, Calif., is making software for ordinary business professionals. They will be able to blend their own corporate data with the large amounts of publicly-available data, in search of new statistical insights. There just aren’t enough statisticians, the company figures, to address the demand corporations will have for Big Data.
“We’re about making data consumable,” says Sharmila Shahani-Mulligan, ClearStory’s founder and chief executive. “The world is talking about the size of Big Data sources, but at the end of the day it will be about the ease of consumption.” Indeed, the McKinsey Global Institute has projected that the U.S. needs up to 190,000 people with “analytical expertise” than it has.
The idea of combining proprietary corporate data with public data is hardly new; fast food companies have long looked at census data to figure out where to put outlets, and what to feed people there. The public data sources for ClearStory, will largely be the new category of “Data Marts,” like Datasift, Factual, the Windows Azure Data Marketplace from Microsoft, and Infochimps. These data marts collectively have perhaps millions of data types, from Twitter feeds to mean temperatures and store locations.
“There aren’t enough experts at companies to handle all these data sources,” Ms. Shahani-Mulligan says. “Brick and mortar stores in the Midwest can’t get the kind of data scientists LinkedIn has. We can give them that.” The company offers a range of algorithms to wrangle the data. Longer term, ClearStory hopes to get companies to start sharing more of their own internal data too, what Ms. Shahani-Mulligan calls “the gold locked up in 30 years of relational databases.”
Drawing off all of the data sources could be powerful, if ClearStory can manage to maintain an even level of quality and reliability among its data sources. That already bedevils many data marts. If ClearStory offers sloppy data to customers, however, it will hurt ClearStory.
Another challenge will be designing an intuitive Web-based user experience, so the average marketer or communications professional isn’t flummoxed by statistics. Good “U.E.” people are among the toughest people to find in tech.
ClearStory is a good trend barometer in other ways. The product, which is expected to be ready for public consumption late this summer, will be offered free to casual users. The idea is that enough customers far away from the corporate information technology department will come to depend on ClearStory, and eventually I.T. will want to incorporate a more sophisticated paid version. It is a plan that worked for Yammer, a corporate messaging tool, and Dropbox, with offers storage in the cloud, as well as the Linux operating system.
Ms. Shahani-Mulligan, a marketing veteran of Netscape, Opsware, and Cloudera, is backed by Google Ventures, Andreessen Horowitz, Khosla Ventures, among others. That makes it one of the more “hot backer”-compliant ventures: Google Ventures is specializing in companies with a statistical bent, Andreessen Horowitz, is assembling an ecology of companies in cloud computing, and Vinod Khosla is a renowned tech visionary.
Even with backers like that, ClearStory is making itself visible a lot earlier than most startups risking it will hit the radar of potential competitors even before it has a product. In another sign of the times, and the boom (or, possibly, the bubble) in Big Data, ClearStory needs the publicity in order to hire in an increasingly competitive talent market, however.
“There is a war for talent,” says Ben Horowitz, ClearStory’s backer at Andreessen Horowitz. “Getting the first five people is easy, but the next 20 is hard, when you are selling on your vision and momentum.” Besides, he says, “big companies don’t react to potential threats. They only react to real threats. There are just too many threats out there.”
Subscribe to:
Posts (Atom)

