Showing posts with label Competitiveness. Show all posts
Showing posts with label Competitiveness. Show all posts

Tuesday, January 22, 2013

OF INTEREST: Max Levchin (Former CTO of Paypal)'s keynot at DLD on Big Data


This is the approximate text of the keynote I gave at DLD13 today in Munich. I felt my delivery of this (admittedly, relatively dense) material was not the best, but the content is crucially important. To that end, I am posting the notes here. They were edited for grammar beyond the basics, so you will have to forgive the occasional fifth-grade prose.

Max Levchin, January 21, 2013
Many of today’s big data companies are trying to tackle problems that just aren’t nearly big enough. 

Most are focused on marginally improving existing, digital businesses, but I believe the next big wave of opportunities exists in centralized processing of data gathered from primarily analog systems.

At PayPal, where I was the CTO, we succeeded because we gained deep understanding of the immense quantities of behavioral data that we captured in processing millions of transactions per day. We learned so much about our customers, that we could predict their intentions, and prevent vast majority of intentional fraud. 

At HVF, the project I began in 2011, we seek to create businesses that improve the analog, real world through deep understanding of data. I will tell you more about that in a bit, but first let me expand on what I mean by “digital sensors over analog data.”

Collaborative consumption is a huge current trend. To me, it started to make sense with Über — idling black limo cars put to better use. Über (where I am now an investor) created a simple piece of software to send the idle ones to price-insensitive consumers in need of a cab with a bit of flair. 

The world of real things is very inefficient: slack resources are abundant, so are the companies trying to rationalize their use. Über, AirBnB, Exec, GetAround, PostMates, ZipCar, Cherry, Housefed, Skyara, ToolSpinner, Snapgoods, Vayable, Swifto…it’s an explosion! What enabled this? Why now? It’s not like we suddenly have a larger surplus of black cars than ever before.

Examine the DNA of these businesses: resource availability and demand requests — highly analog, as this is about cars, drivers, and passengers — is captured at the edge, automatically where possible, then transmitted and stored, then processed centrally. Requests are queued at the smart center, and a marketplace/auction is used to allocate them, matches are made and feedback is given in real time. 

Utilization per dollar goes up, because there is simply less idle time! So efficient is this new approach in fact that limo companies started hiring drivers to drive for Über only.  

I am ignoring some key details, like the need for mutual trust between participants that is at least initially enabled by the presence of a trusted third party, and the feedback loop from consumers of the service, but at its core, these business all look very similar.

A key revolutionary insight here is not that the market-based distribution of resources is a great idea — it is the digitalization of analog data, and its management in a centralized queue to create amazing new efficiencies.

Consider the cab calling experience vs Über. When dialing up a cab, you are managing your own spot in the queue. If you hang up in anger, you go back to the end. If you stay on, listening to the on-call tune, you may be an infinite loop, as there is no feedback! 

And even if you are willing to pay a hundred times more than everyone else waiting ahead of you in line to speak to dispatch, you never get to express that demand. The data exists in an analog-only format, and it moves at analog-only speeds.

With Über, the queue can be managed centrally (because the information is converted to a digital format at the edge) with nearly complete transparency to you — you know when the resources you want become available, you know how long you have to wait for them, and most importantly, you are generally assured through feedback that you have not been forgotten or ignored.

So why is all this now? Cheap digital sensors over analog resources (cars, houses, humans, etc) — AT&T pays for your GPS… so you’d consume more data of course. Mobile broadband, of course. Even more key is the critical mass of pretty smart devices that are clients for real-time participation in queue management with feedback. It’s the smart-enough terminal model. 

As an aside, consumers want sexy sensors that give visual feedback to motivate action (and more data) — like the Nike Fuel band. Machines, on the other hand, just want cheap sensors to get more data from.

I sometimes imagine the low-use troughs of sinusoidal curves utilization of all these analog resources being pulled up, filling up with happy digital usage. 

Private jets spend 1h per day on average in the air. BlackJet promises to make that number closer to the commercial average — 10h/day. And not everyone in a suburban neighborhood needs their own lawn mower — they need an app to schedule one of the local kids to come by and take care of their overgrown grass. 

The defensibility of these businesses lies in their ability to build a network effect — a network effect of data. Once a business understands substantially more than any one of the resources managed in their queue, it’s effectively impossible to compete with them on price — they can always see more of the usage as it happens, and price it more efficiently, pushing any competition out. 

So what other businesses can we expect to emerge in analog-data-driven, central-intelligence queue marketplace businesses? Some interesting ones are probably already being built: a market for private neighborhood security (off-duty cops)? An auction for short-term patent licenses (litigator included)? 

Technology already enables efficient redistribution for your spare change: it’s Kickstarter and AngelList. We will definitely see dynamically-priced queues for confession-taking priests, and therapists!

How about dynamic pricing for brain cycles? We have been maximizing utilization of very high-value, very low-frequency specialists — today you can already rent the brain of a data-mining genius via Kaggle by the hour, tomorrow by brain-hour. Just like the SETI@Home screensaver “steals” CPU cycles to sift through cosmic radio noise for alien voices, your brain plug firmware will earn you a little extra cash while you sleep, by being remotely programmed to solve hard problems, like factoring products of large primes.

There is also a neat symmetry to this analog-to-digtail transformation — enabling centralization of unique analog capacities. As soon as the general public is ready for it, many things handled by a human at the edge of consumption will be controlled by the best currently available human at the center of the system, real time sensors bringing the necessary data to them in real time. The freshest, smartest pilot, most familiar with the particular complicated airport will land your plane — via remote control.   

So what’s after that? This is where it gets really interesting. These new modes of operation — remote controlled cars and planes flown by pilots you can’t see, rides in quasi-cabs with people you have never met, legal advice from lawyers whose license you cannot really check. This is going to add a huge amount of new kinds of risks. 

But as a species, we simply must take these risks, to continue advancing, to use all available resources to their maximum. Yet these risks are real, and they cannot be ignored.

The way to deal with risk is of course some form of insurance. Modeling loss from observed past events is hardly news, but dynamically changing the price of the service to reflect individual risk is a big deal. My expectation is that next decade we will see an explosion of insurance and insurance-like products and services — leveraging those very same network effects of data, providing truly dynamic resource pricing and allocation. 

This is the purpose of my new project, HVF — to bring these new products to those that will benefit from them the most. We see data-driven understanding and pricing of risk as the great opportunity to improve lives. The reason the notion of analog is so important here is because it ultimately also means “human.”

Understanding the changing risk profile of a person can deliver to them amazing opportunities they wouldn’t have in today’s world: inferring that a particular college grad is financially responsible by looking at their tweets could allow them to buy their first house on credit, at 21, without any history, and looking at someone’s heart rate monitor data could make their cardiovascular healthcare cost-free. 

These are not non-controversial topics — privacy, unfair discrimination, built-in biases are all possible, and we must be thoughtful and diligent in how we go about bringing this future. But I believe that what we can enable with data insights greatly outweighs the downsides. 

Here is what I mean, by way of a simple example.

On a Sat morning, I load my two toddlers into their respective child seats, and my car’s in-wheel strain gauges detect the weight difference and reports that the kids are with me in a moving vehicle to my insurance via a secure message through my iPhone. The insurance company duly increases today’s premium by a few dollars. 

My keepHonest app sees this too and immediately offers me up as a customer to a few competing insurance companies in the background, but nobody is willing to charge me less right now, and the phone chirps sadly to let me know I’m now paying a higher premium. Safer, but more expensive. 

But In a few hours, my car’s GPS duly reports to my insurer that I only drove two miles to the park, never sped and, and observed all traffic signs. My phone now chirps happily: not only has my rate been discounted, several companies are offering me a deal on insurance!

So to conclude: I believe that in the next decades we will see huge number of inherently analog processes captured digitally. Opportunities to build businesses that process this data and improve lives will abound. 



Friday, October 19, 2012

Will Big Data decide the election?


A new book traces the recent history of data mining in political campaigns. A review of The Victory Lab, by Sasha Issenberg.

Chip Lebovitz, Fortune, October 19, 2012

FORTUNE -- There's a powerful vignette in Sasha Issenberg's The Victory Lab in which political consultant Alexander Gage presents his new data targeting system to Mitt Romney's 2002 gubernatorial campaign.

Gage has combined consumer records with political voting history to identify potential Romney supporters among nontraditional Republican voting blocks. Gage sees his work as revolutionary -- a first in politics, and potentially a first anywhere. Yet just as he completes his presentation, Romney's deputy campaign manager Alex Dunn raises his hand and deadpans, "You mean you don't do this in politics."

Dunn's surprise was not out of place in 2001. And while much has changed since then, politics remains an analog art in many ways. The Victory Lab charts the recent history of political data mining designed to identify persuadable voters and swing elections.

Politico has called Issenberg's book "Moneyball for politics." There are obvious parallels between political operatives like Gage and the protagonist of Moneyball, Oakland Athletics general manager Billy Beane. 

Both pioneered data-centric techniques that gave their respective organizations a leg up over the competition.

Yet in his 14 years as GM, Billy Beane has yet to make it to the World Series. By contrast, the work of Gage and other data-mining political consultants has translated directly into electoral success. These operatives may not have singlehandedly won the past two presidential elections, but the side that had the most advanced data-driven mobilization efforts went a perfect 2-0.

The George W. Bush campaign's sophisticated data mining helped the president win reelection in 2004. Four years later, Barack Obama defeated John McCain in part because the Obama campaign outclassed McCain's operatives on the data front.

Issenberg seems to have interviewed everyone who's anyone in the field of political statistics, but his subject matter doesn't necessarily lend itself to prose. Numbers have power, but too many of them can addle rather than inform. After reading the book, you'll probably remember the outcome of a few experiments, but be stuck scouring your brain for the results of the rest. Readers are advised to keep a pen and pencil handy when reading The Victory Lab, so they can jot down key stats.

Yet Issenberg's material is sufficiently gripping that you'll want to keep turning the pages, even if it means deciphering the results of yet another randomized trial. For example, micro-targeting efforts on behalf of Senator Michael Bennett (D-Colo.)'s 2010 Senate reelection campaign likely spurred 25,000 additional Bennett votes. That might not seem like a big number, until you consider that the race was decided by 15,000 votes.

Issenberg avoids sweeping predictions about the future of political data mining. That's probably wise, given that political professionals tend to define the future as Election Day. Campaigns at their simplest are a one-day snapshot of personal preference. The Victory Lab reminds us, however, that every individual choice is the result of a thousand factors, many of them subject to manipulation by political operatives.


Wednesday, October 17, 2012

Should High Schools Teach Big Data?



Given the anticipated shortage of data scientists, some high school educators have jumped in to expose students to big data concepts.


Changing when advanced database technology is taught has real-world implications, given the realities of today's job market. Both data analytics and big data skills are in high demand in private industry and government.

But there is a looming shortage of workers with these abilities. McKinsey & Co. sounded this alarm back in 2011 with its seminal report that predicted the U.S. would face ashortage of 140,000 to 190,000 workers with the skills to manage and analyze big data.

The popular technology job board Dice.com has seen a spike in listings for "data scientist," up from just a handful a year ago to more than 35 at the start of October. While still an imprecise job designation, "data scientists" command high salaries compared to other IT job titles. (Separately, the unemployment rate for technology professionals dropped in the third quarter to 3.3%, as compared to 4.2% in the same quarter a year ago, according to the Bureau of Labor Statistics.)

These listings--many of which request a PhD in fields like mathematics, economics, or statistics--today cluster in financial services, retail, and e-commerce. Job listings using the phrase "big data" have increased from around 200 in January to nearly 800 in October.

"But increasingly, every industry is dealing with big data questions," said Alice Hill, managing director of Dice.com and president of Dice Labs.

Preparing the U.S. for future high-tech jobs, specifically ones oriented around data, was also a focus of TechAmerica Foundation's Big Data Commission.

Given the anticipated shortage of data scientists, should students start learning the precepts of big data in high school?Analytics and data science are central to making "business, the global economy, and our society work better," Steve Mills, senior VP and group executive at IBM, and co-chair of the Big Data Commission, said in a statement. "That's why it's critical that our country prepares a new generation of experts who know how to corral today's data deluge for world-changing insights."

The commission's new report, "Demystifying Big Data: A Practical Guide to Transforming the Business of Government," included the following recommendations for skills development: Strengthen and expand public-private partnerships to invest in skills-building initiatives for the federal workforce in the area of big data. These should include formal career tracks for IT managers; an IT leadership academy to provide big data and related training and certification; data-intensive degree programs; and scholarships to prepare a new generation of data scientists.

Big Data In High School?
Among those trying to push big data classes down to the high school level is Alex Philp, PhD. Philp is founder and CTO of TerraEchos, a developer of advanced intelligence and surveillance security systems. Philp has been working with the high schools in Missoula, Mont., and has even created a scholarship program at one--Sentinel High School--to introduce these topics to computer science students.

"[We] cooked up the idea of promoting a merit-based, micro-challenge grant at Sentinel," Philp explains. Currently, four student teams have been awarded grants. "Ultimately, I hope these high school students feed into opportunities at the University of Montana and then ultimately into the most exciting businesses and markets involving big data," Philp said. "This also relates to aspects of the overall economic competitiveness of our country and a renewed commitment to science, technology, engineering, arts, and mathematics (STEAM) competency in our country."

Meanwhile, at the university level, Philp has been working with Eric Tangedahl, IT director for the school of business administration at The University of Montana, to create a multidisciplinary course, now in its first year.

"The course draws students from computer science, management information systems, and math," Tangedahl said. "We felt that to tackle these problems, students have to work together in teams, teams with different skill sets," he said.

The three-hour class, now with 18 students, is half lecture about big data and half lab--specifically around IBM Infostream, a programming language that takes advantage of IBM's DB2 and WebSphere platforms.

Another high school on this path is the science and engineering magnet (SEM) school at Yvonne A. Ewell Townview Center in Dallas. SEM, ranked by Newsweek this year as one of the best high schools in the U.S., recently participated in IBM's annual Master the Mainframe Contest for high school students in the U.S. and Canada.

Marilyn Cadenhead, who teaches Advanced Placement computer science at SEM, has many students participating in the contest.

Cadenhead says her students use technology very effectively, and already have a computer-mediated learning style. "This is a digital age, very different from when I was in school," she wrote in an email. "I learned to type on a manual typewriter. If we as teachers do not engage [students] with the latest technology and teaching styles, we will be 'boring' and the students will not be motivated to learn."

"I want my students to know what is going on in the real world," Cadenhead concludes, explaining her enthusiasm for the IBM mainframe contest, as well the annual IBM Innovation summer camp that some SEM students attended this year.

IBM's Innovation summer camp Facebook page describes the program this way: "The students will gain hands-on experience with visualization, mobile application development, Linux, and DB2 coupled with demos of research underway at UTD in mind control, gaming, and robotics."

Other observers, however, point out that big data analysis is powerful precisely because it is about more than raw technology. The highly valued professionals in this space are those who ask the right questions, who can see business-relevant answers inside the data. That kind of maturity and domain expertise is beyond the capacity of high school students, they say.

The University of Montana's Tangedahl concurs. "I wouldn't throw high school student in and expect them to come up with a [big data] program for the Defense department," he said. But, he adds, "you've absolutely got to put the building blocks in place." He recommends adding data statistics and programming classes in high school to better prepare students for college-level classes, like his, that push teams of students to answer "real-world problems."

Like Tangedahl, Dice.com's Hill isn't sold on the idea of pushing big data instruction to high school students. She thinks students should learn relevant technologies, such as data interpretation, data extraction, and data modeling. "It's never too soon to learn those skills," she said.

But TerraEchos' Philp disagrees, arguing it is essential that educators encourage excellence in students at all ages, and not set limits.

"I've spent many years of my life working with thousands of students at various ages, attempting to raise the bar," he said. "High school kids are proposing [projects] that are as good as I'm seeing in college." With motivation, inspiration, and passion, he said, "my experience is, students rise to the occasion."