Showing posts with label emerging technology. Show all posts
Showing posts with label emerging technology. Show all posts

Wednesday, March 6, 2013

The Google Glass feature no one is talking about

Mark Hurst, Creative Good, February 28, 2013

Google Glass might change your life, but not in the way you think. There's something else Google Glass makes possible that no one – no one – has talked about yet, and so today I'm writing this blog post to describe it.

To read the raving accounts of tech journalists who Google commissioned for demos, you'd think Glass was something between a jetpack and a magic wand: something so cool, so sleek, so irresistible that it must inevitably replace that fading, pitifully out-of-date device called the smartphone.

Sergey Brin himself said as much yesterday, observing that it is "emasculating" to use a smartphone, "rubbing this featureless piece of glass." His solution to that piece of glass, of course, is called Glass. And his solution to that emasculation is – well, as VentureBeat put it, "Sergey Brin calls smartphones 'emasculating' – but dorky Google Glass [is] A-OK."

Like every other shiny innovation these days, Google Glass will live or die solely on the experience it creates for people. The immediate, most visible problem in the Glass experience is how dorky the user looks while wearing it. No one wants to be the only person in the bar dressed like a cyborg from a 1992 virtual-reality movie. It's embarrassing. Early adopters will abandon Google Glass if they don't sense the social approval they seek while wearing it.

Google seems to have calculated this already and recently announced a partnership with Warby Parker, known for its designer glasses favored by the all-important younger demographic. (My own proposal, posted the day before, jokingly suggested that Google look into monocles.)

Except for the awkward physical design, the experience of using Google Glass has won high praise from reviewers. Seeing your bitstreams floating in the air in front of you, it would seem, is an ecstatic experience. Weather! Directions! Social network requests! Email overload! All floating in front of you, never out of your sight! For people who delight in a deluge of digital distractions, this is much more exciting than a smartphone, which forces you back to the boring offline world, every so often, when you put the phone away. Glass promises never to do that. In fact, in a feat of considerable chutzpah, Google is attempting to pitch Glass as an antidote to distraction, since users don't have to look down at a phone. Right, because now the distractions are all conveniently placed directly into your eyeball! (For a more accurate exploration of Glass-enabled distraction, see this darkly comic parody video. Even edgier is this parody – warning, some spicy language.)

As if all that wasn't enough, Google Glass comes with yet another, even more important feature: lifebits, the ability to record video of the people, places, and events around you, at all times. Veteran readers will remember that I predicted this six years ago in my book Bit Literacy. From Chapter 13:

The life bitstream will raise new and important issues. Should it be socially acceptable, for example, to record a private conversation with a friend? How will anyone be sure they're not being recorded, in public or private? … Corporations, police, even friends with 'life recorders' will capture the actions and utterances of everyone in sight, whether they like it or not.

Today, finally, that future has arrived: a major company offering the ability to record your life, store it, and share it – all with a simple voice command.

And this is where our story takes a turn, toward a ramification that dwarfs every other issue raised so far on Google Glass. Yes, the glasses look dorky – Google will fix that. And sure, Glass forces users to be permanently plugged-in to Google's digital world – that's hardly a concern for the company or, for that matter, most users out there. No. The real issue raised by Google Glass, which will either cause the project to fail or create certain outcomes you may not want (which I'll describe), has to do with the lifebits. Once again, it's an issue of experience.

The Google Glass feature that (almost) no one is talking about is the experience – not of the user, but of everyone other than the user. A tweet by David Yee introduces it well:

There is a kid wearing Google Glasses at this restaurant which, until just now, used to be my favorite spot.
The key experiential question of Google Glass isn't what it's like to wear them, it's what it's like to be around someone else who's wearing them. I'll give an easy example. Your one-on-one conversation with someone wearing Google Glass is likely to be annoying, because you'll suspect that you don't have their undivided attention. And you can't comfortably ask them to take the glasses off (especially when, inevitably, the device is integrated into prescription lenses). Finally – here's where the problems really start – you don't know if they're taking a video of you.

Now pretend you don't know a single person who wears Google Glass… and take a walk outside. Anywhere you go in public – any store, any sidewalk, any bus or subway – you're liable to be recorded: audio and video. Fifty people on the bus might be Glassless, but if a single person wearing Glass gets on, you – and all 49 other passengers – could be recorded. Not just for a temporary throwaway video buffer, like a security camera, but recorded, stored permanently, and shared to the world.

Now, I know the response: "I'm recorded by security cameras all day, it doesn't bother me, what's the difference?" Hear me out – I'm not done. What makes Glass so unique is that it's a Google project. And Google has the capacity to combine Glass with other technologies it owns.

First, take the video feeds from every Google Glass headset, worn by users worldwide. Regardless of whether video is only recorded temporarily, as in the first version of Glass, or always-on, as is certainly possible in future versions, the video all streams into Google's own cloud of servers. Now add in facial recognition and the identity database that Google is building within Google Plus (with an emphasis on people's accurate, real-world names): Google's servers can process video files, at their leisure, to attempt identification on every person appearing in every video. And if Google Plus doesn't sound like much, note that Mark Zuckerberg has already pledged that Facebook will develop apps for Glass.

Finally, consider the speech-to-text software that Google already employs, both in its servers and on the Glass devices themselves. Any audio in a video could, technically speaking, be converted to text, tagged to the individual who spoke it, and made fully searchable within Google's search index.

Now our stage is set: not for what will happen, necessarily, but what I just want to point out could technically happen, by combining tools already available within Google.

Let's return to the bus ride. It's not a stretch to imagine that you could immediately be identified by that Google Glass user who gets on the bus and turns the camera toward you. Anything you say within earshot could be recorded, associated with the text, and tagged to your online identity. And stored in Google's search index. Permanently.

I'm still not done.

The really interesting aspect is that all of the indexing, tagging, and storage could happen without the Google Glass user even requesting it. Any video taken by any Google Glass, anywhere, is likely to be stored on Google servers, where any post-processing (facial recognition, speech-to-text, etc.) could happen at the later request of Google, or any other corporate or governmental body, at any point in the future.
Remember when people were kind of creeped out by that car Google drove around to take pictures of your house? Most people got over it, because they got a nice StreetView feature in Google Maps as a result.
Google Glass is like one camera car for each of the thousands, possibly millions, of people who will wear the device – every single day, everywhere they go – on sidewalks, into restaurants, up elevators, around your office, into your home. From now on, starting today, anywhere you go within range of a Google Glass device, everything you do could be recorded and uploaded to Google's cloud, and stored there for the rest of your life. You won't know if you're being recorded or not; and even if you do, you'll have no way to stop it.

And that, my friends, is the experience that Google Glass creates. That is the experience we should be thinking about. The most important Google Glass experience is not the user experience – it's the experience of everyone else. The experience of being a citizen, in public, is about to change.
Just think: if a million Google Glasses go out into the world and start storing audio and video of the world around them, the scope of Google search suddenly gets much, much bigger, and that search index will include you. Let me paint a picture. Ten years from now, someone, some company, or some organization, takes an interest in you, wants to know if you've ever said anything they consider offensive, or threatening, or just includes a mention of a certain word or phrase they find interesting. A single search query within Google's cloud – whether initiated by a publicly available search, or a federal subpoena, or anything in between – will instantly bring up documentation of every word you've ever spoken within earshot of a Google Glass device.

This is the discussion we should have about Google Glass. The tech community, by all rights, should be leading this discussion. Yet most techies today are still chattering about whether they'll look cool wearing the device.

Oh, and as for that physical design problem. If Google Glass does well enough in its initial launch to survive to subsequent versions, forget Warby Parker. The next company Google will call is Bausch & Lomb. Why wear bulky glasses when the entire device fits into a contact lens? And that, of course, would be the ultimate expression of the Google Glass idea: a digital world that is even more difficult to turn off, once it's implanted directly into the user's body. At that point you'll not even know who might be recording you. There will be no opting out.

Monday, January 28, 2013

GE to IBM: Watch your Data, We Are Coming


General Electric, the massive industrial conglomerate, will not be content to let IT leaders like IBM and Google hog all the glory in the internet of things era.


It sure looks like General Electric — the conglomerate that builds stuff ranging from appliances to jet engines — is spending a ton of time and resources to boost its profile in high (as opposed to “low”) tech. In fact it looks like it’s waging a massive PR campaign to show that it is not some grimy industrial relic but a force at the cutting edge of big data and “the internet of things.” If you don’t believe it, just download its November report on the industrial internet, which we covered here.

The latest evidence of this push? An interview with William Ruh, VP of software for GE Research, in ComputerWeekly.com. In the piece, Ruh appeared to take a veiled swipe IBM — which loves to portray itself as the thought leader in bleeding-edge tech and the kingpin in tech patents. (For the record, in 2012 GE came in ninth in patents with a total of 1,652 compared to IBM’s 6,478 — but who’s counting?)

Ruh said the airline industry has gathered tons of data about how jet engines have performed over the past two decades and that historical data should help guide predictive maintenance going forward. Ruh told ComputerWeekly:

“In emerging markets, we are seeing dirt and sandy environments … How are these affecting aero engines? [Business intelligence] cannot answer this. Nor can a supercomputer … Watson cannot tell me when this machine part will break.”

Watson is IBM’s much-hyped computer that boasts human-like thought processes and beat the human champion in Jeopardy a few years back.


GE is banking on the growing acknowledgement that machine data — information generated and collected by the types of industrial gear it makes — gives it an entry into the booming world of big data. That’s probably why GE CEO Jeff Immelt has been cropping up in a lot of interesting venues, including in an interview with Om Malik last month. And why GE came to San Francisco to announce its “Industrial Internet Quests” and tap into the wealth of software and data expertise there. As my colleague Katie Fehrenbacher put it at the time, the quest “calls on developers, data scientists and designers to make algorithms and applications that can increase productivity for the health and aviation sectors” — all sectors where GE plays.
It may be easy for folks in the valley to forget that GE has thousands of its own software developers on staff and builds sophisticated medical imaging and other high-tech gear: it does have credibility. And, at a time when the emphasis on making and building actual products is more valued, GE has lessons to teach.
The conglomerate obviously wants to be seen as a leader in this realm and won’t be content to let the likes of IBM hog all the glory in the internet of things era. After all, it builds an awful lot of those “things.”



Friday, December 28, 2012

What Turned Jaron Lanier Against the Web?


The digital pioneer and visionary behind virtual reality has turned against the very culture he helped create

Ron Rosenbaum, Smithsonian, January 2013 

I couldn’t help thinking of John Le Carré’s spy novels as I awaited my rendezvous with Jaron Lanier in a corner of the lobby of the stylish W Hotel just off Union Square in Manhattan. Le Carré’s espionage tales, such as The Spy Who Came In From the Cold, are haunted by the spectre of the mole, the defector, the double agent, who, from a position deep inside, turns against the ideology he once professed fealty to.

And so it is with Jaron Lanier and the ideology he helped create, Web 2.0 futurism, digital utopianism, which he now calls “digital Maoism,” indicting “internet intellectuals,” accusing giants like Facebook and Google of being “spy agencies.” Lanier was one of the creators of our current digital reality and now he wants to subvert the “hive mind,” as the web world’s been called, before it engulfs us all, destroys political discourse, economic stability, the dignity of personhood and leads to “social catastrophe.” Jaron Lanier is the spy who came in from the cold 2.0.

To understand what an important defector Lanier is, you have to know his dossier. As a pioneer and publicizer of virtual-reality technology (computer-simulated experiences) in the ’80s, he became a Silicon Valley digital-guru rock star, later renowned for his giant bushel-basket-size headful of dreadlocks and Falstaffian belly, his obsession with exotic Asian musical instruments, and even a big-label recording contract for his modernist classical music. (As he later told me, he once “opened for Dylan.” )

The colorful, prodigy-like persona of Jaron Lanier—he was in his early 20s when he helped make virtual reality a reality—was born among a small circle of first-generation Silicon Valley utopians and artificial-intelligence visionaries. Many of them gathered in, as Lanier recalls, “some run-down bungalows [I rented] by a stream in Palo Alto” in the mid-’80s, where, using capital he made from inventing the early video game hit Moondust, he’d started building virtual-reality machines. In his often provocative and astute dissenting book You Are Not a Gadget, he recalls one of the participants in those early mind-melds describing it as like being “in the most interesting room in the world.” Together, these digital futurists helped develop the intellectual concepts that would shape what is now known as Web 2.0—“information wants to be free,” “the wisdom of the crowd” and the like.

And then, shortly after the turn of the century, just when the rest of the world was turning on to Web 2.0, Lanier turned against it. With a broadside in Wired called “One-Half of a Manifesto,” he attacked the idea that “the wisdom of the crowd” would result in ever-upward enlightenment. It was just as likely, he argued, that the crowd would devolve into an online lynch mob.

Lanier became the fiercest and weightiest critic of the new digital world precisely because he came from the Inside. He was a heretic, an apostate rebelling against the ideology, the culture (and the cult) he helped found, and in effect, turning against himself.
***
And despite his apostasy, he’s still very much in the game. People want to hear his thoughts even when he’s castigating them. He’s still on the Davos to Dubai, SXSW to TED Talks conference circuit. Indeed, Lanier told me that after our rendezvous, he was off next to deliver the keynote address at the annual meeting of the Ford Foundation uptown in Manhattan. Following which he was flying to Vienna to address a convocation of museum curators, then, in an overnight turnaround, back to New York to participate in the unveiling of Microsoft’s first tablet device, the Surface.

Lanier freely admits the contradictions; he’s a kind of research scholar at Microsoft, he was on a first-name basis with “Sergey” and “Steve” (Brin, of Google, and Jobs, of Apple, respectively). But he uses his lecture circuit earnings to subsidize his obsession with those extremely arcane wind instruments. Following his Surface appearance he gave a concert downtown at a small venue in which he played some of them.

Lanier is still in the game in part because virtual reality has become, virtually, reality these days. “If you look out the window,” he says pointing to the traffic flowing around Union Square, “there’s no vehicle that wasn’t designed in a virtual-reality system first. And every vehicle of every kind built—plane, train—is first put in a virtual-reality machine and people experience driving it [as if it were real] first.”

I asked Lanier about his decision to rebel against his fellow Web 2.0 “intellectuals.”

“I think we changed the world,” he replies, “but this notion that we shouldn’t be self-critical and that we shouldn’t be hard on ourselves is irresponsible.”

For instance, he said, “I’d been an early advocate of making information free,” the mantra of the movement that said it was OK to steal, pirate and download the creative works of musicians, writers and other artists. It’s all just “information,” just 1’s and 0’s.

Indeed, one of the foundations of Lanier’s critique of digitized culture is the very way its digital transmission at some deep level betrays the essence of what it tries to transmit. Take music.

“MIDI,” Lanier wrote, of the digitizing program that chops up music into one-zero binaries for transmission, “was conceived from a keyboard player’s point of view...digital patterns that represented keyboard events like ‘key-down’ and ‘key-up.’ That meant it could not describe the curvy, transient expressions a singer or a saxophone note could produce. It could only describe the tile mosaic world of the keyboardist, not the watercolor world of the violin.”

Quite eloquent, an aspect of Lanier that sets him apart from the HAL-speak you often hear from Web 2.0 enthusiasts (HAL was the creepy humanoid voice of the talking computer in Stanley Kubrick’s prophetic 2001: A Space Odyssey). But the objection that caused Lanier’s turnaround was not so much to what happened to the music, but to its economic foundation.

I asked him if there was a single development that gave rise to his defection.

“I’d had a career as a professional musician and what I started to see is that once we made information free, it wasn’t that we consigned all the big stars to the bread lines.” (They still had mega-concert tour profits.)
“Instead, it was the middle-class people who were consigned to the bread lines. And that was a very large body of people. And all of a sudden there was this weekly ritual, sometimes even daily: ‘Oh, we need to organize a benefit because so and so who’d been a manager of this big studio that closed its doors has cancer and doesn’t have insurance. We need to raise money so he can have his operation.’

“And I realized this was a hopeless, stupid design of society and that it was our fault. It really hit on a personal level—this isn’t working. And I think you can draw an analogy to what happened with communism, where at some point you just have to say there’s too much wrong with these experiments.”

His explanation of the way Google translator works, for instance, is a graphic example of how a giant just takes (or “appropriates without compensation”) and monetizes the work of the crowd. “One of the magic services that’s available in our age is that you can upload a passage in English to your computer from Google and you get back the Spanish translation. And there’s two ways to think about that. The most common way is that there’s some magic artificial intelligence in the sky or in the cloud or something that knows how to translate, and what a wonderful thing that this is available for free.

“But there’s another way to look at it, which is the technically true way: You gather a ton of information from real live translators who have translated phrases, just an enormous body, and then when your example comes in, you search through that to find similar passages and you create a collage of previous translations.”

“So it’s a huge, brute-force operation?” “It’s huge but very much like Facebook, it’s selling people [their advertiser-targetable personal identities, buying habits, etc.] back to themselves. [With translation] you’re producing this result that looks magical but in the meantime, the original translators aren’t paid for their work—their work was just appropriated. So by taking value off the books, you’re actually shrinking the economy.”

The way superfast computing has led to the nanosecond hedge-fund-trading stock markets? The “Flash Crash,” the “London Whale” and even the Great Recession of 2008?

“Well, that’s what my new book’s about. It’s called The Fate of Power and the Future of Dignity, and it doesn’t focus as much on free music files as it does on the world of finance—but what it suggests is that a file-sharing service and a hedge fund are essentially the same things. In both cases, there’s this idea that whoever has the biggest computer can analyze everyone else to their advantage and concentrate wealth and power. [Meanwhile], it’s shrinking the overall economy. I think it’s the mistake of our age.”

The mistake of our age? That’s a bold statement (as someone put it in Pulp Fiction). “I think it’s the reason why the rise of networking has coincided with the loss of the middle class, instead of an expansion in general wealth, which is what should happen. But if you say we’re creating the information economy, except that we’re making information free, then what we’re saying is we’re destroying the economy.”

The connection Lanier makes between techno-utopianism, the rise of the machines and the Great Recession is an audacious one. Lanier is suggesting we are outsourcing ourselves into insignificant advertising-fodder. Nanobytes of Big Data that diminish our personhood, our dignity. He may be the first Silicon populist.
“To my mind an overleveraged unsecured mortgage is exactly the same thing as a pirated music file. It’s somebody’s value that’s been copied many times to give benefit to some distant party. In the case of the music files, it’s to the benefit of an advertising spy like Google [which monetizes your search history], and in the case of the mortgage, it’s to the benefit of a fund manager somewhere. But in both cases all the risk and the cost is radiated out toward ordinary people and the middle classes—and even worse, the overall economy has shrunk in order to make a few people more.”

Lanier has another problem with the techno-utopians, though. It’s not just that they’ve crashed the economy, but that they’ve made a joke out of spirituality by creating, and worshiping, “the Singularity”—the “Nerd Rapture,” as it’s been called. The belief that increasing computer speed and processing power will shortly result in machines acquiring “artificial intelligence,” consciousness, and that we will be able to upload digital versions of ourselves into the machines and achieve immortality. Some say as early as 2020, others as late as 2045. One of its chief proponents, Ray Kurzweil, was on NPR recently talking about his plans to begin resurrecting his now dead father digitally.

Some of Lanier’s former Web 2.0 colleagues—for whom he expresses affection, not without a bit of pity—take this prediction seriously. “The first people to really articulate it did so right about the late ’70s, early ’80s and I was very much in that conversation. I think it’s a way of interpreting technology in which people forgo taking responsibility,” he says. “‘Oh, it’s the computer did it not me.’ ‘There’s no more middle class? Oh, it’s not me. The computer did it.’

“I was talking last year to Vernor Vinge, who coined the term ‘singularity,’” Lanier recalls, “and he was saying, ‘There are people around who believe it’s already happened.’ And he goes, ‘Thank God, I’m not one of those people.’”

In other words, even to one of its creators, it’s still just a thought experiment—not a reality or even a virtual-reality hot ticket to immortality. It’s a surreality.

Lanier says he’ll regard it as faith-based, “Unless of course, everybody’s suddenly killed by machines run amok.”

“Skynet!” I exclaim, referring to the evil machines in the Terminator films.

At last we come to politics, where I believe Lanier has been most farsighted—and which may be the deep source of his turning into a digital Le Carré figure. As far back as the turn of the century, he singled out one standout aspect of the new web culture—the acceptance, the welcoming of anonymous commenters on websites—as a danger to political discourse and the polity itself. At the time, this objection seemed a bit extreme. But he saw anonymity as a poison seed. The way it didn’t hide, but, in fact, brandished the ugliness of human nature beneath the anonymous screen-name masks. An enabling and foreshadowing of mob rule, not a growth of democracy, but an accretion of tribalism.

It’s taken a while for this prophecy to come true, a while for this mode of communication to replace and degrade political conversation, to drive out any ambiguity. Or departure from the binary. But it slowly is turning us into a nation of hate-filled trolls.

Surprisingly, Lanier tells me it first came to him when he recognized his own inner troll—for instance, when he’d find himself shamefully taking pleasure when someone he knew got attacked online. “I definitely noticed it happening to me,” he recalled. “We’re not as different from one another as we’d like to imagine. So when we look at this pathetic guy in Texas who was just outed as ‘Violentacrez’...I don’t know if you followed it?”

“I did.” “Violentacrez” was the screen name of a notorious troll on the popular site Reddit. He was known for posting “images of scantily clad underage girls...[and] an unending fountain of racism, porn, gore” and more, according to the Gawker.com reporter who exposed his real name, shaming him and evoking consternation among some Reddit users who felt that this use of anonymity was inseparable from freedom of speech somehow.

“So it turns out Violentacrez is this guy with a disabled wife who’s middle-aged and he’s kind of a Walter Mitty—someone who wants to be significant, wants some bit of Nietzschean spark to his life.”

Only Lanier would attribute Nietzschean longings to Violentacrez. “And he’s not that different from any of us. The difference is that he’s scared and possibly hurt a lot of people.”

Well, that is a difference. And he couldn’t have done it without the anonymous screen name. Or he wouldn’t have.

And here’s where Lanier says something remarkable and ominous about the potential dangers of anonymity.
“This is the thing that continues to scare me. You see in history the capacity of people to congeal—like social lasers of cruelty. That capacity is constant.”

“Social lasers of cruelty?” I repeat.

“I just made that up,” Lanier says. “Where everybody coheres into this cruelty beam....Look what we’re setting up here in the world today. We have economic fear combined with everybody joined together on these instant twitchy social networks which are designed to create mass action. What does it sound like to you? It sounds to me like the prequel to potential social catastrophe. I’d rather take the risk of being wrong than not be talking about that.”

Here he sounds less like a Le Carré mole than the American intellectual pessimist who surfaced back in the ’30s and criticized the Communist Party he left behind: someone like Whittaker Chambers.

But something he mentioned next really astonished me: “I’m sensitive to it because it murdered most of my parents’ families in two different occasions and this idea that we’re getting unified by people in these digital networks—”

“Murdered most of my parents’ families.” You heard that right. Lanier’s mother survived an Austrian concentration camp but many of her family died during the war—and many of his father’s family were slaughtered in prewar Russian pogroms, which led the survivors to flee to the United States.

It explains, I think, why his father, a delightfully eccentric student of human nature, brought up his son in the New Mexico desert—far from civilization and its lynch mob potential. We read of online bullying leading to teen suicides in the United States and, in China, there are reports of well-organized online virtual lynch mobs forming...digital Maoism.

He gives me one detail about what happened to his father’s family in Russia. “One of [my father’s] aunts was unable to speak because she had survived the pogrom by remaining absolutely mute while her sister was killed by sword in front of her [while she hid] under a bed. She was never able to speak again.”
It’s a haunting image of speechlessness. A pogrom is carried out by a “crowd,” the true horrific embodiment of the purported “wisdom of the crowd.” You could say it made Lanier even more determined not to remain mute. To speak out against the digital barbarism he regrets he helped create.      



Saturday, November 3, 2012

Google Now: Behind the Predictive Future of Search


How Google learned to un-fragment itself and create the next big thing

Dieter Bohn, The Verge, October 29, 2012


For decades, visions of the future have played with the magical possibilities of computers: they'll know where you are, what you want, and can access all the world's information with a simple voice prompt. That vision hasn't come to pass, yet, but features like Apple's Siri and Google Now offer a keyhole peek into a near future reality where your phone is more "Personal Assistant" than "Bar bet settler." The difference is that the former actually understands what you need while the latter is a blunt search instrument.\

Google Now is one more baby step in that direction. Introduced this past June with Android 4.1 "Jelly Bean," it's designed to ambiently give you information you might need before you ask for it. To pull off that ambitious goal, Google takes advantage of multiple parts of the company: comprehensive search results, robust speech recognition, and most of all Google's surprisingly deep understanding of who you are and what you want to know.

With Android 4.2, launching alongside the Nexus 4 and Nexus 10 on November 13th, Google has updated the feature with new information cards in new categories. And yet, the amount of engineering effort that makes Google Now possible is out of proportion to what it does — it's a massive, cross-company effort for what seems like a relatively small product. That difference is a clue. Google Now isn't important for what it does, well, "now," but the building blocks are there for a radically different kind of platform in the future.
We sat down with the teams responsible for some of the technology that went into Google Now to find out what makes it tick today and discover some hints about what it could be in the future.

A deeper understanding
You may not be familiar with Google Now, primarily because it's only available on the sliver of Android devices running Jelly Bean (and up) — a situation that sadly won't change with the latest version. It's essentially an app that combines two important functions: voice search and "cards" that bubble up relevant information on a contextual basis.

Actually, Google Now technically only refers to the ambient information part of the equation, a branding kerfuffle that distinguishes it from Apple's Siri product yet still causes confusion. Those cards might contain local restaurants, the traffic on your commute home, or when your flight is about to take off. They appear automatically as Google tries to guess the information you'll need at any given moment.

While it seems like a relatively simple service, it's only really possible because of the massive amount of computational power Google can leverage alongside the massive amount of data Google knows about you thanks to your searches. It's "precisely what Google is best at," Android's director of product management, Hugo Barra, tells us. "It really feels like we’ve been working on Google Now for the past ten years. Because Google Now touches every back-end of Google, every different web service that’s been developed over the last ten years or so is part of this service."


The breadth of that backend and the simple cards it enables is what makes Google Now so intriguing as a product. One of Barra’s favorite examples is a voice search for something that pulls from all those multiple sources and turns it into a comprehensible and useful result. Searching for “Directions to the museum with the William Paley exhibition” causes Google to 1) find that exhibition, 2) understand you care about the museum where it is being shown, 3) know your location, and finally 4) present you with a simple map card to the museum itself along with a button to immediately get directions.

Taking all of that complex data and turning it into a relatively simple and useful interface is a gargantuan undertaking, but Google has started with a somewhat small set of categories for the types of cards it shows. With Jelly Bean, you'd see calendar alerts, weather, flight times, sports scores, transit directions, local restaurants, and a few more categories of information.

Even within that limited set of data, Google has to make choices about which cards to show you and when. It uses a few different signals — location, time, and all of your recent searches figuring prominently among them — to decide what to show you in any given moment. "It’s essentially a ranking problem, and it’s a very complicated one," according to Barra, but Google has perhaps more experience at solving ranking problems than any other company after years of delivering search results.

In my experience, Google is able to get you the "right" information you want a relatively small percentage of the time, but that low hit rate doesn't actually hurt the experience all that much. That's mainly thanks to the fairly small number of categories cards fall into, but also to the fact that when Google Now gets it right, it really feels magical. The sort of thing you might manually search for — like your commute time home — is simply waiting for you.

With the latest update, Google is expanding Now into new categories, increasing the different kinds of information it's able to provide. The new additions aren't radically ambitious, but that's in fitting with the overall feel of Google Now. What it shows you is more about serendipitous information than structured data.
The first category involved Gmail integration. With your permission, Google will keep an eye on your inbox and recognize flight confirmations, hotel reservations, restaurant bookings, event tickets, and package tracking emails. It will take that knowledge and give you a relevant card when appropriate — say, giving you your hotel information when you land in the right city or letting you know when it's time to leave for a concert.

The new features are part of Google’s growing efforts to provide relevant results based on the knowledge it’s accumulated about you. As search gets better, so do people’s expectations for what it provides. “Of course Google’s going to access more than just the public information on the web,” Scott Huffman, Engineering Director for Search Quality at Google tells us, “Google’s going to know when my flight is, whether my package has gotten here yet and where my wife is and how long it’s going to take her to get home this afternoon. [...] Of course, Google knows that stuff.” If you’re willing to opt in to letting Google know so much about you — and increasingly, opting in is the default — then Google wants to return the favor by using that information to your benefit. It requires you to trust Google quite a bit, but the company hopes that your trust will be rewarded.

These new cards are actually similar to a feature that Google added to its web search results this past August, both in content and in style. That's probably not an accident — if you assume Google has already won the battle for search, the next battle is giving you information before you even search for it. When it comes to deciding which data to give you, Barra tells us that Google has "a pipeline [...], possibly in the hundreds of cards” from its many engineering teams. Rather than flood users with all of those new cards, Google is taking a slow and steady approach to adding those new features — if only because right now it can only add those cards with a software update.

Some of the other new categories of cards are relatively minor additions: stocks, news, local concerts, movies, and local attractions. It also has a basic exercise tracking card that utilizes the phone’s accelerometer and location data: every month it will let you know how far you've walked or biked and also tell you how it compared to the month previous. Another new card lets you know that you're near a "photo opportunity," as Product Management Director Baris Gultekin told us. It uses data from Google's Panaramio service, noting when you're close to a place that has a "high density of pictures taken at a spot." You can see photos that were taken at the landmark and, Google hopes, take one yourself.

Neural networks
Just as Google Now's ambient information is backed by a massive and unseen engineering effort, Google's voice search is a simple feature that belies the effort that goes behind it. Huffman points out that getting voice search right actually involves more than just turning spoken words into textual queries, "speech recognition, natural language understanding, and understanding entities and knowledge in the world [all] really have to come together."

Voice search is the sort of feature that we take for granted on smartphones — Apple’s Siri and even Windows Phone both use the feature to offer up search results that go beyond basic web searches. What used to be a "hey neat" kind of feature is increasingly becoming an expected feature, and Google is well aware of that, "As you make search better, people’s expectations go up." To meet those expectations, Google is attacking all three of the areas Huffman delineated in equal measure.

Speech recognition is a very difficult problem to solve, as anybody who has dealt with voice search knows all too well. Recently, Google has changed its approach to making it work in a fundamental way, replacing a system that was the result of years of effort with a new framework for understanding the spoken word. Google has shifted to using a neural network that's much more effective at understanding speech.

A neural network is a computer system that behaves a bit like the actual neurons in your brain do. 

Essentially, the computer is designed with layers of software-based "neurons" that do the same thing actual neurons do: take input in and "fire" off to other neurons based on the data they receive. Over the summer, the results of research led by Google Fellow Jeff Dean's on neural networks made some waves: Google had taught a computer to recognize cats in videos. The interesting part is that the neural network essentially created the concept of "cat" on its own without direct human intervention.

Here's how it works: The first layer of neurons looks for very simple things, like angled lines or colors. If it sees something that matches, it fires off a signal. There's then a second layer of neurons, which simply pays attention to sets of neurons firing from the first layer. As you add in more and more layers with the same behavior, you essentially add in layers of conceptual abstraction until, at the very top layer, there's a neuron that has trained itself to recognize cats 15.8 percent of the time.

Of course, that doesn't mean that the computer "understands" cats in a conscious way, but the effect of it being able to recognize something like a cat without direct human training is what's important. "With a lot of other machine learning techniques," Dean explains, "you often have to do a lot of work to hand-engineer exactly the right features [...] Whereas with a neural network you can feed in much rawer forms of data."
Google's research scientists took this method and essentially applied it directly to speech recognition, fellow researcher Vincent Vanhoucke told us. "We picked up the kind of work that Jeff’s team was doing and just changed the input of the system." Google used the neural network at a very basic level of speech recognition: understanding and interpreting the basic sounds of speech: phonemes.

The approach "led to about between 20 to 25 percent reduction in the error rate in our system," according to Vanhoucke. The neural network turned out to be exceptionally good at solving what used to be very thorny problems in speech recognition. Accounting for "different environments, [...] different accents, different tones of voice, different pitches, different background noise, different microphones, [...], people talking in the background, different audio conditions" became much easier because the network was able to automatically learn how to account for each situation.

Knowledge Graph
Just understanding the words you've spoken isn't enough, obviously. Just as a neural network trades in increasing layers of abstraction, Google itself needs to move beyond basic web queries. In a very real way, Google is trying to get its computers to actually understand what it is you're asking them. Part of that comes from a relatively new initiative called the "Knowledge Graph," the company's effort to compile a database of "entities" in the world.

Today, Google's servers are aware of 500 million such entities, and "knowing" those things means that the company is able to act on them in interesting ways. For example, if you search for a “Tom Cruise,” Google knows you’re referring to a person instead of a vacation and can then tell you specific facts about him instead of simply crawling the web for related words. In truth, Google only knows those details because it is so adept at crawling the web — but the additional layer of abstraction created by putting that information into the structured Knowledge Graph means that Google can do more with search results. It "allows voice search, in some sense, [to] give me something to talk about," says Huffman. In Tom Cruise example, returning an “entity” instead of just a search result means you can contextually ask for more information, like “what movies has he been in?” or “how tall is he, really?”

Having something to talk about and talking to somebody are two different things, and with regard to the latter Google is again taking a Google-esque approach. As opposed to Apple's Siri, which you could say has a distinct personality, Huffman says that Google has "shied away from the idea of kind of a human persona for search or for the entity that you’re interacting with and instead tried to go for, in some sense, ‘hey, you’re interacting with all of Google.’"

Wrap-up
The Google that you're interacting with in Google Now is very different than the Google you used even a year ago. The company's products have often felt fragmented, serving small niches and launched without feeling fully thought-through — and then in too many cases simply killed off. That may have been a function of the fact that Google is so large and does so much — but Google Now is a sign that all the different parts of Google are finally working together in a cohesive way.


"Google Now actually started as a twenty percent project," Barra told us. Google famously encourages its employees to work on "side projects" for some portion of their time, and what's interesting about Google Now is that although it started two years ago as one of these side projects, it's become a catalyst for integrating so many different parts of Google. Barra tells us that “we literally have dozens of teams working with us right now,” and the achievement with Google Now is that it feels like those teams are integrated, not fragmented.

In a single app, the company has combined its latest technologies: voice search that understands speech like a human brain, knowledge of real-world entities, a (somewhat creepy) understanding of who and where you are, and most of all its expertise at ranking information. Google has taken all of that and turned it into an interesting and sometimes useful feature, but if you look closely you can see that it's more than just a feature, it's a beta test for the future.

Thursday, October 18, 2012

IBM's Watson Is Learning Its Way To Saving Lives


A few years ago, IBM’s new computer was a game-playing curiosity. Now Watson is poised to change the way human beings make decisions about medicine, finance, and work.

Jon Gertner, Fast Company, October 15, 2012.

The woman was gravely ill. Her name was Ms. Yamato. Thirty-seven years old, born in Osaka, Japan, she had never smoked, and yet there it was anyway: a spot on her lung.
A doctor had already performed a bronchoscopy and had made the diagnosis of cancer. Then he referred the patient to Mark Kris, an oncologist at Memorial Sloan-Kettering Cancer Center in New York. Seated alongside me in his office on the Upper East Side of Manhattan, Kris is showing me Ms. Yamato's electronic medical record on an iPad. "I'm preparing for the first visit," he explains, swiping the screen to show what that entails. He's interested in running at least two tests on the patient. The first is an MRI, to find out if the cancer has spread to her brain. The second involves a deeper diagnostic regimen. Lung cancer tumors are not all the same; there are thousands of variations. So a test that examines the mutations within a tumor will be crucial, he says. It so happens that cancer patients born in East Asia who have never smoked often have a particular mutation that responds well to a medication by the name of Erlotinib. That may be the case here. One can hope.

Over the past year, IBM executives have come to believe that Watson represents the first machine of the third computer age.

The woman is not real. She happens to be a character within an app that IBM has created for Watson, its new computer. Watson's special talent, its reason for being, is a singular ability to grasp the intricacies of human language and answer exceedingly difficult questions. You may have heard about Watson already. Back in 2007, a group of computer engineers at IBM's research labs in upstate New York began building the machine--named for IBM's founder, Thomas J. Watson--with the goal of creating a question-and-answer technology that would be more authoritative and powerful than anything on the planet. The initial objective of the Watson group was simple: to win in the game show Jeopardy!, something Watson famously achieved in February 2011. Yet the group had a far more important goal: to turn Watson into a business, hopefully one of some scale. So starting in late 2009, a business development team at IBM began holding meetings outside the company in an effort to understand the ultimate worth of this new technology. No doubt it could be a business one day. But what kind of business?

"The first thing that hit us about Watson," recalls John Kelly, IBM's chief of research, "was that this thing could be applied almost anywhere." Early on, IBM executives decided to focus on a field in which Watson could have a notable social impact while also proving its ability to master a complex body of knowledge. The team chose medicine. They believed Watson could help doctors make diagnoses and, even more important, select treatments. Specifically, they thought Watson could be the perfect tool to chart the complex decision trees that cancer specialists like Kris negotiate every day as they weigh treatment options that might involve radiation, surgery, and any of countless chemotherapy drugs. Watson can ingest more data in a day than any human could in a lifetime. It can read all of the world's medical journals in less time than it takes a physician to drink a cup of coffee. All at once, it can peruse patient histories; keep an eye on the latest drug trials; stay apprised of the potency of new therapies; and hew closely to state-of-the-art guidelines that help doctors choose the best treatments. Watson never goes on vacation. And it never forgets a fact. On the contrary, it keeps learning.

This fall, after six months of teaching their treatment guidelines to Watson, the doctors at Sloan-Kettering will begin testing the IBM machine on real patients. The Ms. Yamato app shows how it will work. After Kris inputs the results of her medical tests, Watson begins deliberating. "It's going through its algorithms," Kris says as we stare at the iPad. "It's seeing where the data sends it today." On the screen, a colorful globe spins. In a few seconds, Watson offers three possible courses of chemotherapy, charted as bars with varying levels of confidence--one choice above 90% and two above 80%. "Watson doesn't give you the answer," Kris says. "It gives you a range of answers." Then it's up to Kris to make the call. He regards the options on the screen and wonders how they might change if Ms. Yamato happened to develop a common symptom: hemoptysis, or coughing up blood.

"Let's try that," he says. He inputs the information and shows me the result approvingly. Watson has dropped one drug from the top chemo regimen. That's just what Kris would have done.

To make sense of all this--that is, to gauge both the value of Watson to a hospital like Sloan-Kettering and its potential to change forever the worlds of medicine and business--you could follow two different paths. You might consider Watson's evolutionary promise. Watson can almost certainly generate huge administrative benefits. Already, one large health insurer--Indiana-based Wellpoint--has begun using a Watson computer in its Virginia data center to speed along the authorization for medical procedures. Usually, authorizations are evaluated by a team of trained nurses and can sometimes take weeks to come through. Watsonizing the process would speed it up--a boon for a doctor like Kris, who now must wait while assistants exchange faxes with insurers before he can get clearance for any expensive tests.

Kris shows me what happens when Watson's treatment plan calls for an MRI. A button pops up on his screen to ask for preauthorization. "I just click that," he says, and it's done instantly.

I ask him what if Watson's request is denied.

Kris seems amused by the question. Watson has already consulted the latest medical literature, and it's been trained by the best cancer doctors in the world. "Who is the authority that is going to trump that?" he asks. Insurers balk at paying for unnecessary procedures; Watson's expert opinion essentially guarantees the necessity.

But the more intriguing path is the second one--a consideration of Watson's potential to do something revolutionary. This is the trail that captivates Kris. Eventually, he thinks, Watson could provide any doctor anywhere with the world's best second opinion. A physician in a community hospital in the Midwest, or at a remote medical center in China, could have instant access to everything that the medical field's best oncologists--people like Kris and his colleagues at Sloan-Kettering--have taught Watson. What is more, Watson will be able to excavate facts beyond the ken of Sloan-Kettering's current lineup of specialists. As Kris says, "We could ask Watson: What is the best treatment for this rare condition based on all of Sloan-Kettering's records?" It could then go through several years of cancer cases looking for the most successful outcomes. In time, it could even look at hospital records from around the world. As Manoj Saxena, the IBM executive now in charge of commercializing Watson, tells me: "It's like being able to take a knowledge worker--cancer specialist, nurse, bond trader, portfolio manager, whatever--and equip that person with the best knowledge, and have it available at their fingertips." As Watson evolves, Saxena believes, these knowledge banks will significantly alter how, and how well, humans make decisions.

Within a few years, for instance, Watson may be reaching well beyond oncology to assist patients suffering from any chronic disease and help general practitioners make diagnoses in their offices. Ultimately, Saxena believes, Watson could play an essential role in the diagnosis and treatment of mental health; in the financial services industry, where Citibank is testing it now; and in education. It could become the world's smartest dietitian.

How Watson Works
IBM's Watson computer begins trials in the health care industry this fall. The initial goal is to help oncologists make better decisions for cancer treatment; eventually, the computer will also aid in the diagnosis and treatment of other chronic diseases.

1. For well over a year, the Watson computers have been "trained" in science and medicine. Technicians feed Watson medical textbooks and journals, patient histories, and treatment guidelines.

2. At Memorial Sloan-Kettering Cancer Center in New York, doctors have begun using a Watson appon a tablet to access the computer through the cloud. The doctor logs in to Watson and begins to input data and ask questions.

3. When the oncologist queries Watson about a course of treatment for a lung or breast cancer patient, the computer--with its ability to understand natural language--notes keywords in the query, such as the particular type of cancer and the genomic variant of the tumor.

4. Watson then springs into action, using its massively parallel processors to review millions of pages of text in seconds. It explores the patient's medical history, medications, and other existing conditions. It then combines this information with recent data from the patient's medical tests and may comb through studies of patient groups at Sloan-Kettering who have had similar types of cancer. It also reviews doctors' and nurses' notes, recent medical research, journal articles, and treatment guidelines.

5. Watson then generates hypotheses for treatment. On the tablet app, these appear as separate options with varying levels of confidence. For instance, Watson might score one treatment option--a combination of chemotherapy drugs--with a 95% confidence level, suggesting it would be the most sensible path. It might also highlight options with lower scores as alternative treatment courses. The doctor then weighs the options and makes the call.

Saxena now commands a team of about 200 people who are working to adapt Watson's skills for various IBM clients. He and I are discussing his progress over lunch one day near IBM's upstate New York headquarters when he leans back and tells me that after creating two successful tech startups, both of which he sold (the second to IBM), his current job is far and away the most meaningful endeavor of his life. Those startups, he confides, were exciting, important. "But this," he says of the Watson rollout, "this is stuff that is going to change the course of history."

Over the past year, IBM executives have come to believe that Watson represents the first machine of the third computer age, a category now referred to within the company as cognitive computing. As Kelly describes it, the first generation of computers were tabulating machines that added up figures. "The second generation," he says, "were the programmable systems--the mainframe, the first IBM 360, PCs, all the computers we have today." Now, Kelly believes, we've arrived at the cognitive moment--a moment of true artificial intelligence. These computers, such as Watson, can recognize important content within language, both written and spoken. They do not ask us to communicate with them in their coded language; they speak ours. And perhaps most important, they can learn, so they improve without constant human instruction.

Siri, on the iPhone, might be considered an elementary example. Watson is industrial strength. "Computers do numerical calculations, they move data around, and they've been doing that forever," David Ferrucci, the IBM researcher who commanded the team that built the first Watson computer, tells me one day at IBM's research labs. "When I think about Watson, it's interpreting the information in human terms. It's saying: What does this mean to me? And that's a big deal." Also significant is how Watson renders an answer. Unlike its responses in Jeopardy!, in the real world it will perform as it did for Kris at Sloan-Kettering--by giving not a single solution but a range of probable solutions, each backed up by Watson's evidence and ranked by its level of confidence. In the lingo of computer science, that makes the machine probabilistic rather than deterministic. One might say this trait gives Watson a humanizing glow of humility and diminishes concerns that it marks a stride toward a computer-led dystopia. Watson, in IBM's marketing schema, is here to help with our questions, rather than solve them. In the case of medicine, it--for Watson is not really a he--is here to support doctors, not replace them.


The Watson of today is not precisely the same machine that won in Jeopardy! IBM has fine-tuned its software and algorithms for medical applications (or, in the case of Citibank, financial services applications). Watson has shrunk, too, from a row of about a dozen server racks that would have filled a small bedroom to an assemblage about the size of a double-door refrigerator. But for all the concentrated power, it doesn't look like anything special. Its sleek black servers are standard IBM Power 750s. You could wander around Watson and regard its blinking lights, as I did on a quiet midsummer afternoon at IBM's research labs, and not think something unusual is happening inside it. But there is. The way Watson solves problems--or, rather, the way it looks for answers, simultaneously sending out thousands of inquiries in all directions and then scoring the evidence it collects--is different from how other computers work. One person at IBM likens Watson's process to (1) gathering hundreds or thousands of possible solutions from a vast data bank, (2) pouring them into a giant funnel, (3) stirring with a dash of algorithms, and (4) letting only the best drip out of the bottom.

At the moment, a half-dozen Watsons are scattered around the country. Some are on the premises of IBM clients, as with the insurer Wellpoint, while others are cloud based, which is how hospitals such as Sloan-Kettering will access Watson. "Effectively, there's no limit to how many Watsons there can be," Bernie Meyerson, IBM's VP of innovation, tells me. Watson is a creation of software, not hardware. "That's the beauty of it," he says.

Watson is different from big servers and mainframes in other ways, too. The best computers of today have the extraordinary processing power needed to create, say, complex supply chains for building a new automobile or planning a satellite launch. These machines are good at manipulating the vast amounts of clearly defined data--numbers and facts--known as structured information. But most of the world's information is more ambiguous and less precise and lies beyond their reckoning. "We now have this proliferation of what we call Big Data," Saxena, Watson's business manager, tells me, referring to the flood of information created by our computers, our electronic sensors, and ourselves. "Ninety percent of the world's information was created in the last two years," he says. "But 80% of that 90% is unstructured or semistructured information, like doctor's notes or product reviews on Amazon." This near infinitude also includes tweets, blogs, emails--all the noise and scribble of modern life. So any company that aspired to manage the data of all the world's businesses would today be able to analyze only a small part of it. Watson, though, is a genius at reading unstructured information. And it's precisely this facility that explains why IBM sees such a rich business opportunity here.

It likewise explains why medicine is a logical first choice. While some health information is indeed structured--think of blood-pressure readings or cholesterol counts--the vast majority is unstructured. This cache includes textbooks, medical journals, patient records, and nurse and doctor evaluations. In fact, medicine embodies so much unstructured information that its proliferation has, by the account of many medical professionals, far outstripped the ability of doctors to keep up. Neither better training nor continuing education could ever wholly remedy this problem. When I meet with Herbert Chase, a professor of clinical medicine at Columbia University who consulted with IBM during the early stages of the Watson project, he says it is "not humanly possible" for a busy doctor to keep abreast of the current literature.

One result of information overload is a high rate of misdiagnosis and consequently incorrect treatment. By some estimates, Saxena tells me, 20% of initial diagnoses of cancer are eventually altered. "Imagine the implications of cancer care if there is a one in five chance that for the next six months whatever therapy they're giving you is wrong," he says.
Deciding on a course of treatment is even tougher than making a diagnosis. "It's still possible for a doctor to know the ways that people get sick," says Chase, who is also a kidney specialist. "But what is unmanageable, and what has been for decades, is knowing what the best option is today." Some applications now available to doctors are meant to alleviate this problem; one popular web-based tool is named Isabel. But Watson, in Chase's view, reaches a different level of sophistication. "I'll give you an example of a test we thought up for Watson," he tells me one day in his Manhattan office. "A patient was pregnant, had Lyme disease, and was also allergic to penicillin. And Watson came up with a drug. The first thing I thought was, Watson made a mistake. That drug can't be given to someone allergic to penicillin." But Chase was wrong, not Watson. "My knowledge was about five years old," he says. "And in the past couple of years, all the muckety-mucks had reviewed all the studies and had concluded yes, you can give that drug to someone who's allergic to penicillin."
To Chase, this proves a point: If you're a patient, you don't want to believe your doctor doesn't know everything. But he or she doesn't, and can't. At its best, the dispensation of treatment is inefficient today. "At its worst," Chase says, "it's subpar, incorrect, wrong therapy," and doesn't reach the standard of care to which his profession aspires. "As you can imagine," he adds, "this is not something we like talking about."

Last year, IBM turned 100 years old, which sets it apart from West Coast counterparts like Amazon, Apple, Google, HP, and Microsoft--all younger and ostensibly the tech world's leading innovators. To delve into IBM's recent research, though, is to wonder if our perception of technological leadership sometimes suffers from the distortions of branding and familiarity. We use iPhones and search engines and laser printers every day. But IBM's technologies are lodged deeper within the infrastructure of daily life; you're tapping into them whenever you send an email, for instance, or log on to a website. IBM has been granted more patents than any other company in the world for 19 years in a row. Yet since getting out of the laptop business in 2004, it has not produced a single product that it sells directly to the consumer.

If you're a patient, you don't want to believe your doctor doesn't know everything. But he or she doesn't, and can't.

To understand how Watson figures into the company's culture of ideas, or to see how it represents the kind of large-scale innovation that arguably lies beyond the capabilities of any startup, it helps to understand what the company actually does these days. IBM has operations in 172 countries and an organizational chart that resembles a vast Soviet bureaucracy. It employs about 433,000 men and women. Though IBM still sells hardware--big mainframe computers, silicon chips, and supercomputers--mainly it makes money selling software and consulting services to businesses and governments. The company's strategy has been validated of late by its performance: IBM's stock price has been on an upward trek for the past five years, and its winning streak has attracted the likes of Warren Buffett, who last year decided the company merited an investment of $10.7 billion. Meanwhile, as one of the few global titans to invest staggering sums on R&D ($6 billion to $7 billion a year), IBM maintains one of the world's last great industrial laboratories. At its main research center in Yorktown Heights, New York, a jet-age dream of glass curtain walls and rusticated stone designed by the Finnish-American architect Eero Saarinen, IBM employs the bulk of what is likely the world's largest mathematics department, with 300 members. If you're looking for a new PC design, you're out of luck here. But if you're shopping around for a new or better algorithm, IBM can build you one.

Not everyone is impressed by the direction of IBM's management. A relentless focus on earnings and cost cutting has led to a significant offshoring of domestic jobs, and a vocal corps of disillusioned or laid-off IBMers regularly take to the web to lament that the company's best days are behind it. IBM has also had its share of technological stumbles, apparently bungling several high-profile government contracts in recent years (in Texas and Indiana, for example) that left the company embroiled in disagreements with unhappy clients. And though these flare-ups may be uncommon, the company otherwise rarely quickens the pulse, with a long-standing reputation for being slow, steady, reliable, and maybe a little dull. IBM doesn't have big growth spikes or ballyhooed product launches; rather, it has plodding, long-term client contracts built around its ability to help optimize, say, a company's global IT services or a public utility's electrical grid. The corporation moves along like a supertanker. "IBM's annual revenue base is huge--$100 billion," says Toni Sacconaghi, a technology analyst for Sanford C. Bernstein. "So to move the needle is tough. It's hard to find big new products."


The managers and engineers keep looking anyway. One way IBM tries to infuse the troops with a sense of mission is through its periodic attempts to create for itself a Grand Challenge, such as the construction of Deep Blue, a chess-playing computer, or, more recently, Watson. The Grand Challenges are focused and expensive efforts--IBM will not verify Watson's cost, but estimates put the sum between $100 million and $1 billion--to push the company beyond the competition.

Watson's origins can arguably be traced back some years to a more modest annual initiative IBM calls the Global Technology Outlook, or GTO. Anyone at IBM can contribute to the outlook, and most of the results are eventually made public. The GTO tries to identify future business opportunities by putting a spotlight on various technology trends. A while ago, the IBM outlook pointed to analytics as a potentially huge field. Not long after, then-CEO (and current chairman) Sam Palmisano green-lighted IBM's acquisition of about $16 billion in smaller companies that had computer technologies to do this kind of work--essentially, to comb through vast stores of data, both structured and unstructured, and help extract nuggets from the global corporate babel.

Like Big Data or cloud computing, analytics is one of those contemporary catchphrases that everyone talks about but no one pauses to define. Bernie Meyerson, IBM's VP of innovation, argues that the great promise of analytics is not just to spot trends or glean information for boosting sales but to use computers and software to change the future. "Analytics is the capability to see what no human can," he says. Recently, at a public event, Meyerson was asked if IBM missed out by not building a tablet to compete with the iPad. He responded that as part of its Smarter Cities Initiative, IBM had just spent several years gathering all of the data on car transportation in Singapore; it then fed the data into a model it had built to predict the time and location of traffic jams. "We know from history what happens in Singapore if you slow the lights down in one direction by three seconds, and how to tweak the model so the jam never happens," he told his questioner. "And so there will be a traffic jam that never occurs because we can predict what happens 20 minutes from now, because we can take enough Big Data and crunch it, and do analytics on it. So we're predicting the future, and changing it. And you're asking me if I'm worried about a tablet?"

Watson, too, fits into Meyerson's conception of analytics, though it aims to change not the future of a traffic jam but of illness and investing. And by all indications, that tantalizing promise is not lost on the business community. "I have my shoulder against the door," Saxena tells me. He means he is turning clients away--something I heard from several other sources, too--until IBM executives feel confident Watson has proved its credibility at places like Wellpoint and Sloan-Kettering. Saxena seems certain that Watson will be a multibillion-dollar business, though he will only go so far as to say that by 2015, IBM will have annual revenues of about $16 billion from its analytics portfolio, of which Watson will be a part. When I put the question of Watson's potential to John Kelly, IBM's chief of research, he says: "It's like asking, at the very beginning, How big will the PC industry be?"

Kelly notes that the business model for Watson is still to be determined. He isn't sure whether selling Watson as a computer or marketing it as a service will make the most sense. But he feels he has time to decide. None of IBM's competitors, more than a year after the Jeopardy! victory, has announced a Q&A technology like Watson. "I think we have a huge lead," Kelly tells me. "When people realize this is not a one-off game machine but a new era of computing, then you'll see other companies tripling down to catch up."

I asked a number of people, both within IBM and outside of it, whether other organizations could have built this machine first. The consensus was probably not. The reasons did not precisely connect to IBM's technological capabilities--Google and Microsoft have plenty of computer prodigies in their ranks too. Rather, it was the combination of assets at IBM that made the difference. The company had its vast corporate lab, huge sums it was ready to invest, a profound expertise in hardware as well as software, and a collaborative culture that brought in lots of help from academia. And crucially, it had its business clients. In this respect, being a company that doesn't cater to consumers has advantages. Watson is only as bright as its teachers. Without the staff at Sloan-Kettering, where doctors like Mark Kris teach it oncology, Watson would not be nearly so smart. In fact, it might be kinda dumb. Or it might get all sorts of things wrong, like Siri does, except you'll be looking not for a pizza parlor but for a tumor.


From the start, the team that originally built Watson under David Ferrucci has worked out of a big room on the second floor of IBM's Hawthorne Labs in Westchester County, New York. Hawthorne is a large glass cube of a building situated about 30 miles north of New York City. Inside the Watson work space are five fake wood-grained tables, each home to a group of computer engineers who sit around and alternately immerse themselves in their screens or break to discuss coding with a neighbor. The mood here is sober. The staffers bring water bottles, not junk food. These aren't the unlined faces you'll see at a startup. Indeed, Ferrucci, who sits off to the side, is a suburban dad who looks like he'd be just as comfortable standing in front of a grill with a basting brush as he is overseeing his team. The walls here are covered with huge whiteboards crammed with the hieroglyphics of computer science. Overhead lights cast the room in gloomy fluorescence. The place has the neglected feel of a finished basement in a 1970s-era subdivision.

In early fall, the Watson team, now about 45 strong, began moving its work to a gleaming new space in IBM's main Yorktown Heights research laboratory--a promotion that reflects their importance as they support Saxena's much larger business development group while simultaneously working on the next iteration of Watson, known as Watson 2.0. One of the team's goals is to make Watson adaptable enough so that it doesn't require several dozen people spending a year to get it ready for every new application, such as medicine or financial services. But a more immediate project is to help Watson through the U.S. Medical Licensing Examination, the complex test all med-school graduates must take before practicing. If it passes, says Ferrucci, "that doesn't mean I can have a computer be a doctor." But IBM would gain what he calls "a crisp metric" that proves Watson has a real proficiency in medicine. The credential would no doubt help Watson's standing with health insurers, doctors, and patients, too. Passing the licensing exam is a difficult task--far harder than winning at Jeopardy!--but in early September, Ferrucci seemed pleased by the results. The computer is doing "interestingly well," he said. He sounded confident that Dr. Watson will ace the test by year's end.

Harder to intuit is how soon afterward Watson will infiltrate society. When I ask Jaime Carbonell, a computer science professor at Carnegie Mellon, he says he has no doubt the impact of Watson will be significant. "But I don't think there will be one moment of, 'Now we have it and yesterday we didn't,'" Carbonell remarks. "It will take time to permeate. Like cell phones, which were big, clumsy things you could barely carry at first." Was there a year, or month, or day, he asks, when cell phones began to change the world? "I can't think of when that was," he says. "But now we can't do without them."

Such is the course of technology: Electronic tools initially available only to the elite grow ever faster, smaller, cheaper. Kelly tells me he believes that eventually Watson will shrink to the size of a handheld device. Randy Katz, a computer science professor at UC Berkeley, sees a more approachable Watson, too. "Can the person in the street ask Watson a question now? No, he can't," says Katz. "But in five or 10 years, will there be systems like that--like Siri, but much better? I think the answer is yes."

In many of my conversations at IBM, the talk often drifts to applications of Watson. All sorts of intriguing scenarios are presented to me--for instance, that Watson will soon analyze not just words but images, such as MRIs and EKGs. Or it will diagnose a spider bite on a child's arm in a crop field in Africa, transmitted via smartphone by his worried father to a U.S. hospital. One afternoon, Saxena suggests this one: When you think you're coming down with the flu, Watson will be able to discern, before you even arrive at the doctor's office, that it might be a ragweed allergy, based on your medical record (you've had the same symptoms twice before at this time of year); your symptoms (gleaned from the insurance claim and diagnostic information in journals); and recent news (it just read an article in the Austin-American Statesman on a ragweed outbreak near your hometown).

It all sounds amazing. It's also speculative. Watson has not yet saved a life or a dollar of medical costs, or added anything, really, to IBM's bottom line. It has not yet faced its resistors--doctors who may find the technology objectionable and slow its adoption. It has not yet, as Saxena believes it will, changed the course of history. It has only won a television game show.

Still, Saxena predicts the computer will begin to scale up dramatically late next year. "By then," he says, "we will have built the technology, demonstrated it, built the tooling and methods around it. We will have the recipe book, and then we'll just push it out." But he will only have reached the end of Watson's beginning.

A version of this article appears in the November 2012 issue of Fast Company.