Showing posts with label Analytics. Show all posts
Showing posts with label Analytics. Show all posts

Tuesday, September 11, 2012

Erik Brynjolfsson and Andrew McAfee : Big Data's Management Revolution

Big Data's Management Revolution

by Erik Brynjolfsson and Andrew McAfee  |  10:05 AM September 11, 2012

http://blogs.hbr.org/cs/2012/09/big_datas_management_revolutio.html

Big data has the potential to revolutionize management. Simply put, because of big data, managers can measure, and hence know, radically more about their businesses, and directly translate that knowledge into improved decision making and performance. Of course, companies such as Google and Amazon are already doing this. After all, we expect companies that were born digital to accomplish things that business executives could only dream of a generation ago. But in fact the use of big data has the potential to transform traditional businesses as well.

We've seen big data used in supply chain management to understand why a carmaker's defect rates in the field suddenly increased, in customer service to continually scan and intervene in the health care practices of millions of people, in planning and forecasting to better anticipate online sales on the basis of a data set of product characteristics, and so on.

Here's how two companies, both far from Silicon Valley upstarts, used new flows of information to radically improve performance.

Case #1: Using Big Data to Improve Predictions
Minutes matter in airports. So does accurate information about flight arrival times: If a plane lands before the ground staff is ready for it, the passengers and crew are effectively trapped, and if it shows up later than expected, the staff sits idle, driving up costs. So when a major U.S. airline learned from an internal study that about 10% of the flights into its major hub had at least a 10-minute gap between the estimated time of arrival and the actual arrival time — and 30% had a gap of at least five minutes — it decided to take action.

At the time, the airline was relying on the aviation industry's long-standing practice of using the ETAs provided by pilots. The pilots made these estimates during their final approach to the airport, when they had many other demands on their time and attention. In search of a better solution, the airline turned to PASSUR Aerospace, a provider of decision-support technologies for the aviation industry.

In 2001 PASSUR began offering its own arrival estimates as a service called RightETA. It calculated these times by combining publicly available data about weather, flight schedules, and other factors with proprietary data the company itself collected, including feeds from a network of passive radar stations it had installed near airports to gather data about every plane in the local sky.

PASSUR started with just a few of these installations, but by 2012 it had more than 155. Every 4.6 seconds it collects a wide range of information about every plane that it "sees." This yields a huge and constant flood of digital data. What's more, the company keeps all the data it has gathered over time, so it has an immense body of multidimensional information spanning more than a decade. RightETA essentially works by asking itself "What happened all the previous times a plane approached this airport under these conditions? When did it actually land?"

After switching to RightETA, the airline virtually eliminated gaps between estimated and actual arrival times. PASSUR believes that enabling an airline to know when its planes are going to land and plan accordingly is worth several million dollars a year at each airport. It's a simple formula: Using big data leads to better predictions, and better predictions yield better decisions.

Case #2: Using Big Data to Drive Sales
A couple of years ago, Sears Holdings came to the conclusion that it needed to generate greater value from the huge amounts of customer, product, and promotion data it collected from its Sears, Craftsman, and Lands' End brands. Obviously, it would be valuable to combine and make use of all these data to tailor promotions and other offerings to customers, and to personalize the offers to take advantage of local conditions.

Valuable, but difficult: Sears required about eight weeks to generate personalized promotions, at which point many of them were no longer optimal for the company. It took so long mainly because the data required for these large-scale analyses were both voluminous and highly fragmented — housed in many databases and "data warehouses" maintained by the various brands.

In search of a faster, cheaper way, Sears Holdings turned to the technologies and practices of big data. As one of its first steps, it set up a Hadoop cluster. This is simply a group of inexpensive commodity servers whose activities are coordinated by an emerging software framework called Hadoop (named after a toy elephant in the household of Doug Cutting, one of its developers).

Sears started using the cluster to store incoming data from all its brands and to hold data from existing data warehouses. It then conducted analyses on the cluster directly, avoiding the time-consuming complexities of pulling data from various sources and combining them so that they can be analyzed. This change allowed the company to be much faster and more precise with its promotions.

According to the company's CTO, Phil Shelley, the time needed to generate a comprehensive set of promotions dropped from eight weeks to one, and is still dropping. And these promotions are of higher quality, because they're more timely, more granular, and more personalized. Sears's Hadoop cluster stores and processes several petabytes of data at a fraction of the cost of a comparable standard data warehouse.

These aren't just a few flashy examples. We believe there is a more fundamental transformation of the economy happening. We've become convinced that almost no sphere of business activity will remain untouched by this movement.

Without question, many barriers to success remain. There are too few data scientists to go around. The technologies are new and in some cases exotic. It's too easy to mistake correlation for causation and to find misleading patterns in the data. The cultural challenges are enormous, and, of course, privacy concerns are only going to become more significant. But the underlying trends, both in the technology and in the business payoff, are unmistakable.

The evidence is clear: Data-driven decisions tend to be better decisions. In sector after sector, companies that embrace this fact will pull away from their rivals. We can't say that all the winners will be harnessing big data to transform decision making. But the data tell us that's the surest bet.

This blog post was excerpted from the authors' upcoming article "Big Data: The Management Revolution," which will appear in the October issue of Harvard Business Review.
_____________________

BIG DATA INSIGHT CENTER

·       Use Big Data to Find New Micromarkets

·       Integrate Data Into Products, or Get Left Behind

·       How to Avoid the Big Data "Gotcha's"

·       Big Data, Analytics and the Path from Insights to Value

Monday, September 10, 2012

Tech's New Wave, Driven by Data



Steve Lohr, The New York Times, September 8, 2012

From the article: "TECHNOLOGY tends to cascade into the marketplace in waves. Think of personal computers in the 1980s, the Internet in the 1990s and smartphones in the last five years.

Computing may be on the cusp of another such wave. This one, many researchers and entrepreneurs say, will be based on smarter machines and software that will automate more tasks and help people make better decisions in business, science and government. And the technological building blocks, both hardware and software, are falling into place, stirring optimism.

Michael R. Stonebraker, a pioneer in database research, is one of the optimists. Software used by companies and government agencies — in products sold by Oracle, I.B.M., Microsoft and others — descends from research done in the 1970s by Mr. Stonebraker and Eugene Wong, a colleague at the University of California, Berkeley, as well as a team of scientists at I.B.M.

Today, Mr. Stonebraker sees an opportunity for new kinds of ultrafast databases. The new software, he explains, takes advantage of rapid advances in computer hardware to help businesses and researchers find insights in the rising flood of data coming from so many sources, including Web-browsing trails, sensor data, genetic testing and stock trading.

So, at 68, Mr. Stonebraker is a co-founder and chief technology officer of two start-ups in the field of data-driven discovery, VoltDB and Paradigm4.

“Now is the time,” says Mr. Stonebraker, who is an adjunct professor at the Massachusetts Institute of Technology’s computer science and artificial intelligence laboratory. “The economics and the technology are ripe.”

The case for optimism is by no means unqualified. The march of these technologies raises social issues, including privacy concerns, and the timing is uncertain. All of the bold predictions in the 1990s that the Internet would disrupt traditional industries like media, advertising and retailing did come true — a decade later.

But a series of related technologies, scientists and entrepreneurs say, has reached a critical mass — come to a digital boiling point, so to speak — so that new products and capabilities become possible. The technical ingredients, they note, include powerful, low-cost computing and storage spread across thousands of computers. The digital engine rooms of Google and Amazon are prime examples.

Another fast-improving technology involves inexpensive and intelligent sensors, which are crucial to a new breed of automated machines like experimental driverless cars and battlefield drones. Clever software — notably machine-learning algorithms — animates much of the current wave of smarter technology. Two well-known examples are found in Watson, the “Jeopardy”-winning computer from I.B.M., and the movie recommendations on Netflix.

ADVANCES in such underlying technologies are fueling the current excitement in fields like artificial intelligence, robotics and data analysis and prediction. “All parts of the technology pipeline are gearing up at the same time, and that’s how you get this explosion of new applications and uses,” says Jon Kleinberg, a computer scientist at Cornell University.

Behind the seeming explosion, experts say, is a process of technology evolution. Paul Saffo, a technology forecaster, compares the process to the evolutionary biology concept known as “punctuated equilibria” formulated by the paleontologists Stephen Jay Gould and Niles Eldredge. The idea is that species often evolve in periodic spurts.

Yet, they say, there are typically years of progress before a commercial breakthrough in the technological realm.

“Even in Silicon Valley, it takes most technologies 20 years to become overnight successes,” says Mr. Saffo, a consulting professor at Stanford’s school of engineering.

The Internet provides a case study of both technology’s evolutionary progress and its exponential growth. In 1969, there were only four computers connected to the nascent Internet, compared with roughly a billion computing devices today, from laptops to cellphones, says Edward Lazowska, a computer scientist at the University of Washington.

The early increases in connected computers drew scant attention. “But at some point in the late 1990s,” Mr. Lazowska says, “you were going from 4 million to 8 million to 16 million to 32 million to 64 million, and people started to notice that something revolutionary was going on.”

Rocket Fuel is a four-year-old Silicon Valley start-up that uses artificial-intelligence software to place display advertisements for marketers on the Web. The company can not only tailor ads by demographic slices of viewers’ ages, gender and interests, but can also use its predictive algorithms to produce campaigns based on results, says George H. John, the company’s chief executive.

For example, a luxury carmaker might tell Rocket Fuel that it wants to place 100 million ads in the next month, and it will pay the company, say, $80 for generating a sales lead, as evidenced by a potential customer downloading a brochure or filling out an online form.

Rocket Fuel is growing fast, having nearly doubled its work force since the start of the year, to 240. So far in 2012, it has handled campaigns for more than 500 advertisers, including BMW, Duncan Hines, Allstate, Pizza Hut and Ace Hardware. It has raised $76 million in venture funding and debt, and its thousands of computers handle 19 billion bid requests a day on ad exchanges. Each online auction for ad space is typically completed in about 100 milliseconds, a tenth of a second.

Rocket Fuel, Mr. John says, is using some of the ideas he worked on in the 1990s as a doctoral student focusing on artificial intelligence at Stanford — research that was supported with government dollars from the National Science Foundation and other agencies, as is so often the case. In the last few years, building a business around those ideas has become achievable and affordable. “And a lot of it has to do with the underlying technology,” Mr. John says.

FOR Mr. Stonebraker, the hardware advance that opens the door to his start-ups is the striking improvement of solid-state memory, as performance climbs and prices plunge. Solid-state, or flash, memory is most widely known as the lightweight storage technology used in consumer devices like small music players and smartphones.

But increasingly, solid-state memory can be used in big computers, holding a hefty database in memory instead of sending data off to be stored on disk drives. According to Mr. Stonebraker, some data-handling tasks can now be completed 50 times faster than with conventional systems.

“Memory is the new disk,” he says. “The obvious thing to do is to exploit that technology.”

In the yin and yang of computing, it is software that exploits hardware, enabling a computer to do useful things. And machine-learning programs and other data-sifting software are advancing swiftly.

“There is no point in collecting and storing all this data if the algorithms are not able to find useful patterns and insights in the data,” says Mr. Kleinberg at Cornell. “But the software is scaling up to the task.”

A version of this article appeared in print on September 9, 2012, on page BU4 of the New York edition with the headline: Tech’s New Wave, Driven By Data.

Tuesday, September 4, 2012

Atul Gawande (interview): A marriage of data and caregivers gives Dr. Atul Gawande hope for health care


A marriage of data and caregivers gives Dr. Atul Gawande hope for health care
How transparency, real-time feedback, and lessons from the police can improve health outcomes.


Alex Howard, O'Reilly Radar, August 31, 2012

Dr. Atul Gawande (@Atul_Gawande) has been a bard in the health care world, straddling medicine, academia and the humanities as a practicing surgeon, medical school professor, best-selling author and staff writer at the New Yorker magazine. His long-form narratives and books have helped illuminate complex systems and wicked problems to a broad audience.

One recent feature that continues to resonate for those who wish to apply data to the public good is Gawande’s New Yorker piece “The Hot Spotters,” where Gawande considered whether health data could help lower medical costs by giving the neediest patients better care. That story brings home the challenges of providing health care in a city, from cultural change to gathering data to applying it.

This summer, after meeting Gawande at the 2012 Health DataPalooza, I interviewed him about hot spotting, predictive analytics, networked transparency, health data, feedback loops and the problems that technology won’t solve. Our interview, lightly edited for content and clarity, follows.

Given what you’ve learned in Camden, N.J. — the backdrop for your piece on hot spotting — do you feel hot spotting is an effective way for cities and people involved in public health to proceed?

Gawande: The short answer, I think, is “yes.”

Here we have this major problem of both cost and quality — and we have signs that some of the best places that seem to do the best jobs can be among the least expensive. How you become one of those places is a kind of mystery.

It really parallels what happened in the police world. Here is something that we thought was an impossible problem: crime. Who could possibly lower crime? One of the ways we got a handle on it was by directing policing to the places where there was the most crime. It sounds kind of obvious, but it was not apparent that crime is concentrated and that medical costs are concentrated.

The second thing I knew but hadn’t put two and two together about is that the sickest people get the worst care in the system. People with complex illness just don’t fit into 20-minute office visits.

The work in Camden was emblematic of work happening in pockets all around the country where you prioritize. As soon as you look at the system, you see hundreds, thousands of things that don’t work properly in medicine. But when you prioritize by saying, “For the sickest people — the 5% who account for half of the spending — let’s look at what their $100,000 moments are,” you then understand it’s strengthening primary care and it’s the ability to manage chronic illness.

It’s looking at a few acute high-cost, high-failure areas of care, such as how heart attacks and congestive heart failure are managed in the system; looking at how renal disease patients are cared for; or looking at a few things in the commercial population, like back pain, being a huge source of expense. And then also end-of-life care.

With a few projects, it became more apparent to me that you genuinely could transform the system. You could begin to move people from depending on the most expensive places where they get the least care to places where you actually are helping people achieve goals of care in the most humane and least wasteful ways possible.

The data analytics office in New York City is doing fascinating predictive analytics. That approach could have transformative applications in health care, but it’s notable how careful city officials have been about publishing certain aspects of the data. How do you think about the relative risks and rewards here, including balancing social good with the need to protect people’s personal health data?

Gawande: Privacy concerns can sometimes be a barrier, but I haven’t seen it be the major barrier here. There are privacy concerns in the data about households as well in the police data.

The reason it works well for the police is not just because you have a bunch of data geeks who are poking at the data and finding interesting things. It’s because they’re paired with people who are responsible for responding to crime, and above all, reducing crime. The commanders who have the responsibility have a relationship with the people who have the data. They’re looking at their population saying, “What are we doing to make the system better?”

That’s what’s been missing in health care. We have not married the people who have the data with people who feel responsible for achieving better results at lower costs. When you put those people together, they’re usually within a system, and within a system, there is no privacy barrier to being able to look and say, “Here’s what we can be doing in this health system,” because it’s often that particular.

The beautiful aspect of the work in New York is that it’s not at a terribly abstract level. Yes, they’re abstracting the data, but they’re also helping the police understand: “It’s this block that’s the problem. It’s shifted in the last month into this new sector. The pattern of the crime is that it looks more like we have a problem with domestic violence. Here are a few more patterns that might give you a clue about what you can go in and do.” There’s this give and take about what can be produced and achieved.

That, to me, is the gold in the health care world — the ability to peer in and say: “Here are your most expensive patients and your sickest patients. You didn’t know it, but here, there’s an alcohol and drug addiction issue. These folks are having car accidents and major trauma and turning up in the emergency rooms and then being admitted with $12,000 injuries.”

That’s a system that could be improved and, lo and behold, there’s an intervention here that’s worked before to slot these folks into treatment programs, which by and large, we don’t do at all.

That sense of using the data to help you solve problems requires two things. It requires data geeks and it requires the people in a system who feel responsible, the way that Bill Bratton made commanders feel responsible in the New York police system for the rate of crime. We haven’t had physicians who felt that they were responsible for 10,000 ICU patients and how well they do on everything from the cost to how long they spend in the ICU.

Health data is creating opportunities for more transparency into outcomes, treatments and performance. As a practicing physician, do you welcome the additional scrutiny that such collective intelligence provides, or does it concern you?

Gawande: I think that transparency of our data is crucial. I’m not sure that I’m with the majority of my colleagues on this. The concerns are that the data can be inaccurate, that you can overestimate or underestimate the sickness of the people coming in to see you, and that my patients aren’t like your patients.

That said, I have no idea who gets better results at the kinds of operations I do and who doesn’t. I do know who has high reputations and who has low reputations, but it doesn’t necessarily correspond to the kinds of results they get. As long as we are not willing to open up data to let people see what the results are, we will never actually learn.

The experience of what happens in fields where the data is open is that it’s the practitioners themselves that use it. I’ll give a couple of examples. Mortality for childbirth in hospitals has been available for a century. It’s been public information, and the practitioners in that field have used that data to drive the death rates for infants and mothers down from the biggest killer in people’s lives for women of childbearing age and for newborns into a rarity.

Another field that has been able to do this is cystic fibrosis. They had data for 40 years on the performance of the centers around the country that take care of kids with cystic fibrosis. They shared the data privately. They did not tell centers how the other centers were doing. They just told you where you stood relative to everybody else and they didn’t make that information public. About four or five years ago, they began making that information public. It’s now available on the Internet. You can see the rating of every center in the country for cystic fibrosis.

Several of the centers had said, “We’re going to pull out because this isn’t fair.” Nobody ended up pulling out. They did not lose patients in hoards and go bankrupt unfairly. They were able to see from one another who was doing well and then go visit and learn from one and other.

I can’t tell you how fundamental this is. There needs to be transparency about our costs and transparency about the kinds of results. It’s murky data. It’s full of lots of caveats. And yes, there will be the occasional journalist who will use it incorrectly. People will misinterpret the data. But the broad result, the net result of having it out there, is so much better for everybody involved that it far outweighs the value of closing it up.

U.S. officials are trying to apply health data to improve outcomes, reduce costs and stimulate economic activity. As you look at the successes and failures of these sorts of health data initiatives, what do you think is working and why?

Gawande: I get to watch from the sidelines, and I was lucky to participate in Datapalooza this year. I mostly see that it seems to be following a mode that’s worked in many other fields, which is that there’s a fundamental role for government to be able to make data available.

When you work in complex systems that involve multiple people who have to, in health care, deal with patients at different points in time, no one sees the net result. So, no one has any idea of what the actual experience is for patients. The open data initiative, I think, has innovative people grabbing the data and showing what you can do with it.

Connecting the data to the physical world is where the cool stuff starts to happen. What are the kinds of costs to run the system? How do I get people to the right place at the right time? I think we’re still in primitive days, but we’re only two or three years into starting to make something more than just data on bills available in the system. Even that wasn’t widely available — and it usually was old data and not very relevant to this moment in time.

My concern all along is that data needs to be meaningful to both the patient and the clinician. It needs to be able to connect the abstract world of data to the physical world of what really happens, which means it has to be timely data. A six-month turnaround on data is not great. Part of what has made Wal-Mart powerful, for example, is they took retail operations from checking their inventory once a month to checking it once a week and then once a day and then in real-time, knowing exactly what’s on the shelves and what’s not.

That equivalent is what we’ll have to arrive at if we’re to make our systems work. Timeliness, I think, is one of the under-recognized but fundamentally powerful aspects because we sometimes over prioritize the comprehensiveness of data and then it’s a year old, which doesn’t make it all that useful. Having data that tells you something that happened this week, that’s transformative.

Are you using an iPad at work?

Gawande: I do use the iPad here and there, but it’s not readily part of the way I can manage the clinic. I would have to put in a lot of effort for me to make it actually useful in my clinic.

For example, I need to be able to switch between radiology scans and past records. I predominantly see cancer patients, so they’ll have 40 pages of records that I need to have in front of me, from scans to lab tests to previous notes by other folks.

I haven’t found a better way than paper, honestly. I can flip between screens on my iPad, but it’s too slow and distracting, and it doesn’t let me talk to the patient. It’s fun if I can pull up a screen image of this or that and show it to the patient, but it just isn’t that integrated into practice.

What problems are immune to technological innovation? What will need to be changed by behavior?

Gawande: At some level, we’re trying to define what great care is. Great care means being able to provide optimally knowledgeable care in the right time and the right way for people and not wasting resources.

Some of it’s crucially aided by information technology that connects information to where it needs to be so that good decision-making happens, both by patients and by the clinicians who work with them.

If you’re going to be able to make health care work better, you’ve got to be able to make that system work better for people, more efficiently and less wastefully, less harmfully and with much better teamwork. I think that information technology is a tool in that, but fundamentally you’re talking about making teams that can go from being disconnected cowboys in care to pit crews that actually work together toward solving a problem.

In a football team or a pit crew, technology is really helpful, but it’s only a tiny part of what makes that team great. What makes the team great is that they know what they’re aiming to do, they’re very clear about their goals, and they are able to make sure they execute every basic thing that’s crucial for that success.

What do you worry about in this surge of interest in more data-driven approaches to medicine?

Gawande: I worry the most about a disconnect between the people who have to use the information and technology and tools, and the people who make them. We see this in the consumer world. Fundamentally, there is not a single [health] application that is remotely like my iPod, which is instantly usable. There are a gazillion number of ways in which information would make a huge amount of difference.

That sense of being able to understand the world of the user, the task that’s accomplished and the complexity of what they have to do, and connecting that to the people making the technology — there just aren’t that many lines of marriage. In many of the companies that have some of the dominant systems out there, I don’t see signs that that’s necessarily going to get any better.

If people gain access to better information about the consequences of various choices, will that lead to improved outcomes and quality of life?

Gawande: That’s where the art comes in. There are problems because you lack information, but when you have information like “you shouldn’t drink three cans of Coke a day — you’re going to put on weight,” then having that information is not sufficient for most people.

Understanding what is sufficient to be able to either change the care or change the behaviors that we’re concerned about is the crux of what we’re trying to figure out and discover.

When the information is presented in a really interesting way, people have gradually discovered — for example, having a little ball on your dashboard that tells you when you’re accelerating too fast and burning off extra fuel — how that begins to change the actual behavior of the person in the car.

No amount of presenting the information that you ought to be driving in a more environmentally friendly way ends up changing anything. It turns out that change requires the psychological nuance of presenting the information in a way that provokes the desire to actually do it.

We’re at the very beginning of understanding these things. There’s also the same sorts of issues with clinician behavior — not just information, but how you are able to foster clinicians to actually talk to one another and coordinate when five different people are involved in the care of a patient and they need to get on the same page.

That’s why I’m fascinated by the police work, because you have the data people, but they’re married to commanders who have responsibility and feel responsibility for looking out on their populations and saying, “What do we do to reduce the crime here? Here’s the kind of information that would really help me.” And the data people come back to them and say, “Why don’t you try this? I’ll bet this will help you.”

It’s that give and take that ends up being very powerful.



Wednesday, August 22, 2012

Health IT's Next Big Challenge: Comparative Effectiveness Research


Innovative approach to medical data analysis can yield new treatment options at a lower cost.
 
Paul Cerrato,   InformationWeek August 21, 2012

Healthcare providers are being pushed to deliver more cost effective medical care and to improve the health of not just individual patients but large populations. One key to carrying out both mandates is finding more clinically effective treatment options.

Many academic medical thought leaders insist that the best way to find those treatment protocols is to test them in randomized controlled trials. Such RCTs require a large group of control subjects to receive either a placebo or conventional therapy and a large group to receive the experimental treatment in question. The problem is RCTs are outrageously expensive. In today's cost conscious healthcare system, that's a problem.

Enter comparative effectiveness research. CER compares two or more accepted treatments to determine which are most effective. Medical informatics comes into the picture because it's now possible to get these projects off the ground by analyzing huge patient databases. And much of that patient data can now be gleaned from electronic health record systems.

The American Recovery and Reinvestment Act of 2009 has earmarked $1.1 billion for CER. The Agency for Healthcare Research and Quality (AHRQ), the federal agency tasked with improving the quality, safety, efficiency, and effectiveness of health care, has been using part of that money to fund research on data infrastructure so that clinicians can figure how to take advantage of all the patient data in the Medicare system to compare treatment options. Other AHRQ-sponsored research has been looking at how to create an all-payer, all-claims database that clinicians can tap into for the same purpose.

Other CER-related projects include one led by David J. Magid, MD, director of research at the Colorado Permanente Group. His team searched through thousands of the group's EHRs to figure out which anti-hypertensive drugs are most effective when patients don't respond to first-line treatment with diuretics. The team managed to keep its research costs down to $200,000, a small fraction of what a randomized controlled trial would cost, and still came up with useful results, namely that beta blockers and ACE inhibitors work well.

Similarly a consortium of large healthcare systems, including Kaiser Permanente and Mayo Clinic, is capitalizing on the power of tens of millions of e-records to generate research. For example, they recently launched programs to mine their EHRs to compare treatment protocols for diabetes.

"With these large databases and detailed clinical information, we can conduct comparative, effective research in real world settings, with a full range of patients, not just those selected for clinical trials," Joe V. Selby, director of Kaiser's research division, states in a recent issue of Scientific American.

Boston's Beth Israel Deaconess Medical Center, one of the teaching hospitals affiliated with Harvard Medical School, recently entered the CER arena in a big way. Starting this month, the medical center launched Clinical Query, a searchable patient data repository that lets researchers and clinicians look for potential connections between diseases, treatment options, and risk factors, which in turn can become the jumping off point for a research project.

So if a Harvard researcher wants to compare the benefits of diuretics to ACE inhibitors among patients with hypertension, he can use Clinical Query to look at the records of more than 2 million patients and 200 million data points, including diagnoses, medications taken, lab values, and radiology images.

A comparison of data on the two classes of high blood pressure meds might reveal that one is more effective than the other. And while the results of that CER analysis may not carry the same weight as a randomized clinical trial in which groups of patients were actually given the drugs in real time to see which were more effective, the CER results can still guide clinicians on treatment options for their patients.

Given the fact that comparative effectiveness research will likely cost far less than a randomized clinical trial, it's time healthcare stakeholders take a closer look at this approach. The challenge for IT departments is going to be getting searchable patient data repositories up and running. Few hospitals have the resources to create their own version of Clinical Query. But at the very least, they need to start ramping up their data warehousing and data mining initiatives.

EHR systems are now collecting invaluable information that physicians can use to detect disease patterns, clusters of patients exposed to specific toxins, and groups of patients who respond well to various drug regimens. We can't waste this gold mine.

InformationWeek Healthcare brought together eight top IT execs to discuss BYOD, Meaningful Use, accountable care, and other contentious issues. Also in the new, all-digital CIO Roundtable issue: Why use IT systems to help cut medical costs if physicians ignore the cost of the care they provide? (Free with registration.)

Tuesday, August 21, 2012

Three kinds of big data

Looking ahead at big data's role in enterprise business intelligence, civil engineering, and customer relationship optimization.

Alistair Croll,  O'Reilly Radar,  August 21, 2012 
 
In the past couple of years, marketers and pundits have spent a lot of time labeling everything ”big data.” The reasoning goes something like this:

· Everything is on the Internet.

· The Internet has a lot of data.

· Therefore, everything is big data.

When you have a hammer, everything looks like a nail. When you have a Hadoop deployment, everything looks like big data. And if you’re trying to cloak your company in the mantle of a burgeoning industry, big data will do just fine. But seeing big data everywhere is a sure way to hasten the inevitable fall from the peak of high expectations to the trough of disillusionment.

We saw this with cloud computing. From early idealists saying everything would live in a magical, limitless, free data center to today’s pragmatism about virtualization and infrastructure, we soon took off our rose-colored glasses and put on welding goggles so we could actually build stuff.

So where will big data go to grow up?

Once we get over ourselves and start rolling up our sleeves, I think big data will fall into three major buckets: Enterprise BI, Civil Engineering, and Customer Relationship Optimization. This is where we’ll see most IT spending, most government oversight, and most early adoption in the next few years.

Enterprise BI 2.0

For decades, analysts have relied on business intelligence (BI) products like Hyperion, Microstrategy and Cognos to crunch large amounts of information and generate reports. Data warehouses and BI tools are great at answering the same question — such as “what were Mary’s sales this quarter?” — over and over again. But they’ve been less good at the exploratory, what-if, unpredictable questions that matter for planning and decision making because that kind of fast exploration of unstructured data is traditionally hard to do and therefore expensive.

Most “legacy” BI tools are constrained in two ways:

· First, they’ve been schema-then-capture tools in which the analyst decides what to collect, then later capture that data for analysis.

· Second, they’ve typically focused on reporting what Avinash Kaushik (channeling Donald Rumsfeld) refers to as “known unknowns” — things we know we don’t know, and generate reports for.

These tools are used for reporting and operational purposes, usually focused on controlling costs, executing against an existing plan, and reporting on how things are going.

As my Strata co-chair Edd Dumbill pointed out when I asked for thoughts on this piece:

“The predominant functional application of big data technologies today is in ETL (Extract, Transform, and Load). I’ve heard the figure that it’s about 80% of Hadoop applications. Just the real grunt work of log file or sensor processing before loading into an analytic database like Vertica.”

The availability of cheap, fast computers and storage, as well as open source tools, have made it okay to capture first and ask questions later. That changes how we use data because it makes it okay to speculate beyond the initial question that triggered the collection of data.

What’s more, the speed with which we can get results — sometimes as fast as a human can ask them — makes data easier to explore interactively. This combination of interactivity and speculation takes BI into the realm of “unknown unknowns,” the insights that can produce a competitive advantage or an out-of-the-box differentiator.

We saw this shift in cloud computing: first, big public clouds wooed green-field startups. Then, in a few years, incumbent IT vendors introduced their private cloud offerings. Private clouds included only a fraction of the benefits of public clouds, but were nevertheless a sufficient blend of smoke, mirrors, and features to delay the inevitable move to public resources by a few years and appease the business. For better or worse, that’s where most of IT budgets are being spent today according to IDC, Gartner, and others.

In the next few years, then, look for acquisitions and product introductions — and not a little vaporware — as BI vendors that enterprises trust bring them “big data lite”: enough to satisfy their CEO’s golf buddies, but not so much that their jobs are threatened. This, after all, is how change comes to big organizations.

Ultimately, we’ll see traditional “known unknowns” BI reporting living alongside big-data-powered data import and cleanup, and fast, exploratory data “unknown unknown” interactivity.

Civil Engineering

The second use of big data is in society and government. Already, data mining can be used to predict disease outbreaks, understand traffic patterns, and improve education.

Cities are facing budget crunches, infrastructure problems, and a crowding from rural citizens. Solving these problems is urgent, and cities are perfect labs for big data initiatives. Take a metropolis like New York: hackathons; open feeds of public data; and a population that generates a flood of information as it shops, commutes, gets sick, eats, and just goes about its daily life.



I think municipal data is one of the big three for several reasons: it’s a good tie breaker for partisanship, we have new interfaces everyone can understand, and we finally have a mostly-connected citizenry.

In an era of partisan bickering, hard numbers can settle the debate. So, they’re not just good government; they’re good politics. Expect to see big data applied to social issues, helping us to make funding more effective and scarce government resources more efficient (perhaps to the chagrin of some public servants and lobbyists). As this works in the world’s biggest cities, it’ll spread to smaller ones, to states, and to municipalities.

Making data accessible to citizens is possible, too: Siri and Google Now show the potential for personalized agents; Narrative Science takes complex data and turns it into words the masses can consume easily; Watson and Wolfram Alpha can give smart answers, either through curated reasoning or making smart guesses.

For the first time, we have a connected citizenry armed (for the most part) with smartphones. Nielsen estimated that smartphones would overtake feature phones in 2011, and that concentration is high in urban cores. The App Store is full of apps for bus schedules, commuters, local events, and other tools that can quickly become how governments connect with their citizens and manage their bureaucracies.

The consequence of all this, of course, is more data. Once governments go digital, their interactions with citizens can be easily instrumented and analyzed for waste or efficiency. That’s sure to provoke resistance from those who don’t like the scrutiny or accountability, but it’s a side effect of digitization: every industry that goes digital gets analyzed and optimized, whether it likes it or not.

Customer Relationship Optimization

The final home of applied big data is marketing. More specifically, it’s improving the relationship with consumers so companies can, as Sergio Zyman once said, sell them more stuff, more often, for more money, more efficiently.

The biggest data systems today are focused on web analytics, ad optimization, and the like. Many of today’s most popular architectures were weaned on ads and marketing, and have their ancestry in direct marketing plans. They’re just more focused than the comparatively blunt instruments with which direct marketers used to work.

The number of contact points in a company has multiplied significantly. Where once there was a phone number and a mailing address, today there are web pages, social media accounts, and more. Tracking users across all these channels — and turning every click, like, share, friend, or retweet into the start of a long funnel that leads, inexorably, to revenue is a big challenge. It’s also one that companies like Salesforce understand, with its investments in chat, social media monitoring, co-browsing, and more.

This is what’s lately been referred to as the “360-degree customer view” (though it’s not clear that companies will actually act on customer data if they have it, or whether doing so will become a compliance minefield). Big data is already intricately linked to online marketing, but it will branch out in two ways.

First, it’ll go from online to offline. Near-field-equipped smartphones with ambient check-in are a marketer’s wet dream, and they’re coming to pockets everywhere. It’ll be possible to track queue lengths, store traffic, and more, giving retailers fresh insights into their brick-and-mortar sales. Ultimately, companies will bring the optimization that online retail has enjoyed to an offline world as consumers become trackable.

Second, it’ll go from Wall Street (or maybe that’s Madison Avenue and Middlefield Road) to Main Street. Tools will get easier to use, and while small businesses might not have a BI platform, they’ll have a tablet or a smartphone that they can bring to their places of business. Mobile payment players like Square are already making them reconsider the checkout process. Adding portable customer intelligence to the tool suite of local companies will broaden how we use marketing tools.

Headlong into the trough

That’s my bet for the next three years, given the molasses of market confusion, vendor promises, and unrealistic expectations we’re about to contend with. Will big data change the world? Absolutely. Will it be able to defy the usual cycle of earnest adoption, crushing disappointment, and eventual rebirth all technologies must travel? Certainly not.


O'Reilly Radar (http://s.tt/1lhOc)

Saturday, August 18, 2012

Data and Analytics Key to Health Reform, but Challenges Stand in Way

Kate Ackerman, iHealthBeat, August 15, 2012
 
NATIONAL HARBOR, Md. -- At the eHealth Initiative's National Forum on Data and Analytics in Healthcare last week, stakeholders discussed the importance of data and analytics in implementing health reform, as well as the challenges associated with it.

Jennifer Covich Bordenick, CEO of the eHealth Initiative, said, "Our survey, CIOs and members kept telling us that they are concerned about analytics. They don't feel they have the tools necessary to meet the demands of accountable care and meaningful use." She noted that "93% of the CIOs believe it is very important, but 72% don't feel their organizations have what they need to meet the analytical needs."

By convening experts in health data and analytics, the forum aimed to highlight organizations that are leading the way and facilitate conversations around the need for improvement, she said.

Federal Government Touts Data and Analytics To Support Health Reform

Niall Brennan -- director of the Policy and Data Analysis Group at CMS -- told attendees that health data analytics is "absolutely central" to everything that has to do with health reform.

Brennan -- who stepped in for U.S. Chief Technology Officer Todd Park to give the afternoon's keynote speech -- highlighted the federal government's efforts "to make the data more helpful," while not compromising individuals' privacy.

He said that in the past CMS was "overly conservative" in terms of data release. However, in the last few years -- in large part because of the Affordable Care Act -- the agency has made great strides in liberating health data, he said.

"It might not look like it on the outside, but we are literally constantly pushing the envelope," Brennan said.

He cited the Blue Button Initiative and HealthData.gov as examples of the government's efforts to make privacy-protected health information available to the public.

He also noted that Section 10332 of the ACA authorizes the release of Medicare fee-for-service data to qualified entities if they agree to combine the CMS data with claims data from other sources to compile performance reports.

While Brennan touted the federal government's release of more health care data, he acknowledged, "You can have all the data in the world, but if you don't have the right analytics ... it's just a bunch of useless numbers."

Brennan said the federal government -- from CMS to HHS to the White House -- is "very, very committed" to data and analytics.

Survey Finds Industry Still Has Far To Go

At the forum, Jason Goldwater -- vice president of programs and research at eHI -- offered a sneak peek into the results of a survey eHI conducted with the College of Healthcare Information and Management Executives to get a picture of the types of data and analytics being used in health care.

The survey, which was conducted in July, focused on four areas:

· Types of data used;

· Types of analytic functions used;

· Types of functions needed; and

· Challenges to the use of data and analytics.

When asked what data their organizations actively exchange:

· 76.6% of respondents said lab results;

· 74.5% said demographics;

· 70.2% said discharge summaries;

· 46.8% said allergy information;

· 36.2% said continuity of care documents; and

· 36.2% said problem lists.

The survey found that most health care organizations are focusing their resources on retrospective analysis, with 58.3% citing that as the area in which they direct the majority of their analytical resources. According to the survey, 16.7% of respondents said their organizations direct the majority of their analytical resources toward real-time decision support, 13.9% said optimization and efficiency and 2.8% cited predictive analytics.

When asked what type of analytical functions their organizations primarily use:

· 87.5% of respondents said ad-hoc queries;

· 61.1% said data mining;

· 56.9% said data warehousing;

· 34.7% said exploratory data analysis;

· 30.6% said on-line analytical processing; and

· 23.6% said predictive modeling.

Goldwater said that despite health care organizations' interest in data and analytics, respondents cited several challenges, including:

· Lack of standardized data across systems;

· Lack of a system infrastructure to support analytics;

· Cost of analytical software;

· Concerns about privacy and security of the data; and

· Limited utility of the results to the organizations.

Goldwater said that eHI, CHIME and McKesson will host a webinar on Aug. 30 to discuss the survey results in more detail and that eHI will release an issue brief in the fall.

Speakers Highlight Challenges Associated With Health Data and Analytics

Several speakers offered real-life examples of the challenges cited by respondents to the eHI/CHIME survey.

Jason Williams -- vice president of business analytics at RelayHealth -- said that health care cost and quality reforms require inter-stakeholder transparency and that there needs to be "more emphasis" in that area. He also cited the demand for business and technology analysts and the need to find a balanced approach to privacy as areas for improvement.

Micky Tripathi -- president and CEO of the Massachusetts eHealth Collaborative and chair of eHI's Board of Directors -- said that variations in EHR systems can be problematic for analytics.

Brendan Mullen -- senior director of PINNACLE, the American College of Cardiology's outpatient registry -- noted that its system integration tool has to look at 41 locations in the NextGen EHR system to determine if a physician provided patient education on heart failure.

Tripathi said that as vendors are working to address the issues, new measures are coming down the pipeline.

For the Healthcare Information and Management Systems Society Conference in February, MAeHC compared the results of its certified Quality Data Center with the Office of the National Coordinator for Health IT-sponsored, open-source popHealth tool to evaluate meaningful use quality measures. The tools used the same exact data, but they did not produce the same results for any of the 44 measures, Tripathi said.

He explained that further investigation found several reasons for the discrepancies, including the definition of the continuity of care document, coding and mapping, and the interpretation of certain measures, such as age.

Tripathi said, "We have a ton of work to do," adding that the industry needs to keep pushing along.

Covich Bordenick told iHealthBeat, "It was clear by the end of that day that this was just the start of a conversation that is going to take years to explore," adding, "eHealth Initiative is going to help unpack this issue."

MORE ON THE WEB

· National Forum on Data and Analytics in Healthcare

· HealthData.gov

· "Data Analysis and the Future of Health Care" (Tibken, Wall Street Journal, 4/16
).