Thursday, September 27, 2012

New Talk by Clay Shirky -


Clay Shirky’s Ted Talk “How the Internet will (one day) transform government”

The open-source world has learned to deal with a flood of new, oftentimes divergent, ideas using hosting services like GitHub -- so why can’t governments? In this rousing talk Clay Shirky shows how democracies can take a lesson from the Internet, to be not just transparent but also to draw on the knowledge of all their citizens.

Clay Shirky argues that the history of the modern world could be rendered as the history of ways of arguing, where changes in media change what sort of arguments are possible -- with deep social and political implications

Wednesday, September 26, 2012

How "Big-Data-as-a-Service" Can Help Smaller Companies Compete


Brian Proffitt, Read Write Web, September 26, 2012

The common perception of how big data is used centers around giant multi-national enterprises spending millions trying to fine tune their business strategies to eke out every last penny from their customers. But in reality, big data is worming its way into businesses large and small, often as a service instead of on-premises software.


The Evolution Of The Comment Card
Visit a bustling diner in a small town and you may see them tucked in among the bottles of ketchup and sugar packets on the linoleum counter: 3 X 5 comment cards. How was your meal? How was your service? Fill it out and tuck it in the little wooden box by the cash register, please.

Regulars might make their concerns known directly to the staff, but diners such as this - like nearly every other business in the world - need to attract and keep new customers in order to grow. That's the point behind comment cards: get as much feedback as you can so you to improve what needs fixing and keep doing what's working well.

Moving the comment card into the 21st Century is essentially what big data is all about.

The most common form of big data in busines today was created as a technological response to tracking all of the data that was generated by commercial websites. Once marketers and other business execs saw that they could monitor an online customer's responses all the way down to the mouse click, software engineers started figuring out a way to keep that data and mine it for ever more useful information.

The combination of speed and volume needed to catch all of this information is what makes big data tools and technology really necessary. But for the vast majority of businesses that do not have big ecommerce sites, is big data even worth the attempt?

It turns out, yes.

Big Data Tools Work For Small Data, Too
One of the easier ways smaller businesses are taking on big data is seeing the general value of data analytics no matter how big or small the data set is. That message is practically a non-brainer: business owners are scanning the headlines every day and getting excited about applying data to their decision-making process.
Social media is one quick way to implement big data within a smaller business. Used and analyzed properly, the information from social media can give any business instant feedback that's far more robust and immediate than those old-fashioned comment cards.

"Every time we perform a search, tweet, send an email, post a blog, comment on one, use a cell phone, shop online, update our profile on a social networking site, use a credit card, or even go to the gym, we leave behind a mountain of data, a digital footprint, that provides a treasure trove of information about our lifestyles, financial activities, health habits, social interactions, and much more," wrote former Tivoli CEO Frank Moss in his 2011 book The Sorcerers and Their Apprentices.

Big-Data-as-a-Service (BDaaS)
Using big data can go beyond mining social media as a juiced-up form of the comment card. Big data can also be integrated with existing business practices to improve and expand day-to-day operations.

It's becoming pretty well-known that big data and fast data analysis are being used by large hotels and chains to improve their yield management processes. This kind of infrastructure is typically beyond the reach of smaller hotels, inns and bed and breakfasts. It's probably overkill anyway: while a big hotel near Orlando, Fla., area might see 75 room pricing changes per day, a small independent lodging in a less-volatile market might only see a couple of price moves daily.

But that doesn't make the need for a smaller inn to adjust to local market changes any less important, says Erik Hovanec, CEO of LeisureLink, which specializes in providing yield-management as a service to smaller hospitality locations.

One of LeisureLink's clients is a 120-room property outside Myrtle Beach, SC, that "is great at hospitality, not necessarily at IT." Hovanec described. Using his company's service, the property is able to tap into information about local Myrtle Beach hotels and see real-time pricing information on other properties and make adjustments accordingly.

This is exactly what the larger hotels and chains have been doing for a while. But now this service is available to smaller, mid-market establishments. Typically, a large hotel or chain might invest $30-$40 million just to increase their yield management from 90% to 95%. Hovanec boasts that since LeisureLink's service is often the first real step into automated yield management for smaller hotels, their efficiency in yield management can rocket from 20% to 80%.

Can Big Data Work For Every Business?
At some point, any company considering a big-data approach needs to consider the one basic question: is there information out there that will help improve the business? If there's a yes in there, then a search for a big data solution might be worth the effort.

It's not that we really need another another "as-a-Service" acronym, but thanks to the use of Internet-based Big-Data-as-a-Service (BDaaS), you don't have to be a giant enteprise to play any more. These days, there's a good chance that someone out there will have the information you need or can help you find it.

Thursday, September 20, 2012

Big Data for All


Omer Tene, Concurring Opinions, September 20, 2012

Much has been written over the past couple of years about “big data” (See, for example, here and here and here). In a new article, Big Data for All: Privacy and User Control in the Age of Analytics, which will be published in the Northwestern Journal of Technology and Intellectual Property, Jules Polonetsky and I try to reconcile the inherent tension between big data business models and individual privacy rights. We argue that going forward, organizations should provide individuals with practical, easy to use access to their information, so they can become active participants in the data economy. In addition, organizations should be required to be transparent about the decisional criteria underlying their data processing activities.

The term “big data” refers to advances in data mining and the massive increase in computing power and data storage capacity, which have expanded by orders of magnitude the scope of information available for organizations. Data are now available for analysis in raw form, escaping the confines of structured databases and enhancing researchers’ abilities to identify correlations and conceive of new, unanticipated uses for existing information. In addition, the increasing number of people, devices, and sensors that are now connected by digital networks has revolutionized the ability to generate, communicate, share, and access data.

Data creates enormous value for the world economy, driving innovation, productivity, efficiency and growth. In the article, we flesh out some compelling use cases for big data analysis. Consider, for example, a group of medical researchers who were able to parse out a harmful side effect of a combination of medications, which were used daily by millions of Americans, by analyzing massive amounts of online search queries. Or scientists who analyze mobile phone communications to better understand the needs of people who live in settlements or slums in developing countries.

At the same time, the “data deluge” presents formidable privacy concerns. Protecting privacy become harder as information is multiplied and shared ever more widely among multiple parties around the world. As more information regarding individuals’ health, financials, location, electricity use and online activity percolates, concerns arise about profiling, tracking, discrimination, exclusion, government surveillance and loss of control. From a more technical legal angle, big data challenges some of the most fundamental concepts of privacy law, including the definition of “personally identifiable information”, the role of individual control, and the principles of data minimization and purpose limitation.

In our article, we make the case for providing individuals with usable access to their data. The call for transparency is not new, of course. Rather the emphasis is on access to data in usable format, which can work to create value to individuals. Transparency and access alone have not emerged as potent tools because individuals do not care for, and cannot afford to indulge in transparency and access for their own sake (see one oft-cited counterexample here). The enabler of transparency and access is the ability to use the information and benefit from it in a tangible way. This will be achieved through “featurization” or “app-ification” of privacy. Organizations should build as many dials and levers as needed for individuals to engage with their data.

We expect that “featurization” of big data, harnessing its immense force for not only organizational but also individual benefit, will unleash a wave of innovation and create a market for personal data applications. The technological groundwork has already been completed with mash-ups and real-time APIs making it easier for organizations to combine information from different sources and services into a single user experience. Regardless of lingering questions concerning who – if anyone – “owns” the information, we think that fairness dictates that individuals enjoy beneficial use of the data about them.

Our second proposal would require organizations to disclose the decisional criteria underpinning their data analytics machinery. In a big data world, it is often not the data but rather the inferences drawn from them that give cause for concern. Inaccurate, manipulative or discriminatory conclusions may be drawn from perfectly innocuous, accurate data. Much like in quantum physics, the observer in big data analysis can affect the results of her research by defining the data set, proposing a hypothesis or writing an algorithm. At the end of the day, big data analysis is an interpretative process, in which one’s identity and perspective informs one’s results. Like any interpretative process, it is subject to error, inaccuracy and bias. Louis Brandeis, who together with Samuel Warren “invented” the legal right to privacy in 1890, has also written that “[s]unlight is said to be the best of disinfectants”. We trust if the existence and uses of databases were visible to the public, organizations would be more likely to avoid unethical or socially unacceptable uses of data.

Wednesday, September 19, 2012

When data disrupts health care


The convergence of data, privacy and cost have created a unique opportunity to reshape health care.

Mac Slocum, O'Reilly Radar, September 18, 2012

Health care appears immune to disruption. It’s a space where the stakes are high, the incumbents are entrenched, and lessons from other industries don’t always apply.

Yet, in a recent conversation between Tim O’Reilly and Roger Magoulas it became evident that we’re approaching an unparalleled opportunity for health care change. O’Reilly and Magoulas explained how the convergence of data access, changing perspectives on privacy, and the enormous expense of care are pushing the health space toward disruption.

As always, the primary catalyst is money. The United States is facing what Magoulas called an “existential crisis in health care costs” [discussed at the 3:43 mark]. Everyone can see that the current model is unsustainable. It simply doesn’t scale. And that means we’ve arrived at a place where party lines are irrelevant and tough solutions are the only options.

“Who is it that said change happens when the pain of not changing is greater than the pain of changing?” O’Reilly asked. “We’re now reaching that point.” [3:55]

(Note: The source of that quote is hard to pin down, but the sentiment certainly applies.)

This willingness to change is shifting perspectives on health data. Some patients are making their personal data available so they and others can benefit. Magoulas noted that even health companies, which have long guarded their data, are warming to collaboration.

At the same time there’s a growing understanding that health data must be contextualized. Simply having genomic information and patient histories isn’t good enough. True insight — the kind that can improve quality of life — is only possible when datasets are combined.

“Genes aren’t destiny,” Magoulas said. “It’s how they interact with other things. I think people are starting to see that. It’s the same with the EHR [Electronic Health Record]. The EHR doesn’t solve anything. It’s part of a puzzle.” [4:13]

And here’s where the opportunity lies. Extracting meaning from datasets is a process data scientists and Silicon Valley entrepreneurs have already refined. That means the same skills that improve mindless ad-click rates can now be applied to something profound.

“There’s this huge opportunity for those people with those talents, with that experience, to come and start working on stuff that really matters,” O’Reilly said. “They can save lives and they can save money in one of the biggest and most critical industries of the future.” [5:20]

The language O’Reilly and Magoulas used throughout their conversation was telling. “Save lives,” “work on stuff that matters,” “huge opportunity” — these aren’t frivolous phrases. The health care disruption they discussed will touch everyone, which is why it’s imperative the best minds come together to shape these changes.

The full conversation between O’Reilly and Magoulas is available in the following video.

Here are key points with direct links to those segments:

· Internet companies used data to solve John Wanamaker’s advertising dilemma (“Half the money I spend on advertising is wasted; the trouble is I don’t know which half”). Similar methods can apply to health care. [17 seconds in]

· The “quasi-market system” of health care makes it harder to disrupt than other industries. [3:15]

· The U.S. is facing an existential crisis around health care costs. “This is bigger than one company.” [3:43]

· We can benefit from the multiple data types coming “on stream” at the same time. These include electronic medical records, inexpensive gene sequencing, and personal sensor data. [4:28]

· The availability of different datasets presents an opportunity for Silicon Valley because data scientists and technologists already have the skills to manage the data. Important results can be found when this data is correlated: “The great thing is we know it can work.” [5:20]

· Personal data donation is a trend to watch. [6:40]

· Disruption is often associated with trivial additions to the consumer Internet. With an undisrupted market like health care, technical skills can create real change. [7:04]

· “There’s no question this is going to be a huge field.” [8:15]

If the disruption of health care and associated opportunities interests you, O’Reilly has more to offer. Check out our interviews, ongoing coverage, our recent report, “Solving the Wanamaker problem for health care,” and the upcoming Strata Rx conference in San Francisco.


  O'Reilly Radar (http://s.tt/1nGAn)

Thursday, September 13, 2012

Healthcare: Cyber wards

April Dembosky, The Financial Times,  September 12, 2012

Silicon Valley is hoping technology can reform the labyrinthine US medical system

Patients sit quietly in a waiting room at a San Francisco hospital, their expressions concentrated as they answer questions about their medical history. Rather than scribbling on clipboards, they are tapping iPads.

They press icons for male or female – the unmistakable lavatory illustrations of a blue straight-legged man and skirted woman in pink. The zero to 10 rating scale for mental health has a smiley green face at 10, and a sad red face at zero.

Various other visual cues make filling in forms more engaging but, more importantly, extract more accurate and complete information from patients. The electronically culled answers are zipped directly to a central hospital database with all of the patient’s medical records, saving secretaries from making mistakes by having to decipher countless styles of handwriting.

“Just about any problem you can see in healthcare is trying to be addressed by a tech company,” says Sterling Lanier, chief executive of Tonic Health, a Silicon Valley start-up that makes the iPad software.

Mr Lanier is one of many Valley entrepreneurs turning away from the world of social networks, mobile games and digital advertising to develop tools that address the entrenched problems in the lumbering healthcare system.

And it is not just a matter of iPads intended to scythe through bureaucracy. Patients are increasingly measuring their blood pressure and glucose levels from home with remote monitoring devices and iPhone apps. Some computer games now foster healthy behaviour. Enormous databases build up huge case histories to reduce medical mistakes and costs.

Entrepreneurs say their technology could smooth revolutionary reforms of medical care in the US, which spends $2.6tn a year on health, or 17 per cent of gross domestic product. As policy changes roll out over the next few years, insurance companies will be forced to limit their profits, and hospitals will face penalties if patients return to the hospital within 30 days of being discharged. Doctors will no longer be paid for how many X-rays they take or laboratory tests they run but for how well their patients are doing.

However, while the entrepreneurs exude optimism about their ability to streamline the healthcare system, the sprawling industry proved resistant of reforms in the 1990s. It was difficult to translate the vision of a few bright technology experts to the massive healthcare administration sector.

Fears about the proposed technology revolution resonate in several other countries that have hit roadblocks when turning to technology to address healthcare problems. Doctors and other medical professionals around the world have historically been slow to adopt new technology, wary of the costs and the time needed to learn and adjust to new administrative procedures.

Cancer – a case for Doctor Watson

Now scientists at IBM have proved what the Watson supercomputer can do on a hit television game show, they want to see if it can treat cancer.

Since Watson’s victory on Jeopardy! last year, doctors at Memorial Sloan-Kettering Cancer Centre in New York have been training the computer on cancer treatment protocols, hoping it will eventually become a much faster and more reliable diagnostician than human doctors, and in turn reduce costs and mistakes in one of the most expensive sectors of healthcare ...

To read the rest of this article, click here

In England, the National Health Service faltered in its efforts to overhaul its medical records system, switching from paper to electronic records, in part because of its inability to persuade doctors to make changes. The British government wanted to centralise technology purchases from a handful of vendors and rope all the hospitals in one region into the same system. Doctors had to change their habits and did not have a choice about which technology would work for them.

“There was no incentive, and it was a huge disruption to their workflow,” said Harry Greenspun, a doctor and senior healthcare technology adviser at Deloitte.

In the US, where the government has also called for a transition to electronic medical records, only 25 per cent of physicians are “on target” to meet the deadline for the new federal standards, according to a study by the California HealthCare Foundation.

For entrepreneurs building technology tools that go beyond government mandates, such as online appointment registries or mobile phone apps for patients to contact doctors, there are even more obstacles.

Health is a heavily regulated industry in the US and a slow-moving one, particularly when compared with the rocket speed of social media and ecommerce developments. Investors do not want to be stuck in a multiyear approval process at the Food and Drug Administration or tied up in ever-evolving government policies.

While funding for information technology in the health sector is growing fast, according to the National Venture Capital Association, it is still relatively meagre compared with the billions flowing into the next would-be Twitters and Facebooks. Venture capital funding in healthcare IT hit $860m last year, increasing from only $340m in 2002. The biggest single jump came between 2010 and 2011 when funding soared from $499m.

Several entrepreneurs, such as Mr Lanier at Tonic Health, still see a lingering fear in the market and have relied solely on angel investors or friends and family for early financing.

“A lot of venture capital firms may be scared by the category or had a negative experience in the past, namely because of the long sale cycles, or because they could never get the technology to fully work,” he says. “We encountered some hesitation from VCs who said: ‘I love this but the category makes me gun-shy.’ ”

. . .

When Khan Wong got a strange rash on his calf last summer, he didn’t worry. When it swelled and turned purple in the following days, he did. His doctor was baffled and unable to make a diagnosis. Mr Wong chronicled the development of his rash as it got bigger and more purple in a series of photos he took with his iPhone and posted on his Facebook page. Among the series of horrified comments from his friends, a few offered their own diagnoses: cellulitis, spider bite, allergic reaction. In the end, Mr Wong saw three different doctors who proffered three different theories. A few days of ointment and a steroid shot and Mr Wong was fine, but there was never an exact diagnosis.

Technology is already becoming a key feature of healthcare as more and more patients and doctors find medical uses for their everyday consumer technologies. Doctors, too, take pictures of strange rashes to get second opinions from their colleagues. Medical students across the country are given iPads on their first day of learning, loaded with all their textbooks and a list of recommended apps to help with studying.

“It is impossible to go into an operating room anywhere in the US today and not see an iPad,” says Chuck Farkas, a senior partner at Bain & Company’s healthcare practice. And there would be strong support among doctors and nurses for the rumoured 7-inch iPad mini: “the biggest call Apple has had from healthcare professionals is for an iPad that fits in the pocket of a white coat”, he says.

It is the Silicon Valley-style, easy-to-use, consumer-friendly nature of these devices and mobile applications that US government officials are trying to introduce into healthcare. For the past few years, the US Department of Health and Human Services has been courting California engineers to make finding a doctor’s appointment or a good price on a prescription drug as easy as online banking or shopping for an airline reservation.

“We were drawn as a magnet to Silicon Valley, as they’ve solved so many other problems in other sectors, like financial services and travel,” says Greg Downing, executive director for innovation at HHS.

His agency is offering data from other agencies – for example, disease occurrence and death rates from the Centers for Disease Control and Prevention, and clinic and doctor locations from the Health Resources and Services Administration – to engineers to weave into websites and mobile apps.

“We call it the ‘liberation of data’,” he adds.

One result is a start-up called Castlight Health, which makes price comparison software for common medical procedures such as mammograms and colonoscopies. It functions like a travel website, aggregating data from employers’ cache of past insurance claims, and packages it into a searchable price database for consumers.

The goal is to bring transparencyto a field that has historically been opaque, permitting vast discrepancies in the cost for common procedures. With 30m-40m joining insurance rolls in the next few years as a result of healthcare reform, many are expected to be funnelled into “high-deductible health plans” and will pay much closer attention than before to the costs of services their doctors prescribe.

“Healthcare has become a very inefficient market because there’s no correlation between price and quality,” says Peter Isaacson, Castlight Health’s chief marketing officer. “You can have the highest prices for the lowest quality.”

Healthcare technology is being developed in various pockets of the country, in the Midwest and the north-east, where the industry has a strong presence and technology experts frequently interact with insurance and hospital executives. The Silicon Valley renegade approach is sometimes interpreted as hubristic in these circles, but Valley insiders say their way of working has big advantages.

“There’s an unfettered approach to problem-solving here,” says Mark Goines, a partner at Morganthaler Ventures. “The inspiration comes from not being impaired by the way it’s always been done in healthcare.”

Despite the obstacles and past failures, various experts argue that things will work out for Silicon Valley this time. Current government policies and incentives support experts to develop new technology and encourage healthcare providers to buy it, Mr Downing says.

And consumer electronics, particularly smartphones,are reaching record levels of penetration: 81 per cent among doctors and 48 per cent among the general population, according to data from Manhattan Research and Nielsen.

If the Valley has its way, patients will routinely be making appointments with a few taps on their smartphones, exchanging text messages with their doctors and even receiving treatment recommendations based on computer algorithms.

“We really are at a tipping point,” Mr Farkas says.

Copyright The Financial Times Limited 2012. You may share using our article tools.
Please don't cut articles from FT.com and redistribute by email or post to the web.

Tuesday, September 11, 2012

Erik Brynjolfsson and Andrew McAfee : Big Data's Management Revolution

Big Data's Management Revolution

by Erik Brynjolfsson and Andrew McAfee  |  10:05 AM September 11, 2012

http://blogs.hbr.org/cs/2012/09/big_datas_management_revolutio.html

Big data has the potential to revolutionize management. Simply put, because of big data, managers can measure, and hence know, radically more about their businesses, and directly translate that knowledge into improved decision making and performance. Of course, companies such as Google and Amazon are already doing this. After all, we expect companies that were born digital to accomplish things that business executives could only dream of a generation ago. But in fact the use of big data has the potential to transform traditional businesses as well.

We've seen big data used in supply chain management to understand why a carmaker's defect rates in the field suddenly increased, in customer service to continually scan and intervene in the health care practices of millions of people, in planning and forecasting to better anticipate online sales on the basis of a data set of product characteristics, and so on.

Here's how two companies, both far from Silicon Valley upstarts, used new flows of information to radically improve performance.

Case #1: Using Big Data to Improve Predictions
Minutes matter in airports. So does accurate information about flight arrival times: If a plane lands before the ground staff is ready for it, the passengers and crew are effectively trapped, and if it shows up later than expected, the staff sits idle, driving up costs. So when a major U.S. airline learned from an internal study that about 10% of the flights into its major hub had at least a 10-minute gap between the estimated time of arrival and the actual arrival time — and 30% had a gap of at least five minutes — it decided to take action.

At the time, the airline was relying on the aviation industry's long-standing practice of using the ETAs provided by pilots. The pilots made these estimates during their final approach to the airport, when they had many other demands on their time and attention. In search of a better solution, the airline turned to PASSUR Aerospace, a provider of decision-support technologies for the aviation industry.

In 2001 PASSUR began offering its own arrival estimates as a service called RightETA. It calculated these times by combining publicly available data about weather, flight schedules, and other factors with proprietary data the company itself collected, including feeds from a network of passive radar stations it had installed near airports to gather data about every plane in the local sky.

PASSUR started with just a few of these installations, but by 2012 it had more than 155. Every 4.6 seconds it collects a wide range of information about every plane that it "sees." This yields a huge and constant flood of digital data. What's more, the company keeps all the data it has gathered over time, so it has an immense body of multidimensional information spanning more than a decade. RightETA essentially works by asking itself "What happened all the previous times a plane approached this airport under these conditions? When did it actually land?"

After switching to RightETA, the airline virtually eliminated gaps between estimated and actual arrival times. PASSUR believes that enabling an airline to know when its planes are going to land and plan accordingly is worth several million dollars a year at each airport. It's a simple formula: Using big data leads to better predictions, and better predictions yield better decisions.

Case #2: Using Big Data to Drive Sales
A couple of years ago, Sears Holdings came to the conclusion that it needed to generate greater value from the huge amounts of customer, product, and promotion data it collected from its Sears, Craftsman, and Lands' End brands. Obviously, it would be valuable to combine and make use of all these data to tailor promotions and other offerings to customers, and to personalize the offers to take advantage of local conditions.

Valuable, but difficult: Sears required about eight weeks to generate personalized promotions, at which point many of them were no longer optimal for the company. It took so long mainly because the data required for these large-scale analyses were both voluminous and highly fragmented — housed in many databases and "data warehouses" maintained by the various brands.

In search of a faster, cheaper way, Sears Holdings turned to the technologies and practices of big data. As one of its first steps, it set up a Hadoop cluster. This is simply a group of inexpensive commodity servers whose activities are coordinated by an emerging software framework called Hadoop (named after a toy elephant in the household of Doug Cutting, one of its developers).

Sears started using the cluster to store incoming data from all its brands and to hold data from existing data warehouses. It then conducted analyses on the cluster directly, avoiding the time-consuming complexities of pulling data from various sources and combining them so that they can be analyzed. This change allowed the company to be much faster and more precise with its promotions.

According to the company's CTO, Phil Shelley, the time needed to generate a comprehensive set of promotions dropped from eight weeks to one, and is still dropping. And these promotions are of higher quality, because they're more timely, more granular, and more personalized. Sears's Hadoop cluster stores and processes several petabytes of data at a fraction of the cost of a comparable standard data warehouse.

These aren't just a few flashy examples. We believe there is a more fundamental transformation of the economy happening. We've become convinced that almost no sphere of business activity will remain untouched by this movement.

Without question, many barriers to success remain. There are too few data scientists to go around. The technologies are new and in some cases exotic. It's too easy to mistake correlation for causation and to find misleading patterns in the data. The cultural challenges are enormous, and, of course, privacy concerns are only going to become more significant. But the underlying trends, both in the technology and in the business payoff, are unmistakable.

The evidence is clear: Data-driven decisions tend to be better decisions. In sector after sector, companies that embrace this fact will pull away from their rivals. We can't say that all the winners will be harnessing big data to transform decision making. But the data tell us that's the surest bet.

This blog post was excerpted from the authors' upcoming article "Big Data: The Management Revolution," which will appear in the October issue of Harvard Business Review.
_____________________

BIG DATA INSIGHT CENTER

·       Use Big Data to Find New Micromarkets

·       Integrate Data Into Products, or Get Left Behind

·       How to Avoid the Big Data "Gotcha's"

·       Big Data, Analytics and the Path from Insights to Value

Here Comes the Data Economy


New companies are creating services using government data on health care, education, and more.

Alexander B. Howard, Slate, September 10, 2012


We're living in the exabyte age, where the actions of billions of humans using the Web and their mobile devices are creating massive amounts of big data to collect, store, analyze, and put to work.

If big data is a strategic resource, as has been suggested, then many national and state governments have public reserves that can be tapped for the public good in this young century's version of the industrial revolution. Given that the United States economy is still coming out of the worst recession and financial shock since the Great Depression, supporting civic and tech entrepreneurs enjoys political support from both sides of the aisle.

Entrepreneurs, big and small, are mashing up data from the rapidly expanding collection of sources and building new businesses on it or improve their existing services, like Zillow or Google Maps or Consumer Reports or Bloomberg Government. In a time when job creation is critical, using public sector information to create jobs isn’t an aim to dismiss lightly, although the terms and conditions under which such activity occurs must be clear to all actors involved, to avoid the creation of new monopolies based upon artificial scarcity.

My publisher, long-time open source and open government advocate Tim O'Reilly, has asked how government can act as a platform to enable people inside and outside government to innovate on top of it. One answer is certainly releasing open data. In that context, open data and application programming interfaces, more commonly known as APIs, increasingly look like fundamental infrastructure for digital government in the 21st century.

There's good reason to think that open data could have an overall effect on the economy akin to open source and small business. Gartner, the IT research analysis firm, recently highlighted how open data creates value in the public and private sector.

You may not realize it, but services you use on a daily basis have been built upon data released by the government. Weather data collected by the National Oceanic and Atmospheric Association has an annual estimated economic value of $10 billion, according to U.S. Chief Information Officer Steven VanRoekel and U.S. Chief Technology Officer Todd Park. NOAA data sets are used by Weather.com, Weather Underground, and the Weather Channel—and the nation's farmers consult these forecasts to manage both their crops and the risks of loss. VanRoekel and Park estimate the annual economic value of the data from the U.S. global positioning system at some $90 billion. From companies like TomTom or Garmin to dashboard GPS systems to smartphones and associated location-based applications, GPS data sets are baked into an expanding number of services and products.

Now, as Park seeks to scale open data across the federal government, we’re on the verge of the next generation of services driven by open data, which will involve everything from energy to health care to consumer finance to transit sectors. The challenge is that the cities and federal agencies that hold vast amounts of data may not always understand the value of the information they hold or how to create or sustain businesses using it. That's where open innovation in the public sector and the dynamism of entrepreneurs will play an important role in making the people's data more useful to the people.

BrightScope is a notable example of what dogged persistence can create. The California startup made a profitable business using government data to help the American people understand the fees associated with their 401(k)s. Last May, BrightScope went further, launching financial adviser pages based on open government data from the Securities and Exchange Commission and the Financial Industry Regulatory Authority, the largest independent securities regulator in the United States. Previously, financial adviser profiles could only be found through exact queries at an obscure URL on the regulators' websites. Now, information that citizens care about—the records of financial advisers in their geographic region—is available where they're looking for it: in search engine results.

Just as labor and regulatory data fuels BrightScope's business, there's an expanding number of startups that are tapping into other data released so-called “smart disclosure” initiatives. Smart disclosure is when a private company or government agency provides a person with periodic access to his or her own data in open formats that enable them to easily put the information to use. Startups like Billshrink.com and Hello Wallet are already using a combination of private sector and public sector data to enhance consumer finance decisions. The success of such consumer finance startups suggests an important lesson: The most successful apps and services will combine government, industry, and user-generated data.

The key open data story to watch in the federal government, however, centers on health care. McKinsey and Associates estimates the annual economic value of big, open liquid health data at about $350 billion annually. The explosion of mHealth apps are just the beginning of the disruption in health care from open health data. The effort to revolutionize the health care industry by making health data as useful as weather data is still in its infancy—but the early results are promising. iTriage, which was acquired by Aetna, is enabling people to make better mobile health care decisions where and when they need to do so. It uses a combination of government and private sector data to evaluable symptoms or conditions and point users to nearby medical care. Another startup, Castlight, is analyzing health care data to empower patients, acting like Kayak.com for those who want more transparency about costs. In May, Castlight completed a $100 million round of financing.

But for these sorts of initiatives to take off, entrepreneurs and regulators will have to work together to get contextual consent right and inform patients about the reuse of their data. Transparency is crucial to building a health data commons and thriving startup ecosytem based upon it.

If that balance can be struck, there's considerable potential for entrepreneurs to create better civic interfaces for many digital services. If open government data have helped build new tools, open data disclosed by private companies could create even more value for citizens. But currently, few businesses release anonymized data in an open, usable format. It will soon be time for the government to step in, convene stakeholders, and answer some key questions: How can we create uniform standards that will allow entrepreneurs and developers to innovate? When should data be licensed? Most of the big data releases we have seen come from finance, with bank records or stock trades. But there are significant opportunities to help both entrepreneurs and empowered consumers in health care, energy, education, and telecommunications, to name just a few.

Just as the glowing blue dot on the maps in our smartphone screens revolutionized how we navigate the world, similar "blue dots" could emerge for health care, finance, energy, and any product or service that is regulated or cataloged by government and industry. First, however, they'll need to open the data.

Also in the Future Tense package on government and open data: why Yelp and the government should share data; what a burger mob tells us about the future of democracy; and how Mexico is using open data to move beyond its authoritarian past.


Monday, September 10, 2012

Artificial Intelligence, Powered by Many Humans



Crowdsourcing can create an artificial chat partner that's smarter than Siri-style personal assistants.

Tom Simonite, Technology Review, September 10, 2012

Personal assistants such as Apple's Siri may be useful, but they are still far from matching the smarts and conversational skills of a real person. Researchers at the University of Rochester have demonstrated a new, potentially better approach that creates a smart artificial chat partner from fleeting contributions from many crowdsourced workers.


Crowdsourcing typically involves posting simple tasks to a website such as Amazon Mechanical Turk, where Web users complete them for a reward of a few cents. The tasks are often simple, repetitive jobs that are easy for humans but tough for computers, such as categorizing images. Crowdsourcing has become a popular way for companies to handle such tasks, but some researchers, including the group at Rochester, believe it can also be used to take on more complex tasks.

When people talk to the new crowd-powered chat system, called Chorus, using an instant messaging window, they get an experience practically indistinguishable from chatting with a single real person. Yet behind the scenes, each response is the result of tens of people paid a few cents to perform small tasks: including suggesting possible replies and voting for the best suggestions submitted by other workers.

Tests where Chorus was asked for travel advice showed that it could be smarter than any one individual in the crowd, because around seven people were contributing to its responses at any one time. Helpers built this way might also be cheaper than paying a conventional one-on-one assistant. "It shows how a crowd-powered system that is relatively simple can do something that AI has struggled to do for decades," says Jeffrey Bigham, an assistant professor at the University of Rochester, and a member of the research team that created Chorus. Bigham jokes that Chorus is more likely to pass a Turing Test, which challenges an artificial intelligence system to fool someone into thinking it's human, than conventional chat software, although it may not meet most definitions of artificial intelligence.

In trials of the system, people asked Chorus for advice on restaurants to visit in Los Angeles and New York, and quickly received suggestions. Feedback such as "Hmm. That seems pricey," was quickly taken on board by the crowd, which came up with an alternative. AI systems such as Siri typically have difficulty following this kind of back-and-forth conversation, particularly in colloquial language.

Bigham worked with Rochester colleagues Walter Lasecki and Rachel Wesley, and Anand Kulkarni, the cofounder of crowdsourcing company MobileWorks (see "Human Workers, Managed by an Algorithm"). Their goal was to find a new way to increase the power of crowdsourcing, which is typically limited to simple, isolated tasks, such as adding labels to image files. "What we're really interested in is when a crowd as a collective can do better than even a high-quality individual," says Bigham, by combining work on many simple tasks into a coherent, complex whole.

Chorus does that with three simple types of task. First, any new chat updates from the human user are passed along to many crowd workers, who are asked to suggest a reply. Those suggestions are then voted on by crowd workers to determine the one that will be sent back.

A final mechanism creates a kind of working memory that ensures that Chorus's replies reflect the history of the conversation so far, crucial if it is to carry out long conversations—something that is a challenge for apps like Siri and even AI chatbots intended to showcase conversational skills.

For the working memory component, crowd members are asked to maintain a short running list of the eight most important snippets of information under discussion, to be used as a reference when workers suggest replies. This is important, as to allow for the natural turnover of crowdsourcing workers. "A single person may not be around for the duration of the conversation—they come and go, and some may contribute more than others," says Bigham.

Bigham says Chorus has the potential to be more than just a neat demonstration. "We definitely want to start embedding it into real systems," he says. "Perhaps you could help someone with cognitive impairment by having a crowd as a personal assistant."

Another possibility is to combine Chorus with another system previously developed at Rochester, which has crowd workers collaborate to steer a robot. "Could you create a robot this way that can drive around and interact intelligently with humans?" asks Bigham.

Michael Bernstein, an assistant professor at Stanford University who is currently doing research at Facebook, agrees that Chorus could lead to real-world applications (see "Adding Human Intelligence to Software").

"You could go from today where I call AT&T and speak with an individual, to a future where many people with different skills work together to act as a single incredibly intelligent tech support," says Bernstein. He says the Chorus software could become a true expert if it were able to direct incoming questions to members of the crowd with particular knowledge or skills.

However, Bernstein adds that it may be necessary to add more reviewing steps to Chorus in order to filter a crowd's suggestions to prevent it developing a split personality when faced with difficult questions. This is a familiar problem in applying crowdsourcing. Bigham's crowd-steered robot, for example, has been known to crash into obstacles dead ahead because half the crowd workers steering it wanted it to go left, and the other half wanted it to go right.

Tech's New Wave, Driven by Data



Steve Lohr, The New York Times, September 8, 2012

From the article: "TECHNOLOGY tends to cascade into the marketplace in waves. Think of personal computers in the 1980s, the Internet in the 1990s and smartphones in the last five years.

Computing may be on the cusp of another such wave. This one, many researchers and entrepreneurs say, will be based on smarter machines and software that will automate more tasks and help people make better decisions in business, science and government. And the technological building blocks, both hardware and software, are falling into place, stirring optimism.

Michael R. Stonebraker, a pioneer in database research, is one of the optimists. Software used by companies and government agencies — in products sold by Oracle, I.B.M., Microsoft and others — descends from research done in the 1970s by Mr. Stonebraker and Eugene Wong, a colleague at the University of California, Berkeley, as well as a team of scientists at I.B.M.

Today, Mr. Stonebraker sees an opportunity for new kinds of ultrafast databases. The new software, he explains, takes advantage of rapid advances in computer hardware to help businesses and researchers find insights in the rising flood of data coming from so many sources, including Web-browsing trails, sensor data, genetic testing and stock trading.

So, at 68, Mr. Stonebraker is a co-founder and chief technology officer of two start-ups in the field of data-driven discovery, VoltDB and Paradigm4.

“Now is the time,” says Mr. Stonebraker, who is an adjunct professor at the Massachusetts Institute of Technology’s computer science and artificial intelligence laboratory. “The economics and the technology are ripe.”

The case for optimism is by no means unqualified. The march of these technologies raises social issues, including privacy concerns, and the timing is uncertain. All of the bold predictions in the 1990s that the Internet would disrupt traditional industries like media, advertising and retailing did come true — a decade later.

But a series of related technologies, scientists and entrepreneurs say, has reached a critical mass — come to a digital boiling point, so to speak — so that new products and capabilities become possible. The technical ingredients, they note, include powerful, low-cost computing and storage spread across thousands of computers. The digital engine rooms of Google and Amazon are prime examples.

Another fast-improving technology involves inexpensive and intelligent sensors, which are crucial to a new breed of automated machines like experimental driverless cars and battlefield drones. Clever software — notably machine-learning algorithms — animates much of the current wave of smarter technology. Two well-known examples are found in Watson, the “Jeopardy”-winning computer from I.B.M., and the movie recommendations on Netflix.

ADVANCES in such underlying technologies are fueling the current excitement in fields like artificial intelligence, robotics and data analysis and prediction. “All parts of the technology pipeline are gearing up at the same time, and that’s how you get this explosion of new applications and uses,” says Jon Kleinberg, a computer scientist at Cornell University.

Behind the seeming explosion, experts say, is a process of technology evolution. Paul Saffo, a technology forecaster, compares the process to the evolutionary biology concept known as “punctuated equilibria” formulated by the paleontologists Stephen Jay Gould and Niles Eldredge. The idea is that species often evolve in periodic spurts.

Yet, they say, there are typically years of progress before a commercial breakthrough in the technological realm.

“Even in Silicon Valley, it takes most technologies 20 years to become overnight successes,” says Mr. Saffo, a consulting professor at Stanford’s school of engineering.

The Internet provides a case study of both technology’s evolutionary progress and its exponential growth. In 1969, there were only four computers connected to the nascent Internet, compared with roughly a billion computing devices today, from laptops to cellphones, says Edward Lazowska, a computer scientist at the University of Washington.

The early increases in connected computers drew scant attention. “But at some point in the late 1990s,” Mr. Lazowska says, “you were going from 4 million to 8 million to 16 million to 32 million to 64 million, and people started to notice that something revolutionary was going on.”

Rocket Fuel is a four-year-old Silicon Valley start-up that uses artificial-intelligence software to place display advertisements for marketers on the Web. The company can not only tailor ads by demographic slices of viewers’ ages, gender and interests, but can also use its predictive algorithms to produce campaigns based on results, says George H. John, the company’s chief executive.

For example, a luxury carmaker might tell Rocket Fuel that it wants to place 100 million ads in the next month, and it will pay the company, say, $80 for generating a sales lead, as evidenced by a potential customer downloading a brochure or filling out an online form.

Rocket Fuel is growing fast, having nearly doubled its work force since the start of the year, to 240. So far in 2012, it has handled campaigns for more than 500 advertisers, including BMW, Duncan Hines, Allstate, Pizza Hut and Ace Hardware. It has raised $76 million in venture funding and debt, and its thousands of computers handle 19 billion bid requests a day on ad exchanges. Each online auction for ad space is typically completed in about 100 milliseconds, a tenth of a second.

Rocket Fuel, Mr. John says, is using some of the ideas he worked on in the 1990s as a doctoral student focusing on artificial intelligence at Stanford — research that was supported with government dollars from the National Science Foundation and other agencies, as is so often the case. In the last few years, building a business around those ideas has become achievable and affordable. “And a lot of it has to do with the underlying technology,” Mr. John says.

FOR Mr. Stonebraker, the hardware advance that opens the door to his start-ups is the striking improvement of solid-state memory, as performance climbs and prices plunge. Solid-state, or flash, memory is most widely known as the lightweight storage technology used in consumer devices like small music players and smartphones.

But increasingly, solid-state memory can be used in big computers, holding a hefty database in memory instead of sending data off to be stored on disk drives. According to Mr. Stonebraker, some data-handling tasks can now be completed 50 times faster than with conventional systems.

“Memory is the new disk,” he says. “The obvious thing to do is to exploit that technology.”

In the yin and yang of computing, it is software that exploits hardware, enabling a computer to do useful things. And machine-learning programs and other data-sifting software are advancing swiftly.

“There is no point in collecting and storing all this data if the algorithms are not able to find useful patterns and insights in the data,” says Mr. Kleinberg at Cornell. “But the software is scaling up to the task.”

A version of this article appeared in print on September 9, 2012, on page BU4 of the New York edition with the headline: Tech’s New Wave, Driven By Data.

Friday, September 7, 2012

One Woman's Data Trail Diary



Scott Shane, The New York Times, August 31, 2012

As part of The Agenda, The Times’s look at major issues facing the next administration, we have been examining the trade-offs, more than a decade after the Sept. 11 attacks, between security and privacy and civil liberties. Some readers have written in about the electronic data trail that all of us leave as we go about our lives, using the Internet and carrying smartphones.


Heidi Boghosian, a New Yorker and author of a book on surveillance scheduled for publication next year, “Spying on Democracy: A Short History of Government/Corporate Collusion in the Technology Age,” agreed to try to document her own data trail on one recent day. Her account, below, is nothing extraordinary – and that’s the point. It is impossible to live in urban, wired America without leaving clues about ourselves, our movements and our views everywhere. And it is all but impossible to be certain who is looking at the resulting data or video and how much of it is accessible to federal, state or local government.

Ms. Boghosian is executive director of the National Lawyers Guild, a group of self-described radical lawyers and law students founded in 1937, and between her day job and her book research, she thinks far more than most people about surveillance and privacy. But the exercise of documenting her day was nonetheless informative, she said.

“Definitely, for me, going through the process reinforced my sense of the role corporations play in our daily lives,” she said. “And I don’t think most people realize the extent to which corporations cooperate in turning over personal information to the government.”

Here is the record she made:

A Day of Surveillance:

(1) 8:30 a.m.: Closed Circuit Television (CCTV) in hallway permits private landlord to monitor departure of tenant from apartment building at 173 Avenue A, New York, N.Y. A sign is posted alerting tenants that their actions are being monitored.

(2) City-owned video surveillance camera, mounted atop a streetlight pole, records pedestrian and vehicular traffic on corner of Avenue A and 11th Street.

(3) 9:45 a.m.: Internet Protocol-based, closed-circuit television CCTV/video surveillance camera at Chase Bank A.T.M. on Second Avenue and 10th Street records clear image of person withdrawing money. I.P. video surveillance footage probably transmitted to a central monitoring room and digitally stored (allowing for advanced search techniques), or viewed over the Internet. Intelligent I.P. cameras with video analytics such as motion sensors, facial recognition and behavioral recognition are used to identify abnormal activity in and around banking locations.

(4) 10 a.m.: Customer Loyalty Card at East Village coffee shop Café Pick Me Up, Avenue A and 9th Street, likely allow the business to track and predict customer spending habits.

(5) 10:30 a.m.: iPhone (with G.P.S. tracking) in owner’s back pocket allows phone owner’s movement and location to be tracked (by government, if cellphone provider gives access) through day and evening, even if phone is turned off, as phone owner walks to Astor Place subway stop.

(6) 10:45 a.m.: Passed by “smart sign” (digital billboard with cameras that gauges demographics of passers-by) that delivers ads tailored to the demographics of the passer-by.

(7) 10:45 a.m.: CCTV in NY subway system monitor boarding #6 subway line to work.

(8) 11:11 a.m.: CCTV in elevator records ride to ninth-floor office in office building. Building security guard has four cameras behind front desk showing elevator, stairways and hallways.

(9) 11:20 a.m.: Facebook software tracks user activities on sites on Internet after logging in and reading a few comments. “Tag Suggestions” feature employs facial recognition technology and suggests name tags after uploading photos.

(10) 11:30 a.m.: Cookies on Internet monitor all Internet searches throughout day on range of subjects; ads appear on screen related to items purchased on line (athletic shoes, face cream).

(11) 1 p.m.: Credit card at Macy’s Department store, used for quick purchase, has embedded RFID (Radio Frequency ID) chip, tracking consumer spending habits and providing that information to big business.

(12) 1:15 p.m.: Shoes in Macy’s new shoe store all have RFID chips (unique identifier linked to database) in them.

(13) 1:30 p.m. Downloaded iTunes onto iPhone. Online music providers may share personal information with third parties.

(14) 2 p.m.: Video cameras and motion detectors in local supermarket track physical movements of customers (allegedly to aim for improved customer service) as customer drops in to pick up some fruit for lunch.

(15) 2:30 p.m.: Social security number and driver’s license information, required by Fulton Street Verizon cellular store winds up in Verizon’s digital database, as customer switches from AT&T. Allegedly needed Social Security number to conduct credit check even though customer has had a landline account with Verizon for many years.

(16) 3:30 p.m.: I.P. address may have been included on a bar code on a digital coupon while registering for at hotelcoupons.com to get a discount hotel deal.

(17) 4 p.m.: Continuous, systematic desktop monitoring surveillance of personal use of business computer to access site to order shoes could have been conducted had employer installed software to monitor my real time actions, purportedly to avoid discrimination and sexual harassment lawsuits that may result from inappropriate e-mails sent within company.

(18) 6 p.m.: Dropped by anti-fracking protest on West 14th Street. Unmarked police van with tinted windows probably had NYPD Technical Assistance Response Unit (TARU) personnel recording protest activities, especially because several Occupy protesters were present. TARU provides investigative technical equipment and tactical support to all N.Y.P.D. bureaus and also provides assistance to other city, state and federal agencies. The unit employs several forms of computer forensics.

(19) 7:30 p.m.: CCTV cameras inside East Village restaurant while meeting friend for dinner after the protest.

(20) 9:30 p.m.: Surveillance cameras on several buildings passed on way home is captured on tape.

(21) 10 p.m.: Final check of Gmail, and a few Google searches, allow Google to collect even more data on consumer habits and personal interests.

Wednesday, September 5, 2012

REINVENTING SOCIETY IN THE WAKE OF BIG DATA


A Conversation with Alex (Sandy) Pentland      Edge Video (24-Minutes)

ALEX 'SANDY' PENTLAND is a pioneer in big data, computational social science, mobile and health systems, and technology for developing countries. He is one of the most-cited computer scientists in the world and was named by Forbes as one of the world's seven most powerful data scientists. He currently directs MIT's Human Dynamics Laboratory and the MIT Media Lab Entrepreneurship Program, and advises the World Economic Forum, Nissan Motor Corporation, and a variety of start-up firms.

Recently I seem to have become MIT's Big Data guy, with people like Tim O'Reilly and "Forbes" calling me one of the seven most powerful data scientists in the world. I'm not sure what all of that means, but I have a distinctive view about Big Data, so maybe it is something that people want to hear.

I believe that the power of Big Data is that it is information about people's behavior instead of information about their beliefs. It's about the behavior of customers, employees, and prospects for your new business. It's not about the things you post on Facebook, and it's not about your searches on Google, which is what most people think about, and it's not data from internal company processes and RFIDs. This sort of Big Data comes from things like location data off of your cell phone or credit card, it's the little data breadcrumbs that you leave behind you as you move around in the world.

"What those breadcrumbs tell is the story of your life. It tells what you've chosen to do. That's very different than what you put on Facebook. What you put on Facebook is what you would like to tell people, edited according to the standards of the day. Who you actually are is determined by where you spend time, and which things you buy. Big data is increasingly about real behavior, and by analyzing this sort of data, scientists can tell an enormous amount about you. They can tell whether you are the sort of person who will pay back loans. They can tell you if you're likely to get diabetes."

Sandy Pentland's EDGE Profile page: http://edge.org/memberbio/alex_(sandy)_pentland

Permalink: http://www.edge.org/conversation/reinventing-society-in-the-wake-of-big-data

[ED. NOTE: Part of the ongoing series "COMPUTATIONAL SOCIAL SCIENCE @ Edge"

http://edge.org/event/special/computational-social-science]

Brookings Report: Big Data for Education: Data Mining, Data Analytics, and Web Dashboards



Darrell M. West, The Brookings Institution, September 4, 2012

Imagine this scenario: twelve-year-old Susan took a course designed to improve her reading skills. She read short stories and the teacher would give her and her fellow students a written test every other week measuring vocabulary and reading comprehension. A few days later, Susan’s instructor graded the paper and returned her exam. The test showed that she did well on vocabulary, but needed to work on retaining key concepts.

In the future, her younger brother Richard is likely to learn reading through a computerized software program. As he goes through each story, the computer will collect data on how long it takes him to master the material. After each assignment, a quiz will pop up on his screen and ask questions concerning vocabulary and reading comprehension. As he answers each item, Richard will get instant feedback showing whether his answer is correct and how his performance compares to classmates and students across the country. For items that are difficult, the computer will send him links to websites that explain words and concepts in greater detail. At the end of the session, his teacher will receive an automated readout on Richard and the other students in the class summarizing their reading time, vocabulary knowledge, reading comprehension, and use of supplemental electronic resources.

In comparing these two learning environments, it is apparent that current school evaluations suffer from several limitations. Many of the typical pedagogies provide little immediate feedback to students, require teachers to spend hours grading routine assignments, aren’t very proactive about showing students how to improve comprehension, and fail to take advantage of digital resources that can improve the learning process. This is unfortunate because data-driven approaches make it possible to study learning in real-time and offer systematic feedback to students and teachers.

In this report, I examine the potential for improved research, evaluation, and accountability through data mining, data analytics, and web dashboards. So-called “big data” make it possible to mine learning information for insights regarding student performance and learning approaches.[1] Rather than rely on periodic test performance, instructors can analyze what students know and what techniques are most effective for each pupil. By focusing on data analytics, teachers can study learning in far more nuanced ways.[2] Online tools enable evaluation of a much wider range of student actions, such as how long they devote to readings, where they get electronic resources, and how quickly they master key concepts.


[1] James Manyika, Michael Chui, Brad Brown, Jacques Bughin, Richard Dobbs, Charles Roxburgh, and Angela Byers, “Big Data: The Next Frontier for Innovation, Competition, and Productivity,” McKinsey Global Institute, May, 2011.
[2] Felix Castro, Alfredo Vellido, Angela Nebot, and Francisco Mugica, “Applying Data Mining Techniques to e-Learning Problems,” Studies in Computational Intelligence, Volume 62, 2007, pp. 183-221.