Tuesday, March 22, 2011

Open Networking Foundation Pursues New Standards





MOUNTAIN VIEW, Calif. — Acknowledging that so-called cloud computing will blur the distinctions between computers and networks, about two dozen big information technology companies plan to announce on Tuesday a new standards-setting group for computer networking.

The group, to be called the Open Networking Foundation, hopes to help standardize a set of technologies pioneered at Stanford and the University of California, Berkeley, and meant to make small and large networks programmable in much the same way that individual computers are.

The changes, if widely adopted, would have implications for global telecommunications networks and large corporate data centers, but also for small household networks. The benefits, proponents say, would be more flexible and secure networks that are less likely to suffer from congestion. Someday, they say, networks might even be less expensive to build and operate.

The new approach could allow for setting up on-demand “express lanes” for voice and data traffic that is time-sensitive. Or it might let big telecommunications companies, like Verizon or AT&T, use software to combine several fiber optic backbones temporarily for particularly heavy information loads and then have them automatically separate when a data rush hour is over.

For households, the new capabilities might let Internet service providers offer remote services like home security or energy control.

The foundation’s organizers also say the new technologies will offer ways to improve computer security and could possibly enhance individual privacy within the e-commerce and social networking markets. Those markets are the fastest-growing uses for computing and network resources.

While the new capabilities could be crucial to network engineers, for business users and consumers the changes might be no more noticeable than advances in plumbing, heating and air-conditioning. Everything might work better, but most users would probably not know — or care — why or how.

The members of the Open Networking Foundation will include Broadcom, Brocade, Ciena, Cisco, Citrix, Dell, Deutsche Telekom, Ericsson, Facebook, Force10, Google, Hewlett-Packard, I.B.M., Juniper, Marvell, Microsoft, NEC, Netgear, NTT, Riverbed Technology, Verizon, VMWare and Yahoo.

“This answers a question that the entire industry has had, and that is how do you provide owners and operators of large networks with the flexibility of control that they want in a standardized fashion,” said Nick McKeown, a professor of electrical engineering and computer science at Stanford, where his and colleagues’ work forms part of the technical underpinnings, called OpenFlow.

The effort is a departure from the traditional way the Internet works. As designed by military and academic experts in the 1960s, the Internet has been based on interconnected computers that send and receive packets of data, paying little heed to the content and making few distinctions among the various types of senders and receivers of information.

The intelligence in the original Internet was meant to reside largely at the end points of the network — the computers — while the specialized routing computers were relatively dumb post offices of various size, mainly confined to reading addresses and transferring packets of data to adjacent systems.

But these days, when cloud computing means a lot of the information is stored and processed on computers out on the network, there is growing need for more intelligent control systems to orchestrate the behavior of thousands of routing machines. It will make it possible, for example, for managers of large networks to program their network to prioritize certain types of data, perhaps to ensure quality of service or to add security to certain portions of a network.

The designers argue that because OpenFlow should open up hardware and software systems that control the flow of Internet data packets, systems that have been closed and proprietary, it will cause a new round of innovation focused principally upon the vast computing systems known as cloud computers.”This is a pragmatic solution,” said David Farber, a computer scientist at Carnegie Mellon who was one of the pioneers of data networking technology.

“The idea of moving intelligence to the end points of the network was one of the original design points of the Internet,” Mr. Farber said. But he noted that as the network evolved to offer sophisticated advanced services through centralized cloud computers, including the delivery of digital voice and video, it became less feasible to continue relying on decentralized network design.

Mr. Farber noted that there have been other research projects aimed at redesigning the Internet. For example, the National Science Foundation, in addition to supporting the OpenFlow initiative, has financed the Global Environment for Network Innovations, or GENI. Open Flow appears to have generated broad industry support, he said, but it must still prove itself in the market.

A number of networking companies, including Cisco, Hewlett-Packard and Juniper, have already produced prototype systems that support OpenFlow technology. Morever, at least one Silicon Valley start-up, Nicira Networks, is now testing OpenFlow products that are meant to enter the cloud computing market later this year.

“If you look at really big companies like Google and Amazon, they take really smart programmers and give them a problem like search or automating storage and then turn the crank and out pops a great system,” said Martin Casado, Nicira’s chief technologist and one of the members of the OpenFlow research project at Stanford.

But Mr. Casado noted that in the past the one part of the system that could not be programmed was the network. “This customizes the network to the applications that are actually being run.”

Saturday, March 19, 2011

'We're in a critical period for the internet' . . . Tim Wu

The Wu master
The internet as a model of free speech and access is coming to an end, says web expert Tim Wu
     
The internet is under threat. At risk is what's known as "net neutrality", or the principle of free access for each user to every online site, regardless of content. That's the view of the man who coined the above term, Tim Wu, whose new book, The Master Switch, was published yesterday. It argues the internet now runs the risk of not just political censorship – as seen in Libya and Egypt, and in the American reaction to WikiLeaks – but that of commercial censorship, too. Monopolies such as Google and Apple may soon decide to choose which parts of the internet to give us – or switch off – and in some cases have already started to do so.

"We are in a critical period for the internet," Tim Wu, the book's author, says. "What the internet is, is in flux." Wu looks, a colleague suggests, like a cleverer version of Keanu Reeves. In reality, he is a senior adviser to the Obama administration on, fittingly, the competition issues that concern internet and mobile industries. A position which, ironically, makes him a distant colleague of the officials waging war against WikiLeaks and Bradley Manning. An academic lawyer by trade – he has taught at Chicago, Columbia and Stanford – Wu has also long been a respected commentator on internet issues, and writes regularly for Slate magazine. The first to coin the term "net neutrality", Wu is sometimes mentioned in the same breath as social media experts Jay Rosen and Jeff Jarvis, and sceptics Evgeny Morozov, Nicholas Carr and Jaron Lanier. But unlike these five, whose work is mainly concerned with a discussion about the (de)merits of online activity, Wu's book perhaps places him in a critically different category. The Master Switch is less concerned with the rights and wrongs of the internet today, and more concerned with its long-term future.

"The internet is about 15 years into its cycle as an open medium," says Wu, "and at that moment in their cycle, most open media tend to turn to closed media." What Wu means is that the internet might be about to go the same way as the information services of the 20th century: the telephone, radio, cinema and television. "Internet is the descendant of these industries," Wu says, "a 15-year-old teenager." And if we want to know what kind of adult this teenager will become, "the clearest way is to look at its parents, and look what happened to them when they reached their 20s".

What happened, he argues, is that they went from being technologies used by lots of different individuals and companies, to ones controlled by just a few monopolies. He uses the example of AT&T, the great American telephone monopolist, "who went to the American people and said, 'We will be good, we will build the best telephone network in the world: give us the monopoly'." He points to the American film industry, and shows how it quickly was transformed from an industry that was relatively easy to enter, to one mainly controlled by a few Hollywood studios.

Wu's fear is that a similar consolidation of power may be about to happen to the internet. "When we talk about the internet," he says, "we're only talking about three or four companies . . . Amazon, Google, Apple, Facebook." These big four are so large that one or more of them could team up with a mobile network and end net neutrality (a term that Wu coined in 2003) by privileging that network's clients above all others. It's something that is already partly happening. Apple's iPhone was at first only available to AT&T customers in the US and O2 clients in the UK. And since Google teamed up with American phone firm Verizon last summer, it now at least has the option to do something similar. Wu won't comment on these deals directly, due to his position with the US government. But, he says, like AT&T in the 1910s, "Google has similarly promised the world: 'We will be a good company.' And we have essentially conferred it a dominance over the market, and over how we find our information, because we believe it will be good."

However, warns Wu, "the question everyone has is whether one day Google will have its Heart of Darkness moment."  If it does, says Wu, what's at stake is the principle on which the internet was founded. At its inception, "the internet was typically a place where you could put up content without anybody's permission". But partnerships between the bigger web-related companies might squeeze out the smaller ones. Using the example of the online news industry, Wu suggests that if newspapers were to follow the example of Rupert Murdoch's new iPad-based "paper", The Daily, and "become exclusive partners with Apple, it may be easier for them to make money, but we may also end up with a media on the internet that is significantly more closed than it is now." This is because, he says, "You can imagine a future where blogs don't really have a meaningful future, because the content provided on a platform [such as Apple] doesn't create any room for anyone other than its exclusive media partners." So, Wu concludes: "The internet as a forum for speech, as a place where an individual with a talent can compete with a major newspaper – I'm suggesting that model may be passing."

But though the internet was a freer place in its younger days, I ask Wu, wasn't it only available to a privileged few? Big conglomerates may be growing ever powerful, but haven't they at least brought the internet to a much wider audience than the academics and techies of the early 90s? "It shouldn't be a trade-off," Wu replies. "There is some truth to the idea that companies are interested in consumers, and so they bring [the internet] to a broader marketplace – but it is still important to stand up for the original values of the internet." Wu sees these values "as fundamental to a free society", values that we should preserve even as the internet becomes "a mass consumption product. And I think it's possible. You don't have to throw those values out the window just because millions of people are using it."

Wu recommends protecting these values through the maintenance of something that in his book he calls the "separation principle". Just as journalists maintain "a separation of news and opinion", Wu argues, "the people who move and carry information should stay at some distance from the creators of content, because they have a natural conflict of interest." In other words, though Wu does not name names, companies such as Apple and Google should stay well away from mobile networks such as AT&T and Verizon.

He feels government and consumers have a dual responsibility to police this conflict. Legislation should prevent mergers between the carriers of internet content, and the producers of the content. Consumers should boycott any company that threatens net neutrality. Will it work? Wu is undecided. "My big question is whether, five years from now, the big four companies will be even more consolidated, with other companies mattering less and less – or whether the internet will have proved its truly radical nature, and a whole new cast of characters will have emerged . . . And I don't know the answer."

• The Master Switch: The Rise and Fall of Information Empires by Tim Wu is published by Atlantic Books, £19.99

Tuesday, March 15, 2011

This Data Isn't Dull. It Improves Lives.




GOVERNMENTS have learned a cheap new way to improve people’s lives. Here is the basic recipe:
Take data that you and I have already paid a government agency to collect, and post it online in a way that computer programmers can easily use. Then wait a few months. Voilà! The private sector gets busy, creating Web sites and smartphone apps that reformat the information in ways that are helpful to consumers, workers and companies.

Not surprisingly, San Francisco, with its proximity to Silicon Valley, has been a pioneer in these efforts. For some years, Bay Area transit systems had been tracking the locations of their trains and buses via onboard GPS. Then someone got the bright idea to post that information in real time. Thus the delightful app Routesy was born. Install it on a smartphone and the app can tell you that your bus is stuck in traffic and will be 10 minutes late — or it can help you realize that you are standing on the wrong street, dummy. It gives consumers a great new way to find out when and where the bus is coming, and all at minimal government expense.

Another example involves weather data produced by the National Oceanographic and Atmospheric Administration. The forecasts you find on the Weather Channel, or on the evening news or online, use the agency’s information. Again, the government produces and releases raw data, and the private sector transforms it into something useful for the public.

Several other departments in the Obama administration are looking to expand the use of such techniques. On data.gov, you will find huge amounts of downloadable data that had heretofore been inaccessible. As a sign of the importance that President Obama has attached to this approach, he put it on the government’s agenda on Jan. 21, 2009, his second day in office. (Disclosure: My book, “Nudge,” published in 2008, advocated this broad idea; Cass R. Sunstein, co-author of the book, is now administrator of the White House Office of Information and Regulatory Affairs.)

Now the administration is pushing to use this concept as a tool for regulation, and as a method of avoiding more heavy-handed rule making. The idea is that making things more transparent can immediately turn consumers into better shoppers and make markets work better. One might think that such an initiative would receive nearly universal support — after all, who could be against openness and transparency? But it turns out that some people are.

Two cases are under discussion right now.

First, the Department of Transportation is considering a new rule requiring airlines to make all of their prices public and immediately available online. The postings would include both ticket prices and the fees for “extras” like baggage, movies, food and beverages. The data would then be accessible to travel Web sites, and thus to all shoppers.

The airlines would retain the right to decide how and where to sell their products and services. But many of them are insisting that they should be able to decide where and how to display these extra fees. The issue is likely to grow in importance as airlines expand their lists of possible extras, from seats with more legroom to business-class meals served in coach.

Electronic disclosure of all fees can make it much easier for consumers to figure out what a trip really costs, and thus make markets more efficient, without requiring new rules and regulations. (As someone who once bought two tickets on a discount airline from London to Dublin for the advertised price of £1 each, then ended up paying hundreds of dollars for the privilege of bringing along two heavy suitcases, I acknowledge having a sore spot on this issue.)

Another initiative has been proposed by the Consumer Product Safety Commission. In 2008, Congress overwhelmingly passed and President George W. Bush signed legislation mandating an online database of reported safety issues in products, at saferproducts.gov. The Web site ran for a few months in a “soft launch” and went into full operation on Friday.

But a majority in the House of Representatives passed an amendment last month that might have stopped this initiative in its tracks. The amendment, sponsored by Representative Mike Pompeo, a Kansas Republican, would have prohibited the agency from spending any further money to start the site. One goal, of course, was to cut the budget, although proponents of the amendment also argue that the Web site might include information that is erroneous and damaging to the businesses that sell children’s products.

Yet several provisions in the final rules protect manufacturers from false or malicious statements. Consumers have to include identifying information and sign an affidavit testifying to the truth of their complaints. Furthermore, manufacturers will be able to see complaints before they are posted, and can then correct mistakes or add comments.

ALTHOUGH this amendment was passed in the name of deficit reduction, the requested money for the site is a puny $3 million a year. If we want to reduce the cost of government regulation, this is exactly the kind of effort we should be applauding and expanding.

Compared with the tiny costs, the benefits of this program could be enormous. Thirteen years ago, two of my dear friends experienced the nightmare that parents dread most. They were called at work by their child-care provider and told that their 18-month-old son had died in a crib accident. Imagine their anguish when they later learned that other children had died in this model of crib, and that still others had died in cribs with similar design. Yet there was no easy way for any parent or child-care provider to know that.

In a recent three-year span, some 265 children under the age of 5 died in accidents related to nursery products, the government has reported. If this program could reduce that number even slightly, the cost would seem amply justified.

Moving the government into the 21st century should be applauded. In a future column, I will explain how the release of some kinds of data can even help consumers better understand themselves.

Richard H. Thaler is a professor of economics and behavioral science at the Booth School of Business at the University of Chicago.

Wednesday, March 2, 2011

Halamka: Freeing the Data

Wednesday, March 2, 2011

I'm keynoting this year's Intersystems Global Conference on the topic of "Freeing the Data" from the transactional systems we use today such as Enterprise Resource Planning (ERP), Customer Relationship Management (CRM),  Electronic Health Records (EHR), etc.  As I've prepared my speech,  I've given a lot of thought to the evolving data needs we have in our enterprises.

In healthcare and in many other industries, it's increasingly common for users to ask IT for tools and resources to look beyond the data we enter during the course of our daily work.   For one patient, I know the diagnosis, but what treatments were given to the last 1000 similar patients.  I know the sales today, but how do they vary over the week, the month, and the year?   Can I predict future resource needs before they happen?

In the past, such analysis typically relied on structured data, exported from transactional systems into data marts using Extract/Transform/Load (ETL) utilities, followed by analysis with Online Analytical Processing (OLAP) or Business Intelligence (BI) tools.

In a world filled with highly scalable web search engines,  increasingly capable natural language processing technologies, and practical examples of artificial intelligence/pattern recognition (think of IBM's Jeopardy-savvy Watson as a sophisticated data mining tool), there are novel approaches to freeing the data that go beyond a single database with pre-defined hypercube rollups.   Here are my top 10 trends to watch as we increasingly free data from transactional systems.

1.  Both structured and unstructured data will be important

In healthcare, the HITECH Act/Meaningful Use requires that clinicians document the smoking status of 50% of their patients.   In the past, many EHRs did not have structured data elements to support this activity.    Today's certified EHRs provided structured vocabularies and specific pulldowns/checkboxes for data entry, but what do we do about past data?   Ideally, we'd use natural language processing, probability, and search to examine unstructured text in the patient record and figure out smoking status including the context of the word smoking such as "former", "active", "heavy", "never" etc.

Businesses will always have a combination of structured and unstructured data.   Finding ways to leverage unstructured data will empower businesses to make the most of their information assets.

2.  Inference is possible by parsing natural language

Watson on Jeopardy provided an important illustration of how natural language processing can really work.   Watson does not understand the language and it is not conscious/sentient.   Watson's programming enables it to assign probabilities to expressions.     When asked "does he drink alcohol frequently?", finding the word "alcohol" associated with the word "excess" is more more likely to imply a drinking problem than finding "alcohol" associated with  "to clean his skin before injecting his insulin".    Next generation Natural Language Processing tools will provide the technology to assign probabilities and infer meaning from context.

3.  Data mining needs to go beyond single databases owned by a single organization.

If I want to ask questions about patient treatment and outcomes, I may need to query data from hundreds of hospitals to achieve statistical significance.   Each of those hospitals may have different IT systems with different data structures and vocabularies.   How can a query a collection of heterogenous databases?   Federation will possible by normalizing the queries through middleware.   For example, data might be mapped to a common Resource Description Framework (RDF) exchange language using standardized SPARQL query tools.   At Harvard, we've created a common web-based interface called SHRINE that queries all our hospital databases, providing aggregate de-identified answers to questions about diagnosis and treatment of millions of patients.

4.  Non-obvious associations will be increasingly important

Sometimes, it is not enough to query multiple databases.   Data needs to be linked external resources to produce novel information.  For example, at Harvard, we've taken the address of each faculty member, examined every publication they have ever written, geo-encoded the location of every co-author, and created visualizations of productivity, impact, and influence based on the proximity of colleagues.   We call this "social networking analysis"

5.  The President's Council of Advisors on Science and Technology (PCAST) Healthcare IT report will offer several important directional themes to will accelerate "freeing the data".

The PCAST report suggests that we embrace the idea of universal exchange languages, metadata tagging with controlled vocabularies, privacy flagging, and search engine technology with probabilistic matching to transform transactional data sources into information, knowledge and wisdom.    For example, imagine if all immunization data were normalized as it left transactional systems and pushed into state registries that were united by a federated search that included privacy protections.  Suddenly every doctor could ensure that every person had up to date immunizations at every visit.

6.  Ontologies and data models will be important to support analytics

Part of creating middleware solutions that enable federation of data sources requires that we generally know what data is important in healthcare and how data elements relate to each other.   For example, it's important to know that an allergy has a substance, a severity, a reaction, an observer, and an onset data.   Every EHR may implement allergies differently, but by using common detailed clinical model for data exchange and querying we can map heterogeneous data into comparable data.

7.  Mapping free text to controlled vocabularies will be possible and should be done as close to the source of data as possible.

Every industry has its jargon.   Most clinicians do not wake up every morning thinking about SNOMED-CT concepts of ICD-10 codes.   One way to leverage unstructured data is to turn it into structured data as it is entered.   If a clinician types "Allergy to Pencillin", it could become SNOMED-CT concept 294513009 for Pencillins.  As more controlled vocabularies are introduced in medicine and other industries, transforming text into controlled concepts for later searching will be increasingly important.   Ideally, this will be done as the data is entered, so it can be checked for accuracy.  If not at entry, then transformations should be done as close to the source systems as possible to ensure data integrity.   With every transformation and exchange of data from the original source, there is increasing risk of loss of meaning and context.

8.  Linking identity among heterogenous databases will be required for healthcare reform and novel business applications.

If a patient is seen in multiple locations how can we combine their history together so they get the maximum benefit of alerts, reminders, and decision support?    Among the hospitals I oversee, we have persistent linkage of all medical record numbers between hospitals - a master patient index.   Surescripts/RxHub does a realtime probabilistic match on name/gender/date of birth for over 150 million people in real time.   There are other interesting creative techniques such as those pioneered by Jeff Jonas for creating a unique hash of data for every person, then linking data based on that hash.   For example John, Jon, Jonathan, and Johnny are reduced to one common root name John.    The combination of John + lastname + date of birth is hashing using SHA-1.   In this way, records about the person can be aggregated without ever disclosing who the person really is - it's just that hash that is used to find common records.

9.  New tools will empower end users

All users, not just power users, want web-based or simple to use client server tools that allow data queries and visualizations without requiring a lot of expertise.  The next generation of SQL Server and PowerPivot offer this kind of query power from the desktop.    At BIDMC, we've created web-based parameterized queries in our Meaningful Use tools, we're implementing PowerPivot, and we're creating a powerful hospital-based visual query tool using I2B2 technologies.

10.  Novel sources of data will be important

Today, patients and consumers are generating data from apps on smart phones, from wearable devices, and social networking sites.   Novel approaches to creating knowledge and wisdom will source data from consumers as well as traditional corporate transactional systems.

Thus, as we all move toward "freeing the data" it will no longer be ufficient to use just structured transaction data entered by experts in a single organization, then mined by professional report writers.   The speed of business and the need for enhanced quality and efficiency is pushing us toward near real time business intelligence and visualizations for all users.   In a sense this mirrors the development of the web itself, evolving from expert HTML coders, to tools for content management for non-technical designated editors, to social networking where everyone is an author, publisher, and consumer.

"Freeing the data" is going to require new thinking about the way we approach application design and requirements.   Just as security needs to be foundational, analytics need to be built in from the beginning.

I look forward to my keynote in a few weeks.  Once I've delivered it, I'll post the presentation on my blog.

Tuesday, March 1, 2011

World Economic Forum Drives Health Data Initiative

 
The Global Health Data Charter calls for the use of technology to overcome worldwide gaps in health information collection, availability, privacy, and analysis.

By Nicole Lewis, Information Week,  Feb 28, 2011

The World Economic Forum has launched the Global Health Data Charter, an initiative to advance global health through the management and collection of data. The charter aims to enable individuals and patients, health professionals, and policymakers to make more informed decisions through secure access to comprehensive health data.

Officials at the World Economic Forum in Geneva said at the charter's1 unveiling last week that accurate health data is not available across health systems operating in developed and developing countries, and that gaps in data can be overcome through the use of technology, which will be a main driver in the collection, analysis, and application of health information.

In an interview with InformationWeek, Olivier Raynaud, head of global health and healthcare industries at the World Economic Forum, said the charter is a foundation document that can be used by national and individual organizations and clinicians. Healthcare stakeholders, such as health research organizations, academia, providers, insurers, and nongovernmental organizations (NGOs), can play a role in and benefit from the capture, storage, sharing, and use of health data.

"Even though it is 2011, most health data is still captured with pen and paper. The cost of digital support for information is getting closer to zero and this immediately makes information easier to store, retrieve, share, and aggregate," Raynaud said. "There are large-scale programs taking place in the most challenging areas (such as the monitoring of pregnancies by BRAC3, a large NGO in Bangladesh, where midwifes were equipped with PDAs), which have shown that it is perfectly feasible and generates immediate and tangible results.

He also said many nations are at a tipping point as they transition from paper-based systems to capturing health data electronically, and the hope is that the charter will help foster and enable a data-based, digital health era that will address global disparities in health.

"Disadvantaged populations will gain more from improved health data management; similar to what has been seen with mobile communication, digital health information has the potential to enable a dramatic change in the pace of progress towards universal coverage and access to health," Raynaud said.