Monday, August 26, 2013

book review | Radical Abundance: How a Revolution in Nanotechnology Will Change Civilization

book review | Radical Abundance: How a Revolution in Nanotechnology Will Change Civilization
August 24, 2013 by José Luis Cordeiro<http://www.kurzweilai.net/jose-luis-cordeiro>

[http://www.kurzweilai.net/images/Radical-Abundance.png]Eric Drexler, popularly known as "the founding father of nanotechnology," introduced the concept in his seminal 1981 paper in Proceedings of the National Academy of Sciences.

This paper established fundamental principles of molecular engineering and outlined development paths to advanced nanotechnologies.

He popularized the idea of nanotechnology in his 1986 book, Engines of Creation: The Coming Era of Nanotechnology, where he introduced a broad audience to a fundamental technology objective: using machines that work at the molecular scale to structure matter from the bottom up.

He went on to continue his PhD thesis at MIT, under the guidance of AI-pioneer Marvin Minsky, and published it in a modified form as a book in 1992 as Nanosystems: Molecular Machinery, Manufacturing, and Computation<http://e-drexler.com/p/idx04/00/0411nanosystems.html>.

Drexler's new book, Radical Abundance: How a Revolution in Nanotechnology Will Change Civilization<http://www.kurzweilai.net/radical-abundance-how-a-revolution-in-nanotechnology-will-change-civilization>, tells the story of nanotechnology from its small beginnings, then moves quickly towards a big future, explaining what it is and what it is not, and enlightening about what we can do with it for the benefit of humanity.

In his pioneering 1986 book, Engines of Creation, he defined nanotechnology as a potential technology with these features: "manufacturing using machinery based on nanoscale devices, and products built with atomic precision."

In his 2013 sequel, Radical Abundance, Drexler expands on his prior thinking, corrects many of the misconceptions about nanotechnology, and dismisses fears of dystopian futures replete with malevolent nanobots and gray goo.

In 1986 he talked about nanotechnology as the engine of abundance, but now he talks about radical abundance:

Imagine what the world might be like if we were really good at making things — better things — cleanly, inexpensively, and on a global scale. … The global prospect would be, not scarcity, but unprecedented abundance — radical, transformative, and sustainable abundance. We would be able to produce radically more of what people want and at a radically lower cost — in every sense of the word, both economic and environmental.

What if industrial production as we know it can be changed beyond recognition? The consequences would change almost everything else, and this new industrial revolution is visible on the horizon. Imagine a world where the gadgets and goods that run our society are produced not in a far-flung supply chain of industrial facilities, but in compact, even desktop-scale machines.

Imagine replacing an enormous automobile factory and all of its multi-million dollar equipment with a garage-sized facility that can assemble cars from inexpensive, microscopic parts, with production times measured in minutes. Then imagine that the technologies that can make these visions real are emerging — under many names, behind the scenes, with a long road still ahead, yet moving surprisingly fast.

The new nanotech: atomically precise manufacturing

His new book clearly identifies nanotechnology with atomically precise manufacturing (APM). He shows that APM rests on well-understood scientific and engineering principles that will support large-scale, low-cost production of advanced products, enable solutions to seemingly intractable global problems, and facilitate rapid draw-down of atmospheric CO2 levels.

Drexler describes how APM will radically reduce materials and energy costs, since most future devices will be manufactured using very abundant and common elements like carbon, hydrogen, nitrogen, oxygen and silicon, and at a very low cost of less than one dollar per kilogram. He argues that photovoltaic systems using nanotechnology will allow us to power human civilization using only 0.2% of the Earth's land surface, without using any more fossil fuels and helping to improve the environment.

Consider that a "one-gram platform built with advanced technologies could provide teraflops of computational power (and much more, in bursts), together with a million-terabyte data storage capacity and better-than-human sensors, all with a power demand comparable to that of cell phone on standby," he suggests.

He argues that new devices produced with APM will radically use less resources and energy, and will be radically more efficient and stronger. Instead of using milliwatts, future devices will use nanowatts of power. We will produce much more for much less, while preserving the environment and improving it.

Drexler makes many comparisons between the information revolution and what he now calls the "APM revolution." What the first did with bits, the second will do with atoms: "Image files today will be joined by product files tomorrow. Today one can produce an image of the Mona Lisa without being able to draw a good circle; tomorrow one will be able to produce a display screen without knowing how to manufacture a wire."

Civilization, he says, is advancing from a world of scarcity toward a world of abundance — indeed, radical abundance.

José Cordeiro is Director, Venezuela Node of The Millennium Project and Energy Advisor/Faculty, Singularity University

Thursday, August 22, 2013

Data Monetization in the Age of Big Data

Data Monetization in the Age of Big Data

Source: Accenture

The volume and richness of the data now uniquely accessible to mobile providers—whether in the form of transactions, inquiries, text messages or tweets, GPS locations or live video feeds—offers a veritable gold mine of insights and applications. And even as mobile phones have become the primary device through which consumers get their information, those very same devices have begun to facilitate new types of information, including extremely precise, real-time, geolocation information.

Not surprisingly, operators today are talking about when and how to tap into this data and what to do with it. In particular, they want to know how to monetize it: how to sort, analyze and manipulate the data and put it to use. This holds true not only for internal applications, but increasingly for building new revenue streams or collaborating on external applications with third parties as well.

How can mobile operators best approach this new territory? What are the opportunities and challenges? And how can operators shape new business models to monetize their Big Data?

Direct link to document (PDF; 4.6 MB)

Wednesday, August 21, 2013

A Guardian guide to your Metadata

A Guardian guide to your Metadata
http://www.theguardian.com/technology/interactive/2013/jun/12/what-is-metadata-nsa-surveillance#meta=0000100

Metadata is information generated as you use technology, and its use has been the subject of controversy since NSA's secret surveillance program<http://www.guardian.co.uk/world/the-nsa-files> was revealed. Examples include the date and time you called somebody or the location from which you last accessed your email. The data collected generally does not contain personal or content-specific details, but rather transactional information about the user, the device and activities taking place. In some cases you can limit the information that is collected – by turning off location services on your cell phone for instance – but many times you cannot. Below, explore some of the data collected through activities you do every day. On Thursday, June 13 The Guardian's data editor James Ball will answer your questions about the NSA data collection program<http://q-and-a.guardian.co.uk/qanda/859036> in the US from 3pm-4pm EST | 8pm-9pm BST
Choose the services you use in a day

Email
Phone
Camera
Facebook
Twitter
Search
Web browser
Tweet<https://twitter.com/intent/tweet?url=http%3A%2F%2Fwww.theguardian.com%2Ftechnology%2Finteractive%2F2013%2Fjun%2F12%2Fwhat-is-metadata-nsa-surveillance%23meta%3D0000100&text=How%20much%20metadata%20do%20you%20reveal%20every%20day%3F%20Find%20out%20via%20%40GuardianUS%20%23NSAfiles%20INTERACTIVE%3A> Share Reset
Twitter

* your name, location, language, profile bio information and url
* when you created your account
* your username and unique identifier
* tweet's location, date, time and timezone

* tweet's unique ID and ID of tweet replied to
* contributor IDs
* your follower, following and favorite count
* your verification status
* application sending the tweet

What metadata looks like

Below is a tweet from @GuardianUS (right) and a truncated version of its metadata (left). Accessing metadata is often possible through services offered by the provider and can be retrieved in a structured format that could include raw text, XML, or in this example, JSON. An easy way to see some of your own metadata is by looking at your browser's history which provides information about what websites you visited and when

What you can tell using metadata:
A case study of the Petraeus scandal

1. To communicate, Paula Broadwell and David Petraeus shared an anonymous email account

[http://gia.guim.co.uk/2013/06/metadata/img/broadwell_patreaus_frames01.png]

2. Instead of sending emails, both would login to the account, edit and save drafts

[http://gia.guim.co.uk/2013/06/metadata/img/broadwell_patreaus_frames02.png]

3. Broadwell logged in from various hotels' public Wi-Fi, leaving a trail of metadata that included times and locations

[http://gia.guim.co.uk/2013/06/metadata/img/broadwell_patreaus_frames03.png]

4. The FBI crossed-referenced hotel guests with login times and locations leading to the identification of Broadwell

[http://gia.guim.co.uk/2013/06/metadata/img/broadwell_patreaus_frames04.png]
Sources: Microsoft<http://msdn.microsoft.com/en-us/library/gg309512.aspx>, IPTC<http://www.iptc.org/std/photometadata/0.0/documentation/IPTC-PhotoMetadataWhitePaper2007_11.pdf>, Verizon telephone data court order<http://www.guardian.co.uk/world/interactive/2013/jun/06/verizon-telephone-data-court-order>, Photo Metadata<http://www.photometadata.org/meta-tutorials-adobe-photoshop>, Facebook<https://developers.facebook.com/docs/reference/login/extended-profile-properties>, Beyond the Wire<http://blog.newswire.ca/wordpress-mu/2011/06/28/uncovering-the-metadata-in-a-tweet>, Google<https://support.google.com/accounts/answer/54068?hl=en>, ACLU<http://www.aclu.org/blog/technology-and-liberty-national-security/surveillance-and-security-lessons-petraeus-scandal>

Saturday, August 17, 2013

Is Big Data an Economic Big Dud?

August 17, 2013

Is Big Data an Economic Big Dud?

The astounding rate of growth would make any parent proud. There were 30 billion gigabytes of video, e-mails, Web transactions and business-to-business analytics in 2005. The total is expected to reach more than 20 times that figure in 2013, with off-the-charts increases to follow in the years ahead, according to Cisco, the networking giant.

How much data is that? Cisco estimates that in 2012, some two trillion minutes of video alone traversed the Internet every month. That translates to over a million years per week of everything from video selfies and nannycams to Netflix downloads and “Battlestar Galactica” episodes.

What is sometimes referred to as the Internet’s first wave — say, from the 1990s until around 2005 — brought completely new services like e-mail, the Web, online search and eventually broadband. For its next act, the industry has pinned its hopes, and its colossal public relations machine, on the power of Big Data itself to supercharge the economy.

There is just one tiny problem: the economy is, at best, in the doldrums and has stayed there during the latest surge in Web traffic. The rate of productivity growth, whose steady rise from the 1970s well into the 2000s has been credited to earlier phases in the computer and Internet revolutions, has actually fallen. The overall economic trends are complex, but an argument could be made that the slowdown began around 2005 — just when Big Data began to make its appearance.

Those factors have some economists questioning whether Big Data will ever have the impact of the first Internet wave, let alone the industrial revolutions of past centuries. One theory holds that the Big Data industry is thriving more by cannibalizing existing businesses in the competition for customers than by creating fundamentally new opportunities.

In some cases, online companies like Amazon and eBay are fighting among themselves for customers. But in others — here is where the cannibals enter — the companies are eating up traditional advertising, media, music and retailing businesses, said Joel Waldfogel, an economist at the University of Minnesota who has studied the phenomenon.

“One falls, one rises — it’s pretty clear the digital kind is a substitute to the physical kind,” he said. “So it would be crazy to count the whole rise in digital as a net addition to the economy.”

Robert J. Gordon, a professor of economics at Northwestern University, said comparing Big Data to oil was promotional nonsense. “Gasoline made from oil made possible a transportation revolution as cars replaced horses and as commercial air transportation replaced railroads,” he said. “If anybody thinks that personal data are comparable to real oil and real vehicles, they don’t appreciate the realities of the last century.”

Other economists believe that Big Data’s economic punch is just a few years away, as engineers trained in data manipulation make their way through college and as data-driven start-ups begin hiring. And of course the recession could be masking the impact of the data revolution in ways economists don’t yet grasp. Still, some suspect that in the end our current framework for understanding Big Data and “the cloud” could be a mirage.

“I think it’s conceivable that the data era will be a bust for the things people expect it to be useful for,” said Scott Wallsten, a senior fellow at the Technology Policy Institute and the Georgetown Center for Business and Public Policy. Some entirely new use will have to turn up for data to fulfill its economic potential, he added.

There is no disputing that a wide spectrum of businesses, from e-marketers to pharmaceutical companies, are now using huge amounts of data as part of their everyday business.

Josh Marks is the chief executive of one such company, masFlight, which helps airlines use enormous data sets to reduce fuel consumption and improve overall performance. Although his first mission is to help clients compete with other airlines for customers, Mr. Marks believes that efficiencies like those his company is chasing should eventually expand the global economy.

For now, though, he acknowledges that most of the raw data flowing across the Web has limited economic value: far more useful is specialized data in the hands of analysts with a deep understanding of specific industries. “The promises that are made around the ability to manipulate these very large data sets in real time are overselling what they can do today,” Mr. Marks said.

Some economists argue that it is often difficult to estimate the true value of new technologies, and that Big Data may already be delivering benefits that are uncounted in official economic statistics. Cat videos and television programs on Hulu, for example, produce pleasure for Web surfers — so shouldn’t economists find a way to value such intangible activity, whether or not it moves the needle of the gross domestic product?

In addition, infrastructure investments often take years to pay off in a big way, said Shane Greenstein, an economist at Northwestern University. He cited high-speed Internet connections laid down in the late 1990s that have driven profits only recently. But he noted that in contrast to the Internet’s first wave, which created services like the Web and e-mail, the impact of the second wave — the Big Data revolution — is harder to discern above the noise of broader economic activity.

“It could be just time delay, or it could be that the value just isn’t there,” said Mr. Greenstein, who has studied the competitive success of online businesses in media, advertising and retailing.

Perhaps surprisingly, the parallel most tightly embraced by digital futurists — the rise of the electricity grid — is largely dismissed by those who have studied the history of the subject. The idea is that a ubiquitous Internet will make data and “cloud” computing available anywhere, like electricity through a socket.

The numerical comparisons are tantalizing. As illustrated in “The Electric City,” by Harold L. Platt, the booming quantity and adoption rates of electricity flowing on the Chicago grid in the late 19th and early 20th centuries instantly bring to mind those charts showing data growth today.

Despite those similarities, Mr. Platt, a professor emeritus of history at Loyola University Chicago, said it was unlikely that the revolutions unleashed in manufacturing, domestic life, transportation and high and low society by electricity could ever be matched by the data era. “I’d be hard pressed to quickly draw comparisons,” he said.

But even as Mr. Platt, 68, spoke by cellphone from Chicago, fragments of today’s inescapable data flood found him as he received messages from his grown children. “I have to text them or else they won’t answer me back,” Mr. Platt said gamely. “I’m going with the flow.”

James Glanz is an investigative reporter for The New York Times.

Friday, August 16, 2013

Five myths about big data

Five myths about big data

By Samuel Arbesman, Friday, August 16, 11:29 AM

Samuel Arbesman, an applied mathematician and network scientist, is a senior scholar at the Ewing Marion Kauffman Foundation and the author of “The Half-Life of Facts.” Follow him on Twitter: @Arbesman.

by Samuel Arbesman Big data holds the promise of harnessing huge amounts of information to help us better understand the world. But when talking about big data, there’s a tendency to fall into hyperbole. It is what compels contrarians to write such tweets as “Big Data, n.: the belief that any sufficiently large pile of s--- contains a pony.” Let’s deflate the hype.

1. “Big data” has a clear definition.

The term “big data” has been in circulation since at least the 1990s, when it is believed to have originated in Silicon Valley. IBM offers a seemingly simple definition: Big data is characterized by the four V’s of volume, variety, velocity and veracity. But the term is thrown around so often, in so many contexts — science, marketing, politics, sports — that its meaning has become vague and ambiguous.

There’s general agreement that ranking every page on the Internet according to relevance and searching the phone records of every Verizon customer in the United States qualify as applications of big data. Beyond that, there’s much debate. Does big data need to involve more information than can be processed by a single home computer? If so, marketing analytics wouldn’t qualify, and neither would most of the work done by Facebook. Is it still big data if it doesn’t use certain tools from the fields of artificial intelligence and machine learning? Probably.

Should narrowly focused industry efforts to glean consumer insight from large datasets be grouped under the same term used to describe the sophisticated and varied things scientists are trying to do? There’s a lot of confusion, and industry experts and scientists often end up talking past one another.

2. Big data is new.

By many accounts, big data exploded onto the scene quite recently. “If wonks were fashionistas, big data would be this season’s hot new color,” a Reuters report quipped last year. In a May 2011 report, the McKinsey Global Institute declared big data “the next frontier for innovation, competition, and productivity.”

It’s true that today we can mine massive amounts of data — textual, social, scientific and otherwise — using complex algorithms and computer power. But big data has been around for a long time. It’s just that exhaustive datasets were more exhausting to compile and study in the days when “computer” meant a person who performed calculations.

Vast linguistic datasets, for example, go back nearly 800 years. Early biblical concordances — alphabetical indexes of words in the Bible, along with their context — allowed for some of the same types of analyses found in modern-day textual data-crunching.

The sciences also have been using big data for some time. In the early 1600s, Johannes Kepler used Tycho Brahe’s detailed astronomical dataset to elucidate certain laws of planetary motion. Astronomy in the age of the Sloan Digital Sky Survey is certainly different and more awesome, but it’s still astronomy.

Ask statisticians, and they will tell you that they have been analyzing big data — or “data,” as they less redundantly call it — for centuries. As they like to argue, big data isn’t much more than a sexier version of statistics, with a few new tools that allow us to think more broadly about what data can be and how we generate it.

3. Big data is revolutionary.

In their new book, “Big Data: A Revolution That Will Transform How We Live, Work, and Think,”Viktor Mayer-Schonberger and Kenneth Cukier compare “the current data deluge” to the transformation brought about by the Gutenberg printing press.

If you want more precise advertising directed toward you, then yes, big data is revolutionary. Generally, though, it’s likely to have a modest and gradual impact on our lives.

When a phenomenon or an effect is large, we usually don’t need huge amounts of data to recognize it (and science has traditionally focused on these large effects). As things become more subtle, bigger data helps. It can lead us to smaller pieces of knowledge: how to tailor a product or how to treat a disease a little bit better. If those bits can help lots of people, the effect may be large. But revolutionary for an individual? Probably not.

4. Bigger data is better.

In science, some admittedly mind-blowing big-data analyses are being done. In business, companies are being told to “embrace big data before your competitors do.” But big data is not automatically better.

Really big datasets can be a mess. Unless researchers and analysts can reduce the number of variables and make the data more manageable, they get quantity without a whole lot of quality. Give me some quality medium data over bad big data any day.

And let’s not forget about bias. There’s a common misconception that throwing more data at a problem makes it easier to solve. But if there’s an inherent bias in how the data are collected or examined, a bigger dataset doesn’t help. For example, if you’re trying to understand how people interact based on mobile phone data, a year of data rather than a month’s worth doesn’t address the limitation that certain populations don’t use mobile phones.

Many interesting questions can be explored with little datasets. Big data has refined our idea of six degrees of separation: Facebook has shown that it’s actually closer to four degrees. But the first six-degrees study was done by psychologist Stanley Milgram using a lot of cleverness and a small number of postcards.

Furthermore, although it’s exciting to have massive datasets with incredible breadth, too often they lack much in the way of a temporal dimension. To really understand a phenomenon, such as a social one, we need datasets with large historical sweep. We need long data, not just big data.

5. Big data means the end of scientific theories.

Chris Anderson argued in a 2008 Wired essay that big data renders the scientific method obsolete: Throw enough data at an advanced machine-learning technique, and all the correlations and relationships will simply jump out. We’ll understand everything.

But you can’t just go fishing for correlations and hope they will explain the world. If you’re not careful, you’ll end up with spurious correlations. Even more important, to contend with the “why” of things, we still need ideas, hypotheses and theories. If you don’t have good questions, your results can be silly and meaningless.

Having more data won’t substitute for thinking hard, recognizing anomalies and exploring deep truths.

Read more from Outlook, friend us on Facebook, and follow us on Twitter

Saturday, August 10, 2013

Too much information

<http://www.aeonmagazine.com/about/>

Our instincts for privacy evolved in tribal societies where walls didn't exist. No wonder we are hopeless oversharers

by Ian Leslie<http://www.aeonmagazine.com/author/ian-leslie/> 2,000 words

http://www.aeonmagazine.com/living-together/do-we-have-a-privacy-instinct-or-are-we-wired-to-share/

inShare20
<mailto:?subject=Too%20much%20information&body=Too%20much%20information%0AOur%20instincts%20for%20privacy%20evolved%20in%20tribal%20societies%20where%20walls%20didn%27t%20exist.%20No%20wonder%20we%20are%20hopeless%20oversharers%0A%0ARead%20now:%20http://www.aeonmagazine.com/living-together/do-we-have-a-privacy-instinct-or-are-we-wired-to-share/>

In October 2012 a woman from Massachusetts called Lindsey Stone went on a work trip to Washington DC, and paid a visit to Arlington National Cemetery, where American war heroes are buried. Crouching next to a sign that said 'Silence and Respect', she raised a middle finger and pretended to shout while a colleague took her photo. It was the kind of puerile clowning that most of us (well me, anyway) have indulged in at some point, and once upon a time, the resulting image would have been noticed only by the few friends or family to whom the owner of the camera showed it. However, this being the era of sharing, Stone posted the photo to her Facebook profile.

Within weeks, a 'Fire Lindsey Stone' page had materialised, populated by commentators frothing with outrage at a desecration of hallowed ground. Anger rained down on Stone's employer, a non-profit that helps adults with special needs. Her employers decided, reluctantly, that Stone and her colleague would have to leave.

More recently, Edward Snowden's revelations about the panoptic scope of government surveillance have raised the hoary spectre of 'Big Brother'. But what Prism's fancy PowerPoint decks and self-aggrandising logo suggest to me is not so much an implacable, omniscient overseer as a bunch of suits in shabby cubicles trying to persuade each other they're still relevant. After all, there's little need for state surveillance when we're doing such a good job of spying on ourselves. Big Brother isn't watching us; he's taking selfies and posting them on Instagram like everyone else. And he probably hasn't given a second thought to what might happen to that picture of him posing with a joint.

Walls are a relatively recent innovation. Members of pre-modern societies happily coexisted while carrying out almost all of their lives in public view

Stone's story is hardly unique. Earlier this year, an Aeroflot air hostess was fired from her job after a picture she had taken of herself giving the finger to a cabin full of passengers circulated on Twitter. She had originally posted it to her profile on a Russian social networking site without, presumably, envisaging it becoming a global news story. Every day, embarrassments are endured, jobs lost and individuals endangered because of unforeseen consequences triggered by a tweet or a status update. Despite the many anxious articles about the latest change to Facebook's privacy settings, we just don't seem to be able to get our heads around the idea that when we post our private life, we publish it.

At the beginning of this year, Facebook launched the drably named 'Graph Search', a search engine that allows you to crawl through the data in everyone else's profiles. Days after it went live, a tech-savvy Londoner called Tom Scott started a blog in which he posted details of searches that he had performed using the new service. By putting together imaginative combinations of 'likes' and profile settings he managed to turn up 'Married people who like prostitutes', 'Single women nearby who like to get drunk', and 'Islamic men who are interested in other men and live in Tehran' (where homosexuality is illegal).

Scott was careful to erase names from the screenshots he posted online: he didn't want to land anyone in trouble with employers, or predatory sociopaths, or agents of repressive regimes, or all three at once. But his findings served as a reminder that many Facebook users are standing in their bedroom naked without realising there's a crowd outside the window. Facebook says that as long as users are given the full range of privacy options, they can be relied on to figure them out. Privacy campaigners want Facebook and others to be clearer and more upfront with users about who can view their personal data. Both agree that users deserve to be given control over their choices.

But what if the problem isn't Facebook's privacy settings, but our own?

A few years ago George Loewenstein, professor of behavioural economics at Carnegie Mellon University in Pittsburgh, set out to investigate how people think about the consequences of their privacy choices on the internet. He soon concluded that they don't.

In one study, Loewenstein and his collaborators asked two groups of students to fill out an online survey about their lives. Everyone received the same questions, ranging from the innocuous to the embarrassing or potentially incriminating. One group was presented with an official-looking website that bore the imprimatur of their university, and were assured that their answers would remain anonymous. The other group filled out the questions on a garishly coloured website on which the question 'How BAD Are U???' was accompanied by a grinning devil. It featured no assurance of anonymity.

Bizarrely, the 'How BAD Are U???' website was much more likely to elicit revealing confessions, like whether a student had copied someone else's homework or tried cocaine. The first set of respondents reacted cautiously to the institutional feel of the first website and its obscurely concerning assurances about anonymity. The second group fell under the sway of the perennial youthful imperative to be cool, and opened up, in a way that could have got them into serious trouble in the real world. The students were using their instincts about privacy, and their instincts proved to be deeply wayward. 'Thinking about online privacy doesn't come naturally to us,' Loewenstein told me when I spoke to him on the phone. 'Nothing in our evolution or culture has equipped us to deal with it.'

When a boy hit puberty, he disappeared into the jungle, returning a man. In today's digital culture this is precisely the stage at which we make our lives most exposed to the public gaze

We might be particularly prone to disclosing private information to a well-designed digital interface, making an unconscious and often unwise association between ease-of-use and safety. For example, a now-defunct website called Grouphug.us solicited anonymous confessions. The original format of the site was a masterpiece of bad font design: it used light grey text on a dark grey background, making it very hard to read. Then, in 2008, the site had a revamp, and a new, easier-to-read black font against a white background was adopted. The cognitive scientists Adam Alter and Danny Oppenheimer gathered a random sample of 500 confessions from either side of the change. They found that the confessions submitted after the redesign were generally far more revealing than those submitted before: instead of minor peccadilloes, people admitted to major crimes. (Facebook employs some of the best web designers in the world.)

This is not the only way our deeply embedded real-world instincts can backfire online. Take our rather noble instinct for reciprocity: returning a favour. If I reveal personal information to you, you're more likely to reveal something to me. This works reasonably well when you can see my face and make a judgment about how likely I am to betray your confidence, but on Facebook it's harder to tell if I'm trustworthy. Loewenstein found that people were much readier to answer probing questions if they were told that others had already answered them. This kind of rule-of-thumb — when in doubt, do what everyone else is doing — works pretty well when it comes to things such as what foods to avoid, but it's not so reliable on the internet. As James Grimmelmann, director of the intellectual property programme at the University of Maryland, puts it in his article 'Facebook and the Social Dynamics of Privacy' (2008): 'When our friends all jump off the Facebook privacy bridge, we do too.'

Giving people more control over their privacy choices won't solve these deeper problems. Indeed, Loewenstein found evidence for a 'control paradox'. Just as many people mistakenly think that driving is safer than flying because they feel they have more control over it, so giving people more privacy settings to fiddle with makes them worry less about what they actually divulge.

Then again, perhaps none of this matters. Facebook's founder Mark Zuckerberg is not the only tech person to suggest that privacy is an anachronistic social convention about which younger generations care little. And it's certainly true that for most of human existence, most people have got by with very little private space, as I found when I spoke to John L Locke, professor of linguistics at Ohio University and the author of Eavesdropping: An Intimate History (2010). Locke told me that internal walls are a relatively recent innovation. There are many anthropological reports of pre-modern societies whose members happily coexisted while carrying out almost all of their lives in public view.

You might argue, then, that the internet is simply taking us back to something like a state of nature. However, hunter-gatherer societies never had to worry about invisible strangers; not to mention nosy governments, rapacious corporations or HR bosses. And even in the most open cultures, there are usually rituals of withdrawal from the arena. 'People have always sought refuge from the public gaze,' Locke said, citing the work of Paul Fejos, a Hungarian-born anthropologist who, in the 1940s, studied the Yagua people of Northern Peru, who lived in houses of up to 50 people. There were no partitions, but inhabitants could achieve privacy any time they wanted by simply turning away. 'No one in the house,' wrote Fejos, 'will look upon, or observe, one who is in private facing the wall, no matter how urgently he may wish to talk to him.'

The need for privacy remains, but the means to meet it — our privacy instincts — are no longer fit for purpose

From the 1960s onwards, Thomas Gregor, professor of anthropology at Vanderbilt University in Nashville, studied an indigenous Brazilian tribe called the Mehinaku, who lived in oval huts with no internal walls, each housing a family of 10 or 12. Mehinaku villagers were expected to remove themselves altogether from the life of the village at important stages of life, such as adolescence. When a boy hit puberty, he disappeared into the jungle, returning a man. In today's digital culture, of course, this is precisely the stage at which we make our lives most exposed to the public gaze.

Grimmelmann thinks the suggestion that we are voluntarily waving goodbye to privacy is nonsense: 'The way we think about privacy might change, but the instinct for it runs deep.' He points out that today's teenagers retain as fierce a sense of their own private space as previous generations. But it's much easier to shut the bedroom door than it is to prevent the spread of your texts or photos through an online network. The need for privacy remains, but the means to meet it — our privacy instincts — are no longer fit for purpose.

Over time, we will probably get smarter about online sharing. But right now, we're pretty stupid about it. Perhaps this is because, at some primal level, we don't really believe in the internet. Humans evolved their instinct for privacy in a world where words and acts disappeared the moment they were spoken or made. Our brains are barely getting used to the idea that our thoughts or actions can be written down or photographed, let alone take on a free-floating, indestructible life of their own. Until we catch up, we'll continue to overshare.

A long-serving New York Times journalist who recently left his post was clearing his desk when he came across an internal memo from 1983 on computer policy. It said that while computers could be used to communicate, they should never be used for indiscreet or potentially embarrassing messages: 'We have typewriters for that.' Thirty years later, and the Kremlin's security agency has concluded that The New York Times IT department was on to something: it recently put in an order for electric typewriters. An agency source told Russia's Izvestiya newspaper that, following the WikiLeaks and Snowden scandals, and the bugging of the Russian prime minister Dmitry Medvedev at the G20 summit in London, 'it has been decided to expand the practice of creating paper documents'.

Its invention enabled us to capture and store our thoughts and memories but, today, the best thing about paper is that it can be shredded.

Published on 7 August 2013

<http://www.aeonmagazine.com/living-together/do-we-have-a-privacy-instinct-or-are-we-wired-to-share/#top><http://www.aeonmagazine.com/living-together/do-we-have-a-privacy-instinct-or-are-we-wired-to-share/#top>

Tuesday, July 30, 2013

Esther Dyson: 3D Fantasies

3D Fantasies

24 July 2013

http://www.project-syndicate.org/print/how-3d-printing-will-change-the-world-by-esther-dyson


NEW YORK – How will 3D printing change the world? Today, you can read about jewelry and custom can openers, much as three decades ago you could have read that the personal computer would enable people to keep their recipes organized. Of course, PCs became much more useful than that. Many entrepreneurs and small businesspeople can now run their entire operations on a computer, and people keep their recipes not only organized, but also online. They also track their workouts, monitor their babies, and amass huge collections of digital friends (for better or worse).

The Internet changed the balance of power between individuals and institutions. It enabled millions of people to have jobs without having bosses. Instead, they have agents – such as TaskRabbit or Amazon Web Services or Uber – who match providers and customers.

I think we will see a similar story with 3D printing, as it grows from a novelty into something useful and disruptive – and sufficiently cheap and widespread to be used for (relatively) frivolous endeavors as well. We will print not just children's playthings, but also human prostheses – bones and even lungs and livers – and ultimately much machinery, including new 3D printers.

So, even as custom-manufactured goods become cheaper and people talk about local manufacture as well as local foods, other goods may get more expensive if we do it right. "Juan got his wife Alice a real wooden chair for her birthday!" you might hear. But their daughter Mika got a reprinted chair made from the same old materials plus a little more, marking her growth from her last birthday. Only the size and the filigree on the back are different, reflecting her new interest in space travel; last year, it was horses.

Like computers and the Internet, 3D printing will affect business and behavior around the world and across industries. Already, there is a growing number of shared 3D printing services, enabling you to print something of your own design or use (a customized version of) designs that you can find in online catalogues or order through 3D design shops.

Over time, these print shops will replace thousands of stores carrying millions of items, some of which sit around for months waiting to be bought. They will print goods using designs from online services that offer designs for both open-source, free-design goods and branded goods that may not seem very distinct except for a logo.

Indeed, branding and intellectual property issues will become increasingly "complicated" for hardware, just as they are now for software and content. Many people will have to shift from controlling design to offering better services to make money, or perhaps band together under a particular brand known for some other quality.

Materials may come to be one such differentiator, as illustrated by a startup called Emerging Objects<http://www.emergingobjects.com/>. As in the world of content and software, new design brands are likely to emerge and die more quickly; the pace of change will increase and it will be harder to stay on top for long.

Outside the world of manufacturing, where mass-produced goods may still have a substantial cost advantage over custom-printed ones, 3D printing will have far greater impact downstream, in the market for spare parts and replacements, where demand is less predictable but more precise. (If you want a widget, any widget will do, but once you have widget 94303, only part 94303A will satisfy you.)

One early example is KeyMe.net, which makes house keys on demand. The user needs one original, which he registers by inserting into the KeyMe kiosk; he can then store that design anonymously in KeyMe's database, with unique access to it via his fingerprints and email (but with no reference to a physical address). Then, when he loses the key or needs a spare, he can get a new copy at any location with a kiosk – of which there will soon be many, the company hopes.

The cost in money (let alone convenience) is a fraction of that for going to an ordinary key maker – especially at the hours when such emergencies usually occur. KeyMe does not actually use 3D printing; it cuts them out of blanks the "traditional" way, but uses the same kind of electronic design representation that a 3D printer would. In fact, I consider it a brilliant forerunner of the overall impact of 3D printing – making the occasional production of cheap copies of a specific item easy and available anywhere, anytime.

Today, for example, many businesses are devoted to managing and storing spare parts. Each location needs to carry thousands of different spare parts because it is not clear which ones will be needed where. But, in the future, if something breaks, you will be able to take it over to the 3D print shop to be reprinted. Better yet, the shop may be able to reuse the materials in your broken part – saving the costs and environmental burden of throwing things away, shipping them somewhere, and so forth.

Consider Apple power cords (the item that I lose most often), which are a huge source of profit for all involved. That will change - hallelujah! Of course, my reduced cost will be someone else's reduced revenue – and not just Apple's.

One big loser in this world will be the freight business (along with junkyards, logistics companies, and centralized recycling operations). When things can be made, used, broken/worn out, and recycled closer to home, the need for transport is reduced dramatically. Recycled materials do not need to be delivered to centralized processing centers and then forwarded to factories. Products will not need to be made in those factories and then shipped to customers or to inventory centers.

Right now, US inventories held by manufacturers, wholesalers, and retailers are valued at around $1.7 trillion<http://www.census.gov/mtis/www/data/pdf/mtis_current.pdf> – or about 10% of annual GDP. This includes many things that cannot be 3D-printed (anytime soon, at least), but it does hint at how much stuff is just sitting around.

In the short run, this means greater efficiency and more and speedier recycling, happening locally rather than centrally. In the long run, 3D printing will allow more efficient use of physical resources and faster diffusion of the best designs, boosting living standards around the world.

Wednesday, July 24, 2013

Metadata Liberation Movement

Politicians don't lose their jobs from accusations of softness on food poisoning or lightning strikes, but some lose their jobs from being accused of softness on terrorism (former Georgia Sen. Max Cleland comes to mind). Hence the massive and expensive exercise in barn-door closing known as airport security. Let us apply this tendency to the national surveillance debate.

"It is not rational to give up massive amounts of privacy and liberty to stay marginally safer from a threat that, however scary, endangers the average American far less than his or her daily commute," writes Conor Friedersdorf in the Atlantic, expressing a common view.

Another kind of loss of liberty comes when our tax dollars are spent on useless programs.

These considerations will be in the background as President Obama's newly resuscitated Privacy and Civil Liberties Oversight Board, an agency whose existence previously was almost a secret itself, prepares recommendations on how and whether to relieve some of the secrecy around now-controversial national surveillance efforts. But one question still isn't being asked: With respect to the famous metadata surveillance, why is it conducted by a secret agency at all?

This is the program that anonymously collects data about electronic transactions—from the time, duration and numbers involved in a phone call, to email connections, to financial transactions—everything but the actual content. It may be that Americans have nothing to fear from computers raking through piles of anonymous data. The threat to liberty and privacy comes only when computers kick up red flags for an actual human being to look at, which now means an employee of an agency necessarily removed from popular oversight.

The biggest problem, then, with metadata surveillance may simply be that the wrong agencies are in charge of it. One particular reason why this matters is that the potential of metadata surveillance might actually be quite large but is being squandered by secret agencies whose narrow interest is only looking for terrorists.

Highway serial killers are enough of a problem that the FBI formed a task force devoted to them, its Highway Serial Killers Initiatives. Instead of finding a suspect and trying to tie him to bodies, could metadata help us quickly find suspects based on the locations of bodies?

Could metadata be used to alert us to the troubled recluse who suddenly starts buying guns and ammunition? Could it be used to raise the cost of organized crime, such as drug smuggling or product counterfeiting or identity theft, which obviously requires elements of organization, which means lots of electronic "transactions"?

One peculiarly bad argument that assails every crime-prevention strategy is that, if someone is determined to commit a crime, he will find a way. People differ in their degree of determination. Economics wouldn't exist if raising the cost of engaging in an activity—whether planning a shooting spree, organizing a drug ring or being a serial killer—didn't lessen our supply of these activities.

"Big data" is only as good as the algorithms used to find out things worth finding out. The efficacy and refinement of big-data techniques are advanced by repetition, by giving more chances to find something worth knowing. Bringing metadata out of its black box wouldn't only be a way to improve public trust in what government is doing. It would be a way to get more real value for society out of techniques that are being squandered on a fairly minor threat.

Bringing metadata out of the black box would open up new worlds of possibility—from anticipating traffic jams to locating missing persons after a disaster. It would also create an opportunity to make big data more consistent with the constitutional prohibition of unwarranted search and seizure. In the first instance, with the computer withholding identifying details of the individuals involved, any red flag could be examined by a law-enforcement officer to see, based on accumulated experience, whether the indication is of interest.

If so, a warrant could be obtained to expose the identities involved. If not, the record could immediately be expunged. All this could take place in a reasonably aboveboard, legal fashion, open to inspection in court when and if charges are brought or—this would be a good idea—a court is informed of investigations that led to no action.

Our guess is that big data techniques would pop up way too many false positives at first, and only considerable learning and practice would allow such techniques to become a useful tool. At the same time, bringing metadata surveillance out of the shadows would help the Googles, Verizons and Facebooks defend themselves from a wholly unwarranted suspicion that user privacy is somehow better protected by French or British or (heavens) Chinese companies from their own governments than U.S. data is from the U.S. government.

Most of all, it would allow these techniques to be put to work on solving problems that are actual problems for most Americans, which terrorism isn't.

A version of this article appeared July 24, 2013, on page A13 in the U.S. edition of The Wall Street Journal, with the headline: Metadata Liberation Movement.

Tuesday, July 16, 2013

GRAND BARGAINS FOR BIG DATA: THE EMERGING LAW OF HEALTH INFORMATION


FRANK PASQUALE
*
ABSTRACT
Health information technology can save lives, cut costs, and expand access to care. But its full promise will only be realized if policymakers broker a "grand bargain" between providers, patients, and administrative agencies. In exchange for subsidizing systems designed to protect intellectual property and secure personally identifiable information, health regulators should have full access to key data those systems collect.
Successful data-mining programs at the Centers for Medicare & Medicaid Services ("CMS") provide one model. By requiring standardized collection of billing data and hiring private contractors to analyze it, CMS pioneered innovative techniques for punishing fraud. Now it must move beyond deterring illegal conduct and move toward data-driven promotion of best practices.
With this aim in mind, CMS is already subsidizing technology but more than money is needed to optimize the collection, analysis, and use of data. Policymakers need to navigate intellectual property and privacy rights skillfully.
They must condition current (and future) government support for providers and insurers on better collection and dissemination of health information.If they succeed, the law of health information might better incorporate public values than information law generally

Saturday, July 13, 2013

Open Data Tools: Turning Data into ŒActionable Intelligence¹

Open Data Tools: Turning Data into ‘Actionable Intelligence’

10 July 2013 by Shannon Bohle, posted in Open Access, Uncategorized

http://www.scilogs.com/scientific_and_medical_libraries/open-data-tools-turning-data-into-actionable-intelligence/

My previous two articles were on open access and open data. They conveyed major changes that are underway around the globe in the methods by which scientific and medical research findings and data sets are circulated among researchers and disseminated to the public. I showed how E-science and ‘big data’ fit into the philosophy of science though a paradigm shift as a trilogy of approaches: deductive, empirical, and computational, which was pointed out, provides a logical extenuation of Robert Boyle's tradition of scientific inquiry involving “skepticism, transparency, and reproducibility for independent verification” to the computational age.

The Honourable Robert Boyle 1627–1691, Experimental Philosopher
(Image credit: Wellcome Library, London). Published with written permission.

First published in 1661,
The Sceptical Chymist: or Chymico-Physical Doubts & Paradoxes, Touching the Spagyrist's Principles Commonly call'd Hypostatical; As they are wont to be Propos'd and Defended by the Generality of Alchymists. Whereunto is præmis'd Part of another Discourse relating to the same Subject was written by Robert Boyle and is the source of the name of the modern field of 'chemistry.'
Image credit: Project Gutenberg.
(Click image to enlarge).

There has been a strong support in the belief that information should be freely available, a tradition that libraries have advocated since the days of Andrew Carnegie’s campaign to build free public libraries in the United States, Canada, the United Kingdom, and elsewhere, beginning in 1883. That policy is now being backed through legislative policy in the US, UK, and the EU mandating that scientific and medical articles and their data be freely accessible to the public when they are a result of taxpayer’s funding. The control over the dissemination of this scientific information seems up for grabs and challenges the traditional model of subscription-based journals as the primary mode by which this information is circulated. Libraries, publishers, commercial entities, as well as government agencies, are all competing against one another for the unfettered access and publication of these “free” materials for the potential revenue streams they will bring in. Supposedly, these revenues will occur through memberships and advertising in the case of journal publishers, continued state and grant funding in the case of libraries, advertising and new product creation in the case of commercial enterprises, as well as job continuation or shifted positions with new duties for civil servants and government contractors through compliance with legislative mandates, Congressional requests, and Presidential directives. Competition, as usual, is also occurring within these sectors amongst one another. Hopefully, all this disruption and competition will lead to valuable new economic growth and benefits for society. Making the published articles and credited data openly available is just the first step. How can that data be cleaned and merged with similar data perhaps from disparate sources? How will it be visualized to make the data clearer? How will the data be stored and preserved so that it is not corrupted over time?

Video credit: Digital Curation Centre. "Managing Research Data."
Produced by Piers Video Production, (Duration: 0:12:36).

This third article on open access and open data evaluates new and suggested tools when it comes to making the most of the open access and open data OSTP mandates. According to an article published in The Harvard Business Review’s “HBR Blog Network,” this is because, as its title suggests, “open data has little value if people can't use it.” Indeed, “the goal is for this data to become actionable intelligence: a launchpad for investigation, analysis, triangulation, and improved decision making at all levels.” Librarians and archivists have key roles to play in not only storing data, but packaging it for proper accessibility and use, including adding descriptive metadata and linking to existing tools or designing new ones for their users. Later, in a comment following the article, the author, Craig Hammer, remarks on the importance of archivists and international standards, “Certified archivists have always been important, but their skillset is crucially in demand now, as more and more data are becoming available. Accessibility—in the knowledge management sense—must be on par with digestibility / 'data literacy' as priorities for continuing open data ecosystem development. The good news is that several governments and multilaterals (in consultation with data scientists and - yep! - certified archivists) are having continuing 'shared metadata' conversations, toward the possible development of harmonized data standards...If these folks get this right, there's a real shot of (eventual proliferation of) interoperability (i.e. a data platform from Country A can 'talk to' a data platform from Country B), which is the only way any of this will make sense at the macro level.”

Image credit: Rock Health. Rock Health is a business incubator started by four Harvard graduates, and has an impressive lineup of partners such as Harvard Medical School, Mayo Clinic, Genentech, and others. They offer many resources, like these free videos, and a four month program for those wishing to create new tools for mobile and health 2.0 initiatives. (Click image to enlarge).

From a business perspective, the management of open data in the health sciences, for example, holds both the potential to reduce losses and increase profits. Preserving, storing, and retrieving data in a manageable fashion, then, will affect not only data consumers but also data producers. According to a National Law Review article published in late June this year, the “increased availability of health care data means more oversight and more litigation” because “data is the lifeblood of health care fraud enforcement efforts” which affects the overall cost structure of service provision. According to a report by the U.S. Federal Bureau of Investigation, “health care fraud costs the country an estimated $80 billion a year…[so] rooting out health care fraud is central to the well-being of both our citizens and the overall economy.” At the same time as cutting fraud and waste, data tools developed by U.S. “digital health startups net[ted] $849M in investments in first half of 2013," and $78M of that was attributed to analytics and 'big data.'

What Percentage of Caregivers
Conduct the Following Online Health-Related Activities?

Additionally, a report by issued in June 2013 by The Pew Research Center and the California HealthCare Foundation indicated that as many as "39% of U.S. adults are caregivers and many navigate health care with the help of technology," but of those, "39% of caregivers manage medications for a loved one; few use tech to do so." Similarly, it reported, "Most caregivers say the internet is helpful to them" and "nine in ten caregivers own a cell phone and one-third have used it to gather health information." It seems, then, that a viable window is open for new open data tools in the area of internet and mobile technologies to provide caregivers with more information about medical tests, medications, and clinical trials using metadata descriptors.

Nature will launch Scientific Data in 2014.

"Scientific Data is a new open-access, online-only publication for descriptions of scientifically valuable datasets. It introduces a new type of content called the Data Descriptor, which will combine traditional narrative content with curated, structured descriptions of research data, including detailed methods and technical analyses supporting data quality. Scientific Data will initially focus on the life, biomedical and environmental science communities, but will be open to content from a wide range of scientific disciplines. Publications will be complementary to both traditional research journals and data repositories, and will be designed to foster data sharing and reuse, and ultimately to accelerate scientific discovery."
Video credit: Nature, "Scientific Data."

In addition to financial rewards, there are also financial incentives sponsored by government agencies, non-profit charities, publishers, and for-profit businesses to develop tools and to create successful commercial projects engaged in data re-use. The Obama administration is offering a “Big Data Research Initiative” backed by $200M for new projects routed through six departments: DARPA, DOE, DOD, HHS/NIH, NSF, and USGS. The deadline to reply to this year’s call for projects involving big data collaborations is September 2, 2013. Interested individuals can send proposals to the Networking and Information Technology Research and Development (NITRD) program (BigDataprojects@nitrd.gov). A detailed description of the requirements can be found on their website. According to a White House press release, Dr. John P. Holdren, Assistant to the President and Director of the White House’s Office of Science and Technology Policy, stated that,

John P. Holdren, Director of the Office of Science and Technology Policy. (Image credit: Wikipedia).

“In the same way that past Federal investments in information-technology R&D led to dramatic advances in supercomputing and the creation of the Internet, the initiative we are launching...promises to transform our ability to use Big Data for scientific discovery, environmental and biomedical research, education, and national security,” This is the second year of the Big Data Initiative, and “the Administration is encouraging multiple stakeholders including federal agencies, private industry, academia, state and local government, non-profits, and foundations, to develop and participate in Big Data innovation projects across the country.”

Victoria Costello (vcostello@plos.org) of PLOS I OPEN FOR DISCOVERY manages the ASAP awards program, which is sponsored by 27 global organizations including Google, PLOS, and the Wellcome Trust. “The ASAP Program will award three top awards of $30,000 each” in October 2013 to recognize the best in creative data “reuse, remixing, [and] repurposing—which enables countless clinical translations and subsequent discoveries based on previously published (OA) research.” Similarly, Eli Lilly will be accepting submissions until October 2, 2013, for a "Clinical Trial Revisualization Design" competition, and is giving away $75K in cash and prizes. Their goal is to encourage "designers and developers to re-imagine clinical trial information in a patient-centric way [because] clinical trial information can often be dense and difficult to digest from a patient’s perspective."

Going back to 2011, before the recent open access and open data mandates, the National Library of Medicine sponsored their own contest, "NLM Show Off Your Apps: Innovative Uses of NLM Information", with 35 entries using their biomedical data. Wondering if databases such as PubMed Central might make use of added metadata from the field of health informatics, and open data elsewhere, I contacted Betsy L. Humphreys (blh@nlm.nih.gov), Deputy Director of the National Library of Medicine, to discuss the feasibility of adding health information metadata tags found in electronic health records (EHRs) to the records in their database, either on their side or on the entrepreneurial side of things.

Betsy L. Humphreys,
Deputy Director,
National Library
of Medicine
(Image credit:
National Library
of Medicine).
Published with
written permission
from Betsy L.
Humphreys.

According to Humphreys, my concept of “connecting EHRs with NLM databases is very sensible,” but directly adding metadata tags is not a very practical approach. Part of the problem in doing this lay in the sheer number of PubMed records, more than 22 million of them, and the fact that they are updated nightly. Another major problem is that whenever the Systematized Nomenclature of Medicine—Clinical Terms (SNOMED CT) or the Logical Observation Identifiers Names and Codes (LOINC) (used to identify lab tests and clinical observations) is updated, some of the records with those tags would need updating. Perhaps the best reason why this approach should not be implemented by the NLM, Humphries informed me, is that within PubMed Central, there is a lot of interoperability, including a correspondence table that includes SNOMED CT using the Unified Medical Language System (UMLS) Metathesaurus, and whenever possible, records are mapped with their synonyms. But, on the entrepreneurial side of things, she said, "there are indeed many EHR vendors who use the MedlinePlus Connect API, which enables the use of SNOMED CT, LOINC, or RxNorm in search arguments, in order to integrate NLM’s data into an electronic health record through a patient portal.” The interview with Humphreys left me with hope that there are indeed ample opportunities for entrepreneurs to create new data tools or physical products that create something new through mashups that combine electronic health records, published medical research records in existing databases like Medline Plus, PubMed, and PubMed Central, as well as other data.

Pictured above is Matthew D. Carmichael delivering a talk about data integrity for archival administration of digital data and other e-records. Information critical to the safekeeping of digital data was discussed at the Fourth Annual Virtual Center for Archives and Records Administration (VCARA) Conference on Digital Stewardship & Knowledge Dissemination in the 21st Century held on May 22, 2013, and hosted in the virtual world Second Life by San Jose State University’s School of Library and Information Science. The conference promoted “the impact of archivists, records administrators and digital curators on the future of knowledge dissemination in the context of cultural heritage institutions.” (Click image to enlarge).

Rating the Quality of Open Data

When working with government data it may be helpful to keep a few key guidelines in mind. The problem is, there are many guidelines. A working group within OpenGovData.org developed "8 Principles of Open Government Data" which are: "1. Data Must Be Complete... 2. Data Must Be Primary... 3. Data Must Be Timely... 4. Data Must Be Accessible... 5. Data Must Be Machine processable... 6. Access Must Be Non-Discriminatory... 7. Data Formats Must Be Non-Proprietary... 8. Data Must Be License-free." This is very similar to the Sunlight Foundation's "Ten Principles for Opening Up Government Information"— "1. Completeness... 2. Primacy... 3. Timeliness 4. Ease of Physical and Electronic Access... 5. Machine readability... 6. Non-discrimination... 7. Use of Commonly Owned Standards...8. Licensing... 9. Permanence... 10. Usage Costs." Open government data initiatives could also be held up to a 5-star rating method, which has been proposed by Tim Berners-Lee, the British computer scientist credited with inventing the World Wide Web:

★          Available on the web (whatever format), but with an open licence to be Open Data
★★        Available as machine-readable structured data (e.g. Excel instead of image scan of a table)
★★★      As (2), plus non-proprietary format (e.g. CSV instead of Excel)
★★★★    All the above, plus use W3C open standards (RDF and SPARQL)
★★★★★  All the above, plus link your data to other people’s data to provide context

The Open Data Institute has created an "Open Data Certificate" for data and rates it against a checklist. Certificates awarded grade data as "Raw: A great start at the basics of publishing open data, Pilot: Data users receive extra support from, and can provide feedback to the publisher, Standard: Regularly published open data with robust support that people can rely on, and Expert: An exceptional example of information infrastructure." In 2009, The White House created a scorecard by which open data can be evaluated according to 10 criteria: "high value data, data integrity, open webpage, public consultation, overall plan, formulating the plan, transparency, participation, collaboration, and flagship initiative." The U.S. government's simple stoplight-like rating system was as follows: green for data that "meets expectations," yellow for data that demonstrates "progress toward expectations," and red for data that "fails to meet expectations." At the other end of the spectrum, there is an exceptionally complex checklist offered by OPQUAST. On May 9, 2013, President Obama issued an Executive Order "Making Open and Machine Readable the New Default for Government Information" wherein a new "Open Data Policy" has just been established and being newly implemented through "Project Open Data" in which there are seven key principles: "public, accessible, described, reusable, complete, timely, and managed post-release." There does not seem to be an associated rating system, however, to evaluate how well the data complies with the principles. Finally, Nature has has set up three criteria for data: firstly, "experimental rigor and technical data quality," secondly, "completeness of the description," and lastly, "integrity of the data files and repository record."

Finding Solutions with Data Analysis and Visualization

Personally, while at Cambridge, I veered from historical preservation a bit and spent some time studying about the preservation of modern science data. As a librarian and archivist, knowing how to preserve scientific and and medical history meant learning about some of the standard file formats in which science data is typically stored, the software tools used to create and analyze scientific data (which may also be required for reproducibility and long-term access to preserved data sets), along some hands-on training to actually use those software tools.

I completed computer training useful in scientific computing through the UCS service, such as "Programming Concepts for Beginners," "Python: Introduction for Absolute Beginners," "Unix Intro," "Unix: Simple Shell Scripting for Scientists," "Programming Concepts-Pattern Matching," "Emacs," "mySQL," and "Condor and CamGRID" used in parallel, distributed, and grid computing. I did not have the opportunity to work with Cambridge’s COSMOS, “the world's first national cosmology supercomputer,” which was “founded in January 1997 by a consortium of leading UK cosmologists, brought together by Stephen Hawking.” But it would have been pretty exciting to use this data-intensive system.

The truth is, many scientists are routinely employed in hacking their own programs to link measuring equipment, data analysis, and visualization together and my best guess is that they would benefit from more end-to-end integrative open source software systems development. While by no means exhaustive, I compiled a list linking to 349 subject specific tools and 123 general tools that are useful for all of this newly available open data. Certainly, some if not all of the tools in this list might have its quality rated according to the eight aforementioned standards.

More than 349 Subject Specific Open Data Tools

• Earth, geology, ecology, climate, and weather sciences tools include Microsoft’s SciScope. Metadata standards include EML (The Ecological Metadata Language) and the ISO 19115:2003 International Standard for Geographic Information. Polymaps is a tool for making dynamic, interactive maps.

Weather Modeling (Image Credit: Wikimedia Commons)

• GoGeo is a UK-based site that provides a listing of 161 free software products, 50 data services, and 10 search portals relating to geographical information. A standard for this field is the CSDGM (Content Standard for Digital Geospatial Metadata).
• Berkeley compiled a good listing of molecular biology resources including databases and tools for protein and nucleotide sequencing as well as model organisms. Foldit is a game that enables players to solve real problems in protein folding through its simulation software. Some of the metadata standards in the biological fields include: ABCD (Access to Biological Collections Data) Schema developed by the (Taxonomic Databases Working Group (TDWG) of Australia and with the International Union of Biological Sciences, Darwin Core, developed in Australia by TDWG for natural history specimens and observations, Genome Metadata, MIGS/MIMS (Minimum Information About a (Meta)Genome Sequence), and FGDC. GenBank Flat File Format is a standardized format for biological data.
• In astronomy, NASA maintains the FITS file format standard and documentation. A new metadata astronomy thesaurus to provide better linking among scholarly astronomy articles, called the Unified Astronomy Thesaurus (UAT), was recently released. The World Wide Telescope (WWT) “is an application that runs in Windows that utilizes images and data stored on remote servers enabling you to explore some of the highest resolution imagery of the universe available in multiple wavelengths.” An international consortium comprised of the Spitzer Science Center, ESA/Hubble, California Academy of Sciences, IPAC/IRSA, and the University of Arizona established a metadata standard for astronomy called the Astronomy Visualization Metadata Standard.
• CERN developed ROOT, a tool for ‘big data’ analysis in physics, and CASTOR (Cern Advanced STORage Manager) that is written in Scientific Linux for storage management of ‘big data’ physics files.
• Chemistry tools include PubChemChemSpider, Chemical Markup Language (CML), as well as eBank UK, a digital repository for crystallographic data.

Neurological Modeling (Image credit: By Polygon data were generated by Database Center for Life Science(DBCLS)[2]. (Polygon data are from BodyParts3D[1]) [CC-BY-SA-2.1-jp via Wikimedia Commons].

• In medicine, several online tools allow patients to better manage their health care through a Personal Health Record (PHR) including: Microsoft HealthVault, PatientsLikeMe, getHealtZ, onpatient, WebMD Health Record, and Patient Ally. Other tools include EM data analysis and visualization of the brain like NeuroTrace. CARMEN (Code, Analysis, Repository and Modeling for e-Neuroscience) in the UK is a virtual neurophysiology laboratory. Amira was designed for data analysis and visualization in the life sciences along with MeVisLab for medical image processing and visualization. ClinicalTrials.gov Protocol Data Element Definitions and Infobuttons are useful when working with data in the health sciences. Standard tools include PubMed, PubMed Central, MedlinePlus, and GenBank.

123 General Open Data Tools (Organized by Topic)

Big Data: BigQuery, BOINC, Hadoop, HTCondor (formerly Condor), MapReduce, mrjob, NoSQL database, OGSA (Open Grid Services Architecture), Platfora. (9 Tools).

Data Analysis and Calculation, Data Mining, Graphics Plotting, Image Processing, Data Visualization and Simulation, Reports: Avizo, Cave5D, ELKI, Excel, Feko, Fiji, ggplot2, GIMP, Gnuplot, GoogleVis, Graph, IDL, ImageJ, JMP, KaleidaGraph, MATLAB, Multiphysics, NumPy, Omniscope, OpenCV, OpenGL, OpenGL Vizserver, Open Office, OriginLab, ParaView, PCA/PLS plots, Photoshop, RGGobi, SCIDAVIS, SCILAB, SciPy, SenseWeb, Simulink, Stata, Tableplot, Tabplots, VisIt, Wolfram Alpha, Word. (39 tools).

Data Integrity, Data Recovery, Digital Forensics:  AIR (Automatic Image and Restore), Autopsy, CRC-32, dc3dd, dcfldd, FITools, FTK Imager, Guymager, LOCKSS Box, MagicRescue, OSFClone, OSForensic, PhotoRec, PyFlag, SCALPEL (Source Code Analysis, Libre and Portable Library), Sleuthkit. (16 tools).

Data Management Planning and Audit Assessment: CARDIO (Collaborative Assessment of Research Data Infrastructure and Objectives), DMP Online, DRAMBORA (Digital Repository Audit Method Based On Risk Assessment), DROID (Digital Record Object Identification). (4 tools).

Data Storage or Repository: arXiv, CEPH, D-Space: RTI and ControlDesk, DataFlow / DataBank, Dryad. (5 tools).

File Sharing, Identification, and Management: CollateBox, FigShare, Libmagic, TrlD File Identifier for .NET. (4 tools).

Instrument Control: LabVIEW. (1 tool).

Metadata Standards, Protocols, Preservation Formats, and Registries: AGLS, AGRkMS (Australian Government Recordkeeping Metadata Standard), DCAT (Data Catalog Vocabulary),DOI (Digital Object Identifier), ISO 15386 DCMI (The Dublin Core Metadata Initiative), EAC-CPF (Encoded Archival Context -Corporate Bodies, Persons, and Families), EAD (Encoded Archival Description),
e-GMS (e-Government Metadata Standard) 1.0-2002, GDFR (Global Digital Format Registry), ISO 8:1977 Publishing, ISO 215:1986 Publishing, ISO 23081, ISO International Standards on Archives and Records Management, ISO IT Applications in Science, ISO Medical Science and Health Care Facilities, ISO Natural Sciences (07), JHOVE, JHOVE2, METS (Metadata Encoding and Transmission Standard), MODS (Metadata Object Description Schema), NISO Z39.87-2006 Metadata for Images in XML, OAI-PMH (Open Archives Initiative Protocol for Metadata Harvesting), OData (Open Data Protocol ), PREMIS (Preservation Metadata: Implementation Strategies), PRONOM, RAD (Rules for Archival Description), RDF(Resource Description Framework), RIR (Representation Information Registry Repository), SKOS (Simple Knowledge Organisation System) Core, Standard Archive Format for Europe (SAFE), UDFR (Universal Digital Format Registry),
XML Formatted Data Unit (XFDU), XMP. (33 tools).

Programming: C, C++, d3.js, Fortran, Java, JSON (JavaScript Object Notation), Octave, PROLOG, Python, R, SQL, VBA. (12 tools).

Will ‘Actionable Intelligence’ Ultimately
be a Result of ‘Artificial Intelligence’
?

Despite the fact that my research proposal to enter the computer science graduate program at Cambridge (below) was not accepted, as a proposal, I thought it raised some relevant points about the future applications of artificial intelligence and machine learning for integration and analysis of science data across platforms or in the cloud:

Methodological Approaches in Artificial Intelligence
for the Reuse of "Big Data" in Scientific Computing

Preservation of the science cyberinfrastructure requires an understanding of long-term and archival digital storage formats, data provenance, metadata schemas, and storage repositories essential for both reproducibility and accountability of scientific results and reuse of existing science data sets. These two factors drive mandated availability of prepublication data sets[1] in scientific research projects generated from governmentally-funded agencies like the NSF[2], NIH[3], MRC[4] and Wellcome Trust[5]. Rules mandating data reuse hold the potential to optimize research fiscal spending, spur unprecedented new scientific discoveries in digital repositories by finding relationships across differently funded projects, and promote greater scientific collaboration amongst geographically separated institutions. Yet, in a 2005 symposium on digital biology, additional problems cited “bottlenecks” that “occur owing to our limited capacity to control quality and integrate data from myriad sources, to share data across multiple tasks and to exchange data among different people and organizations...Problems with data integration affect all data tasks, including semantic interpretation, data representation, modeling, data storage and query."[6] Despite the fact funders require data set policies, they leave standards to individual institutions; concurrency of metadata standards within and across disciplines has not occurred, though some advocate RDF for semantic interoperability. Legacy data provide substantial difficulty in updating to new standards and backlogs exist: "Even if standards to facilitate data discoverability, access, and use were to be introduced in the future ... huge problems ... applying ... standards retrospectively to the huge amounts of existing data that do not conform to them [was forseen]."[7] Another problem is that achieving scalability requires supervised, semi-supervised or autonomous automation. I would like to explore the use of artificial intelligence, establishing and monitoring probabilistic relationships across legacy and new linked and/or unlinked datasets housed within distributed or cloud-based data repositories, as a more robust way to search legacy data and data with missing or minimal metadata. This approach may hold the potential for innovative breakthrough discoveries in fields such as basic biological research, clinical medical research, genomics, proteomics, climate modeling, and computational astrophysics.

Sources:

1. Toronto International Data Release Workshop Authors. (9 September 2009). Nature 461, 168-170 doi:10.1038/461168a
2.National Science Foundation. (January 2011). Chapter II - proposal preparation instructions. Retrieved from http://www.nsf.gov/pubs/policydocs/pappguide/nsf11001/gpg_2.jsp#dmp
3. National Institutes of Health. (April 17, 2007). NIH data sharing policy. Retrieved from http://grants.nih.gov/grants/policy/data_sharing/
4.Medical Research Council. (September 2011). MRC policy on research data-sharing. Retrieved from http://www.mrc.ac.uk/Ourresearch/Ethicsresearchguidance/datasharing/Policy/index.htm
5. Wellcome Trust. (August 2010). Policy on data management and sharing. Retrieved from http://www.wellcome.ac.uk/About-us/Policy/Policy-and-position-statements/WTX035043.htm
6. Morris, R. W., et al. (2006). Digital biology: an emerging and promising discipline. Trends in Biotechnology, doi:10.1016/j.tibtech.2005.01.005
7. Smithsonian Institution. Office of Policy Analysis. (March 2011) Sharing Smithsonian digital scientific research data from biology. Retrieved from http://www.si.edu/content/opanda/docs/Rpts2011/DataSharingFinal110328.pdf

The future of AI lies in the imagination. Fortunately, ideas and concepts can be simulated even if it is not yet possible to actualize them in the real world. In my award-winning Mars environment simulation, Curiosity AI, an embodied artificial intelligent agent could search through academic databases in various subjects like astronomy and provide answers, and robot equipment could perform in-situ analysis of data. The humanoid robot could command and operate expert AI systems and launch a swarm designed to take temperature and pressure readings in the environment. The data I used in my expert system for the Phoenix Lander were raw archived data from the NASA mission (because that was what I had available through the Mars Data Archive), but the concept was designed with the idea in mind that AI would enable the processing and calculation of real-time in-situ mission data directly from instruments on Mars.

The amount of raw scientific and medical data is swelling beyond what can be viewed and interpreted by individuals without the intermediary aid of machines filtering and computing that data deluge for us. As such, the paradigm shift introducing the computational approach mentioned at the beginning of this article may become Kuhnian, sweeping away and gaining domination over the two older scientific approaches.

I can only conclude this article on open data tools by saying that when attempting to be knowledgeable about many things, it usually results in becoming an expert at nothing. Find the tools that work best for you and stick with them; however, every now and again it is good to see what new tools might be out there, and now is that time.

If you enjoyed this three-part series and would like more information, NISO will be conducting two webinars later this year on “Research Data Curation:” Part 1: E-Science Librarianship (September 11, 2013) and Part 2: Libraries and Big Data (September 18, 2013).

Disclaimer: I own a startup company (not mentioned here) related to computer technology, primarily engaged in library and archives consulting, computer technology R&D, 3-D modeling and simulation, and artificial intelligence research.