November 12, 2013 - Fact Sheet: Progress by Federal Agencies: Data to Knowledge to Action
November 12, 2013 - Fact Sheet: New Announcements: Data to Knowledge to Action
November 12, 2013 - Press Release: Data to Knowledge to Action Event
November 12, 2013 - Fact Sheet: Progress by Federal Agencies: Data to Knowledge to Action
November 12, 2013 - Fact Sheet: New Announcements: Data to Knowledge to Action
November 12, 2013 - Press Release: Data to Knowledge to Action Event
IT can be disconcerting to learn what, not to mention how much, marketers know about us. Consider a consumer like Scott E. Howe.
The Acxiom Corporation, a marketing technology company that has amassed details on the household makeup, financial means, shopping preferences and leisure pursuits of a majority of adults in the United States, knows that Mr. Howe is 45, married with children, the owner of a house in the 2,500-square-foot range, and is interested, among other things, in tennis, domestic travel, cooking, crafts, sweepstakes and contests. Those intimate details, Mr. Howe says, are entirely accurate.
“I am crazy about that stuff,” he says of the sweepstakes and contests.
Mr. Howe is one of the first Americans to get a detailed glimpse of his own marketing profile because he happens to be the chief executive of Acxiom. But most consumers never learn the specific pieces of information that have been compiled about them by marketers.
That is about to change. Acxiom, one of the most secretive and prolific collectors of consumer information, is embarking on a novel public relations strategy: openness. On Wednesday, it plans to unveil a free Web site where United States consumers can view some of the information the company has collected about them, just as Mr. Howe did.
The data on the site, called AbouttheData.com, includes biographical facts, like education level, marital status and number of children in a household; homeownership status, including mortgage amount and property size; vehicle details, like the make, model and year; and economic data, like whether a household member is an active investor with a portfolio greater than $150,000. Also available will be the consumer’s recent purchase categories, like plus-size clothing or sports products; and household interests like golf, dogs, text-messaging, cholesterol-related products or charities.
Each entry comes with an icon that visitors can click to learn about the sources behind the data — whether self-reported consumer surveys, warranty registrations or public records like voter files. The program also lets people correct or suppress individual data elements, or to opt out entirely of having Acxiom collect and store marketing data about them.
With about $1.1 billion in revenue in its 2013 fiscal year, Acxiom is a leading player in an industry called data brokerage. The company collects, stores, analyzes and sells consumer data with the aim of helping its clients — including well-known banks, credit card issuers, insurance companies, department stores and carmakers — tailor marketing to their most valuable current customers or identify new customers.
A credit card issuer, for instance, could ask Acxiom to help aim a campaign for elite-level cards with concierge services at people above a certain income who live in certain suburbs or drive luxury cars. To do that, Acxiom, like many of its competitors, often uses its own proprietary classification system to segment consumers into socioeconomic marketing categories, like “Frugal Families” or “McMansions and Minivans.”
Some federal regulators and privacy advocates warn that this kind of data-mining could be used to aim at consumers vulnerable to predatory lending practices, for instance, or to favor certain high-value consumers with instant, attentive customer service while relegating other people to interminable wait time.
Mr. Howe says he wants to counter such fears by making industry practices more transparent. A former Microsoft executive, he came to Acxiom as C.E.O. in 2011, bringing the online industry’s enthusiasm for data sharing to what had been a hermetic company.
“We are not going to get anywhere by hiding,” he said in a recent interview at Acxiom’s headquarters in Little Rock, Ark. “You have to make things visible.”
But AbouttheData.com is as much ruthlessly pragmatic as idealistic. Mr. Howe recognizes that regulation of his industry may be coming and that it’s better for Acxiom to be seen as a part of the solution than a part of the problem.
ONE afternoon in late August, Mr. Howe sat in an executive conference room at Acxiom’s headquarters overlooking the Arkansas River, demonstrating a version of AbouttheData.com that was still a work in progress. Having filled out an identity verification form that asked for his name, birth date, address and the last four digits of his Social Security number, he landed on a page that gave him a choice of six data categories to examine.
Visitors who log in may be surprised at the volume of information that may be available and the detailed picture it can give of their personal lives. The household interest section, for instance, listed Mr. Howe as interested in health and medical issues (he subscribes to health industry trade journals and founded a site called Health123.com); crafts (he periodically works with stained glass); woodworking (he paid for his undergraduate education at Princeton in part by working as an apprentice carpenter); tennis (he was on his high school team); gardening (his wife subscribes to Fine Gardening magazine); and “religious/inspirational.”
“I don’t know how inspirational I am,” Mr. Howe said. “I am Methodist. My uncle is a Methodist preacher. I go to church very regularly.”
But consumers, he said, should not expect all information to be current or correct. For instance, the site listed Mr. Howe as the father of two; in fact, he is the father of three. It had also pegged him as Italian, but he is actually of Norwegian descent. (The system predicts likely ethnicity based on surname and is clearly imperfect.)
The home section, meanwhile, which listed such details as the year his house was built and its estimated market value, had incorrect information about his mortgage. “I don’t have a loan on my house anymore. It’s drawing on old data,” Mr. Howe explained. “That’s one I would absolutely go in and change.”
If a personal detail is corrected on the site, the new entry will appear with an aside noting the previous, incorrect entry, letting consumers see what they amended. Mr. Howe acknowledged that the system was fallible because Acxiom obtains information from many different suppliers, and the latest data is not always available in its databases. He said he couldn’t predict how Acxiom’s clients might react to a system that lets consumers update profiles and perhaps fictionalize them, or opt out altogether from Acxiom’s marketing database.
“What happens if a flock of people who are 45 decide to be 39?” Mr. Howe asked. “What happens if 20 percent of the American population decides to opt out? It would be devastating for our business.”
If past consumer behavior is any sign, the number of opt-outs isn’t likely to be high. For instance, Forrester Research reported this year that only 18 percent of Web users had activated an option in their browsers, called Do Not Track, that informs sites and ad networks that they don’t want their browsing histories tracked for marketing purposes.
“It’s a little bit of a risk,” Mr. Howe added of the opt-out provision. “But I feel it’s the right thing to do.”
It may be both the right and the timely thing. The new consumer site should help Acxiom get out in front of potential regulation — at a time when the company is about to introduce a more powerful consumer-targeting engine for its corporate clients that may well set off an outcry from privacy advocates. Last year, some members of Congress opened inquiries into the business practices of data brokers in light of an investigation by The New York Times.
Unlike consumer reporting agencies, which are required by the federal Fair Credit Reporting Act to show consumers free copies of their credit reports every 12 months and let them correct errors, information resellers like Acxiom aren’t required to share marketing data with consumers and allow corrections. But some legislators and regulators believe that they should be.
“Citizens don’t know what of our personal information is on file or how it is being used,” Julie Brill, a member of the Federal Trade Commission, wrote in an op-ed article in The Washington Post in August, asking companies like Acxiom to make their practices more transparent. She added: “This frames the fundamental challenge to consumer privacy in the online marketplace: our loss of control over our most private and sensitive information.”
Industry representatives say it is unnecessary to show consumers their marketing data because, they say, reputable companies limit that data’s use to innocuous purposes; they also contend that such consumer services would be too technically challenging and expensive to develop. In an open letter to Ms. Brill responding to her article, the Direct Marketing Association, a trade group, said consumer access programs “would lead to more fraud and limit the efficacy of companies and data.”
Acxiom’s new site seems calculated to allay regulators’ concerns, at least in part, and to challenge the industry status quo at the same time.
“You may be surprised to know that we are in favor of heightened industry regulation, but we want to make sure we have a voice in the process,” Mr. Howe said. Aboutthedata.com is Acxiom’s bid to have a say in any legislative or regulatory developments. “If we are on our front foot, if we innovate and we are learning,” he said, “we think that earns us a seat at the table.”
MR. HOWE calls himself a “data geek” and believes that every business decision could be made better with “the intelligent application of data.” He uses that strategy to inform his own pursuits.
When he was new to Little Rock and wanted to find great barbecue, he canvassed Acxiom employees by e-mail and received several hundred responses. He visited the 10 most recommended places and later sent his rankings, with the reasons behind them, to employees. (The winner was the Whole Hog Cafe.)
As Mr. Howe envisioned an Acxiom consumer portal, he used a similar method, researching how other industries have responded to consumer concerns about lack of transparency. He saw how Subway, the sandwich chain, made nutrition information clearer and introduced a customer, Jared Fogle, to personify healthier menu options. Mr. Howe also consulted executives at credit reporting bureaus for advice on how a business-to-business industry could create a service for consumers.
“Everybody said the hardest thing was changing the culture,” he said. “Take Acxiom. For 40 years, we talked about our three constituencies: our shareholders, our associates and our clients. But there’s a fourth constituency: consumers.”
Even so, Aboutthedata.com is a self-serving endeavor, promoting Acxiom’s take on data-powered marketing to consumers. In fact, Mr. Howe sees the potential for developing a consumer business at Acxiom in which people could customize the kinds of advertising they want to see by selecting the activities and brands that interest them via AbouttheData.com. Consumers are going to receive ads no matter what, he said, so they may as well elect to receive pitches for stuff they enjoy.
Acxiom plans to add more data elements to the site regularly. Eventually, Mr. Howe said, the site may even ask people for more information about themselves in exchange for special services, online subscriptions or discounts. “Could this be a new utility for consumers, having a user agent where they can choose the information they want to share?” he said.
Although the site shows visitors a few facts that some might consider sensitive, like race and ethnicity, it initially omits, at least in the version I saw, intimate references — like “gambling,” “senior needs,” “smoker in the household” and “adult with wealthy parent” — that Acxiom markets to corporate clients but that might discomfit consumers if they knew they were for sale. (Acxiom said that the site includes the “core” facts it has collected about consumers, but that it might add “derived” data, like propensity for gambling, at a later date.)
This kind of anodyne presentation of data-mining, says Joseph Turow, a professor at the Annenberg School for Communication at the University of Pennsylvania, could prompt people to collude in their own surveillance by perfecting their profiles. That would improve the quality and resale value of the data for Acxiom, he says, perhaps to consumers’ detriment.
“Bits of Acxiom data that seem totally benign on their own could be matched with other data and used in ways that consumers don’t want it to be used,” said Professor Turow, who had not yet seen the new Acxiom site.
Privacy advocates like Professor Turow warn that this kind of system could influence consumers to provide more information than is good for them. “It’s just a game they are playing for marketing purposes and to make regulators feel better,” he said of the Acxiom project.
But Mr. Howe said Acxiom planned to solicit and respond to feedback once the site opens. “Some people are going to say, ‘That’s not nearly good enough,’ ” he said. “This is a first step.”
Data Monetization in the Age of Big Data
Source: Accenture
The volume and richness of the data now uniquely accessible to mobile providers—whether in the form of transactions, inquiries, text messages or tweets, GPS locations or live video feeds—offers a veritable gold mine of insights and applications. And even as mobile phones have become the primary device through which consumers get their information, those very same devices have begun to facilitate new types of information, including extremely precise, real-time, geolocation information.
Not surprisingly, operators today are talking about when and how to tap into this data and what to do with it. In particular, they want to know how to monetize it: how to sort, analyze and manipulate the data and put it to use. This holds true not only for internal applications, but increasingly for building new revenue streams or collaborating on external applications with third parties as well.
How can mobile operators best approach this new territory? What are the opportunities and challenges? And how can operators shape new business models to monetize their Big Data?
Direct link to document (PDF; 4.6 MB)
The astounding rate of growth would make any parent proud. There were 30 billion gigabytes of video, e-mails, Web transactions and business-to-business analytics in 2005. The total is expected to reach more than 20 times that figure in 2013, with off-the-charts increases to follow in the years ahead, according to Cisco, the networking giant.
How much data is that? Cisco estimates that in 2012, some two trillion minutes of video alone traversed the Internet every month. That translates to over a million years per week of everything from video selfies and nannycams to Netflix downloads and “Battlestar Galactica” episodes.
What is sometimes referred to as the Internet’s first wave — say, from the 1990s until around 2005 — brought completely new services like e-mail, the Web, online search and eventually broadband. For its next act, the industry has pinned its hopes, and its colossal public relations machine, on the power of Big Data itself to supercharge the economy.
There is just one tiny problem: the economy is, at best, in the doldrums and has stayed there during the latest surge in Web traffic. The rate of productivity growth, whose steady rise from the 1970s well into the 2000s has been credited to earlier phases in the computer and Internet revolutions, has actually fallen. The overall economic trends are complex, but an argument could be made that the slowdown began around 2005 — just when Big Data began to make its appearance.
Those factors have some economists questioning whether Big Data will ever have the impact of the first Internet wave, let alone the industrial revolutions of past centuries. One theory holds that the Big Data industry is thriving more by cannibalizing existing businesses in the competition for customers than by creating fundamentally new opportunities.
In some cases, online companies like Amazon and eBay are fighting among themselves for customers. But in others — here is where the cannibals enter — the companies are eating up traditional advertising, media, music and retailing businesses, said Joel Waldfogel, an economist at the University of Minnesota who has studied the phenomenon.
“One falls, one rises — it’s pretty clear the digital kind is a substitute to the physical kind,” he said. “So it would be crazy to count the whole rise in digital as a net addition to the economy.”
Robert J. Gordon, a professor of economics at Northwestern University, said comparing Big Data to oil was promotional nonsense. “Gasoline made from oil made possible a transportation revolution as cars replaced horses and as commercial air transportation replaced railroads,” he said. “If anybody thinks that personal data are comparable to real oil and real vehicles, they don’t appreciate the realities of the last century.”
Other economists believe that Big Data’s economic punch is just a few years away, as engineers trained in data manipulation make their way through college and as data-driven start-ups begin hiring. And of course the recession could be masking the impact of the data revolution in ways economists don’t yet grasp. Still, some suspect that in the end our current framework for understanding Big Data and “the cloud” could be a mirage.
“I think it’s conceivable that the data era will be a bust for the things people expect it to be useful for,” said Scott Wallsten, a senior fellow at the Technology Policy Institute and the Georgetown Center for Business and Public Policy. Some entirely new use will have to turn up for data to fulfill its economic potential, he added.
There is no disputing that a wide spectrum of businesses, from e-marketers to pharmaceutical companies, are now using huge amounts of data as part of their everyday business.
Josh Marks is the chief executive of one such company, masFlight, which helps airlines use enormous data sets to reduce fuel consumption and improve overall performance. Although his first mission is to help clients compete with other airlines for customers, Mr. Marks believes that efficiencies like those his company is chasing should eventually expand the global economy.
For now, though, he acknowledges that most of the raw data flowing across the Web has limited economic value: far more useful is specialized data in the hands of analysts with a deep understanding of specific industries. “The promises that are made around the ability to manipulate these very large data sets in real time are overselling what they can do today,” Mr. Marks said.
Some economists argue that it is often difficult to estimate the true value of new technologies, and that Big Data may already be delivering benefits that are uncounted in official economic statistics. Cat videos and television programs on Hulu, for example, produce pleasure for Web surfers — so shouldn’t economists find a way to value such intangible activity, whether or not it moves the needle of the gross domestic product?
In addition, infrastructure investments often take years to pay off in a big way, said Shane Greenstein, an economist at Northwestern University. He cited high-speed Internet connections laid down in the late 1990s that have driven profits only recently. But he noted that in contrast to the Internet’s first wave, which created services like the Web and e-mail, the impact of the second wave — the Big Data revolution — is harder to discern above the noise of broader economic activity.
“It could be just time delay, or it could be that the value just isn’t there,” said Mr. Greenstein, who has studied the competitive success of online businesses in media, advertising and retailing.
Perhaps surprisingly, the parallel most tightly embraced by digital futurists — the rise of the electricity grid — is largely dismissed by those who have studied the history of the subject. The idea is that a ubiquitous Internet will make data and “cloud” computing available anywhere, like electricity through a socket.
The numerical comparisons are tantalizing. As illustrated in “The Electric City,” by Harold L. Platt, the booming quantity and adoption rates of electricity flowing on the Chicago grid in the late 19th and early 20th centuries instantly bring to mind those charts showing data growth today.
Despite those similarities, Mr. Platt, a professor emeritus of history at Loyola University Chicago, said it was unlikely that the revolutions unleashed in manufacturing, domestic life, transportation and high and low society by electricity could ever be matched by the data era. “I’d be hard pressed to quickly draw comparisons,” he said.
But even as Mr. Platt, 68, spoke by cellphone from Chicago, fragments of today’s inescapable data flood found him as he received messages from his grown children. “I have to text them or else they won’t answer me back,” Mr. Platt said gamely. “I’m going with the flow.”
James Glanz is an investigative reporter for The New York Times.
Samuel Arbesman, an applied mathematician and network scientist, is a senior scholar at the Ewing Marion Kauffman Foundation and the author of “The Half-Life of Facts.” Follow him on Twitter: @Arbesman.
by Samuel Arbesman Big data holds the promise of harnessing huge amounts of information to help us better understand the world. But when talking about big data, there’s a tendency to fall into hyperbole. It is what compels contrarians to write such tweets as “Big Data, n.: the belief that any sufficiently large pile of s--- contains a pony.” Let’s deflate the hype.
1. “Big data” has a clear definition.
The term “big data” has been in circulation since at least the 1990s, when it is believed to have originated in Silicon Valley. IBM offers a seemingly simple definition: Big data is characterized by the four V’s of volume, variety, velocity and veracity. But the term is thrown around so often, in so many contexts — science, marketing, politics, sports — that its meaning has become vague and ambiguous.
There’s general agreement that ranking every page on the Internet according to relevance and searching the phone records of every Verizon customer in the United States qualify as applications of big data. Beyond that, there’s much debate. Does big data need to involve more information than can be processed by a single home computer? If so, marketing analytics wouldn’t qualify, and neither would most of the work done by Facebook. Is it still big data if it doesn’t use certain tools from the fields of artificial intelligence and machine learning? Probably.
Should narrowly focused industry efforts to glean consumer insight from large datasets be grouped under the same term used to describe the sophisticated and varied things scientists are trying to do? There’s a lot of confusion, and industry experts and scientists often end up talking past one another.
2. Big data is new.
By many accounts, big data exploded onto the scene quite recently. “If wonks were fashionistas, big data would be this season’s hot new color,” a Reuters report quipped last year. In a May 2011 report, the McKinsey Global Institute declared big data “the next frontier for innovation, competition, and productivity.”
It’s true that today we can mine massive amounts of data — textual, social, scientific and otherwise — using complex algorithms and computer power. But big data has been around for a long time. It’s just that exhaustive datasets were more exhausting to compile and study in the days when “computer” meant a person who performed calculations.
Vast linguistic datasets, for example, go back nearly 800 years. Early biblical concordances — alphabetical indexes of words in the Bible, along with their context — allowed for some of the same types of analyses found in modern-day textual data-crunching.
The sciences also have been using big data for some time. In the early 1600s, Johannes Kepler used Tycho Brahe’s detailed astronomical dataset to elucidate certain laws of planetary motion. Astronomy in the age of the Sloan Digital Sky Survey is certainly different and more awesome, but it’s still astronomy.
Ask statisticians, and they will tell you that they have been analyzing big data — or “data,” as they less redundantly call it — for centuries. As they like to argue, big data isn’t much more than a sexier version of statistics, with a few new tools that allow us to think more broadly about what data can be and how we generate it.
3. Big data is revolutionary.
In their new book, “Big Data: A Revolution That Will Transform How We Live, Work, and Think,”Viktor Mayer-Schonberger and Kenneth Cukier compare “the current data deluge” to the transformation brought about by the Gutenberg printing press.
If you want more precise advertising directed toward you, then yes, big data is revolutionary. Generally, though, it’s likely to have a modest and gradual impact on our lives.
When a phenomenon or an effect is large, we usually don’t need huge amounts of data to recognize it (and science has traditionally focused on these large effects). As things become more subtle, bigger data helps. It can lead us to smaller pieces of knowledge: how to tailor a product or how to treat a disease a little bit better. If those bits can help lots of people, the effect may be large. But revolutionary for an individual? Probably not.
4. Bigger data is better.
In science, some admittedly mind-blowing big-data analyses are being done. In business, companies are being told to “embrace big data before your competitors do.” But big data is not automatically better.
Really big datasets can be a mess. Unless researchers and analysts can reduce the number of variables and make the data more manageable, they get quantity without a whole lot of quality. Give me some quality medium data over bad big data any day.
And let’s not forget about bias. There’s a common misconception that throwing more data at a problem makes it easier to solve. But if there’s an inherent bias in how the data are collected or examined, a bigger dataset doesn’t help. For example, if you’re trying to understand how people interact based on mobile phone data, a year of data rather than a month’s worth doesn’t address the limitation that certain populations don’t use mobile phones.
Many interesting questions can be explored with little datasets. Big data has refined our idea of six degrees of separation: Facebook has shown that it’s actually closer to four degrees. But the first six-degrees study was done by psychologist Stanley Milgram using a lot of cleverness and a small number of postcards.
Furthermore, although it’s exciting to have massive datasets with incredible breadth, too often they lack much in the way of a temporal dimension. To really understand a phenomenon, such as a social one, we need datasets with large historical sweep. We need long data, not just big data.
5. Big data means the end of scientific theories.
Chris Anderson argued in a 2008 Wired essay that big data renders the scientific method obsolete: Throw enough data at an advanced machine-learning technique, and all the correlations and relationships will simply jump out. We’ll understand everything.
But you can’t just go fishing for correlations and hope they will explain the world. If you’re not careful, you’ll end up with spurious correlations. Even more important, to contend with the “why” of things, we still need ideas, hypotheses and theories. If you don’t have good questions, your results can be silly and meaningless.
Having more data won’t substitute for thinking hard, recognizing anomalies and exploring deep truths.
Read more from Outlook, friend us on Facebook, and follow us on Twitter.
Politicians don't lose their jobs from accusations of softness on food poisoning or lightning strikes, but some lose their jobs from being accused of softness on terrorism (former Georgia Sen. Max Cleland comes to mind). Hence the massive and expensive exercise in barn-door closing known as airport security. Let us apply this tendency to the national surveillance debate.
"It is not rational to give up massive amounts of privacy and liberty to stay marginally safer from a threat that, however scary, endangers the average American far less than his or her daily commute," writes Conor Friedersdorf in the Atlantic, expressing a common view.
Another kind of loss of liberty comes when our tax dollars are spent on useless programs.
These considerations will be in the background as President Obama's newly resuscitated Privacy and Civil Liberties Oversight Board, an agency whose existence previously was almost a secret itself, prepares recommendations on how and whether to relieve some of the secrecy around now-controversial national surveillance efforts. But one question still isn't being asked: With respect to the famous metadata surveillance, why is it conducted by a secret agency at all?
This is the program that anonymously collects data about electronic transactions—from the time, duration and numbers involved in a phone call, to email connections, to financial transactions—everything but the actual content. It may be that Americans have nothing to fear from computers raking through piles of anonymous data. The threat to liberty and privacy comes only when computers kick up red flags for an actual human being to look at, which now means an employee of an agency necessarily removed from popular oversight.
The biggest problem, then, with metadata surveillance may simply be that the wrong agencies are in charge of it. One particular reason why this matters is that the potential of metadata surveillance might actually be quite large but is being squandered by secret agencies whose narrow interest is only looking for terrorists.
Highway serial killers are enough of a problem that the FBI formed a task force devoted to them, its Highway Serial Killers Initiatives. Instead of finding a suspect and trying to tie him to bodies, could metadata help us quickly find suspects based on the locations of bodies?
Could metadata be used to alert us to the troubled recluse who suddenly starts buying guns and ammunition? Could it be used to raise the cost of organized crime, such as drug smuggling or product counterfeiting or identity theft, which obviously requires elements of organization, which means lots of electronic "transactions"?
One peculiarly bad argument that assails every crime-prevention strategy is that, if someone is determined to commit a crime, he will find a way. People differ in their degree of determination. Economics wouldn't exist if raising the cost of engaging in an activity—whether planning a shooting spree, organizing a drug ring or being a serial killer—didn't lessen our supply of these activities.
"Big data" is only as good as the algorithms used to find out things worth finding out. The efficacy and refinement of big-data techniques are advanced by repetition, by giving more chances to find something worth knowing. Bringing metadata out of its black box wouldn't only be a way to improve public trust in what government is doing. It would be a way to get more real value for society out of techniques that are being squandered on a fairly minor threat.
Bringing metadata out of the black box would open up new worlds of possibility—from anticipating traffic jams to locating missing persons after a disaster. It would also create an opportunity to make big data more consistent with the constitutional prohibition of unwarranted search and seizure. In the first instance, with the computer withholding identifying details of the individuals involved, any red flag could be examined by a law-enforcement officer to see, based on accumulated experience, whether the indication is of interest.
If so, a warrant could be obtained to expose the identities involved. If not, the record could immediately be expunged. All this could take place in a reasonably aboveboard, legal fashion, open to inspection in court when and if charges are brought or—this would be a good idea—a court is informed of investigations that led to no action.
Our guess is that big data techniques would pop up way too many false positives at first, and only considerable learning and practice would allow such techniques to become a useful tool. At the same time, bringing metadata surveillance out of the shadows would help the Googles, Verizons and Facebooks defend themselves from a wholly unwarranted suspicion that user privacy is somehow better protected by French or British or (heavens) Chinese companies from their own governments than U.S. data is from the U.S. government.
Most of all, it would allow these techniques to be put to work on solving problems that are actual problems for most Americans, which terrorism isn't.
A version of this article appeared July 24, 2013, on page A13 in the U.S. edition of The Wall Street Journal, with the headline: Metadata Liberation Movement.
My previous two articles were on open access and open data. They conveyed major changes that are underway around the globe in the methods by which scientific and medical research findings and data sets are circulated among researchers and disseminated to the public. I showed how E-science and ‘big data’ fit into the philosophy of science though a paradigm shift as a trilogy of approaches: deductive, empirical, and computational, which was pointed out, provides a logical extenuation of Robert Boyle's tradition of scientific inquiry involving “skepticism, transparency, and reproducibility for independent verification” to the computational age.
The Honourable Robert Boyle 1627–1691, Experimental Philosopher
(Image credit: Wellcome Library, London). Published with written permission.
First published in 1661,
The Sceptical Chymist: or Chymico-Physical Doubts & Paradoxes, Touching the Spagyrist's Principles Commonly call'd Hypostatical; As they are wont to be Propos'd and Defended by the Generality of Alchymists. Whereunto is præmis'd Part of another Discourse relating to the same Subject was written by Robert Boyle and is the source of the name of the modern field of 'chemistry.'
Image credit: Project Gutenberg.
(Click image to enlarge).
There has been a strong support in the belief that information should be freely available, a tradition that libraries have advocated since the days of Andrew Carnegie’s campaign to build free public libraries in the United States, Canada, the United Kingdom, and elsewhere, beginning in 1883. That policy is now being backed through legislative policy in the US, UK, and the EU mandating that scientific and medical articles and their data be freely accessible to the public when they are a result of taxpayer’s funding. The control over the dissemination of this scientific information seems up for grabs and challenges the traditional model of subscription-based journals as the primary mode by which this information is circulated. Libraries, publishers, commercial entities, as well as government agencies, are all competing against one another for the unfettered access and publication of these “free” materials for the potential revenue streams they will bring in. Supposedly, these revenues will occur through memberships and advertising in the case of journal publishers, continued state and grant funding in the case of libraries, advertising and new product creation in the case of commercial enterprises, as well as job continuation or shifted positions with new duties for civil servants and government contractors through compliance with legislative mandates, Congressional requests, and Presidential directives. Competition, as usual, is also occurring within these sectors amongst one another. Hopefully, all this disruption and competition will lead to valuable new economic growth and benefits for society. Making the published articles and credited data openly available is just the first step. How can that data be cleaned and merged with similar data perhaps from disparate sources? How will it be visualized to make the data clearer? How will the data be stored and preserved so that it is not corrupted over time?
Video credit: Digital Curation Centre. "Managing Research Data."
Produced by Piers Video Production, (Duration: 0:12:36).
This third article on open access and open data evaluates new and suggested tools when it comes to making the most of the open access and open data OSTP mandates. According to an article published in The Harvard Business Review’s “HBR Blog Network,” this is because, as its title suggests, “open data has little value if people can't use it.” Indeed, “the goal is for this data to become actionable intelligence: a launchpad for investigation, analysis, triangulation, and improved decision making at all levels.” Librarians and archivists have key roles to play in not only storing data, but packaging it for proper accessibility and use, including adding descriptive metadata and linking to existing tools or designing new ones for their users. Later, in a comment following the article, the author, Craig Hammer, remarks on the importance of archivists and international standards, “Certified archivists have always been important, but their skillset is crucially in demand now, as more and more data are becoming available. Accessibility—in the knowledge management sense—must be on par with digestibility / 'data literacy' as priorities for continuing open data ecosystem development. The good news is that several governments and multilaterals (in consultation with data scientists and - yep! - certified archivists) are having continuing 'shared metadata' conversations, toward the possible development of harmonized data standards...If these folks get this right, there's a real shot of (eventual proliferation of) interoperability (i.e. a data platform from Country A can 'talk to' a data platform from Country B), which is the only way any of this will make sense at the macro level.”
Image credit: Rock Health. Rock Health is a business incubator started by four Harvard graduates, and has an impressive lineup of partners such as Harvard Medical School, Mayo Clinic, Genentech, and others. They offer many resources, like these free videos, and a four month program for those wishing to create new tools for mobile and health 2.0 initiatives. (Click image to enlarge).
From a business perspective, the management of open data in the health sciences, for example, holds both the potential to reduce losses and increase profits. Preserving, storing, and retrieving data in a manageable fashion, then, will affect not only data consumers but also data producers. According to a National Law Review article published in late June this year, the “increased availability of health care data means more oversight and more litigation” because “data is the lifeblood of health care fraud enforcement efforts” which affects the overall cost structure of service provision. According to a report by the U.S. Federal Bureau of Investigation, “health care fraud costs the country an estimated $80 billion a year…[so] rooting out health care fraud is central to the well-being of both our citizens and the overall economy.” At the same time as cutting fraud and waste, data tools developed by U.S. “digital health startups net[ted] $849M in investments in first half of 2013," and $78M of that was attributed to analytics and 'big data.'
What Percentage of Caregivers
Conduct the Following Online Health-Related Activities?
"A survey by the Pew Research Center and the California HealthCare Foundation finds that 72% of caregivers reported going online to find health information, while 52% have participated in online social activities related to health and 46% have gone online for a diagnosis." (Image credit: iHealthBeat, Pew Research Center / California HealthCare Foundation.
Additionally, a report by issued in June 2013 by The Pew Research Center and the California HealthCare Foundation indicated that as many as "39% of U.S. adults are caregivers and many navigate health care with the help of technology," but of those, "39% of caregivers manage medications for a loved one; few use tech to do so." Similarly, it reported, "Most caregivers say the internet is helpful to them" and "nine in ten caregivers own a cell phone and one-third have used it to gather health information." It seems, then, that a viable window is open for new open data tools in the area of internet and mobile technologies to provide caregivers with more information about medical tests, medications, and clinical trials using metadata descriptors.
Nature will launch Scientific Data in 2014.
"Scientific Data is a new open-access, online-only publication for descriptions of scientifically valuable datasets. It introduces a new type of content called the Data Descriptor, which will combine traditional narrative content with curated, structured descriptions of research data, including detailed methods and technical analyses supporting data quality. Scientific Data will initially focus on the life, biomedical and environmental science communities, but will be open to content from a wide range of scientific disciplines. Publications will be complementary to both traditional research journals and data repositories, and will be designed to foster data sharing and reuse, and ultimately to accelerate scientific discovery."
Video credit: Nature, "Scientific Data."
In addition to financial rewards, there are also financial incentives sponsored by government agencies, non-profit charities, publishers, and for-profit businesses to develop tools and to create successful commercial projects engaged in data re-use. The Obama administration is offering a “Big Data Research Initiative” backed by $200M for new projects routed through six departments: DARPA, DOE, DOD, HHS/NIH, NSF, and USGS. The deadline to reply to this year’s call for projects involving big data collaborations is September 2, 2013. Interested individuals can send proposals to the Networking and Information Technology Research and Development (NITRD) program (BigDataprojects@nitrd.gov). A detailed description of the requirements can be found on their website. According to a White House press release, Dr. John P. Holdren, Assistant to the President and Director of the White House’s Office of Science and Technology Policy, stated that,
John P. Holdren, Director of the Office of Science and Technology Policy. (Image credit: Wikipedia).
“In the same way that past Federal investments in information-technology R&D led to dramatic advances in supercomputing and the creation of the Internet, the initiative we are launching...promises to transform our ability to use Big Data for scientific discovery, environmental and biomedical research, education, and national security,” This is the second year of the Big Data Initiative, and “the Administration is encouraging multiple stakeholders including federal agencies, private industry, academia, state and local government, non-profits, and foundations, to develop and participate in Big Data innovation projects across the country.”
Victoria Costello (vcostello@plos.org) of PLOS I OPEN FOR DISCOVERY manages the ASAP awards program, which is sponsored by 27 global organizations including Google, PLOS, and the Wellcome Trust. “The ASAP Program will award three top awards of $30,000 each” in October 2013 to recognize the best in creative data “reuse, remixing, [and] repurposing—which enables countless clinical translations and subsequent discoveries based on previously published (OA) research.” Similarly, Eli Lilly will be accepting submissions until October 2, 2013, for a "Clinical Trial Revisualization Design" competition, and is giving away $75K in cash and prizes. Their goal is to encourage "designers and developers to re-imagine clinical trial information in a patient-centric way [because] clinical trial information can often be dense and difficult to digest from a patient’s perspective."
Going back to 2011, before the recent open access and open data mandates, the National Library of Medicine sponsored their own contest, "NLM Show Off Your Apps: Innovative Uses of NLM Information", with 35 entries using their biomedical data. Wondering if databases such as PubMed Central might make use of added metadata from the field of health informatics, and open data elsewhere, I contacted Betsy L. Humphreys (blh@nlm.nih.gov), Deputy Director of the National Library of Medicine, to discuss the feasibility of adding health information metadata tags found in electronic health records (EHRs) to the records in their database, either on their side or on the entrepreneurial side of things.
Betsy L. Humphreys,
Deputy Director,
National Library
of Medicine
(Image credit:
National Library
of Medicine).
Published with
written permission
from Betsy L.
Humphreys.
According to Humphreys, my concept of “connecting EHRs with NLM databases is very sensible,” but directly adding metadata tags is not a very practical approach. Part of the problem in doing this lay in the sheer number of PubMed records, more than 22 million of them, and the fact that they are updated nightly. Another major problem is that whenever the Systematized Nomenclature of Medicine—Clinical Terms (SNOMED CT) or the Logical Observation Identifiers Names and Codes (LOINC) (used to identify lab tests and clinical observations) is updated, some of the records with those tags would need updating. Perhaps the best reason why this approach should not be implemented by the NLM, Humphries informed me, is that within PubMed Central, there is a lot of interoperability, including a correspondence table that includes SNOMED CT using the Unified Medical Language System (UMLS) Metathesaurus, and whenever possible, records are mapped with their synonyms. But, on the entrepreneurial side of things, she said, "there are indeed many EHR vendors who use the MedlinePlus Connect API, which enables the use of SNOMED CT, LOINC, or RxNorm in search arguments, in order to integrate NLM’s data into an electronic health record through a patient portal.” The interview with Humphreys left me with hope that there are indeed ample opportunities for entrepreneurs to create new data tools or physical products that create something new through mashups that combine electronic health records, published medical research records in existing databases like Medline Plus, PubMed, and PubMed Central, as well as other data.
Pictured above is Matthew D. Carmichael delivering a talk about data integrity for archival administration of digital data and other e-records. Information critical to the safekeeping of digital data was discussed at the Fourth Annual Virtual Center for Archives and Records Administration (VCARA) Conference on Digital Stewardship & Knowledge Dissemination in the 21st Century held on May 22, 2013, and hosted in the virtual world Second Life by San Jose State University’s School of Library and Information Science. The conference promoted “the impact of archivists, records administrators and digital curators on the future of knowledge dissemination in the context of cultural heritage institutions.” (Click image to enlarge).
Rating the Quality of Open Data
When working with government data it may be helpful to keep a few key guidelines in mind. The problem is, there are many guidelines. A working group within OpenGovData.org developed "8 Principles of Open Government Data" which are: "1. Data Must Be Complete... 2. Data Must Be Primary... 3. Data Must Be Timely... 4. Data Must Be Accessible... 5. Data Must Be Machine processable... 6. Access Must Be Non-Discriminatory... 7. Data Formats Must Be Non-Proprietary... 8. Data Must Be License-free." This is very similar to the Sunlight Foundation's "Ten Principles for Opening Up Government Information"— "1. Completeness... 2. Primacy... 3. Timeliness 4. Ease of Physical and Electronic Access... 5. Machine readability... 6. Non-discrimination... 7. Use of Commonly Owned Standards...8. Licensing... 9. Permanence... 10. Usage Costs." Open government data initiatives could also be held up to a 5-star rating method, which has been proposed by Tim Berners-Lee, the British computer scientist credited with inventing the World Wide Web:
★ Available on the web (whatever format), but with an open licence to be Open Data
★★ Available as machine-readable structured data (e.g. Excel instead of image scan of a table)
★★★ As (2), plus non-proprietary format (e.g. CSV instead of Excel)
★★★★ All the above, plus use W3C open standards (RDF and SPARQL)
★★★★★ All the above, plus link your data to other people’s data to provide context
The Open Data Institute has created an "Open Data Certificate" for data and rates it against a checklist. Certificates awarded grade data as "Raw: A great start at the basics of publishing open data, Pilot: Data users receive extra support from, and can provide feedback to the publisher, Standard: Regularly published open data with robust support that people can rely on, and Expert: An exceptional example of information infrastructure." In 2009, The White House created a scorecard by which open data can be evaluated according to 10 criteria: "high value data, data integrity, open webpage, public consultation, overall plan, formulating the plan, transparency, participation, collaboration, and flagship initiative." The U.S. government's simple stoplight-like rating system was as follows: green for data that "meets expectations," yellow for data that demonstrates "progress toward expectations," and red for data that "fails to meet expectations." At the other end of the spectrum, there is an exceptionally complex checklist offered by OPQUAST. On May 9, 2013, President Obama issued an Executive Order "Making Open and Machine Readable the New Default for Government Information" wherein a new "Open Data Policy" has just been established and being newly implemented through "Project Open Data" in which there are seven key principles: "public, accessible, described, reusable, complete, timely, and managed post-release." There does not seem to be an associated rating system, however, to evaluate how well the data complies with the principles. Finally, Nature has has set up three criteria for data: firstly, "experimental rigor and technical data quality," secondly, "completeness of the description," and lastly, "integrity of the data files and repository record."
Finding Solutions with Data Analysis and Visualization
Personally, while at Cambridge, I veered from historical preservation a bit and spent some time studying about the preservation of modern science data. As a librarian and archivist, knowing how to preserve scientific and and medical history meant learning about some of the standard file formats in which science data is typically stored, the software tools used to create and analyze scientific data (which may also be required for reproducibility and long-term access to preserved data sets), along some hands-on training to actually use those software tools.
I completed computer training useful in scientific computing through the UCS service, such as "Programming Concepts for Beginners," "Python: Introduction for Absolute Beginners," "Unix Intro," "Unix: Simple Shell Scripting for Scientists," "Programming Concepts-Pattern Matching," "Emacs," "mySQL," and "Condor and CamGRID" used in parallel, distributed, and grid computing. I did not have the opportunity to work with Cambridge’s COSMOS, “the world's first national cosmology supercomputer,” which was “founded in January 1997 by a consortium of leading UK cosmologists, brought together by Stephen Hawking.” But it would have been pretty exciting to use this data-intensive system.
The truth is, many scientists are routinely employed in hacking their own programs to link measuring equipment, data analysis, and visualization together and my best guess is that they would benefit from more end-to-end integrative open source software systems development. While by no means exhaustive, I compiled a list linking to 349 subject specific tools and 123 general tools that are useful for all of this newly available open data. Certainly, some if not all of the tools in this list might have its quality rated according to the eight aforementioned standards.
More than 349 Subject Specific Open Data Tools
• Earth, geology, ecology, climate, and weather sciences tools include Microsoft’s SciScope. Metadata standards include EML (The Ecological Metadata Language) and the ISO 19115:2003 International Standard for Geographic Information. Polymaps is a tool for making dynamic, interactive maps.
Weather Modeling (Image Credit: Wikimedia Commons)
• GoGeo is a UK-based site that provides a listing of 161 free software products, 50 data services, and 10 search portals relating to geographical information. A standard for this field is the CSDGM (Content Standard for Digital Geospatial Metadata).
• Berkeley compiled a good listing of molecular biology resources including databases and tools for protein and nucleotide sequencing as well as model organisms. Foldit is a game that enables players to solve real problems in protein folding through its simulation software. Some of the metadata standards in the biological fields include: ABCD (Access to Biological Collections Data) Schema developed by the (Taxonomic Databases Working Group (TDWG) of Australia and with the International Union of Biological Sciences, Darwin Core, developed in Australia by TDWG for natural history specimens and observations, Genome Metadata, MIGS/MIMS (Minimum Information About a (Meta)Genome Sequence), and FGDC. GenBank Flat File Format is a standardized format for biological data.
• In astronomy, NASA maintains the FITS file format standard and documentation. A new metadata astronomy thesaurus to provide better linking among scholarly astronomy articles, called the Unified Astronomy Thesaurus (UAT), was recently released. The World Wide Telescope (WWT) “is an application that runs in Windows that utilizes images and data stored on remote servers enabling you to explore some of the highest resolution imagery of the universe available in multiple wavelengths.” An international consortium comprised of the Spitzer Science Center, ESA/Hubble, California Academy of Sciences, IPAC/IRSA, and the University of Arizona established a metadata standard for astronomy called the Astronomy Visualization Metadata Standard.
• CERN developed ROOT, a tool for ‘big data’ analysis in physics, and CASTOR (Cern Advanced STORage Manager) that is written in Scientific Linux for storage management of ‘big data’ physics files.
• Chemistry tools include PubChem, ChemSpider, Chemical Markup Language (CML), as well as eBank UK, a digital repository for crystallographic data.
Neurological Modeling (Image credit: By Polygon data were generated by Database Center for Life Science(DBCLS)[2]. (Polygon data are from BodyParts3D[1]) [CC-BY-SA-2.1-jp via Wikimedia Commons].
• In medicine, several online tools allow patients to better manage their health care through a Personal Health Record (PHR) including: Microsoft HealthVault, PatientsLikeMe, getHealtZ, onpatient, WebMD Health Record, and Patient Ally. Other tools include EM data analysis and visualization of the brain like NeuroTrace. CARMEN (Code, Analysis, Repository and Modeling for e-Neuroscience) in the UK is a virtual neurophysiology laboratory. Amira was designed for data analysis and visualization in the life sciences along with MeVisLab for medical image processing and visualization. ClinicalTrials.gov Protocol Data Element Definitions and Infobuttons are useful when working with data in the health sciences. Standard tools include PubMed, PubMed Central, MedlinePlus, and GenBank.
123 General Open Data Tools (Organized by Topic)
Big Data: BigQuery, BOINC, Hadoop, HTCondor (formerly Condor), MapReduce, mrjob, NoSQL database, OGSA (Open Grid Services Architecture), Platfora. (9 Tools).
Data Analysis and Calculation, Data Mining, Graphics Plotting, Image Processing, Data Visualization and Simulation, Reports: Avizo, Cave5D, ELKI, Excel, Feko, Fiji, ggplot2, GIMP, Gnuplot, GoogleVis, Graph, IDL, ImageJ, JMP, KaleidaGraph, MATLAB, Multiphysics, NumPy, Omniscope, OpenCV, OpenGL, OpenGL Vizserver, Open Office, OriginLab, ParaView, PCA/PLS plots, Photoshop, RGGobi, SCIDAVIS, SCILAB, SciPy, SenseWeb, Simulink, Stata, Tableplot, Tabplots, VisIt, Wolfram Alpha, Word. (39 tools).
Data Integrity, Data Recovery, Digital Forensics: AIR (Automatic Image and Restore), Autopsy, CRC-32, dc3dd, dcfldd, FITools, FTK Imager, Guymager, LOCKSS Box, MagicRescue, OSFClone, OSForensic, PhotoRec, PyFlag, SCALPEL (Source Code Analysis, Libre and Portable Library), Sleuthkit. (16 tools).
Data Management Planning and Audit Assessment: CARDIO (Collaborative Assessment of Research Data Infrastructure and Objectives), DMP Online, DRAMBORA (Digital Repository Audit Method Based On Risk Assessment), DROID (Digital Record Object Identification). (4 tools).
Data Storage or Repository: arXiv, CEPH, D-Space: RTI and ControlDesk, DataFlow / DataBank, Dryad. (5 tools).
File Sharing, Identification, and Management: CollateBox, FigShare, Libmagic, TrlD File Identifier for .NET. (4 tools).
Instrument Control: LabVIEW. (1 tool).
Metadata Standards, Protocols, Preservation Formats, and Registries: AGLS, AGRkMS (Australian Government Recordkeeping Metadata Standard), DCAT (Data Catalog Vocabulary),DOI (Digital Object Identifier), ISO 15386 DCMI (The Dublin Core Metadata Initiative), EAC-CPF (Encoded Archival Context -Corporate Bodies, Persons, and Families), EAD (Encoded Archival Description),
e-GMS (e-Government Metadata Standard) 1.0-2002, GDFR (Global Digital Format Registry), ISO 8:1977 Publishing, ISO 215:1986 Publishing, ISO 23081, ISO International Standards on Archives and Records Management, ISO IT Applications in Science, ISO Medical Science and Health Care Facilities, ISO Natural Sciences (07), JHOVE, JHOVE2, METS (Metadata Encoding and Transmission Standard), MODS (Metadata Object Description Schema), NISO Z39.87-2006 Metadata for Images in XML, OAI-PMH (Open Archives Initiative Protocol for Metadata Harvesting), OData (Open Data Protocol ), PREMIS (Preservation Metadata: Implementation Strategies), PRONOM, RAD (Rules for Archival Description), RDF(Resource Description Framework), RIR (Representation Information Registry Repository), SKOS (Simple Knowledge Organisation System) Core, Standard Archive Format for Europe (SAFE), UDFR (Universal Digital Format Registry),
XML Formatted Data Unit (XFDU), XMP. (33 tools).
Programming: C, C++, d3.js, Fortran, Java, JSON (JavaScript Object Notation), Octave, PROLOG, Python, R, SQL, VBA. (12 tools).
Will ‘Actionable Intelligence’ Ultimately
be a Result of ‘Artificial Intelligence’?
Despite the fact that my research proposal to enter the computer science graduate program at Cambridge (below) was not accepted, as a proposal, I thought it raised some relevant points about the future applications of artificial intelligence and machine learning for integration and analysis of science data across platforms or in the cloud:
Methodological Approaches in Artificial Intelligence
for the Reuse of "Big Data" in Scientific Computing
Preservation of the science cyberinfrastructure requires an understanding of long-term and archival digital storage formats, data provenance, metadata schemas, and storage repositories essential for both reproducibility and accountability of scientific results and reuse of existing science data sets. These two factors drive mandated availability of prepublication data sets[1] in scientific research projects generated from governmentally-funded agencies like the NSF[2], NIH[3], MRC[4] and Wellcome Trust[5]. Rules mandating data reuse hold the potential to optimize research fiscal spending, spur unprecedented new scientific discoveries in digital repositories by finding relationships across differently funded projects, and promote greater scientific collaboration amongst geographically separated institutions. Yet, in a 2005 symposium on digital biology, additional problems cited “bottlenecks” that “occur owing to our limited capacity to control quality and integrate data from myriad sources, to share data across multiple tasks and to exchange data among different people and organizations...Problems with data integration affect all data tasks, including semantic interpretation, data representation, modeling, data storage and query."[6] Despite the fact funders require data set policies, they leave standards to individual institutions; concurrency of metadata standards within and across disciplines has not occurred, though some advocate RDF for semantic interoperability. Legacy data provide substantial difficulty in updating to new standards and backlogs exist: "Even if standards to facilitate data discoverability, access, and use were to be introduced in the future ... huge problems ... applying ... standards retrospectively to the huge amounts of existing data that do not conform to them [was forseen]."[7] Another problem is that achieving scalability requires supervised, semi-supervised or autonomous automation. I would like to explore the use of artificial intelligence, establishing and monitoring probabilistic relationships across legacy and new linked and/or unlinked datasets housed within distributed or cloud-based data repositories, as a more robust way to search legacy data and data with missing or minimal metadata. This approach may hold the potential for innovative breakthrough discoveries in fields such as basic biological research, clinical medical research, genomics, proteomics, climate modeling, and computational astrophysics.
Sources:
1. Toronto International Data Release Workshop Authors. (9 September 2009). Nature 461, 168-170 doi:10.1038/461168a
2.National Science Foundation. (January 2011). Chapter II - proposal preparation instructions. Retrieved from http://www.nsf.gov/pubs/policydocs/pappguide/nsf11001/gpg_2.jsp#dmp
3. National Institutes of Health. (April 17, 2007). NIH data sharing policy. Retrieved from http://grants.nih.gov/grants/policy/data_sharing/
4.Medical Research Council. (September 2011). MRC policy on research data-sharing. Retrieved from http://www.mrc.ac.uk/Ourresearch/Ethicsresearchguidance/datasharing/Policy/index.htm
5. Wellcome Trust. (August 2010). Policy on data management and sharing. Retrieved from http://www.wellcome.ac.uk/About-us/Policy/Policy-and-position-statements/WTX035043.htm
6. Morris, R. W., et al. (2006). Digital biology: an emerging and promising discipline. Trends in Biotechnology, doi:10.1016/j.tibtech.2005.01.005
7. Smithsonian Institution. Office of Policy Analysis. (March 2011) Sharing Smithsonian digital scientific research data from biology. Retrieved from http://www.si.edu/content/opanda/docs/Rpts2011/DataSharingFinal110328.pdf
The future of AI lies in the imagination. Fortunately, ideas and concepts can be simulated even if it is not yet possible to actualize them in the real world. In my award-winning Mars environment simulation, Curiosity AI, an embodied artificial intelligent agent could search through academic databases in various subjects like astronomy and provide answers, and robot equipment could perform in-situ analysis of data. The humanoid robot could command and operate expert AI systems and launch a swarm designed to take temperature and pressure readings in the environment. The data I used in my expert system for the Phoenix Lander were raw archived data from the NASA mission (because that was what I had available through the Mars Data Archive), but the concept was designed with the idea in mind that AI would enable the processing and calculation of real-time in-situ mission data directly from instruments on Mars.
The amount of raw scientific and medical data is swelling beyond what can be viewed and interpreted by individuals without the intermediary aid of machines filtering and computing that data deluge for us. As such, the paradigm shift introducing the computational approach mentioned at the beginning of this article may become Kuhnian, sweeping away and gaining domination over the two older scientific approaches.
I can only conclude this article on open data tools by saying that when attempting to be knowledgeable about many things, it usually results in becoming an expert at nothing. Find the tools that work best for you and stick with them; however, every now and again it is good to see what new tools might be out there, and now is that time.
If you enjoyed this three-part series and would like more information, NISO will be conducting two webinars later this year on “Research Data Curation:” Part 1: E-Science Librarianship (September 11, 2013) and Part 2: Libraries and Big Data (September 18, 2013).
Disclaimer: I own a startup company (not mentioned here) related to computer technology, primarily engaged in library and archives consulting, computer technology R&D, 3-D modeling and simulation, and artificial intelligence research.