Showing posts with label Information Infrastructure. Show all posts
Showing posts with label Information Infrastructure. Show all posts
Wednesday, December 12, 2012
Digital Universe is Expanding
Quote from FT below: “Only 0.4 per cent of that data are actually being analysed.”
Paul Taylor, The Financial Times, December 11, 2012
The “digital universe” continues to expand rapidly and will double every two years through 2020, but only a small fraction of this ‘big data’ is being analysed, which leaves a huge untapped opportunity for companies and other organisations.
According to a report published on Tuesday by IDC, the market research firm, a large ‘big data gap’ is opening up between the rapidly expanding digital universe – a measure of all the digital data created, replicated and consumed in a single year. That data holds potential analytic value that could have major business, health and societal impact.
Only 0.4 per cent of that data are actually being analysed.
The reason, IDC suggests, is that the majority of new data being generated is unstructured – for example video surveillance footage, audio files or social media – meaning we know little about the data until it is characterised and tagged. The industry currently faces a shortage of the talent and technology needed to tag and analyse unstructured data.
IDC’s sixth annual report on the “Digital Universe,” sponsored by EMC, the storage market leader, notes that today the digital universe comprises “the images and videos on mobile phones uploaded to YouTube, digital movies populating the pixels of our high-definition TVs, banking data swiped in an ATM, security footage at airports and major events such as the Olympic Games, subatomic collisions recorded by the Large Hadron Collider at CERN, transponders recording highway tolls, voice calls zipping through digital phone lines, and texting as a widespread means of communications.
“With the rise of big data awareness and analytics technology, the digital universe in 2012 has taken on the feel of a tangible geography,” notes the report, “vast, barely charted place full of promise and danger.”
This untapped value could be found in patterns in social media usage, correlations in scientific data from discrete studies, medical information intersected with sociological data, faces in security footage, and so on.
However, even with a generous estimate, the amount of information in the digital universe that is “tagged” accounts for only about 3 per cent of the digital universe in 2012, and that which is analysed is half a per cent of the digital universe. “Herein is the promise of “big data” technology – the extraction of value from the large untapped pools of data in the digital universe,” says the report.
The reports authors conclude that “our digital universe in 2020 will be bigger than ever, more valuable than ever, and more volatile than ever.”
Not surprisingly, IDC predicts that big data will be a big boon for the IT industry. “Web sites that gather significant data need to find ways to monetise this asset. Data scientists must be absolutely sure that the intersection of disparate data sets yields repeatable results if new businesses are going to emerge and thrive. Further, companies that deliver the most creative and meaningful ways to display the results of big data analytics will be coveted and sought after.”
Among the report’s main findings it suggests:
n From 2005 to 2020, the digital universe will grow by a factor of 300 – more than 5,200 gigabytes for every man, woman, and child in 2020. From now until 2020, the digital universe will about double every two years.
n The investment in managing, containing, studying and storing the bits in the digital universe will only grow by 40 per cent between 2012 and 2020. As a result, the investment per gigabyte during that same period will drop from $2.00 to $0.20.
n Between 2012 and 2020, emerging markets’ share of the expanding digital universe will grow from 36 per cent to 62 per cent.
n A majority of the information in the digital universe, 68 per cent in 2012, is created and consumed by consumers – watching digital TV, interacting with social media, sending camera phone images and videos between devices and around the Internet, and so on. Yet enterprises have liability or responsibility for nearly 80 per cent of the information in the digital universe.
n Only a small fraction of the digital universe has been explored for analytic value. IDC estimates that by 2020, as much as 33 per cent of the digital universe will contain information that might be valuable if analysed, compared to 23 per cent today.
n By 2020, nearly 40 per of the information in the digital universe will be “touched” by cloud computing providers – meaning that a byte will be stored or processed in a cloud somewhere in its journey from originator to disposal.
n The proportion of data in the digital universe that requires protection is growing faster than the digital universe itself, from less than a third in 2010 to more than 40 per in 2020.
n The amount of information individuals create themselves – for example writing documents, taking pictures, downloading music – is far less than the amount of information being created about them in the digital universe.
See results of EMC study at http://www.emc.com/leadership/digital-universe/index.htm
Making Dollars and Sense of the Open Data Economy
Is the push to free up government data resulting in economic activity and startup creation?
Alex Howard, O'Reilly Radar, December 11, 2012
Over the past several years, I’ve been writing about how government data is moving into the marketplaces, underpinning ideas, products and services. Open government data and application programming interfaces to distribute it, more commonly known as APIs, increasingly look like fundamental public infrastructure for digital government in the 21st century.
What I’m looking for now is more examples of startups and businesses that have been created using open data or that would not be able to continue operations without it. If big data is a strategic resource, it’s important to understand how and where organizations are using it for public good, civic utility and economic benefit.
Sometimes government data has been proactively released, like the federal government’s work to revolutionize the health care industry by making health data as useful as weather data or New York City’s approach to becoming a data platform.
In other cases, startups like Panjiva or BrightScope have liberated government data through Freedom of Information Act requests and automated means. By doing so, they’ve helped the American people and global customers understand the supply chain, the fees associated with 401(k) plans and the history of financial advisors.
I’ve hypothesized that open data will have an overall effect on the economy akin to that of open source and small business. Gartner’s research has posited that open data creates value in the public and private sector. If government acts as a platform to enable people inside and outside government to innovate on top of it, what are the outcomes?
Over the past four years, the world has heard a rising chorus for raw data from voices like the creator of the World Wide Web, Tim Berners-Lee, and the chief technology officer of the United States, Todd Park. Park, in particular, has been working to scale open data across the federal government as the nation’s “entrepreneur in residence.”
McKinsey and Associates estimated the annual economic value of big, open liquid health data at some $350 billion annually. While that number is eye-opening, which companies and startups stand to change health care using open health data?
Some examples are clear, from mobile apps like iTriage (now owned by Aetna) to Castlight, but they aren’t sufficient to understand what’s happening out there.
Other promising startups are in the consumer finance space, where so-called “smart disclosure” initiatives are enabling people to put their personal data to use. Startups like Billshrink.com and Hello Wallet now are enabling people to make smarter financial decisions.
I know there are more stories out there, and in sectors beyond health care and consumer finance — including transit, energy, education and media. Over the next several months, I’ll be identifying and profiling more civic startups, such as those from the first class in the Code for America accelerator, like Captricity, to specialized search engines, like Zillow, Panjiva and DataMarket.
In the course of that work, I hope to answer some big questions. What are the sustainable business models that successful civic startups are using, whether they use legislative data or other reuse of public sector information? What are the real costs associated with opening up government data to make it usable, both for government and entrepreneurs? And how does it balance against what datasets, at the federal, state or local levels, are the most valuable? Are they open and usable? If so, who’s using them and to what effect? If not, why not?
At the end of this particular project, in February, we’ll publish a report on what I’ve found. In it, I hope to be able to share some answers to several core questions on the topic. Where I need your help is in identifying new startups that are using or consuming government data or in highlighting how existing companies use it in their operations, good or services. Who is doing the most interesting work — and where? If you have research and evidence to share on the questions I posed above, feel free to ring in on that count as well.
Please weigh in through the comments or drop me a line at alex@oreilly.com or at @digiphile on Twitter.
Wednesday, October 10, 2012
Harnessing Data as a New Source of Growth: Big Data Analytics and Policies
OECD Headquarters, Paris, France, October 22, 2012
Introduction
The confluence of several key socio-economic and technological trends is resulting in the generation of huge streams of data every day. The major trends include:
· The increasing migration of social and economic activities on line: Social network site Facebook, for example, now counts over 900 million active participants around the world generating together more than 1 500 status updates every second about their interests and whereabouts. In 2011, e-commerce platform eBay collected data on more than 100 million active users including the 6 million new goods they offered every day.
· The strong decline in the cost of data collection, storage, transportation, and processing: The average cost of consumer hard disk drives (HDDs) per gigabyte, for example, dropped on average by almost 40% per year between 1998 (USD 56 per gigabyte) and 2012 (USD 0.05 per gigabyte). In 1995, as another example, consumers in France paid USD 75 equivalent per month for a dial-up (56 Kb/s) connection, while in 2011 they paid the equivalence of USD 33 per month for a broadband (51 Mb/s) connection, which was almost 1 000 times faster.
· The increasing deployment of “smart” ICT applications such as smart grids and smart transportations based on machine-to-machine (M2M) communication: Connecting one million homes to a smart grid may produce as much as 11 gigabytes of data per day. In order to accommodate for hourly readings through smart meters, a network with a minimum capacity of up to 1 Mbit/s dedicated to M2M communication is needed.
· The continued expansion of mobile communication: In 2011, there were 780 million smart phones worldwide capable of collecting and transmitting geo-location data, which generated more than 600 petabytes (millions of gigabytes) of data every month. It is estimated that the global data traffic generated by mobile communication (including M2M enabled smart devices) will almost double every year to reach 11 exabyte (billions of gigabytes) per month by 2016.
The collection and exploitation of these large data flows through data analytics is leading to a shift towards a data-driven socio-economic model commonly referred to as “big data”. In this model, data is a core asset that provides a huge resource for new industries, processes, services and goods leading to significant competitive advantage. In business, for example, data analytics are increasingly being used in a wide number of operations ranging from optimising the value chain and manufacturing production to more efficiently using labour and improving customer relationships. It is estimated that firms, which adopt data-driven decision-making, for example, have output and productivity that is 5-6% higher than what would be expected given their other investments and information technology usage. These firms also perform better in terms of asset utilisation, return on equity and market value.
To unlock the potential of big data and data analytics (big data analytics), OECD countries need to ensure the development of coherent policies and practices around the collection, transportation, storage and use of data, most prominently in areas related to privacy protection. New data sources, new actors and the increasing ease with which personal data can be collected, linked and processed, seriously challenge the effective implementation of current frameworks for privacy protection. But the potential implications for policy also spill over into many other domains; including among others open access to data, intellectual property rights, competition, skills and employment, infrastructure, and measurement.
Objective
First launched in 2005, Technology Foresight Forums are an annual event organised by the OECD Committee for Information, Computer, and Communications Policy (ICCP) to help identify opportunities and challenges for the Internet Economy posed by technical developments. The 2012 Technology Foresight Forum will focus on the potential of big data analytics as a new source of growth, which could help generate significant economic and social benefits. It will put big data analytics in the context of emerging trends discussed at the last three Foresight Forums, namely mobile communications (2011), smart ICTs (2010), and cloud computing (2009), to highlight the confluence of trends leading towards a data-driven economy (see Figure below).
The confluence of major trends enabling data as a new source of growth
The 2012 Foresight Forum will then discuss the social and economic issues related to big data analytics that may put at risk fundamental values and warrant a review of current policy frameworks, most prominently those aimed at ensuring the protection of privacy, intellectual property, and competition. Other policy areas that will be addressed include health care, science and research, government administration and labour markets.
This year’s Foresight Forum will contribute to OECD’s ongoing horizontal project entitled New Sources of Growth: Intangible Assets (NSG) as well as to the follow-up horizontal project on Knowledge-Based Capital: Seizing the Benefits of New Sources of Growth, which will be launched in 2013. Both projects aim at (i) providing structured evidence of the economic value of intangible assets, including data, as a new source of growth, and (ii) improving understanding of current and emerging challenges for policies related to the increasing relevance of these intangible assets.
Modalities
The Foresight Forum represents a collaborative effort of policy makers, business, civil society, and the Internet technical community and thus provides opportunities to liaise with experts in the fields covered. To increase interactivity with a broader public, the Foresight Forum will be supported by various participative web technologies, including a webcast, and Twitter feeds. Further details will be announced closer to the event.
Agenda
The 2012 Foresight Forum will include two morning and three afternoon sessions: The first (morning) session will introduce data analytics in the context of the key technological and socio-economic trends, namely i) cloud computing; ii) smart ICT applications; and iii) the Internet of Things. The following session will then focus on the socio-economic implications of harnessing data as a new source for growth. The afternoon sessions will take a more in-depth look at specific areas that will include: i) science and research (including public health); ii) marketing and competition; and iii) public administration. Each of these sessions will be taken as a mean to discuss specific potential policy opportunities and challenges ranging from i) privacy and consumer protection; ii) intellectual property rights; iii) open access to data; iv) competition; and v) skills and employment. A concluding session will then highlight the main policy implications to be examined by the OECD in the context of its programme of work for 2013-14.
Subscribe to:
Posts (Atom)