Monday, January 23, 2012

Personalized Medicine-Humanity's Ultimate Big Data Challenge

Robert T Fassett, iHealth Connections,  December 2011;1(2):90–5

Abstract
At the heart of personalized medicine lie big data, really big data—from rapidly accelerating genomics research, from deployments of electronic medical records, and soon from social networking, telemedicine, and the ‘Internet of Things’ remote sensors. The impact personalized medicine will have in transforming healthcare will depend not only on how well we gather and analyze these big data, but also on how effectively we transmit their derivative insights and interventions out to clinicians and ultimately their patients.

Acknowledgment: Editorial assistance was provided by Touch Briefings.

Support: The publication of this article was funded by Oracle Health Sciences.

Disclosure The author has no conflicts of interest to declare.
Correspondence: robert.fassett@oracle.com
We have reached an inflection point between the insular ‘sickcare’ non-system of the past and the collaborative, proactive, true ‘health and wellness’ system of the future. To overcome the inertia of our current ‘system’, disruptive forces are being applied—access reform, value-based reimbursement, evidence-based clinical guidelines, quality reporting, medical homes, and accountable care, among others.1,2

High-definition Healthcare
Another important vector for change has grown out of our massive collective investment in basic biomedical and clinical research. For example, the US alone has funded its National Institutes of Health (NIH) with $484 billion since 1950, with a current annual budget of over $30 billion.3 As we have come to better understand the phenotypic, genotypic, environmental, and lifestyle factors that determine our health, it has become clear that disease and wellness are inherently personal. Any two persons have 99.6 % of their DNA in common. In the remaining set of 24,000,000 base pairs that we each call our own lies humanity’s diversity, our individual predilection for disease, and the potential for truly personalized medicine.4

The US National Cancer Institute defines personalized medicine as “a form of medicine that uses information about a person’s genes, proteins, and environment to prevent, diagnose, and treat disease.”5 This is not to imply that, heretofore, the practice of medicine has been somehow impersonal. Hippocrates already recommended cold foods for ‘phlegmatic patients.’ Two millennia later, we understand that African-Americans respond differently to antihypertensives and prescribe accordingly. What is compelling about this new definition is its resolution. We are now capable of tailoring health and wellness at the molecular level—healthcare in its highest possible definition.6,7

In eight short years, we have progressed from a single human genome to the HapMap, and now to inexpensive whole-genome sequencing and the 1000 Genome Project.8 Genome-wide association studies have identified hundreds of genotype–disease linkages, some of which have strong clinical implications.9


We have begun to appreciate the non-linearity of the old DNA–RNA–protein central dogma and now see phenotype as the result of a complex network of interactions that include DNA structural modifications, novel transcriptional regulation via microRNA and short interfering RNA, post-translational modifications, etc.10 Astonishingly, we have vision into the transcriptome and proteome at the single-cell level.11 All this will soon result in the almost overwhelming growth of the fundamental substrate for personalized medicine: data.

Medicine’s Deep Space Objects
Individualizing treatment for a given patient is a truly daunting, data-driven task. It is not just a matter of wading through three billion base pairs to find a sequence variant that correlates with a particular disease. It is the multivariate ripple effects these polymorphisms have across the DNA–RNA–protein network, and then their interactions with the person’s environmental and lifestyle history, that must be understood. Given the magnitude of this endeavor, the resources committed, and the global cooperation that is needed, this is biology’s version of ‘big science.’

Finding a treatment based on a patient’s genes, proteins, and environment is essentially a signal-detection exercise. Gene defects (i.e., ‘signals’) that are relatively common and have a high penetrance—such as sickle cell anemia—are relatively easy to pick up. (Linus Pauling discovered the causative sickling protein sequence defect in 1949, four years before Watson and Crick determined the structure of DNA.) Gene defects that have a low prevalence and a low penetrance are much harder to detect and understand. They are the biologic equivalents of deep space objects. Many conditions, even common and deadly ones such as obesity, diabetes, and atherosclerosis, are thought to be influenced by many different sets of signals, some easy to detect and others that will require medicine’s equivalent of the Hubble space telescope.

This will ultimately require data volumes and manipulation techniques unprecedented in information science and technology. Detecting rare and variably expressed mutations and correlating them with fine-grained clinical observations and environmental factors in a large population will require massive amounts of high-resolution data. This may seem like a daunting thousand-mile march, but the longest road still lies ahead.