Showing posts with label systems biology. Show all posts
Showing posts with label systems biology. Show all posts

Monday, March 22, 2010

Mining Evolution

Can a worm get breast cancer? And how would you know if it did (since it doesn't have breasts)?

Biomedical research has made great use of "disease models": conditions in lab organisms that resemble the human diseases that the researchers really want to learn about. By seeing how a model condition develops and how it responds to drugs or other changes, researchers can make better guesses about what might help people. But finding such disease models usually requires some obvious similarity between the outward manifestations of the disease in humans and the animal subjects.

In the Proceedings of the National Academy of Sciences, a team from the University of Texas in Austin led by Edward Marcotte use the underlying molecular relationships to find connections between disorders with no such obvious relationship. In addition to a worm analog of breast cancer, they found an amazing connection between plant and human disorders. The analysis of plants' failure to respond to gravity led them to human genes related to Waardenburg syndrome. This syndrome includes an odd constellation of syndromes resulting from defects in the development of neural crest cells.

I saw Marcotte speak about this fascinating work at the conference I attended last December in Cambridge, Massachusetts. My writeup should be posted soon by the New York Academy of Sciences.

Biologists have repeatedly found that the networks of interacting molecules are organized into modules. Over the course of evolution, these modules can be re-used, often for purposes quite different from their original function. This much is well known, although seeming the persistence of modules over the vast evolutionary separation of plants and people is very dramatic.

What the Austin team did was to devise a methodology to identify related molecular modules in different species even without relying on similar outward manifestations, or phenotypes. They combed the known molecular networks of different species for modules that had a lot of "orthologous" genes: those that had retained similarity--and similar relationships--through evolution. They call the particular traits associated with these genes "orthologous phenotypes," or "phenologs." "We're identifying ancient systems of genes that predate the split of these organisms, that in each case retain their functional coherence," Marcotte said at the conference.

The importance of this scheme is that many molecular networks are poorly mapped, especially in humans. But if a particular gene is part of the network underlying a phenotype in another species--such as poor response of a plant to gravity--it's a good guess that the corresponding gene may be active in the orthologous phenotype in people. The researchers in fact confirmed many of these predicted relationships. Some of these genes were previously known to relate to disease, while others were new. The researchers created a list of hundreds more that they still hope to check.

These genes could give researchers many potential new targets for drugs or other interventions in diseases. So evolution is not just helping us to understand the biologic world we live in, but helping us devise ways to improve human health.

Thursday, February 25, 2010

Targeting Cancer

Amy Harmon of The New York Times had an excellent three-part series this week called "Target Cancer." She follows one clinician/researcher as he pursues a "targeted" treatment for melanoma, which aims at the protein produced by a gene called B-Raf that is mutated more than half of the time in this skin cancer.

The series does a great job in following the emotional roller-coaster ride of the doctor, and of course his patients. One early targeted drug doesn't work at all, perhaps because it also attacks normal cells and the side effects become intolerable before the dose is high enough to affect the cancer. A new drug seems not to do anything, but then the team decides to wait for the drug company to reformulate it to deliver higher effective doses.

The results are spectacular: the new formulation causes a virtually unheard of remission in the cancer, and raises hopes in formerly hopeless patients and in the doctors. The excitement and the potential are palpable as some patients dare to hope and others can't bear to. But within a few months, the patients are dying again.

The new drug is an example of personalized medicine, since it is effective only for patients with a particular mutation. There are a few other examples of therapy tuned to patients with a particular genetic profile, such as the breast-cancer drug erbitux and the anticoagulant warfarin (Coumadin).

But this treatment is actually for cancers with a particular mutation--a mutation the normal cells of the patient don't have. Cancers cells generally have more and more of mutations as the disease progresses, because it disrupts the normal quality-control mechanisms in the cell. A study announced last week (registration required) showed that the specific pattern of mutations could be used to monitor the ebb and flow during treatment, although it doesn't look practical yet for tailoring treatment.

Unfortunately, as described in this series, even when a drug targets a mutation in a particular patient's cancer, cancers often develop alternate routes to proliferation. Harmon alludes to one approach to this problem: a multi-pronged "cocktail" that attacks many possible mutations at once. Such cocktails are standard, for example, in treating HIV/AIDS.

Without vilifying the drug companies, she explains some challenges for these profit-oriented companies in pursuing this approach. In particular, even if the cocktail may ultimately be more effective, getting approval might delay or threaten their profits from the drug they have in hand, even if it only extends life for a few months. This is especially true if other drugs in the cocktail are owned by competing companies. In any case, the difficulties in testing multiple drugs make it much harder to know what is effective and what side effects may appear.

The idea of analyzing molecular networks and attacking them at many points simultaneously is a recurring theme in systems biology. But sometimes it seems very far in the future.

Thursday, January 28, 2010

Mix and Match

It's easier to reconfigure a complex system to do new things if it is built from simpler, independent modules. But in biology, modules may be useful for more immediate reasons.

Biological systems, ranging from communities to molecular networks, often feature a modular organization, which for one thing makes it easier for a species to evolve in response to changes in the environment. In some cases, this flexibility might have been selected during prior changes. But modularity can also make life easier for a single organism during its lifetime, and be selected for this reason.

In their book The Plausibility of Life, Marc Kirschner and John Gerhart include modularity as one aspect of "facilitated variation." In particular, they say, genetic changes that affect the "weak linkages" between modules can cause major changes in the resulting phenotype. As long as the modules, the "conserved core processes," are not disrupted, the resulting organism is likely to be viable, and possibly an improvement on its predecessors.

In describing facilitated variation, Kirschner and Gerhart defer the question of whether facilitating rapid future evolution alone causes these features to be selected. Perhaps it does, in some circumstances. But in any case, we can regard the presence of features that enable rapid reconfiguration as an observational fact.

Moreover, the practical challenges of development demand the same sort of robust flexibility that encourages rapid evolutionary change. Over the development of a complex organism, various cells are exposed to drastically different local environments. In addition, genetic or other changes pose unpredictable challenges to the molecular and other systems of the cells. Throughout these changes, critical processes, like metabolism and DNA replication, need to keep working reliably.

To survive and reproduce in the face of these variations organisms need a robust and flexible organization. Features that allow such flexibility should be selected, if only because they improve individual fitness. These same features may then increase evolutionary adaptability, whether or not that adaptability is, by itself, evolutionarily favored.

Whether adaptability is selected for its evolutionary potential or only for helping organisms thrive in a chaotic environment, it has a profound effect on subsequent evolution. A flexible organization including modularity and other features allows small genetic alterations to be leveraged into large but nonfatal changes in the developing creature, so the population can rapidly explore possible innovations.

Wednesday, January 27, 2010

Changing Times

One way to explain the modularity that is seen in biology is that it helps species to evolve quickly as their environment changes.

But the notion that "evolvability" can be selectively favored is tricky intellectual territory, and people can get drawn in to sloppy thinking. Just as group selection must favor more than just "the good of the species," selection for flexibility cannot be grounded in future advantages to the species. To be effective, evolutionary pressure must influence the survival of individuals in the present.

Whether this happens in practice depends on a lot of specific details. Some simulations of the effect of changing environments have not shown any effect. But at a meeting I covered last year for the New York Academy of Sciences, Uri Alon showed one model system that evolves modularity in response to a changing environment.

Alon is well known for describing of "motifs" in networks of molecular interactions. A motif is a regulatory relationship between a few molecules, for example a feed-forward loop, that is seen more frequently in real networks than would be expected by chance. It can be regarded as a building block for the network, but it is not necessarily a module because its action may depend on how it connects with other motifs.

Alon's postdoc Nadav Kashtan simulated a computational system consisting of a set of NAND gates, which perform a primitive logic function. He used an evolutionary algorithm to explore different ways to wire the gates. Wiring configurations that came closest to a chosen overall computation result were rewarded by giving making future generations more likely to resemble them. "The generic thing you see when you evolve something on the computer", Alon said, "is that you get a good solution to the problem, but if you open the box, you see that it's not modular." In general, modules cannot achieve the absolute best performance.

Kashtan then periodically changed the goal, rewarding the system for a different computational output. Over time, the structure of the surviving systems came to have a modular structure. One interesting surprise was that in response to changing goals, the simulated systems evolved much more rapidly than those exposed to a single goal.

But Alon emphasized that this was not a general feature. Instead, the different goals needed to have sub-problems in common. Evolution would then favor the development of dedicated modules to deal with these problems. It is easy to imagine the challenges facing organisms in nature also contain many recurrent tasks, such as the famous "four Fs" of behavior: feeding, fighting, fleeing, and reproducing.

So some biological modularity may reflect the evolutionary response to persistent tasks within a changing environment. But does this explain the wide prevalence of modules? In a future post, I will examine another explanation: that modularity is one of the tools that helps individual organisms adapt to the changing conditions of development and survival during their own lifetimes.

Tuesday, January 26, 2010

Modules

When people design a complex system, they use a modular approach. But why should biology?

For us, modularity is a way to limit complexity. Breaking a big problem into a series or hierarchy of smaller ones makes it more manageable and comprehensible, which is especially important if it is assembled by many people--or one person over an extended time.

The key to a successful module is that its "guts"--the way its parts work together--doesn't depend on how it connects to other stuff. The module can be thought of as a "black box," that just does its job. You don't have to think about it again. For this to work, the connections between modules must be weak, limited to well-defined inputs and outputs that don't directly affect its internal workings.

When it's done right, a module can be easily re-used in new situations. For example, the part of a computer operating system that offers a help menu is tapped by lots of programs without worrying about how it works.

But biology is not designed. Biological systems emerge from an evolutionary process that rewards only survival and reproduction, with no regard for elegance or comprehensibility. Re-usability sounds like a good thing in the long run, but doesn't help an individual survive in the here and now.

Nonetheless, modularity seems to be widespread feature of biological systems. Your gall bladder, for example, is a well-defined blob that receives and releases fluids like blood and bile, but otherwise keeps its own counsel. It can even be removed if necessary. And it does much the same thing in other people and animals.

At a smaller scale, many of the basic components of cells are the same for all eukaryotes. They have the discrete nuclei that define them, and they also have other organelles that perform essential functions, like mitochondria that generate energy. These modules work the same way, whether they happen to appear in a brain cell or a skin cell.

Even at the molecular level, re-usable modules abound. For example, although ribosomes don't have a membrane delimiting them, they consist of very similar bundles of proteins and RNA for all eukaryotic cells, and only modestly different bundles for bacteria. In addition to such complexes, many "pathways," or chains of molecular interactions, recur in many different species.

We have to be careful, of course: simply because we represent complex biological systems as modules doesn't mean they are there. The modules we think we see could simply reflect our limited capacity to understand messy reality. But when researchers have looked at this question carefully, they found that modules really exist in biology, much more than they would in a random system of similar complexity.

But why should evolution favor modular arrangements? And how does a modular structure change the way organisms evolve?

Tuesday, January 19, 2010

The Plausibility of Life

When creationists, or, as they would have it, advocates of "intelligent design," talk about the "weaknesses" of evolutionary theory, knowledgeable people generally roll their eyes and ignore them. This is appropriate, as these advocates only raise the questions in a disingenuous attempt to promote a religious agenda, under the pretense of open-mindedness and "teaching the controversy." In truth, there is no controversy in the scientific community about the dominant role of natural selection (evolution, the theory) in shaping the observed billions of years of change (evolution, the fact) of life on this planet.

But this response obscures the fact that very interesting issues in evolution remain poorly understood.

I'm not referring to the direct exchange of genetic material between single-cell organisms, although that does call into question the tree-like structure of relationships between these simple species. But at the level of complex, multi-cellular creatures like ourselves, this "horizontal gene transfer" is unimportant compared the "vertical" transfer from parents to offspring. The tree metaphor is still intact.

But even for complex creatures-- especially for complex creatures--there are important open questions about how evolution works in detail. The insightful (and cheap!) 2006 book, The Plausibility of Life, by Marc Kirschner and John Gerhart, began to frame some answers to these questions.

The fundamental ingredients of evolution by natural selection were laid out by Darwin: heritable natural variations lead some individuals to be more likely to survive and thus to pass on these variations.

We now know in great detail how cells use some genes in the DNA as a blueprint for proteins, and how these proteins and other parts of the DNA in turn regulate when various genes are active. And we know, as Darwin could only imagine, how that DNA is copied and mixed between generations, only occasionally developing mutations at single positions or in larger chunks. We understand heritable variation.

We also understand the arithmetic of natural selection, which confirms Darwin's intuition: a mutation that improves the chances that its host will survive to be reproduced will spread through a population, while a deleterious mutation will die out (although evolution is indifferent to most mutations). This all takes many generations, but the history of life on earth is long.

But there is something missing, what Kirschner and Gerhart call the third leg of the stool: how does the variability at the DNA level translate into variability at the level of the organism? Selection must occur at this higher level, the level of phenotype, but can only be passed on at the level of the genotype. How do we close this loop?

It would be easy if a creature's fitness were some average of the fitness of each of the three billion bases in the DNA, but it's not that simple. For example, if two proteins work together as a critical team, a mutation in one can kill the organism, even if they could be an even better team if they both mutated in a coordinated way.

This sounds disturbingly reminiscent of the neo-creationist argument that life is so "irreducibly complex" that there must have been a creator--er, designer. But Kirschner and Gerhart don't believe that for a second. What they argue instead is that organisms are constructed so that genetic change can dramatically alter phenotype without sacrificing key functions--in a process they call facilitated variation.

In future posts I will discuss clues that this construction--I'm avoiding the word "design"-- is present in organisms today, and some of the principles it follows.

Friday, January 15, 2010

Sys Devo



Lawrence Berkeley Labs

My latest eBriefing for the New York Academy of Sciences, Growth Networks: Systems Biology Meets Developmental Biology, is now up (the direct link should work only for Academy members; others may get to it through the NYAS page of my website.)

The symposium was very interesting, but, as often happens, it was challenging to present the three talks as a coherent unit. In this case, the overall message (provided by the visionary organizer, Andrea Califano) is that the sweeping and irreversible changes that occur during early development, which are often driven by a relatively few molecular events (perhaps dozens), can provide stringent and useful tests for understanding molecular regulation. This is quite a different way to learn about networks than by gently poking ("perturbing," for example with stress or drugs or RNA interference) a mature animal, in which various molecules are generally cooperating to keep things stable.

The hope is that the overlap between development and systems biology, can have the sort of powerful synergy that have enriched evolutionary and developmental biology in Evo Devo, as popularized by Sean B. Carroll and others. But I suspect the final synthesis will be more of a three way combination, SysEvoDevo.

Angela DePace of Harvard, for example, described her nascent efforts to exploit evolutionary comparisons between related species from the fly genus Drosophila, which have been a playground for development (once called embryology) for nearly a century. In the past couple of decades researchers have learned how to modify particular genes so they produce fluorescent molecules of various colors along with their normal protein products. The results have shown in living color how various transcription factors interact to generate that spatial patterns and compartments that ultimately shape the segmented body of the fly. DePace and her former colleagues at Lawrence Berkeley Labs refined the technique to let them measure the quantitative changes in gene expression at thousands of individual cells in the early embryo (see the figure), which let them test the models of gene activity (and the differences between species) in fascinating detail.

Stanislav Shvartsman of Princeton also looked at early Drosophila development, but he showed that the transcription factors alone don't explain everything. Instead, some of the patterning depends on protein phosphorylation, which is a half-century old process that among other things carries signals from a cell's outer membrane to its nucleus, but is rarely considered in development. Antonio Iavarone of Columbia studies the development of the early nervous system in mice from stem cells to differentiated neurons. This is a process that is subverted by brain cancers, which re-activate this cellular program to grow and nourish themselves.

Pulling these three diverse talks together was a bit of a shoe-horning exercise, but they were all fascinating.

Wednesday, December 2, 2009

Massachusetts Dreaming

Today I'm taking Amtrak to Cambridge--our fair city--MA, for an exciting back-to-back-to-back trio of conferences at the MIT/Harvard Broad (rhymes with "road") Center.

Two of the conferences are described as satellites to RECOMB (Research in Computational Molecular Biology), even though that meeting was in Tucson in May. One of these is on regulatory genomics and the other on systems biology. The third is the fourth meeting of the DREAM assessment of methods for modeling biological networks, a series I've covered since its organizational meeting at the New York Academy of Sciences in 2006.

There's a lot in common between these conferences, so it's not always easy to notice the boundaries. The most tightly focused is DREAM--Dialog on Reverse-Engineering Assessment and Methods. The goal is simple to state: what are the best ways to construct networks that mimic real biological networks, and how much confidence should we have in the results. In practice, things are not so straightforward, and border on the philosophical question of how to distinguish models and "reality." The core activity of DREAM is a competition to build networks based on diverse challenges.

The Regulatory Genomics meeting covers detailed mechanisms of gene regulation, often focusing on more formal and algorithmic aspects than would be expected in a pure biology meeting. The Systems Biology meeting addresses techniques, usually based on high-throughput experimental tools, for attacking large networks head on, rather than taking the more traditional pathway-by-pathway approach.

I'll be writing synopses of the invited talks and the DREAM challenges for an eBriefing at NYAS, but I'll be free to relax and enjoy the contributed talks and posters. This promises to be a rich and exhausting five days.

Tuesday, December 1, 2009

Packing DNA Beads

The dense packing of DNA in the nucleus of eukaryotes strongly affects how genes within it are expressed, with some regions much more accessible to the transcription machinery than others. At the shortest scales, the accessibility of the DNA double helix is reduced where it is wound around groups of eight histone proteins to form nucleosomes, and the precise position of the nucleosomes in the sequence affects which genes are active.

At a slightly larger scale, the nucleosomes are rather closely packed along the DNA. They can remain floppy, like beads on a string, or they can fold into rods of densely packed beads, which further reduces the accessibility of their DNA. Other proteins in the nucleus, notably the histone H1, help to bind together this dense packing. These rods can pack further, with the help of other proteins.

The histone proteins that form the core of the nucleosome, two copies each of H2A, H2B, H3, and H4, have stray "tails" extending from the core. Small chemical changes at particular positions along these tails can have surprisingly large influence on the expression of the associated DNA. For example, the modification H3K27me3 (three methyl groups attached to the lysine at position 27 on the tail of histone H3) represses expression, while acetylation of the same amino acid, H3K27ac activates expression. There is also a more substantial modification, in which histone H2A is replaced by a variant called H2A.Z also modifies expression.

The detailed mechanisms by which the modifications affect expression, such as changing the wrapping of nucleosomes, the packing of nucleosomes, or recruiting of other proteins in the nucleus, are areas of active research.

Since there are dozens of possible histone tail modifications, there are vast numbers of possible combinations of modifications. Some researchers have proposed that these combinations could each prescribe different expression patterns, for example during development. However, the evidence for a combinatorial "histone code" analogous to the three-base codons of the genetic code remains weak.

Nonetheless, proteins that can modify the tails, either adding or removing a chemical group, can have lasting effects on the activity of the underlying genes. The sirtuin proteins that are candidates for longevity-extending drugs, for example, are best known for their role as histone deacetylases.

Some histone modifications can be passed down through cell division or reproduction, so they qualify as epigenetic changes. In contrast to the natural replication of the mirror-image DNA sequence, replicating histone modifications requires a much more complicated process.

Changes in the pattern of histone modifications are found in many basic biological processes, including development, stem-cell maintenance, and cancer. Particular modification patterns have been used to find specific functional sequences within the DNA, such as transcription start sites and enhancers. For these reasons, the ENCODE project mapped modifications as part of their survey of a select part of the human genome for intense study.

Understanding the mechanisms and roles of DNA organization and how it is changed will be essential to a complete picture of gene regulation.


 

Thursday, November 12, 2009

Guilt by Association

Many of the molecular transformations in cells occur inside of complexes, each containing many protein molecules and often other molecules like RNA.

Determining which molecules are in each complex is a critical experimental challenge for unraveling their function.

Ideally, biologists would identify not just the components, but the way they intertwine at an atomic level, for example using x-ray crystallography. The Nobel-prize-winning analysis of the ribosome showed that this detailed structural information also illuminates how the pieces of the molecular machine interact to carry out its biochemical task.

But growing and analyzing crystals takes years of effort. In many cases researchers are happy just to know which molecules are in which complexes. As a first step, biologists have developed several clever techniques to survey thousands of proteins to see which pairs interact, and to confirm whether those interactions really happen in cells.

Identifying additional protein members of complexes requires chemical analysis like chromatography and increasingly powerful mass spectrometry techniques. In contrast, to explore how DNA and RNA act in complexes, researchers can take advantage of the sequence information available for humans and most lab organisms.

To find out which DNA regions bind with a particular protein transcription factor, for example, biologists use Chromatin ImmunoPrecipitation, or ChIP. Bound proteins from a batch of cells are chemically locked to the DNA with a cross-linker like formaldehyde.

This technique then requires an antibody that binds only to the protein (and its bound DNA), and which is sooner or later tethered to a particle. After breaking apart the DNA, the particle precipitates to the bottom of the solution carrying ("pulling down") its bound molecules, which are then separated and analyzed.

A related technique identifies proteins bound to an antibody-targeted partner. Ideally, the methods identify components that were already bound just before the cells are broken up to begin the experiment, rather than all possible binding sites, so they flag only biologically relevant pairings.

For DNA, the state of the art until recently was "ChIP-chip," which uses microarrays to try to match the pulled down DNA to one of perhaps a million complementary test fragments on an analysis "chip." The advent of high-throughput sequencing has allowed "ChIP-seq," in which the sequence of the bound DNA is directly measured and compared by software to the known genome to look for a match. This was the method used recently to find enhancer sequences by their association with a known enhancer-complex protein.

A similar method can identify the RNA targets of RNA-binding proteins, as discussed by Scott Tenenbaum of the University at Albany at a 2007 meeting that I covered for the New York Academy of Sciences (available through the "Going for the Code" link on my website's NYAS page).

Once the fragments are identified, researchers can try to dissect the elements of the sequence that make them prone to binding by a particular protein. When successful, this procedure allows them to identify other targets for interaction with proteins using only computer analysis of sequence information. These bioinformatics techniques are a critical time saver, because the experiments show that each protein can bind to many different molecules in the cell.

Experiments like these are revealing many of the intricate details of cellular regulation, but also how much more there is to learn.

Wednesday, November 11, 2009

Enhancers, Insulators, and Chromatin

Some DNA regions affect the activity of genes that are amazingly far away in the linear sequence of the molecule.

The best known way that genes turn on and off--and thus determine a cell's fate--is when special proteins bind to target DNA sequences right next to different genes--within a few tens of bases. Together with other DNA-binding proteins, these sequence-specific transcription factors promote or discourage transcription of the sequence into RNA. This mechanism is an example of what's called cisregulatory action, because the gene and the target sequence are on the same molecule.

There's another type of sequence that affects genes on the same DNA molecule, but these can be so far away--tens of thousands of bases--that it seems odd to call them cis-regulatory elements. These "enhancers" can be upstream or downstream of the gene they regulate, or even inside of it, in an intron that doesn't code for amino acids. In fact, some of them affect genes on entirely different chromosomes.

Enhancers have important roles in regulating the activity of genes during development, "waking up" in specific tissues at specific times.

The flexible location makes it hard to find enhancers in the genome, and researchers have also struggled to find clear sequence signatures for them. In ongoing work that I described last year for the New York Academy of Sciences, Eddy Rubin and his team at Lawrence Berkeley Labs instead looked for sequences that were extremely conserved during evolution. Although evolutionary conservation is not a perfect indicator, when they attached a dye near these sequences in mice, they often found telltale coloration appearing in particular tissues as the mouse embryos developed.

Years of effort have uncovered many important clues about how enhancers exert their long-distance effects, but still no complete picture. Most researchers envision that the DNA folds into a loop, bringing the enhancer physically close to the promoter region at the start of a gene. Proteins bound to the enhancer region of the DNA, including sequence-specific proteins that can also be called transcription factors, can then directly interact with the proteins bound near the gene, and enhance transcription of the DNA.

But the enhancement effect can be turned off, for example, when researchers insert certain sequences in the DNA sequence between the gene and the enhancer. These "insulator" sequences seem to restrict the influence of the enhancer to specific territories of the genome. Naturally occurring insulators serve the same restrictive function, but they can also be turned off, for example by chemical modification, providing yet another way to regulate gene activity.

If enhancers work by looping, it seems surprising that an intermediate sequence could have such a profound effect. Researchers have proposed various other explanations as well.

In addition to stopping the influence of enhancers, many insulators restrict the influence of chromatin organization. Biologists have long recognized that the histone "spools" that carry DNA can pack in different ways that affect their genetic activity. In a simplistic view, tight packing makes it hard for the transcription machinery to get at the DNA. This chromatin packing can be modified in the cell, and is one important mechanism of epigenetic effects that persistently affect gene expression even through cell division.

As a further confirmation of the close relationship, Bing Ren of the Ludwig Institute and the University of California at San Diego has successfully used known chromatin-modifying proteins to guide him to enhancers, in work that I summarized from the same meeting last year.

One model that combines some of these ingredients says that loops of active DNA are tethered to some fixed component of the nucleus, and that enhancers can only affect genes on the same loop. If insulators act as tethers, this naturally explains how it limits interactions to particular regions (which are then lops). There is still much to be learned, but enhancers and the chromatin packing appear to be tightly coupled.

I suspect that enhancers have been somewhat neglected both because their action mechanism is so confusing and because definitive experiments have been difficult. But recent experiments done in a collaboration between Rubin's and Ren's teams have used a protein called p300, which binds to the enhancer complex, to identify new enhancers with very high accuracy. Moreover, the binding changes with tissue and development just as the enhancer activity does. These and other experiments are opening new windows into these important regulatory elements.


 

Thursday, November 5, 2009

Evolution of a Hairball

Do DNA sequences evolve more slowly if they play important biological roles?

For many genomics researchers, the answer is so self-evidently "yes" that the question is hardly worth asking. Indeed, they often regard sequence conservation between species, or the evolutionary constraint that it implies, as a clear indication of biological function.

And sometimes it is a good indication. But in general, as I described in my story this summer in Science, this connection is only weakly supported by experiments, such as the exhaustive exploration of 1% of the human genome in the pilot phase of the ENCODE project. That work found that roughly 40% of constrained sequences had no obvious biochemical function, while only half of biochemically active sequences seemed to be constrained.

One reason for this is that many important functions, such as markers for alternative splicing, 3D folding of transcribed RNA, or DNA structure that affects binding by proteins, may have an ambiguous signature in the DNA base sequence.

But another reason (there's always more than one!) is that evolutionary pressure depends on biological context.

Several of my sources for the Science story emphasized that redundancy can obscure the importance of a particular region of DNA. For example, deleting one region may not kill an animal, if another region does the same thing. By the same token, redundant sequences may be less visible to evolution, and therefore freer to change over time. Biologists know many cases of important new functions that have evolved from a duplicate copy of a gene.

But redundancy, in which two sequences play interchangeable roles, is only one of many ways that genetic regions affect each other, and a very simple one at that. As systems biologists have been revealing, the full set of interactions between different molecular species forms a rich, complex network, affectionately known as "the hairball."

For some biologists, the importance of context on evolution is obvious. When I spoke on this subject on Tuesday at The Stowers Institute, for example, Rong Li pointed to the work of Harvard systems biologist Mark Kirschner. Kirschner, notably in the 2005 book The Plausibility of Life that he coauthored with John Gerhart, describes biological systems as comprising rigid core components controlled by flexible regulatory linkages.

The conserved core processes generally consist of many complex, precisely interacting pieces. They may be physical structures, like ribosome components, or systems of interactions like signaling pathways. Their structure and relationships are so finely tuned that any change is likely to disrupt their function, so their evolution will be highly constrained.

In contrast, the flexible regulatory processes that link these core components can fine-tune the timing, location, or degree of activity of the conserved core processes. Such changes are at the heart of much biological innovation. For example, the core components that lead to the segmented body plan of insects are broadly similar, at a genetic level, to those that govern our own development. Our obvious differences arise from the way these components are arranged in time and space during development.

The weak predictive power of conservation is particularly relevant as researchers comb the non-protein-coding 98.5% of the genome for new functions. Many of these non-coding DNA sequences are regulatory, so they may evolve faster. Indeed, Mike Snyder of Yale University observed a rapid loss of similarity in regulatory RNA between even closely-related species in deep sequencing studies he described at a symposium at the New York Academy of Sciences (nonmembers can get to my write-up by following the "Go Deep" link at the NYAS section of my website).

Quantifying how evolutionary pressures depend on the way genes interact is likely to keep theorists busy for years to come. But it is clear that the significance of evolutionary constraint in a DNA sequence--or its absence--depends very much on where it fits in the larger biological picture.

Wednesday, October 28, 2009

ENCODE

The mapping of the human genome in draft form in 2000 was a turning point in biology. But the announcement really marked a start, rather than an end, of the practical use of genomic information.

Several large-scale projects in the intervening years have mined particular aspects of the genome and combined it with other sources of information. The HAPMAP project, for example, looked at how common single-nucleotide polymorphisms, or SNPs--changes in a single base--varied among a few selected populations. These studies formed the basis for genome-wide association studies to identify DNA regions associated with various diseases.

Another project, called ENCODE, for ENCyclopedia Of DNA Elements, focused on cataloguing the various types of protein-coding and regulatory structures in the genome. By correlating these with biochemical measures of functional activity and the way the DNA is organized in the nucleus, the researchers got a broad view of how the expression of genes is regulated. They also compared the DNA sequences with those of closely and distantly related organisms, to illuminate how the function of the DNA is related to its evolution.

With some 200 co-authors from 80 different institutions, the ENCODE project rivals some big particle-physics experiments for scope and complexity. In fact, the pilot phase of ENCODE selected "only" 1% of the genome--around 3 million bases--for detailed study. An overview of the results appeared in Nature in June 2007. The data are publicly available, and researchers continue to publish papers on aspects of the work. In addition, follow-up work is aimed at analyzing the entire genome.

Among the profound conclusions from the pilot phase that most of the genome is transcribed into RNA, even though only 1.5% or so codes for protein and only about 5% seems is clearly functional. In other words, much of the regulation of genetic activity may be occurring, not at the level of transcription, but at the level of RNA.

The researchers also found that the organization of the chromosomes in the nucleus, in particular the wrapping of the DNA around histones to form nucleosomes, predicts the locations where transcription begins. These results emphasize the known importance of the positions of nucleosomes in regulating genetic activity at different positions.

Some of the researchers looked at various measures of biochemical activity along the DNA, such as binding to proteins that are known to be active in regulation. Their hope was that these assays would serve to identify regions with a biological function in the cell.

Other ENCODE researchers compared the sequences with corresponding sequences from other organisms--both close relatives like mice and distant eukaryotic relatives like yeast. According to a longstanding assumption, the degree of similarity of these sequences, showing how resistant they are to changes from neutral evolution, should also reflect their biological importance.

These studies revealed two surprises. First, not all biochemically active sequences are evolutionarily constrained. This might mean that the biochemical tests don't measure things that are important to the cell after all. Second, and more puzzling, not all of the constrained sequences had any obvious function.

I wrote a story for Science this summer (subscribers only, sorry) discussing possible reasons why evolution and importance don't always track one another.

ENCODE and other large-scale studies will continue to supply us with extensive, detailed information about the genome. The story is only just beginning.

Thursday, October 15, 2009

A Missing Layer

Models of biological networks have always had gaps, but they are bigger than most researchers realized.

Only in the past few years have biologists begun to recognize the extensive regulatory role of naturally occurring small RNAs. The best known of these endogenous RNAs are chains of 21-23 nucleotides with the rather unfortunate designation of "microRNA" (whose abbreviation, miRNA, is awkwardly similar to the mRNA used for messenger RNA). MicroRNAs arise from sections of DNA whose RNA transcripts contain nearly complementary mirror-image sequences, and so naturally fold back on themselves to form a "stem-loop" structure. Processing by a series of protein complexes liberates one strand from the overlapping section and incorporates it into specialized RNA-protein complexes in the cytoplasm that modify protein production.

Traditionally, systems biologists aiming to unravel gene-regulation networks have relied on the wealth of data from microarrays that measure the mRNA precursors of proteins. By regarding the mRNA abundance as a proxy for the corresponding protein, and looking at how various mRNA levels change with cellular conditions, the researchers construct hypothetical networks of interacting genes. In the graphical representations of these networks, genes are connected by a line or "edge" if the protein product of one seems to act as a transcription factor to change the activity of the other.

MicroRNAs complicate this picture dramatically, although many researchers don't yet incorporate them. Improved tools, often directly sequencing of millions of fragments rather than matching pre-chosen sequences in microarrays, let researchers survey the small RNAs in the cell. In many cases, as in the work of Frank Slack of Yale University described in my latest eBriefing for the New York Academy of Sciences, microRNAs control cellular processes in much the same way as traditional protein transcription factors--in Slack's case extending the lifespan of worms. These regulatory RNAs are a previously unsuspected layer of genetic regulation.

Some of the RNA-protein complexes promote degradation of messenger RNA that is complementary to the bound miRNA. In this case the remaining mRNA could still be a good indicator of a gene's activity, although not of its original transcription. Sometimes accounting for the miRNA might require only a change in the mathematical relationship between genes.

Many miRNAs, however, as well as some transcription factors, act as "master regulators," generating coordinated activity among scores of genes. As a result, these master regulators can effectively change one genetic network into an entirely different one--adding or removing edges. For example, researchers have constructed networks for cancerous cells in which the connections differ markedly from those for their healthy counterparts. Such context-dependent networks may be simple and accurate in specific situations, but they obviously lack important ingredients.

In other cases, a miRNA can continuously vary the activity of genes, rather than being a simple on/off switch. Again, such coordinated response will be hard to capture unless the hidden factor is explicitly identified.

A second type of RNA-protein complex slows (or less often speeds) the translation of complementary messenger RNA. One profound implication is that the measured levels of mRNA may no longer be a good proxy for the levels of the protein produced from it. Indeed, in the few cases where researchers have done the experiments, they have found only weak correlations between the levels of mRNA and the corresponding proteins. Any procedure that depends on these levels being equivalent is on thin ice.

In addition to these quantitative effects, qualitatively new behavior appears when molecules are connected in feedback loops. Systems biologists have catalogued the action of many interesting "motifs." Even two molecules, for example, can act to stabilize concentration--if the feedback around the loop is negative--or act as a bistable switch--if the loop feedback is positive. Clearly, if such motifs are acting in the hidden mRNA layer, no tweaking of the gene-gene interactions will replicate their effects.

All of this reinforces the need for researchers to develop and use as many high-throughput techniques as possible to measure different types of RNA as well as the different states of proteins in cells. The "reverse engineering" of networks never seemed easy. Now it's clear that it's even harder than it seemed.

Friday, September 25, 2009

The Map and the Territory

The map is not the territory. Alfred Korzybski

I confused things with their names: that is belief. Jean-Paul Sartre

Ceci n'est pas une pipe. René Magritte

In fields ranging from economics to climate to biology, scientists build representations of collections of interacting entities. Everyone knows that the real systems have so many moving parts, influencing each other in poorly known ways, that any representation or model will be flawed. But even though they understand the limitations, experts routinely talk about these systems using words that come from the models, rather than from reality. Climate scientists talk of the "troposphere," economists talk of "recessions," and biologists talk of "pathways." Such concepts help us organize our thinking, but they are not the same as the real thing.

Sometimes the difference between the "map" and the "territory" is manageable. Roads and rivers are not lines on a piece of paper, but they clearly exist. Similarly, the frictionless pulleys and massless ropes of introductory physics have a simplified but clear relationship to their real-world counterparts (at least after you've spent a semester learning the rules). Still, it's easy to get sucked into thinking of these well-behaved theoretical entities as the essence, the Platonic ideal, even as one learns to decorate them with friction and mass and other real-world "corrections."

For many interesting and important problems, however, the conceptual distance between the idealizations and the boots-on-the-ground reality is much larger. You might think that experts would recognize the cartoonish nature of their models and treat them as crude guides or approximations, rather than fundamental principles partially obscured by noisy details. Judging from the never-ending debates in economics, however, the more obscure the reality, the more compelling the abstractions become.

Even in less contentious fields, experts can mistake the models for reality. For example, the fascinating field of systems biology aspires to map networks containing hundreds or thousands of molecules using high-throughput experiments like microarrays and computer analysis. Although one might like to describe all these interactions using coupled partial differential equations, researchers would often be happy simply to list which molecules interact. This information is often represented as a graph--sometimes called a "hairball"--which represents each molecule as a dot or node, and interactions as a line or edge connecting them.

Finding such graphs or networks is a major goal of systems biology. In principle, an exhaustive map is more useful than the traditional painstaking focus on particular pathways, which are presumably a small piece of the entire network. But to yield benefits, researchers need to understand how "accurate" the models are.

A few years ago, a group of systems biologist decided the time was ripe to critically evaluate this accuracy. They established the "Dialogue on Reverse Engineering Assessment and Methods," or DREAM to compare different ways of "inferring" biological networks from experiments. (I covered the organizational meeting, as well as meetings in 2006, 2007, and 2008, under the auspices of the New York Academy of Sciences. A fourth meeting, which like the third will be held in conjunction with the RECOMB satellite meetings on Systems Biology and Regulatory Genomics, is scheduled for December in Cambridge, Massachusetts.) These meetings, including competitions to "reverse engineer" some known networks, have been very productive.

Nonetheless, one thing the DREAM meetings made clear is that "inferring" or "reverse engineering" the "real" networks is simply not a realistic goal. Once the networks get reasonably complicated, it's essentially impossible to take enough measurements to clearly define the network. The ambiguity even applies to networks that actually have been engineered, that is, created by people on computers. The "inferred" networks are a useful computational device, but they are not "the" network. And they never will be.

For these reasons, many researchers think the only proper way to assess the results is by comparing to experiments. If the models are good, they should not only match observed data, but should extrapolate to accurately predict what happens in a novel situation, such as the response to a new drug. Interestingly, the most recent DREAM challenges included tasks of this type. Disappointingly, however, the methods that best predicted the novel responses simply generalized from other responses: they did not include any network representation at all!

It seems reasonable to expect that a model that tries to mimic the internal network, even if it is flawed, would better predict truly novel situations. But it's hard to know what it will take for the system to hit a tipping point where it does something completely different, which was never observed before or included in the modeling. Often, we won't recognize the limitations of our complex models--in biology, climate, or economics--until they break.

Wednesday, September 23, 2009

Epigenetics

When it comes to inheritance, there's no beating the DNA sequence for storing and passing on complex information. But other, "epigenetic" mechanisms also bequeath information to subsequent cells or offspring, sometimes in response to environmental changes.

In principle, the word "epigenetics" could apply to any inheritance outside of the genetic sequence. For example, when a cell divides, its contents are divided among the daughter cells. Any transcription factors or other chemicals that alter gene expression are therefore passed on independently of the DNA (along with the mitochondria, which have their own DNA). In recent years, however, "epigenetics" has come to be used mainly to describe two types of chemical changes directly associated with DNA in the nucleus, other than its sequence.

These changes modify how active various genes are in a particular cell. They are particularly important for enforcing the "no turning back" feature of differentiation from versatile stem cells to specialized cells, helping to shut off cellular programs that were active in the early embryo. Moreover, epigenetic changes are passed on during cell division, so that the differentiated cells and all cells made from them lose their ability to become other types of cell. It should not be surprising that many cancers subvert the epigenetic programming to help them re-activate embryonic programs to help them survive and spread. Researchers have identified many epigenetic modifications in cancer cells.

Epigenetic changes can also pass between generations. Biologists have long known of cases of "imprinting," in which the mother's or the father's DNA is inactive in the offspring. Even in people, researchers have found that food shortages in Holland at the end of World War II resulted in changes in the metabolism of the children of women conceived during that period. Such effects are unusual, but profound.

This sounds disturbingly like inheritance of acquired characteristics, as in Kipling's "Just-So Stories." This concept, often misleadingly associated with early 1800's evolution pioneer Jean-Baptiste Lamarck, was supplanted by Darwin's notion of natural selection of random variations. But persistently activating or suppressing pre-existing genes for a few generations, even in response to environmental pressures, is not the same thing as creating novel properties. Some scientists, notably Eva Jablonka of Tel Aviv University, maintain that epigenetic effects can be permanently enshrined in the sequence, but that remains a minority view. Equating epigenetics with Lamarckism is misleading, despite having a grain of truth.

The two best known epigenetic mechanisms are chemical changes that alter the transcription of DNA. One mechanism modifies the DNA itself, while the other modifies the packaging of the DNA in the nucleus.




(Click to open in new window. Source: NIH)

In DNA methylation, methyl (-CH3) groups are chemically bonded to a base in the DNA sequence, usually a cytosine (C) next to a guanine (G), together called CpG. The presence of the methyl group suppresses translation of the DNA sequence that contains it. In addition, the cell contains enzymes that recognize methylation of one chain of DNA and methylate the other chain, helping to propagate the information.

The second mechanism affects the packing of the DNA into the compact structure known as chromatin. The paired DNA chains wrap tightly around a cluster of proteins called histones to form a nucleosome. Nucleosomes strung along the DNA chain themselves pack into compact arrangements that make it hard for the transcription machinery to get at them.

The details of this process are only partially understood. One thing that is known is that free "tails" of the histone proteins straggle out of the nucleosomes, and that chemical modifications of these tails modifies transcription. The modifications include single or multiple methylation or acetylation (adding -COCH3) of particular amino acids positions in the tail, as well as binding of other factors. The details matter: particular modifications either increase or decrease transcription.

In recent years researchers have developed techniques for mapping both DNA methylation and chromatin modification over large regions of the genome. Using these techniques and others, biologists are beginning to understand when and where these epigenetic modifications occur in normal and diseased cells, how nutrition and other environmental influences change them, and how specific modifications are actively regulated to modulate gene expression.

Thursday, September 17, 2009

Pathways to Disease

Most common diseases, including the big killers like heart disease, are "complex": they can't be blamed on single causes like a particular gene. Instead, they result from a complicated interaction of factors that may include lifestyle, environmental exposures, or infection, as well as genetic effects. Moreover, large-scale surveys of genetic influences have confirmed that, in many cases, lots of different genes contribute to disease, each in a small way.

These generalizations also apply to cancer. Cancer differs from the other diseases because most of the genetic changes in cancer cells aren't present in the rest of the patient's cells. Instead, mutations, copy number variations, and large-scale chromosome anomalies accumulate as the disease progresses. These alterations are often abetted by early disruptions of the usual mechanisms for maintaining genome quality during cell division and for executing damaged cells. In spite of these differences, the first major results last fall from The Cancer Genome Atlas comparing the genetics of glioblastomas (deadly and virtually untreatable brain cancers) found no specific mutation was present in all of the tumors. The huge team of researchers did a comprehensive analysis including gene expression, copy number changes and epigenetic changes. But although some changes happened rather frequently, there was no single "smoking gun."

Nonetheless, these studies, in both cancer and other diseases, find clear patterns among the genes whose activity is altered in one way or another. When researchers put the changes in the context of the complex network of molecular interactions in the cell, most of the changes cluster along clear "pathways." As Todd Golub told a meeting I covered last year, just after the glioblastoma results were published: "What was gratifying about this was that this was not just a sprinkling of mutations randomly across the genome, which were difficult to decipher in the context of any kind of mechanistic understanding, but rather these were falling together in a set of pathways that were increasingly well understood in cancer."

I regard the word "pathway" is a bit of a misnomer, since it suggests a linear sequence in which each molecule affects the next one in a chain. In the early days, that was about all that experiments could get at, but researchers have long recognized that networks are messier than this. For example, there may be multiple, parallel influences of one molecule on another, and there are almost always feedback paths in which the final outcome comes back to modify the early steps.

Nonetheless, although they are complex and interconnected, these pathways give researchers a useful shorthand for navigating the rich networks of interactions and for communicating with others. In fact, many researchers specialize in particular pathways, getting to know each molecular member "personally," as well as the effects they have on one another.

Results like the glioblastoma study also show that the pathway level may be a more useful level of "granularity" for thinking about disease than the individual molecules are. Focusing on pathways (or "modules," or "motifs," or whatever) gives us simple-minded humans a better intuitive understanding of a disease, which is important. Moreover, in treatment, researchers can be led astray by focusing on molecular-level changes such as individual genetic variants, since these are not the same for everyone. Targeting specific pathways, for example with combination therapies that attack several "nodes" of the network at once, may prove to be more effective against diverse groups of patients.

But the most important benefit of isolating pathways may be that many of them are shared by different diseases, which is leading to new insights into the relationships between diseases.

Tuesday, September 15, 2009

Targeting Cancer

If personalized medicine ever becomes widespread--and I hope it does--it will probably start with cancers.

In fact, it already has. More than ten years, ago, in 1998, the FDA approved the Genentech monoclonal antibody Herceptin (trastuzumab) as part of treatment for metastatic breast cancer--but only for patients who overexpress the membrane receptor ErbB-2 (also called HER2). For these patients, the extra copies of ErbB-2 generate signals that make the cancer spread more aggressively. The antibody binds to the receptor and diminishes this effect. But Herceptin was only shown to be effective in people who, as shown by laboratory tests, have an excess of the receptor. The approval was conditional on positive test results.

Cancers ought to be the best case for personalized medicine because treatment decisions are made by experts in the disease, based on medical tests and observations. These experts recognize that different tumors respond differently, and they are accustomed to adjusting treatment accordingly. In contrast, for many other diseases, such as mental illnesses, doctors often depend on more subjective symptoms, and patients are susceptible to the default "one size fits all" advertising of pharmaceutical companies. Cancer treatment is still the province of experts.

But it's important to ask whether those experts are doing what they need to, to get the drug to the people who will benefit, and not to the people who will not. A new article in Cancer addresses this question, and the answers are troubling.

The main complaint of the article is that there's not enough data to know. Kathryn Phillips, of the Center for Translational and Policy Research on Personalized Medicine at UCSF, and her colleagues find that in many cases there is no documentation that patients are receiving the right tests to guide their treatment. They also cite other results that

  • Perhaps two thirds of patients who could get the test to see if Herceptin would be appropriate may not get it (at least it's not recorded). By implication, many patients aren't getting a treatment that might help them.
  • A fifth of patients who do get the drug have no record of having gotten the test. This means that patients may be taking a drug, and suffering its cost and side effects, without any evidence that it will help them.
  • A fifth of the test results may be incorrect.

The argument for approving the drug was that it would make treatment cheaper and more effective. That only makes sense if the tests are given, are accurate, and are used to guide treatment. The success of personalized medicine depends on new, reliable procedures for ensuring that treatment is coupled with validated tests. If it can't be done with cancer treatment, it's hard to believe that it's a realistic goal for other diseases.

By the way, the researchers get funding for their research (said to be unrestricted) from the foundations of major health insurance companies. I'm not sure what to make of that.

Saturday, September 12, 2009

E pluribus unum

A few diseases can be traced to specific genetic variants. The nerve degeneration of Huntington's Disease, for example, arises exclusively from alterations of either copy of a gene on chromosome 4. This gene specifies a protein that is now called huntingtin. Such diseases are referred to as Mendelian, since they follow the simple rules of inheritance that Gregor Mendel observed in his pea plants.

For most diseases, though, it has proved difficult to find individual genes that explain much of the risk. Instead, the growing evidence from large-scale studies is that many variants contribute, each contributing only weakly. Even then, the genes alone do not condemn a person to the disease, which may also depend on microbes or non-living elements of the environment or on lifestyle. These "complex" diseases include all of the biggies, like heart disease and stroke, cancers, and many mental illnesses.

In some ways, the failure of the "one-gene/one-disorder" hypothesis shouldn't be too surprising. After all, a gene that reliably causes a fatal disease should have been largely weeded out by natural selection. Huntington's disease avoids this fate because it usually appears late in life, often after people have already had children (including, fortunately for us, Arlo Guthrie). Sickle-cell anemia persists because people with a single variant gene are resistant to malaria, although two copies cause the disease.

Nonetheless, lots of other diseases have an import genetic component, which can be determined by comparing the disease rate for close relatives. For example, if pairs of "identical" twins are more likely to both get a disease than are fraternal twins, the difference presumably arises because they share their entire genome, rather than only half.

For simple Mendelian diseases, researchers have extended this approach to locate where the disease gene resides in the chromosomes. This "linkage" analysis looks at which close relatives inherited a disease, and what known chromosome features they also inherited. This technique was applied in the 1980s to locate the Huntington's gene by testing dozens of residents of a Venezuelan village that had unusually many cases.

But human populations aren't particularly well suited for linkage studies. People don't have a lot of children, and they resist attempts at controlled breeding. As a result, it's hard to see weak genetic effects.

To get more subjects, researchers use association studies, which compare the genetics of unrelated individuals. Historically, you really had to know where to look to make associations studies work. But in the past few years researchers have done dozens of "genome-wide association studies," or GWAS, that look without prejudice across the entire human genome.

These studies are tricky. For one thing, since they monitor perhaps a million genetic markers at once, the chances are good that a marker will correlate with the disease by dumb (bad) luck. In individual experiments, researchers traditionally ignore a result if the probability of it arising by chance isn't less than 5% (P<0.05). For testing a million markers, they might need to ignore a result unless the effect is so strong that the probability that it arose by chance is less than perhaps 5x10-8. To get such a convincing effect requires a lot of human subjects, generally hundreds or thousands. Even so, GWAS results often fail to recur when someone else tries the experiment.

Nonetheless, some genome-wide studies, like two for Alzheimer's I wrote about recently, have uncovered genes repeatedly associated with disease. In addition to variations of the DNA sequence, these studies often include structural variants such as copy-number variations, as well as "epigenetic" tags that change the expression of particular DNA regions. In spite of finding some likely genetic suspects, though, the total effect of all of the known variants is generally less than the known genetic component of these complex diseases. Researchers are actively debating the causes of this discrepancy; probably part of it comes because there are other contributions that are too weak to be seen in these studies.

Because complex diseases depend on the small contributions of many genetic variants, as well as the environment, buying your personal genome often won't tell you much definitive. But by studying these variants, and the way their effects interact in cells, researchers are learning a great deal about the nature of the diseases, including potential strategies for treating them.

Monday, September 7, 2009

Alzheimer's and Inflammation

Two online letters, just out in Nature Genetics (here and here), found three genes that had a statistically significant correlation with late-onset Alzheimer's disease. For the past 16 years, only one gene, APOE, had been connected with this common form of the disease, explaining about half of its genetic heritability. In contrast, the rare, early-onset form has a more classic "Mendelian" genetic pattern, in which, if you have a mutation in one of three genes, you have a high probability of getting the disease.

The new results have the common disappointments of the last few years of genome-wide association studies, or GWAS, of complex diseases: (1) The effects of any particular genetic variant are weak, so they can only be seen by studying thousands of subjects. (2) Because half a million candidate mutations are tested simultaneously, it's hard to assess the significance of something that looks like an association. It's easy to pick up false positives by chance alone, so researchers need to apply big corrections. (3) Different studies identify different variants. In this case, the studies agree on a gene called CLU, but each of them also finds another gene that the other study doesn't. (4) The cumulative effect of all the variants found is not enough to explain the observed heritability of the disease. It seems that there must be many other, unidentified contributors, each having only a small effect.

Confirming this weakness, co-author Michael Owen, of Cardiff University in Wales, noted in a supplementary statement on the Nature Genetics website that "the current genes on their own are not strong predictors of risk and are not suitable for risk testing." I'll have a lot more to say about GWAS and disease in future posts.

But although the genes aren't very useful for predicting risk, they do give clues about the biological mechanisms of the disease. Most previous discussions of Alzheimer's, including these two papers, concerns two types of protein deposits in brain cells: "plaques" of β-amyloid protein and "tangles" of tau protein. Both of these deposits are often seen in the brains of Alzheimer's patients after they die. Clusterin, which is the protein coded by CLU, may help clean up the plaques.

But in her supplementary statement, Julie Williams, also of Cardiff, noted that "clusterin has a role in dampening down inflammation in the brain. Up until now increased inflammation seen in the brains of Alzheimer's sufferers had been viewed as a secondary effect of disease. Our results suggest the possibility that inflammation may be primary to disease development."

This reminded me of Paul Ewald's talk at a January 2007 symposium at Hunter College, "Evolution, Health, and Disease," which I covered on behalf of the New York Academy of Sciences. Ewald, of the University of Louisville, noted that inflammation of arterial plaques is a common feature of the atherosclerosis that often leads to heart disease. (The test for inflammation using the C-reactive protein (CRP) is often used to predict heart-attack risk.) But he also noted that the troublesome ε4 variant of the EPOE gene "is the major risk factor, not only for atherosclerosis and stroke, but also for sporadic Alzheimer's and multiple sclerosis," even though the fat transport that influences atherosclerosis is a completely different chemical property than the formation of protein plaques in Alzheimer's or the myelin-sheath destruction in multiple sclerosis. "The idea that ε4 would be bad in all of these different ways," Ewald said, "is really stretching it."

Instead, Ewald suspects that the common element in these various diseases is infection, perhaps by Chlamydia pneumonia. I imagine that his view remains on the fringe, and perhaps it will remain there. But 25 years ago, the idea that many ulcers are caused by a bacteria was also a fringe idea. Barry Marshall and Robin Warren won the 2005 Nobel Prize in Physiology or Medicine for tracing ulcers to Helicobacter pylori. Maybe in 25 years we will find it natural to associate Alzheimer's with infection, too.