Wednesday, September 30, 2009

Never Fold Alone

Predicting the structure of a protein--the three dimensional pattern that a particular sequence of amino acids folds into to become biologically active--is a perennial challenge of biology.

Researchers have long recognized major drivers of the final shape, such as exposing hydrophilic amino acids to the aqueous environment, while keeping hydrophobic amino-acids tucked away safely inside or in regions that will lie inside of membranes. Chemists back to Linus Pauling have also recognized recurring structural motifs such as alpha helices and beta sheets that allow somewhat regular packing. But even with these constraints, a chain of hundreds of amino acids can arrange in an astronomical number of ways. Exploring these configurations one by one would take virtually forever, so how do real proteins find the few configurations that will let them do their biological job?

One answer is that they don't always succeed. Generally tens of percent of the molecules get mangled along the way and have to be disposed of. But this just reduces the astronomical challenge by a small factor.

Another important fact is that proteins don't fold in a vacuum--or even in a water environment. Even as it is being translated from messenger RNA by the ribosome, a growing polypeptide is joined by proteins called chaperones. These key proteins help to ensure that the new chain folds properly, and also keeps it from aggregating with others (which is another way to tuck away hydrophobic amino acids).

These molecular chaperones are the best-known members of the family of heat-shock proteins (denoted hsp), which are produced in large quantities by cells that have been stressed. Heat, for example, tends to disrupt protein folding, and the chaperones can help put them back together again. In addition to the small hsp70 chaperone that binds to the growing protein, another protein called hsp60 forms a kind of dressing room where the still-folding protein can assemble itself in privacy.

This activity of these chaperones is driven by ATP, the cell's energy currency. Here is a movie of both processes. I'm afraid it didn't help me much, though.

The important point is that protein folding in a cell, like the processing of DNA and RNA, involves the close coordination of other biological macromolecules. This may be part of the reason that, although researches have made a lot of progress in structure prediction from sequence, in part by draw analogies with similar sequences in proteins with known structure, they still struggle with completely novel sequences.

Folding is only one step in the processing of proteins. They also may be acetylated or phosphorylated, crosslinked with sulfur, and combined with metals like iron, zinc, or manganese. They will be decorated with sugars that can, for example, serve as address labels for their final destinations. Those proteins headed for membranes will not be sent out into the cell to fend for themselves, but will bound with membrane and directly handed off. Much of this activity happens in the endoplasmic reticulum, where proteins that have been mangled are identified and recycled.

Even after processing, many proteins will be further modified chemically, for example by adding or removing phosphate groups to modify their activity. Moreover, many proteins do their work as part of complexes with other proteins, either in pairs or other small groups or in larger complexes that may include RNA.

Biology (the reality, as well as the science) is a team sport.

Tuesday, September 29, 2009

Science/Journalism

I've gotten many good insights from Chris Mooney. In a 2004 story in the Columbia Journalism Review called Blinded by Science, for example (oddly unlinkableposted here), he criticized the journalistic tradition of "balance," as it applied to climate change. He explained that although including diverse points of view gives an impression of objectivity, this habit was giving undeserved credibility to the rare deniers of the consensus on climate. In the intervening years, journalists have become more aware of this problem and more frank in distinguishing the mainstream from the fringe (supported by the increasingly dire predictions of the mainstream view).

In one small section of their recent book, Unscientific America: How Scientific Illiteracy Threatens Our Future, Chris and his coblogger at The Intersection, Sheril Kirshenbaum, expand on this and other ways that journalistic traditions obscure scientific realities. Chief among the disconnects is the news focus on, well, news: what's happening now that we didn't know yesterday? Such event-driven coverage serves poorly many ongoing trends in science (as well as in other areas) that develop continuously or incrementally. The need for a "hook" drives reporters to focus on specific articles in the big journals, rather than the accumulating evidence that they are merely an example of.

Journalists are also prone to framing stories around human elements, especially conflict. There are good reasons for this: people read these stories. But the focus on personalities or revolutions often distracts from the real issues. Biobloggers Larry Moran and T. Ryan Gregory, for example, routinely complain about the misleading narrative that "scientists used to think most of the genome was 'junk," but now they've realized it's good for something." (Scientists have long known that much of it was good for something. Much of it is still junk.)

These differences--driven largely by the business of journalism--are important. Scientists who can't follow Mooney and Kirshenbaum's dictum to transform into public communicators would do well to appreciate what happens to their message when it leaves their hands.

Nonetheless, as someone who has morphed from one to the other, I think the similarities between scientists and journalists are greater than the differences. At a fundamental level, both are professionals dedicated to uncovering reality, wherever it lies. Both groups rely on evidence, and treat personal opinions and popular fads with suspicion, as much as they can recognize them. In each profession, there is a strong social obligation that transcends any loyalty to one's employer or even to one's own prejudices. It is an obligation, as best one can, to speak the truth.

Monday, September 28, 2009

We Did All We Could

As you read this on your computer screen, it's easy to take for granted the billions of transistors--driving the screen, running the programs, storing the data, and bringing it to you over the internet--that make it all possible.

This embarrassment of transistors is affordable because they're made in a parallel process that produces vast numbers of similar devices at once, combined into integrated circuits (ICs). Making sure that they each behave the way they're supposed to demands extraordinarily clean and reproducible manufacturing processes. In fact, after inventing the transistor, Bell Labs was late to the IC party because they didn't think anybody could get them all to work at once.

Later, Bell Labs' parent, AT&T, did get good at ICs. Towards the end of my time in semiconductor device research at Bell Labs I worked with the excellent developers of the upcoming integrated circuit generations for what was then AT&T Microelectronics, who had moved to Orlando, Florida.

One benefit of visiting Orlando and learning about their challenges was that I managed to design some test structures that they included on the photomasks they used to develop their process. It took some convincing for them give up even a tiny piece (about 0.002 square centimeters!) of their very precious real estate. They also need to be sure that my devices wouldn't flake off and mess up other structures that they needed to do their real work.

Months later, it was a real rush to get the first silicon wafers with my devices on them.

First, the structures looked exactly like what I designed. Instead of looking at multicolored rectangles in a CAD program on a computer screen, though, I was looking at multicolored rectangles in a microscope: real semiconductor devices.

Second, there were lots of them. Even though the entire array of test structures was over a square centimeter in area, there were dozens of repetitions on each eight-inch-diameter silicon wafer.

Third, they were all the same. They didn't just look the same: on the unfortunate occasions when I blew one up with too much voltage, I learned that its repeated version would have very much the same electrical behavior.

I also made friends with people who did testing, robotically stepping across the wafer to measure each repetition. So for simple measurements, after a lot of up-front planning, I could sit back and let the data roll in. Whenever development ran a lot of 25 wafers through the several hundred steps it took to get finished ICs, they also made me hundreds of test structures, and measured them, too.

Compared to what I was used to in the physics labs, where you might work weeks to get a sample or two, this was heaven.

With lots of people helping out, we also did something more challenging, which was to explore new ways to process the wafers. For example, my research colleague Joze Bevk devised a scheme to improve the addition of electrical dopants into the narrow poly-crystalline-silicon ribs that formed the gates of the transistors. Our development colleagues helped track the wafers through step after step of the modified process.

One day, when Joze and I were visiting Orlando, our colleague Steve Kuehne approached us. In the matter of a surgeon telling waiting relatives "I'm sorry. We did all we could," Steve gave us the bad news: "The gates are falling off." Joze and I were very disappointed at this failure, since from Steve's grave expression it was clear that the result was a disaster.

Over the next hour or so, as we discussed what sort of stresses in the materials might cause these terrible problems, an interesting fact emerged. Out of many millions of gates on the test wafer, perhaps 20 had fallen off! Only the high-throughput measurement tools in the development line, which scan the entire wafer looking for anomalies, could even detect them. This is what Steve meant when he said the gates were falling off. For him, a process with even that many broken devices was a non-starter.

I don't doubt that the developers could have devised modifications of the process that reduced the problem, it if had seemed worthwhile--or if they had invented it themselves. Nonetheless, it was a powerful reminder of the degree of reproducibility that IC manufacturing demands.

When I see a news story about some new technique that's going to change the way ICs are made (like this one or this one--not to pick on IBM), I remember how few failures are deadly. If you can see variation in a handful of devices, then someone is going to have to do an awful lot of work before they can be made by the billions.

Now go back to taking them for granted.

Friday, September 25, 2009

The Map and the Territory

The map is not the territory. Alfred Korzybski

I confused things with their names: that is belief. Jean-Paul Sartre

Ceci n'est pas une pipe. René Magritte

In fields ranging from economics to climate to biology, scientists build representations of collections of interacting entities. Everyone knows that the real systems have so many moving parts, influencing each other in poorly known ways, that any representation or model will be flawed. But even though they understand the limitations, experts routinely talk about these systems using words that come from the models, rather than from reality. Climate scientists talk of the "troposphere," economists talk of "recessions," and biologists talk of "pathways." Such concepts help us organize our thinking, but they are not the same as the real thing.

Sometimes the difference between the "map" and the "territory" is manageable. Roads and rivers are not lines on a piece of paper, but they clearly exist. Similarly, the frictionless pulleys and massless ropes of introductory physics have a simplified but clear relationship to their real-world counterparts (at least after you've spent a semester learning the rules). Still, it's easy to get sucked into thinking of these well-behaved theoretical entities as the essence, the Platonic ideal, even as one learns to decorate them with friction and mass and other real-world "corrections."

For many interesting and important problems, however, the conceptual distance between the idealizations and the boots-on-the-ground reality is much larger. You might think that experts would recognize the cartoonish nature of their models and treat them as crude guides or approximations, rather than fundamental principles partially obscured by noisy details. Judging from the never-ending debates in economics, however, the more obscure the reality, the more compelling the abstractions become.

Even in less contentious fields, experts can mistake the models for reality. For example, the fascinating field of systems biology aspires to map networks containing hundreds or thousands of molecules using high-throughput experiments like microarrays and computer analysis. Although one might like to describe all these interactions using coupled partial differential equations, researchers would often be happy simply to list which molecules interact. This information is often represented as a graph--sometimes called a "hairball"--which represents each molecule as a dot or node, and interactions as a line or edge connecting them.

Finding such graphs or networks is a major goal of systems biology. In principle, an exhaustive map is more useful than the traditional painstaking focus on particular pathways, which are presumably a small piece of the entire network. But to yield benefits, researchers need to understand how "accurate" the models are.

A few years ago, a group of systems biologist decided the time was ripe to critically evaluate this accuracy. They established the "Dialogue on Reverse Engineering Assessment and Methods," or DREAM to compare different ways of "inferring" biological networks from experiments. (I covered the organizational meeting, as well as meetings in 2006, 2007, and 2008, under the auspices of the New York Academy of Sciences. A fourth meeting, which like the third will be held in conjunction with the RECOMB satellite meetings on Systems Biology and Regulatory Genomics, is scheduled for December in Cambridge, Massachusetts.) These meetings, including competitions to "reverse engineer" some known networks, have been very productive.

Nonetheless, one thing the DREAM meetings made clear is that "inferring" or "reverse engineering" the "real" networks is simply not a realistic goal. Once the networks get reasonably complicated, it's essentially impossible to take enough measurements to clearly define the network. The ambiguity even applies to networks that actually have been engineered, that is, created by people on computers. The "inferred" networks are a useful computational device, but they are not "the" network. And they never will be.

For these reasons, many researchers think the only proper way to assess the results is by comparing to experiments. If the models are good, they should not only match observed data, but should extrapolate to accurately predict what happens in a novel situation, such as the response to a new drug. Interestingly, the most recent DREAM challenges included tasks of this type. Disappointingly, however, the methods that best predicted the novel responses simply generalized from other responses: they did not include any network representation at all!

It seems reasonable to expect that a model that tries to mimic the internal network, even if it is flawed, would better predict truly novel situations. But it's hard to know what it will take for the system to hit a tipping point where it does something completely different, which was never observed before or included in the modeling. Often, we won't recognize the limitations of our complex models--in biology, climate, or economics--until they break.

Wednesday, September 23, 2009

Epigenetics

When it comes to inheritance, there's no beating the DNA sequence for storing and passing on complex information. But other, "epigenetic" mechanisms also bequeath information to subsequent cells or offspring, sometimes in response to environmental changes.

In principle, the word "epigenetics" could apply to any inheritance outside of the genetic sequence. For example, when a cell divides, its contents are divided among the daughter cells. Any transcription factors or other chemicals that alter gene expression are therefore passed on independently of the DNA (along with the mitochondria, which have their own DNA). In recent years, however, "epigenetics" has come to be used mainly to describe two types of chemical changes directly associated with DNA in the nucleus, other than its sequence.

These changes modify how active various genes are in a particular cell. They are particularly important for enforcing the "no turning back" feature of differentiation from versatile stem cells to specialized cells, helping to shut off cellular programs that were active in the early embryo. Moreover, epigenetic changes are passed on during cell division, so that the differentiated cells and all cells made from them lose their ability to become other types of cell. It should not be surprising that many cancers subvert the epigenetic programming to help them re-activate embryonic programs to help them survive and spread. Researchers have identified many epigenetic modifications in cancer cells.

Epigenetic changes can also pass between generations. Biologists have long known of cases of "imprinting," in which the mother's or the father's DNA is inactive in the offspring. Even in people, researchers have found that food shortages in Holland at the end of World War II resulted in changes in the metabolism of the children of women conceived during that period. Such effects are unusual, but profound.

This sounds disturbingly like inheritance of acquired characteristics, as in Kipling's "Just-So Stories." This concept, often misleadingly associated with early 1800's evolution pioneer Jean-Baptiste Lamarck, was supplanted by Darwin's notion of natural selection of random variations. But persistently activating or suppressing pre-existing genes for a few generations, even in response to environmental pressures, is not the same thing as creating novel properties. Some scientists, notably Eva Jablonka of Tel Aviv University, maintain that epigenetic effects can be permanently enshrined in the sequence, but that remains a minority view. Equating epigenetics with Lamarckism is misleading, despite having a grain of truth.

The two best known epigenetic mechanisms are chemical changes that alter the transcription of DNA. One mechanism modifies the DNA itself, while the other modifies the packaging of the DNA in the nucleus.




(Click to open in new window. Source: NIH)

In DNA methylation, methyl (-CH3) groups are chemically bonded to a base in the DNA sequence, usually a cytosine (C) next to a guanine (G), together called CpG. The presence of the methyl group suppresses translation of the DNA sequence that contains it. In addition, the cell contains enzymes that recognize methylation of one chain of DNA and methylate the other chain, helping to propagate the information.

The second mechanism affects the packing of the DNA into the compact structure known as chromatin. The paired DNA chains wrap tightly around a cluster of proteins called histones to form a nucleosome. Nucleosomes strung along the DNA chain themselves pack into compact arrangements that make it hard for the transcription machinery to get at them.

The details of this process are only partially understood. One thing that is known is that free "tails" of the histone proteins straggle out of the nucleosomes, and that chemical modifications of these tails modifies transcription. The modifications include single or multiple methylation or acetylation (adding -COCH3) of particular amino acids positions in the tail, as well as binding of other factors. The details matter: particular modifications either increase or decrease transcription.

In recent years researchers have developed techniques for mapping both DNA methylation and chromatin modification over large regions of the genome. Using these techniques and others, biologists are beginning to understand when and where these epigenetic modifications occur in normal and diseased cells, how nutrition and other environmental influences change them, and how specific modifications are actively regulated to modulate gene expression.

Tuesday, September 22, 2009

Forbidden Questions

To navigate the quantum world, you have to know what questions not to ask.

In the everyday world, we get along fine assuming that a baseball, for example, had a certain momentum even before we whacked it and felt the effects. But at the quantum level, an observable effect like the push on a bat does not give us permission to regard the ball's earlier momentum as having been a "real" quantity, independent of the swinging bat. Talking about such unobserved properties is a recipe for trouble.

This takes a lot of getting used to.

I ran head on into this problem in my latest story for Physical Review Focus. I first titled the story "How Long is a Photon?" and described the experiments as measuring the "duration of individual photons." That description was wrong, and came from asking forbidden questions.

Optics experts often measure the duration of pulses that are only a few femtoseconds (10-15 seconds) long. This is much too fast for direct electronic measurements, so they do it by making two similar pulses and measuring whether they overlap. Delaying one of the pulses by more than their length stops them from overlapping. Actually the researchers repeat the experiment with millions of pairs of pulses, each with a particular delay, to build up a picture of how the overlap varies with delay. For pulses consisting of many photons, it is natural to regard the overlap time as reflecting the length of the underlying pulses.

The new experiments look a lot like this. But the difference is critical.

Kevin O'Donnell, at CICESE in Baja California, built on earlier experiments from the Weizmann Institute in Israel. Instead of pairs of pulses, however, these groups measure pairs of photons. They get the pairs by shining a steady green laser into a special crystal, which splits about one green photon in ten million into two infrared photons. Because these two photons are created as a pair in a single quantum-mechanical process, they are "entangled": properties deduced from measurements on one will always be related to properties deduced from measurements on the other, even if the measurements are done far apart.

The nature of this connection is one of the central oddities of quantum mechanics. In fact, we could save a lot of trouble by not talking about the individual photons at all, because in a profound sense they do not exist as separate entities, even after "they" move away from each other. But our language makes it hard to talk about a pair without think of it as a pair of something.

As in the experiments on pulses, O'Donnell delays one photon with respect to the other and measures their overlap. (I can say that without saying photon, but it gets a lot more complicated.) But as he explained to me, it is not meaningful to relate this overlap to the "length" of the photons. Instead, the result of the overlap experiment at different delays is a property of the combined state of the two photons. The final story, "The Overlap of Two Photons," takes pains to describe that correctly, at the cost of clunkier language and probably losing some readers.

As another example, a researcher who measures the energy of one photon in a pair can be assured that the energy of the other will be just right, so that their combined energy equals that of the original green photon. But that doesn't mean that the photon "had" that energy before the measurement was made. More complex experiments, in fact, show that the unmolested photon does not have any particular energy.

In the current experiment, you don't go too far wrong by imagining (incorrectly) that it measures the length of a photon. But this "bad habit," as described by David Mermin in Physics Today (subscribers only, or you can google the title), of conferring reality on properties that aren't or can't be measured, is the root of much confusion. More importantly, thinking (and talking) precisely about what actually exists is key to understanding the nature of the quantum world we inhabit.


 


 

Monday, September 21, 2009

Photo Finish for the Netflix Prize

It's not every day you see seven computer scientists grinning awkwardly behind one of those goofy six-foot checks they give out to lottery winners.

That was the scene in Manhattan's Four Seasons Hotel this morning as Netflix announced the winners of their $1 million competition to improve the system that underlies the movie recommendations that they make to their customers. I wrote about these "recommender systems" and the Netflix Prize in the August issue of Communications of the Association for Computing Machinery. Just as that article was going to press, someone beat the 10%-improvement threshold for bringing home the award.

But only today did Netflix announce that the winning team was BellKor's Pragmatic Chaos, a longtime leader, some of whose members appeared in my story. Although this team was the first to break the barrier in late June, other teams had subsequently passed them, and they submitted their winning entry only 20 minutes before the deadline (30 days after the barrier breaking). In fact, another submission matched their 10.6% improvement --but because the Ensemble team submitted their entry ten minutes later, they spent the presentation clapping politely from the audience.

Most researchers will say the prize money is only a part of the excitement of this competition--and in any case the winning members from AT&T Labs will be handing their winnings over to their corporate sponsor. A major draw for researchers was access to Netflix's enormous database of real-world data. The company also maintained an academic flavor by requiring that winners publish their findings in the open literature, and by maintaining a discussion board where competitors discussed results and strategy.

We don't know how many other companies have taken advantage of these open results, but Netflix certainly has. Netflix's Chief Product Officer Neil Hunt says the company said the company has already incorporated the two or three most effective algorithms from the interim "progress prizes." Moreover, "we've measured a retention improvement" among customers, Hunt said. The company is still evaluating which of the several hundred algorithms that were blended together to win the prize will be incorporated into future recommendations, since the results need to be generated very rapidly. "We still have to assess the complexity of introducing additional algorithms," Hunt said.

At the ceremony, the company didn't talk much about the other features, beyond predicting "star" rankings, that define good recommendation systems. As discussed in my CACM story, these include aspects of the user interface, such as the way users are encouraged to enter data and the way the results are presented. In addition, a good recommendation needs to go beyond predictable satisfaction to include serendipitous choices that a customer would not find on their own.

Rather than take on these more psychological challenges, the "Second Netflix Prize" will address the more algorithmic challenge of making predictions for customers who haven't ranked many movies, for example those who just signed up or who don't feel like providing ratings. To augment this "sparse" data, Netflix will provide competitors with various other tidbits of data, including demographic information like zip code and data about prior movie orders. But not, Hunt hastened to add, names or credit-card numbers.

As my earlier story discussed, such "implicit" user information is of growing importance for recommender systems. For one thing, it's harder to distort this kind of input by pumping up certain products with fake ratings. In addition, although Netflix can easily cajole customers to take the time to enter ratings, many commercial sites are more limited and have only implicit data to work with.

The new prize doesn't set any explicit performance goals. Instead, Netflix plans to award $500,000 each to the best performers as or April, 2010 and April, 2011. But most of the winners today weren't sure they were going to sign on to the new challenge. They were too tired.