Wednesday, November 4, 2009

The Permanent Record

The history of science is documented in carefully crafted publications, which contain data the authors have selected and analyzed to best support their claims. But how should day-to-day scientific research be documented?

Once upon a time, measurements were taken by an experimenter reading an instrument, and transcribing the result by hand into a lab notebook. Written in pen, numbered sequentially, and dated, this notebook provided a time-stamped record of the raw data of the research.

Those days are gone. Modern instruments take massive amounts of data under computer control and often save it in files with inscrutable names that can be later modified at will and leaving no trace.

When I served on a committee investigating possible fraud by investigation of Hendrik Schön in 2002, we found numerous discrepancies in his published data. The committee hoped that we might resolve the problems by examining Schön's raw data. We were sorely disappointed.

We requested supporting information for only six of the eventual 25 papers under investigation. Schön (who was still working at Bell Labs) gave us a two-inch-thick stack of printouts. Essentially none of them met the traditional notion of an archival record of the raw data with document provenance. All were processed data--and some of these were clearly manipulated. He said that storage limitations on the computer forced him to delete some of the data.

Schön also gave us a CD-ROM or two full of "raw" data files. Many of these proved to be files produced by the plotting program, Origin--also not raw data. These files were at least time-stamped, with dates corresponding to the original acquisition of the data. However, they had a curious property that, although the dates on the files varied over a couple of years, the creation times on the files showed a steady progression, one every few minutes, even as the dates changed. Real raw data could have saved Schön, but what he produced only made things worse.

Bill Neaves, the Chief Executive Officers at the Stowers Institute for Medical Research in Kansas City, understands the need for an archival record, from investigations scientific misconduct cases in his former position at the Southwestern Medical Center in Dallas. When he was helping to set up the Stowers Institute almost a decade ago, he says, he decided to take advantage of the lack of institutional history to require that researchers maintain lab notebooks. Moreover, every week the notebooks are scanned and stored in an unalterable, time-stamped format.

Would this procedure have stopped Schön? Maybe not. During the investigation he showed that he could fabricate "raw" data as well as he could publishable data. But it would have slowed him down a lot, and it would have ensured that he knew that people were watching. And if he had been honest, the archive would have absolved him, as Neaves says the Stowers system has already done for one falsely-accused researcher.

But these days, the written record is only a shadow of the activities of science. I'd also like to see all data-acquisition software unalterably configured to create read-only files with clear time stamps. With the low price of digital memory today, there is no excuse for not archiving all measurements.

Unfortunately, even this is not enough. Many of the great biology frauds have involved researchers altering samples to get the expected readout--for example spiking them with radioactive compounds that will show up at the right place on a gel. Without mechanisms to track laboratory materials and their manipulation, and connect them to the eventual measurements, there will always be room for chicanery.

But the Stowers efforts are step in the right direction. They send a strong message to researchers that the integrity of the scientific process is paramount, and that ensuring integrity will not be left to chance or to the reputation of any individual.


 

Tuesday, November 3, 2009

Spoiler Alert

On this election day, many New Jersey voters faced a common, unsatisfying, and ultimately unnecessary choice: vote for the best candidate, or vote for a candidate that the polls suggest could win?

Why unnecessary? In a reasonable electoral system, voting for the person you prefer would not increase the chances for someone you don't. But that's not how we do it.

A typical vote today awards the election to the candidate with the most votes--a plurality, not necessarily a majority. If a third candidate enters a two-person race, he or she is more likely to draw votes from the more similar candidate. The result is that an extra candidate with a particular viewpoint can reduce the chances of that viewpoint prevailing. This is the "spoiler effect."

The best known example is Ralph Nader's entry in the 2000 election in Florida. Although Nader got only a few percent of the votes, if everyone who voted for him had voted for Gore, Gore would have beaten Bush--in Florida, and in the country.

But there's nothing liberal or conservative about the problem. Any candidate can lose because of a strong third-party candidate with similar views.

This is just wrong.

It should be changed.

And there is a simple way to change it.

In instant runoff voting, voters rank all the candidates, rather than just voting for their first choice. Not too hard, right?

If no one gets a majority of first-choice votes, the candidate with the fewest is eliminated. Those ballots are redistributed to the second choice on each ballot. No votes are "wasted," but voters get to express their true preference.

The results are exactly to what would happen in an ideal runoff election, except that there's no need for the expense and low turnout of a second election.

In 2000 Florida, for example, Nader would have been eliminated, and the ballots that listed him as first choice would have been allocated between Bush and Gore, depending on who people listed as their second choice. In 2009 New Jersey, independent Chris Daggett was expected to draw more votes from the Republican candidate, which could tip the balance to the Democrat.

Many individual cities, like Oakland and Memphis as well as some national elections use instant runoff voting now. There's nothing particularly difficult about it, except that the two dominant parties may see it as a threat to their exclusive right to power.

Instant runoff voting isn't perfect. Pretty much all voting systems have some quirks that sometimes give results that seem obviously wrong.

But it's not nearly as bad as what we have now.


 


 

Monday, November 2, 2009

Kansas City Here I Come

I'm going to Kansas City.

I've been invited to give a talk Tuesday at the Stowers Institute for Medical Research, based on my story this summer in Science about the surprisingly weak connection between the apparent biological importance of a DNA sequence and the preservation of it sequence through evolution.

I thought at first that Stowers had mistaken me for a real researcher. But they assured me that a writer can sometimes do a better job of providing perspective than a research who is immersed in day-to-day technical details. Since hosting Matt Ridley in 2001, the Institute has periodically included science writers among their speakers.

Preparing the slides for the talk has been a lot of work, but it has reminded me of some big differences between communicating science as a spectator and as a participant.

The most obvious difference is the thoroughness of the discussion. For one researcher to convince another requires data, covering the entire logical chain as well as possible alternative explanations. In contrast, a journalist rarely gives a complete description of the evidence. Instead, as David Ehrenstein, the Physical Review Focus editor, likes to say, we are happy to convince the reader of the plausibility of a conclusion.

This breezier discussion of the evidence gives journalists a freedom to convey the big picture. Ordinary researchers rarely get this opportunity, until their reputations reach a level where others are happy to hear their opinions for their own sake. This freedom is a terrific luxury for science writers.

Ideally, however, a journalist is not expressing a single opinion, however wise, but synthesizing or contrasting the range of opinions in a field. When this is done well, it conveys the entirety of a field more accurately than any single view. Unfortunately, in active, contentious fields, it's easy to get bogged down in the disagreements, obscuring instead of illuminating the big picture. The common journalistic focus on conflict doesn't help.

Especially when covering disagreements, the journalist needs to convey a sense of authority about the key issues are, if not their ultimate resolution. Without being able to rely on the detailed technical results, this authority often comes from the researchers ("sources") interviewed for the story. Of course, a well-written story can suggest authority even when these sources are not representative of the field, which is one reasons science writers are not created equal.

To convey this authority, journalists often use direct quotes. Again, the reader is dependent on the writer to choose truly representative quotes from a much longer interview with each scientist. Still, this appeal to authority (other than the speaker's) rarely happens in scientific presentations.

Finally, the nature of visuals is strikingly different in technical talks and science writing. In a talk, a scientist might use a cartoon or other light material as a diversion, but the meat will be data: descriptions of procedures, photographs of representative results, graphs summarizing the results, and perhaps an abstract representation or cartoon to convey the concept. In a journal article, in fact, many researchers reading a journal article will skip the text and go straight to the figures.

For much science writing, the only one of these elements that is at all likely to survive is the cartoon conveying the concept. The supporting details get short shrift. For many outlets, the figures won't have any meat at all, and may have only a tangential relation to the subject. As a result, the heavy lifting for science journalism is all done by the written word. Clearly this approach does not translate to a presentation, unless it is a classic speech without visuals.

The PowerPoint presentation I'll be presenting at the Stowers Institute will be a kind of mutant hybrid. I don't plan to use any direct quotes on the slides, for example. But I expect I'll do some name dropping and invoke my interviews for authority, especially to convince the audience that there is a puzzle to be solved in the relationship between evolutionary constraint and biological function. Hopefully I'll be able to get that bigger puzzle across, as well as the intriguing possibilities that may arise by solving it.

As Isaac Asimov said, "the most exciting phrase to hear in science, the one that heralds new discoveries, is not 'Eureka!' (I've found it!), but 'That's funny....'"


 

Friday, October 30, 2009

Conductance is Transmission

My latest story in Physical Review Focus describes measurements of electrical conduction between two "buckyballs," or C60
molecules. This sort of characterization is a prerequisite for the sort of understanding and control that would be needed for future "molecular electronics."

The electrical conductance (the inverse of the resistance) in such tiny systems is limited to values of the order of 2e2/h, where e is the electron charge and h is Planck's constant, which sets the scale for quantum phenomena. This combination goes by the name of "conductance quantum," or G0.

Unlike other quanta like photons, however, the conductance is often not generally required to come in discrete packets. Under special experimental circumstances, however, such as in "quantum point contacts," the conductance can take on reasonably stable values that are simple multiples of G0.

Still, the idea that conductance has special value was quite jarring when it became popular in the 1980s. Most materials have a well-defined conductivity determined by number of electrons and how frequently they scatter from imperfections of atomic motion. The conductance, which is just the total current divided by the voltage, is then calculated from the conductivity by multiplying by the cross-sectional area of a piece of material, and dividing by its length.

In very small devices, however, electrons move as a wave from one end to the other. The conductance is then determined by the likelihood that they propagate to the far end. The visionary IBM researcher Rolf Landauer laid the groundwork for this view in a 1957 article in the IBM Journal of Research and Development.

Only a quarter-century later in the 1980s, however, did experiments start to catch up. Researchers had been doing experiments at very low temperatures in clean semiconductor systems, where the electrons propagate cleanly as waves over many microns. Lithographic patterning can easily create structures that are smaller than this distance, and comparable to the wavelength of the electrons themselves (typically a few hundred angstroms, or a few hundredths of a micron).

In the semiconductor samples, electrons are free to flow only in a thin sheet near the surface. Researchers can apply a voltage a metal film on top of a semiconductor so that the electrons have to avoid the region under the metal. If there is a small gap in a line of metal, electrons can squeeze through this quantum point contact between them. This is the situation where the quantum effects become important.

The usual derivation goes like this (feel free to skip over this long paragraph): on the two sides, electrons fill up the available states equally, so filled states on one side face filled states on the other and have no way to move across. Applying a voltage V raises the energy of electrons on one side, so the top ones now face empty states on the other. The number of such exposed states is the energy change, eV, times the number of states in each energy interval. Here's the magic: the number of states, for the special case where they are one dimensional waves, is determined by how their energy E varies with wave vector k: dE/dk. Their group velocity--the rate at which they impinge on the contact--is (1/h) times dk/dE. Each carries a charge of e, and there are two electrons in each state because there are two spin states. Presto: the total current is eV(dE/dk)(1/h)(dk/dE)2e = 2e2/h x V.

To me it's rather unsatisfying to go through these shenanigans to get a simple answer like 2e2/h. It doesn't seem right that we have to introduce all these extra quantities just to have them cancel out. Is there an easier way to get to this answer?

In any case, it is now clearly established that each quantum state has an overall conductance of G0=2e2/h, multiplied by the transmission coefficient, which is the probability of a particular wavelike electron making it to the other side. This result applies to any quantum transmission, whether it's in an engineered semiconductor or a single C60 molecule.

Thursday, October 29, 2009

Junk or No Junk?

The phrase "junk DNA" is a hot button. Authors of press releases, news stories, and even some journal articles seem powerless to resist casting any new discovery of function in non-protein-coding DNA as overthrowing a cherished belief that most of the DNA is junk.

In contrast, bloggers including T. Ryan Gregory and Larry Moran regularly gripe that this framing, like many "people used to think x, but now…" stories, is misleading: biologists have known for decades that non-coding DNA contains important regulatory and other functional sequences. Nobody seriously thought it was all junk: that's just a myth that makes the story seem more exciting.

Still, most biologists agree that DNA is mostly junk.

John Mattick is not so sure.

In a 2007 paper in Genome Research, Mattick and his co-author Michael Pheasant, both of the University of Queensland in Australia, suggested that evolution could be sheltering much more than the 5% of the genome estimated by the ENCODE project and others. Those researchers estimate the background rate of neutral evolution by looking at sequences that they assume to be useless, such as "ancient repeats" left behind from long-ago genomic invasions.

Instead, if these sequences are a little bit useful, perhaps because they are occasionally drafted by the cell for other uses, then they would be slightly preserved during evolution. If this is the case, then other sequences that are also slightly preserved may be useful, too.

I discussed this issue with Mattick for my story for Science (subscribers only), but it was hard enough to capture the issues for strongly selected sequences, so this subject didn't make the cut. Mattick didn't claim that the issue is settled, only that the logic was in danger of being circular: "It is basically an open question. We have no good idea how much of the genome is conserved, except for that which is dependent on questionable assumptions about the nonconservation of reference sequence."

For him, the extensive transcription of the genome seen by ENCODE may not be a sign that RNA production is unselective, but that a large fraction of the DNA is serving some useful, if so far unknown, function.

Mattick described himself as a "minor author" among the scores on the ENCODE project. Ewan Birney, of the European Bioinformatics Institute in the U.K., played a coordinating role. But he doesn't strongly dispute Mattick's observations. "Ancient repeats provide a marker of evolution. They may very well be under some selection," Birney said. But "the striking thing," he stressed, "regardless about where the line is between selection and no selection, is that a lot of the functional regions are at the absolute lowest end of what we see across the human genome." They may be important, but not very important.

Or maybe the biochemical assays don't measure biological importance at all. "Many people instinctively feel," Birney said, "that all the functional elements really must be selected for in some sense. But there's alternative view, which is that there's just a big set of cases which are generated randomly, are perfectly assayable, when you assay them they're always there in that species, but in fact evolution doesn't select either for or against them. They are truly neutral elements: they are selected neither for nor against."

More philosophically, Birney doesn't find the evolutionary question to be central to short-term concerns about human health. "For disease biology we're interested in understanding the disease. We're not so interested in proving whether they're weakly under selection or something like that."

In fact, for both Birney and Mattick, the extensive biochemical activity of the weakly selected 95% of the DNA suggests its potential as a reservoir of "spare parts." Whether or not the long-term potential of that reservoir puts evolutionary pressure on its components, so that they are conserved, may not be the key issue. The important thing is that those partially-assembled genetic tools are ready to be called into action for future innovations.


 


 

Wednesday, October 28, 2009

ENCODE

The mapping of the human genome in draft form in 2000 was a turning point in biology. But the announcement really marked a start, rather than an end, of the practical use of genomic information.

Several large-scale projects in the intervening years have mined particular aspects of the genome and combined it with other sources of information. The HAPMAP project, for example, looked at how common single-nucleotide polymorphisms, or SNPs--changes in a single base--varied among a few selected populations. These studies formed the basis for genome-wide association studies to identify DNA regions associated with various diseases.

Another project, called ENCODE, for ENCyclopedia Of DNA Elements, focused on cataloguing the various types of protein-coding and regulatory structures in the genome. By correlating these with biochemical measures of functional activity and the way the DNA is organized in the nucleus, the researchers got a broad view of how the expression of genes is regulated. They also compared the DNA sequences with those of closely and distantly related organisms, to illuminate how the function of the DNA is related to its evolution.

With some 200 co-authors from 80 different institutions, the ENCODE project rivals some big particle-physics experiments for scope and complexity. In fact, the pilot phase of ENCODE selected "only" 1% of the genome--around 3 million bases--for detailed study. An overview of the results appeared in Nature in June 2007. The data are publicly available, and researchers continue to publish papers on aspects of the work. In addition, follow-up work is aimed at analyzing the entire genome.

Among the profound conclusions from the pilot phase that most of the genome is transcribed into RNA, even though only 1.5% or so codes for protein and only about 5% seems is clearly functional. In other words, much of the regulation of genetic activity may be occurring, not at the level of transcription, but at the level of RNA.

The researchers also found that the organization of the chromosomes in the nucleus, in particular the wrapping of the DNA around histones to form nucleosomes, predicts the locations where transcription begins. These results emphasize the known importance of the positions of nucleosomes in regulating genetic activity at different positions.

Some of the researchers looked at various measures of biochemical activity along the DNA, such as binding to proteins that are known to be active in regulation. Their hope was that these assays would serve to identify regions with a biological function in the cell.

Other ENCODE researchers compared the sequences with corresponding sequences from other organisms--both close relatives like mice and distant eukaryotic relatives like yeast. According to a longstanding assumption, the degree of similarity of these sequences, showing how resistant they are to changes from neutral evolution, should also reflect their biological importance.

These studies revealed two surprises. First, not all biochemically active sequences are evolutionarily constrained. This might mean that the biochemical tests don't measure things that are important to the cell after all. Second, and more puzzling, not all of the constrained sequences had any obvious function.

I wrote a story for Science this summer (subscribers only, sorry) discussing possible reasons why evolution and importance don't always track one another.

ENCODE and other large-scale studies will continue to supply us with extensive, detailed information about the genome. The story is only just beginning.

Tuesday, October 27, 2009

Climate Cover-Up

In their new book, Climate Cover-Up: The Crusade to Deny Global Warming, James Hoggan and Richard Littlemore waste little time wringing their hands about the reality or seriousness of the global warming threat. They dispense with this question quickly, showing that the essential features of carbon-dioxide induce warming have been known for over a century.

Although a devil might lurk in the details, the recent state of the science is captured by Naomi Oreskes' 2004 literature study in Science, which found zero dissenters from the consensus among 928 journal articles referencing "global climate change." Similarly, the 2007 Fourth Report of the Intergovernmental Panel on Climate Change, whose political charter leads it to avoid poorly understood possibilities like collapsing ice sheets, nonetheless states that "most of the observed increase in global average temperatures since the mid-20th century is very likely due to the observed increase in anthropogenic [greenhouse gas] concentrations."

In contrast, a Pew Survey released last week concludes that only 36% of Americans think there is solid evidence that the Earth is warming because of human activity, down from an already low 47% over the past few years. Climate Cover-Up explores how it has come to pass that the public still thinks that this is an open scientific question. Hoggan and Littlemore describe the extensive, organized efforts to make it appear open, largely funded by corporations with much to lose from effective climate actions.

Hoggan is a public-relations professional who says that PR people have a duty to serve the public good. In 2005 he founded DeSmogBlog to highlight just the sorts of systematic distortions that the book catalogs, and Littlemore is editor at the site. Their highly readable book describes these efforts, and the funding behind them, with journalistic precision and documentation.

Their laundry list of deception includes "astroturf" groups that use sanitized corporate funds to present a "grass roots" appearance; "think tanks" that increasingly eschew analysis for promotion of policies that favor their sponsors, and petitions signed by scientists who are not or who have little or no expertise in climate. Hoggan and Littlemore systematically discuss these and other programs to frame the "debate" as one in which huge uncertainties remain--as if that should be a source of comfort.

Sowing doubt is a tried-and-true strategy for delaying government response. David Michaels' excellent 2008 book, Doubt Is Their Product, for example, and Devra Davis' 2007 The Secret War on Cancer related how the tobacco industry perfected this technique to delay serious government action against their product for decades. Some of the same firms are coordinating the skeptical response to climate change, and some of the same scientists, like Frederick Seitz and S. Fred Singer, have played roles in both controversies (despite having expertise in neither cancer nor climate).

Hoggan and Littlemore describe how these omnipresent figures benefit financially from their support of corporate needs, and they reveal the irrelevance of many of the "thousands" of signatories on some highly publicized petitions. But they don't address why other scientists--many intelligent and sincere--sign on to such statements. Why are researchers who have no expertise in climate, as well as members of the public, so willing to question those who do, when they would never presume to second guess articles about cancer treatments or particle theory?

In the end, though, these efforts have achieved their goal: keeping the journalistic treatment of global warming "balanced," unlike the clear trend in expert opinion. This problem was captured in a 2004 study by Maxwell T. Boykoff and Jules M. Boykoff, "Balance as bias: global warming and the US prestige press" which was published in 2004 in the journal Global Environmental Change, and in Chris Mooney's story, "Blinded by science: how 'balanced' coverage lets the scientific fringe hijack reality" in Columbia Journalism Review later that year.

This book is not likely to convince true skeptics of the seriousness of global warming. For those who understand the stakes, however, the book is a powerful inoculation to help recognize the conspiracy-theory talking points, most recently regurgitated by the authors of SuperFreakonomics, for the misinformation it is.

Where there's smoke, there's not always fire. Sometimes, it's a massive corporate-sponsored campaign to blow smoke.