Saturday, May 31, 2008

Sunday Book Review: Analysis of Phylogenetics and Evolution with R by Emmanuel Paradis (Springer 2006)

We've talked a bit about the use of R for phylogenetic comparative analyses. We love it. One of the biggest complaints with this software has always been its the steep learning curve. Paradis' book takes us big step toward the elimination of this obstacle.

By introducing the R package known as ape, Paradis has done already done more than anyone to bring R to phylogeneticists. His books now makes it possible for even a novice R user to get their feet wet with a broad range of R applications related to phylogenetics and evolution. Although many of the R applications to evolutionary biology introduced by Paradis remain relatively primitive (e.g., direct analysis of sequence data, reconstruction of phylogenetic trees), the phylogenetic applications he discusses are at the bleeding edge. The flexibility R offers for graphical output of phylogenetic trees are also unrivaled.

There are really only two problems worth noting with this book, one ironic, the other tragic. The irony is that book espousing the benefits of free software (and even written using the free typesetting software LaTeX) is anything but free itself. The bloodsuckers at Springer are actually trying to extract >$50 for this slim 211 page paperback. If you're short on cheddar you might try contacting Paradis, word on the street is that he's a good dude. The tragic problem is that the book is already out of date and incorrect in places, due in part to some untimely, and seemingly unnecessary, revisions of ape's code by Paradis himself. I was tearing my hair out for a good long while before I figured out that he had changed the format of the $edge portion of ape's default tree format. Internal nodes were previously labeled with negative integers, but are not coded with positive integers.

R will change the way you do science. Paradis will help. The good will of the community will get you the rest of the way.

Software Review: Figtree a graphical tree viewer from Andrew Rambaut's Research Group


In spite of numerous noble efforts (e.g., TreeView, TreeEdit), the phylogenetics community has always lacked a simple, fully-featured application for viewing trees and their associated features. Figtree comes closer to meeting this need than anything that has come before. This program focuses exclusively on displaying trees and producing "publication-ready figures" and does not actually conduct any analyses (even things as simple as generating consensus topologies). Nevertheless, it has quickly become one of the most important tools in the phylogeneticist's box. It's easy-to-use graphical user interface permits users to do everything from selectively shading branches to displaying support values or branch lengths. Trees can then be exported to the PDF format, whether it be for publication or subsequent revision in a program like Adobe Illustrator. Although some users might have been hoping for exports in JPEG format, I'm glad this isn't included. JPEG files, of course, result from conversion of high quality vector graphics to compressed, rasterized graphics that are invariable of a lower quality than the originals. Nothing has done more to contribute to the hideous pixelated images that grace the pages of your favorite journals than the use of JPEG files (or other similar formats).

Saturday, May 24, 2008

Serious BiSSEness

In what will surely become one of the most influential papers of 2007, Maddison and colleagues propose a new model to resolve an important and often unrecognized problem in ancestral state reconstructions and studies of key innovations. 

Specifically, the effect of character states on the process that generates the observed phylogeny (speciation and extinction rates differ depending on whether the lineage is in state 0 or 1, for example) frequently made it almost impossible for previous models used in reconstruction of ancestry to infer the correct ancestral states and transition rates. The opposite is also true--the inaccurate inference of ancestry made it impossible to infer correct state-associated speciation and extinction rates associated with each character state.   The new model, named BiSSE (binary state speciation and extinction), is implemented in Mesquite. 

Plainly stated, if one wishes to analyze the evolution of a character with two states, each of which is associated with different speciation and extinction rates, the use of the old Mk-family of models is likely inadequate. Equal net diversification rates for alternate states are unlikely, as are equal transition rates. 

And who wants to study characters that do not affect net diversification rates?

SIMMAP Back On-Line

Stochastic mapping is a method for inferring the position of mutational changes or shifts in morphological or ecological traits on phylogenetic trees. It's a cool method with lots of potentially interesting applications. Although basic stochastic mapping can now be implemented in Mesquite, Jonathan Bollback's program SIMMAP permits one to expand the basic methodology a bit further (one can use SIMMAP, for example, to test character correlations). Unfortunately, this program was removed from its original web-site (the one referenced in Bollback's BMC Bioinformatics application note) and isn't readily accessible via internet searches. Fortunately, he has just provided a link to the program's new page. Check it out. Looks like a new version is on the horizon...

The CIPRES Portal

Supercomputers or computing clusters are now a popular solution to the computational challenges posed by increasingly large phylogenetic datasets. By using a cluster, you can speed a typical MrBayes run up by at least eight times (by running each of the eight chains required by the default MCMCMC settings on a different processor). The obvious problem with these resources, of course, is that many users don't have access to a cluster. Fortunately, this is beginning to change. One emerging resource is the CIPRES portal, which offers public access to computing resources at the San Diego Supercomputing Center. Although some have complained that this massively multi-PI, NSF-funded resource has been slow to develop, there has been tangible and important progress over the past few years. At this point, users can implement some of the most popular applications in phylogenetics (e.g., PAUP*, MrBayes, RaxML) through a web interface. In most cases, unfortunately, this interface is limiting; for example, some of the most popular options in MrBayes (e.g., parititioning) and PAUP* (e.g., multiple randomized sequence addition replicates in a heuristic search) are remain unavailable. Nevertheless, they're aware of these limitations and I've been told that improvements are on the horizon. This is an important resource and I want very much for it to succeed.

Friday, May 23, 2008

Molecular Phylogenetics has Plunged Into Crisis!

We're going to try to avoid the whole evolution versus creationism thing here. There are enough other blogs and resources covering this topic. It is fun sometime though, to see how phylogenetic uncertainty has gotten tangled up in the debate. Some of these fanatics are actually reading our papers and talking about discordance among markers! It's a good read, they cite lots of the papers from the 90s about reconstructing the early history of the tree of life. Who wants to explain the conceptual basis of gene tree/species tree conflict to somebody who thinks the Flintstones are a documentary? Love that quote from the Moonie Jonathan Well's 2000 book: "Inconsistencies among trees based on different molecules, and the bizarre trees that result from some molecular analyses, have now plunged molecular phylogeny into a crisis."

I wonder if they're going to revise in light of Dunn et al.?

Thursday, May 22, 2008

Push-button science

It is difficult to grasp the pace of scientific progress these days. One benchmark is how we do our work. The most common programs people use to carry out their analyses were unavailable when I started my PhD research, and many of the key methods had not been invented yet. This could be because we're in a particularly fruitful time for research, but I tend to think it's more of a sign of things to come. We need to be prepared for the fact that the next batch of scientists, 5-10 years from now, will be applying techniques that have not even been invented yet on data sets that we can hardly imagine. This perspective is expressed well by Ken Robinson. (Thanks to Larry Forney for pointing me to that video).

What does this mean? To me, it suggests that there's a serious lack in training of graduate students. In other math-intensive fields (like physics), students are required to take a variety of math courses to prepare them to deal with complex data and equations. In biology, students sometimes take these courses, but it's usually not required. I think it's a key ingredient for success in an uncertain and fast-moving future.

What classes are the most valuable? To me, these have had the most pay-off:

1. Probability (something more advanced than a basic stats course)
2. Calculus
3. Matrix algebra

Take a math course or two, it won't kill you.

Wednesday, May 21, 2008

Hot Off the Press: Evolution 62(5)

The new editors will stop at nothing to spice up this journal; the May number's cover features two chickens doing the nasty. If you're interested in plant mating systems and genetics you're going to love this number. If you're looking for phylogenies, the reading material is less fertile. Hedtke et al. do some fairly standard phylogenetic analyses with Garli and MrBayes in their analysis of androgenesis (presence of father-only nuclear chromosomes in offspring) in the clam genus Corbicula. They use Brown & Lemmon's program MrConverge to diagnose convergence of their Bayesian analyses, but provide the same link for this program that has been down for weeks. There's also an intriguing paper by Egan et al. on the identification of host-specific loci using a genome scan. The approach uses >400 AFLPs to identify loci under selection in populations of beetles specialized for different host trees. It seems like a reasonable, if crude, option for identifying loci under selection in natural populations when candidate gene or QTL studies are not feasible.

Monday, May 19, 2008

Sunday Book Review: Evolution: What the Fossils Say and Why it Matters by Donald Prothero (Columbia University Press, 2007)

Yes, I know it's Monday, but I was busy with graduation yesterday...

Prothero takes an impressively comprehensive approach to debunking the claims that creationists have made about the fossil record. The most useful part of the book debunks the creationist's claims that the fossil record contradicts Darwinian theory. He hits all the creationist's favorites, from the Cambrian explosion and its implications to the proposed absence of transitional forms between ungulates and whales. Prothero rarely minces words in delivering a major smack-down to ignoramuses like Duane Gish. Although the details may leave some hard core evolutionary biologists a bit unsatisfied, Prothero provides all the references to the primary literature that are needed to fill in the gaps. This will be an important reference work.

I do wish he had made a bit more of an effort to integrate the stunning new conclusions revealed by molecular phylogenetic analyses, which serve to further reinforce the validity of Darwin's theory. On a related point, I also can't stand to see so many phylogenetic trees without one iota of support. In some cases, failure to consider molecular phylogenetic studies and phylogenetic uncertainty results in presentation of potentially outdated relationships, like the repeated depiction of turtles as the outgroup to all other extant reptiles and birds (Fig. 5.4 & Fig. 11.1 [which also appears to suggest that snakes are the sister taxon to lizards]). The position of turtles remains controversial, but numerous molecular phylogenetic analyses suggest that they may be closely related to archosaurs rather than branching off at the base of the reptile lineage.

Sunday, May 18, 2008

Lizard Porn

As part of our Saturday ritual to boost our google hits I've scoured the internet (i.e., did one google image search for "lizard porn") for the finest examples of lizard porn. I don't advise anybody else to do the same: people are fucking sick.

One thing is clear from my search: Anolis carolinensis is a porn star. The internet's offerings are dominated by this species. Perhaps its self confidence has been buoyed by the recent sequencing of its complete genome?