I just showed up for a latte at my favorite coffee shop and my buddy Ian was there waiting with a question: "What's more prehistoric, the Robin or the Blue Jay"? He was thinking it was the Blue Jay due to overall physical appearance and their prehistoric squawking calls. It's impossible to answer this question in a manner that's going to satisfy the serious phylogeneticist, but because I like to think of myself as a phylogeneticist of the people I'm going take stab at this one. To avoid troublesome inference about which species is more primitive or more advanced we should focus simply on which extant species has been around for longer. If we look at Hackett et al.'s recent phylogenomic analysis of birds, we find that the Robin's genus (Turdus) is included, but the closest thing to a blue jay is the con-familial crow (Corvus). Let's approach this question from the family level by contrasting the phylogenetic position of the crows and jays (Corvidae) with that of the thrushes (Turdidae). It's clear that the Corvidae branched off from a clade including the Turdidae and a range of other families relatively deep in the Oscine radiation. This pattern certainly fails to reject Ian's hypothesis, but its unclear that anything shy of a comprehensive species-level, time-calibrated phylogeny would be able to do more. Any ornithologists or paleontologists care to weigh in on this important topic?
Thursday, April 16, 2009
Coffee Shop Phlogenetics #2: What's More Prehistoric the Robin or the Blue Jay?
I just showed up for a latte at my favorite coffee shop and my buddy Ian was there waiting with a question: "What's more prehistoric, the Robin or the Blue Jay"? He was thinking it was the Blue Jay due to overall physical appearance and their prehistoric squawking calls. It's impossible to answer this question in a manner that's going to satisfy the serious phylogeneticist, but because I like to think of myself as a phylogeneticist of the people I'm going take stab at this one. To avoid troublesome inference about which species is more primitive or more advanced we should focus simply on which extant species has been around for longer. If we look at Hackett et al.'s recent phylogenomic analysis of birds, we find that the Robin's genus (Turdus) is included, but the closest thing to a blue jay is the con-familial crow (Corvus). Let's approach this question from the family level by contrasting the phylogenetic position of the crows and jays (Corvidae) with that of the thrushes (Turdidae). It's clear that the Corvidae branched off from a clade including the Turdidae and a range of other families relatively deep in the Oscine radiation. This pattern certainly fails to reject Ian's hypothesis, but its unclear that anything shy of a comprehensive species-level, time-calibrated phylogeny would be able to do more. Any ornithologists or paleontologists care to weigh in on this important topic?
Wednesday, April 15, 2009
Why I love the American Museum....
I love the American Museum of Natural History. I was in NYC this past weekend and spent some time on the fourth floor of the AMNH. For those of you who haven’t visited, this floor hosts what may be the most awe-inspiring and beautiful collection of mineralized bone ever displayed. I could spend hours just wandering through the Hall of Vertebrate Origins. If there is anything, anywhere, that better illustrates the shockingly bizarre diversity of vertebrate body plans through time, I haven’t seen it. I really like the fact that the AMNH still believes that bone and stone are preferable to the interactive “discovery center” exhibits that dominate the majority of natural history museums these days. When I go to a museum, I want to see disarticulated ichthyosaurs that speak of rotting flesh on the bottom of a Kansan ocean. Give me wrinkled duck-bill mummies, blocks of dead fish from Eocene swamps, or tangled Coelophysis skeletons from Ghost Ranch. Call me a purist, but I don’t like my fossils soiled by dinomation and artistic reconstruction. When I was an aspiring young paleontologist, I found inspiration in the fossils themselves, and I find it a bit sad that so many museums have moved away from this in favor of the sound and fury of faux dinosaurs. I can't be the only one who feels this way...(?)As an aside, I also think the AMNH has done a fabulous job of grounding this paleodiversity in a phylogenetic framework. Trees are everywhere. My suspicion is that most museum visitors take away very little from this, but I found it to be wonderful.
Dechronization Interviews Jack Sullivan, Editor-in-Chief of Systematic Biology
I have decided to conduct a series of interviews of prominent evolutionary biologists who work with trees, and post them on this blog. For the first of these, I interviewed Jack Sullivan, Editor-in-Chief of Systematic Biology and a professor in my department at the University of Idaho (photo at left, in his natural habitat). I asked Jack a few questions about the field of systematics and some related issues. It is probably worth noting that I didn’t have a tape recorder or anything like that with me, so Jack’s answers are paraphrased. Thanks to Jack for being my guinea pig, Jack; if there are errors below they are probably mine.Question: What are the most exciting recent developments in systematics?
I think there are three. First, there are second-order statistical analyses that can now be applied across a sample of trees from a Bayesian posterior distribution. These include biogeography, comparative methods, and macroevolutionary tests. We used to have to rely on a single tree for our analyses; now we can do the same analyses accounting for phylogenetic uncertainty by sampling from the posterior distribution of trees. Second, the explicit accommodation of incongruence in analyses of multilocus data through the use of the coalescent. I think it will be really cool when we can use these approaches to differentiate between incongruence caused by coalescent stochasticity from that caused by nonvertical transmission such as horizontal gene transfer or hybridization. Third, the development of phylogenomics. I remember a symposium debate at the Evolution meetings when I was a graduate student in the early 1990s. The debate was about total evidence approaches versus other methods. During the debate, someone raised the question of, “If we could sequence every single nucleotide in the genome, would we then get the best possible estimate of the phylogeny?” I think that emerging datasets demonstrate that the answer to this question might be, “not necessarily.”
Question: What is the role of Editor-in-Chief of prominent journals?
It really depends on how heavy-handed you want to be. In our journal, Systematic Biology, content is really meant to be driven by the Society for Systematic Biology (SSB). Because of this, I have tried to be less heavy-handed in the journal’s direction. The direction of the journal should be driven by members of the society as reflected by submissions. There are some topics that I wish we had less submissions (for example, phylocode and DNA barcoding). When papers are submitted and go through review with positive results, I am very reluctant to reject them based on the subject matter.
Question: So you view the editors role as more of a service to the society rather than an opportunity to shape the field?
Both. The editor can shape the field by insisting on maintaining the highly rigorous standards for data analysis that Systematic Biology is known for, especially for empirical papers. Particular things that I require as EIC might differ from my predecessors.
Question: What is the difference between a good and a bad review of a paper?
The primary characteristic of an excellent review is that the reviewer has assumed the role of silent partner - this comes from Dick Olmstead when he was the editor. Reviewers do this because it has been done for them at the journal. We have an incredibly valuable tradition of rigorous yet constructive feedback in reviews.
Question: Do you have any advice for the next generation of systematists?
As early as possible, find your niche that differentiates you from all of your peers that are doing great work. You cannot just do “comparative biology of (fill in the blank)” or “molecular phylogeography of (fill in the blank).” Probably the easiest way to think about this is to imagine yourself on an airplane next to an intelligent layperson. Convey to them what is important about what you do in a manner that is unique. This is critical for the job search - it is a rare situation when the audience [of a job talk] is just phylogeneticists or even evolutionary biologists.
Question: OK now I’m going to ask you about two controversial groups. What is your take on the cladists?
The view that statistics are anathema to systematics is dead. All the vitality in the discipline is in statistical approaches.
Question: And how do you think we should respond to the creationists?
Fighting court battles require very different tactics than changing public opinion. To affect public opinion, there are two things we can do:
1. Publicly deconstruct the false dichotomy between macroevolution and microevolution.
2. Engage in a strong public outreach campaign over the importance of evolution in day to day life.
Snakes on a Plane
Everyone who's seen Rain Man knows that Quantas is the safest airline in the world (but see Quantas fatal accidents). Today's news out of Australia, however, suggests that they may also be the first airline to ground a flight because of a harmless python. The Age is reporting that a flight from Melbourne to Sydney was cancelled after four juvenile Stimson's pythons (Antaresia stimsoni) went missing during a previous flight from Alice Springs to Melbourne. I hope they cancelled this flight because of concern about the snake gumming up the plane's mechanics, because Stimson's pythons are among the most docile snakes on the planet. As someone who's had ~50% success tracking down escaped snakes, I wish Quantas the best of luck in their efforts to track down the little beasts. My advice: If you look at something and say to yourself "I bet a snake couldn't get into that", you're wrong.
Thursday, April 9, 2009
When We Fail MrBayes…
A recent Dechronization post highlighted the unsuccessful attempts at Bayesian estimation of a large-scale bird phylogeny based on a multi-locus data set by Hackett et al. The apparent failure of MrBayes in this particular case (and under similarly challenging inference scenarios associated with large and/or complex data sets, e.g., Soltis et al., 2007; Moore et al., 2008) appears to raise serious concerns regarding our ability to estimate large-scale phylogeny using Bayesian methods.Existing implementations, such as MrBayes, approximate the joint posterior probability density of phylogeny and model parameters using some form of MCMC sampling (typically based on the Metropolis-Hastings algorithm). These methods quietly specify a means of updating the value of each parameter (the proposal mechanisms), the probability of invoking each proposal mechanism (the proposal probability), and the magnitude of the proposed change issued by each proposal mechanism (the tuning parameters). Proposal mechanism design is an art form (there are no hard rules that ensure valid and efficient MCMC sampling for all problems). For this reason, for many (non-phylogenetic) Bayesian inference methods, it is the responsibility of the investigator to explore a range of proposal probabilities and tuning parameterizations that deliver acceptable MCMC performance.
Accordingly, most researchers familiar with Bayesian inference would consider it extremely naïve to expect that any specific MCMC sampling design would perform well for all (or even most) empirical data sets, especially in the very difficult case of phylogeny estimation. Nevertheless, the default settings of existing Bayesian phylogeny estimation programs are so successful that we are actually “shocked, shocked to find that MrBayes does not solve all of our problems!!”. Without going into detail (as doing so would constitute an entirely separate post), the analyses detailed in the supporting material of the Hackett et al. study reads like a recipe for failure, and I would venture that the putative 'impossibility' of obtaining a reliable estimate with MrBayes in this case falls squarely under the third inference scenario defined above.
What does this mean for our phylogenetic community? First, I would argue that researchers interested in Bayesian estimation of phylogeny need to become much, much more sophisticated about diagnosing MCMC performance, carefully assessing convergence (ensuring that the chain has reached the stationary distribution, which is the joint posterior probability density of interest), mixing (assessing movement of the chain over the stationary distribution in proportion to the posterior probability of the parameter space), and sampling intensity (assessing adequacy of the number of independent samples used to approximate the posterior probability). Second, I believe that developers of Bayesian methods need to encourage and facilitate more vigorous and nuanced exploration of MCMC performance among users of these methods.
Tuesday, April 7, 2009
DNA *and* pinned specimen
In a paper from last week's PLoS One (also highlighted in today's Science Times), Thomsen et al. describe a method that may be used to extract DNA from insect specimens collected as far back as 188 years ago. Remarkably, this method also avoids destruction of the pinned insect (see photo - this beetle is post-extraction). Twenty of twenty museum specimens examined yielded good, if short (~200 bp), sequences from mitochondrial genes. There was also some limited success on insects that were even older from non-frozen conditions. One big caveat is that all of their tests were done on beetles, which obviously are some of the most durable of insects, but it was nonetheless exciting. Although they didn't test this, sounds like their DNA might also have been useful for other short fragments, i.e. microsatellites - opening up doors for tracking lots of interesting population biology of insects, including pests and maybe even vectors. Suddenly those cabinets and cabinets of pinned insects I'm surrounded by seem all the more interesting!
Monday, April 6, 2009
Complexity in Crustaceans: A Driven Trend in Organismal Design
One of my favorite papers from last year was the analysis by Adamowicz et al (PNAS) of an apparent trend towards increasing complexity within the Crustacea. Evolutionary biologists have long been fascinated by trends, and many have at least some familiarity with Cope’s Rule, the proposed trend towards increasing body size within lineages. Adamowicz and colleagues looked at complexity within multiple lineages of crustaceans, from the Cambrian to the present. They found that many indices of complexity, including number and disparity of limb types as well as disparity of limb form, generally showed parallel increases in complexity through time.The very issue of whether complexity is a trend has been controversial (e.g., McShea 1996). Moreover, some might complain that this smacks too much of orthogenesis or progressionism for their tastes. And what do we mean by complexity, anyway? These issues have been and will continue to be debated in the literature. But one of the neatest things about the Adamowicz paper is that they provide a possible mechanism for a trend in complexity. They found that newly originated higher taxa had greater limb differentiation than their contemporaries, and that taxa going extinct had lower degree of limb differentiation. Moreover, limb complexity turns out to be one of the strongest predictors of species richness in extant crustacean clades. Together, these suggest the possibility that the trend in complexity might be driven in part by differential speciation and extinction of lineages based on complexity. What if lineages with higher complexity diversified at greater rates than lineages with reduced complexity? Over time, traits associated with complexity might increase simply because of this connection to diversification. This research thus raises some intriguing levels of selection issues, because there is – in principle – no reason why complexity could only be favored by selection at the individual level.
How might limb complexity fuel the diversification process? The authors speculate that increased limb complexity might increase ‘evolvability’ (the meaning of which is even more fun to discuss than ‘complexity’!) and possibly promoting niche specialization. They also note that new limb types might amplify the intensity of sexual selection, possibly serving indirectly to enable that supposed ‘engine of speciation.’ Anyway, don’t expect this paper to end with a case-closed feeling – after all, questions like these are on par with the biggest unresolved issues in biology. But there are lots of things to think about here!
New Insight on Size Free Morphometrics
An interesting discussion on how to remove size from phylogenetic comparative analyses of morphological data is playing out on the R-sig-phylo list-serv (a forum about the use and development of phylogenetic and comparative methods within the R platform). This topic has been contentious for a some time and the ongoing discussion should be of interest to anyone looking to analyze morphological data in a phylogenetic context. Several heavy hitters (Ted Garland & Joe Felsenstein) and Dechronization bloggers (Dan Rabosky & Liam Revell) have already weighed in with insightful remarks.
Coffee Shop Phylogenetics #1: Is the Guinea Pig a Rodent?
Confusion is the typical reaction when you tell somebody you're a phylogeneticist . Sometimes though, us phylogeneticists chance upon someone who's been waiting weeks, months, or even years to consult someone from our profession. For me this tends to happen when I strike up conversations with strangers while waiting in line at my local coffee shop. My most recent experience with coffee shop phylogenetics involved a guinea pig lover with a pressing question about where her beloved pets fall in the tree of life. She was particularly eager to get my thoughts the a rumor circulating among fellow afficionados that the guinea pig is not a rodent (apparently, some guinea pig fans would like to distance themselves from the less-ruputable mouse and rat lovers). I laughed and said "Of course, the guinea pig is a rodent. How else would you explain all of their stunning similarities, like the precence of constantly growing upper incisors?" When I tried to track down the source of her information, however, I was stunned to learn of the guinea pig battles that raged among phylogeneticists in the early to mid-1990s.Things kicked off in 1991 with a parsimony analysis of amino acid sequence data published in Nature by Graur et al. suggesting that mouse-like rodents (myomorphs) were more closely related to primates than they were to guinea pigs (hystricomorphs). Hasegawa et al. responded immediately, showing that monophyly of rodents (myomorphs + hystricomorphs) was supported by maximum likelihood-based analyses and suggesting that the unusual myomorphs + primates inference was due to parsimony's inability to deal with unequal evolutionary rates. In a '92 response, Graur's group stood their ground, arguing that Hasegawa et al.'s results were an anomaly resulting from maximum likelihood analyses of highly divergent, "nonconservative" proteins. Graur continued to discuss the distinctness of guinea pigs and lobbied to have the Hystricomorpha recognized as a distinct order representing "one of the most ancient branches in eutherian evolutionary history." In a '93 PNAS paper, Martignetti and Brosius used the presence of a neural specific small cytoplasmic RNA (BC1 RNA) in guinea pigs and other rodents - but not in other mammals - to argue for inclusion of guinea pigs with rodents. Additional phylogenetic analyses of DNA sequence data by Hasegawa's group in '94 and Frye and Hedges in '95 further supported the guinea pigs as rodents hypothesis. By '96, even Graur had changed his tune and was considering the Hystricognathi a suborder of Rodentia (my knowledge of this history if obviously incomplete and he may have addressed this point more direclty elsewhere).
Just when the dust had settled, things blew up again with the publication of a Nature paper by D'Erchia titled simply "The Guinea Pig is Not a Rodent". Although this paper rejected rodent monophyly, it suggested a rather different tree than that of Graur et al. (1991). This time, the New York Times even got involved. Of course, Hasegawa's group rallied once again to dismiss the guinea pig is not a rodent argument, arguing that, at the very least, there simply wasn't enough support to overturn the traditional classification (a point that was reinforced by similar conclusions from Philippe). Where do things stand now? Suffice to say that nearly everything published over the last 10 years has strongly supported inclusion of gunea pigs in a monophyletic rodentia (e.g., Prasad et al.'s recent phylogenomic analysis fo mammals). In any case, the guinea pig wars represent an interesting historical anecdote and a powerful example of the symptoms that can result when systematists are engaged in intense debate over the value of different types of data (morphological versus molecular) and different types of phylogenetic methods (parsimony versus maximum likelihood).
Sorry guinea pig lovers: you're living with a rodent whether you like it or not.
Friday, April 3, 2009
When MrBayes Fails...
Last summer, Hackett et al. published a widely-read study of phylogenetic relationships among major bird lineages based on 19 independent loci sampled from 169 species (see also Tom Near's previous post). Their study confirmed some patterns suggested by previous phylogenetic studies (e.g., ratites + tinamous as sister to remaining bird species) while also recovering some novel patterns (e.g., passerines sister to parrots [albiet with low support]). One of the more interesting results from their analyses, however, was relegated to the on-line supplement. In this supplement, we learn that all eight of the 10 million generation partitioned analyses they ran in MrBayes apparently failed to reach stationarity (see figure; note that the first 2 million generations are inexplicably trimmed from each analysis as 'burnin-in'). Unpartitioned analyses fared even worse, resulting in immediate crashes "regardless of the memory capacity of the computers used."Among the partitioned analyses, some continued to shift to new areas of the likelihood surface until relatively late in the analysis. Perhaps even more troubling though was the fact that analyses that did appear to reach a stable plateau sampled significantly different likelihood scores (e.g., -lnL -861,000 v. -lnL 859,500). Is this problem unavoidable in analyses of large datasets?
The most obvious solution would be to simply run the analyses for more than 10 million generations. I've certainly had analyses that required more than 10 million generations to reach stationarity. Perhaps this wasn't done because it took two months on a super computer to run the 10 million generation analyses (anybody know if Hackett et al. or others have implemented longer runs since their paper was published?). Another possibile solution to their problems is to modify the parameters of the MC3 analyses implemented by MrBayes (recall that the MrBayes default is to run two independent MC3 analyses with one cold chain and three heated chains). Hackett et al. explored this possibility by running six analyses with one heated chain and one cold chain (B1-B6) and two analyses with six heated chains and one cold chain (A1-A2). The analyses run with multiple heated chains performed significantly better than those with a single heated chain, perhaps due to the fact that multiple chains are incrementally heated by MrBayes (meaning that the fourth of six heated chains has a flatter likelihood surface than the first). Hackett et al. do not discuss the temperatures used for the heated chains in their analyses, but their results suggest that running multiple heated chains in a single analysis is superior to repeatedly running analyses with only one heated chain.
In any case, Hackett et al.'s ultimate solution was to discard all of their Bayesian analyses and rely instead on parsimony and the fast maximum likelihood methods implemented by GARLI and RAxML. Is this shift away from Bayesian inference in favor of fast maximum likelihood searches for computational reasons a sign of things to come (or has this shift already occurred)? Are the fast maximum likelihood methods ready for prime time, or do people remain uncomfortable with the shortcuts they use to acheive their apparent computational efficiency?