Showing posts with label bayesian inference. Show all posts
Showing posts with label bayesian inference. Show all posts

Saturday, June 20, 2009

Are You Down with DPP?

What's on the horizon for phylogenetic analyses? One answer was nicely illustrated by two presentations in the Ernst Mayr awards symposium at Evolution 2009. Jamie Oaks and Charles Linkem - both graduate students in Rafe Brown's group at the University of Kansas - teamed up on two papers about the Dirichlet Process Prior (DPP) and its application to Bayesian phylogenetic inference.

DPP, which is sometimes referred to as the "Chinese restaurant" process, is a nonparametric Bayesian approach used in clustering problems where the data elements are assigned to a set of discrete clusters. Importantly, under the DPP, the number of clusters and the assignment of data elements to clusters are treated as random variables. Phylogenetic applications of the DPP include detecting positive selection in protein-coding DNA sequences (by assigning nucleotide sites to various dn/ds classes: Huelsenbeck et al., 2006), and accommodating among-site variation in substitution rates (by assigning nucleotide sites to various substitution-rate classes Huelsenbeck & Suchard, 2007).

This latter application was the subject of the talks by Oaks and Linkem, who presented results on the ability of the DPP approach for accommodating among-site substitution rate variation (ASRV) using both empirical and simulated data sets. Their results demonstrate that accommodating ASRV under the DPP can significantly improve the marginal likelihood scores relative to conventional methods that assign data partitions to substitution rate classes a priori (such as those based on codon position, etc.). In some cases, estimates under the DPP surpassed those of conventional approaches by hundreds or thousands of log likelihood units! As is well known, the studies by Oaks and Linkem also confirmed that more adequately capturing ASRV can lead to substantially different topologies.

Given that this is the case, why aren't more people using DPP? One issue emphasized by Oaks and Linkem relates to the substantial computational burden associated with inference under the DPP. For example, their analysis of two ~30 taxon data sets required more than two months of run time! Computational expense not withstanding, John Huelsenbeck (and affiliated phylo-geeks at Berkeley...Yeah, I'm looking at you Brian Moore) are working to provide faster, more user-friendly implementation of DPP-based methods that are bound to change our lives forever. Stay tuned for the latest...

Thursday, May 7, 2009

When We Fail MrBayes, Part II

For this (unauthorized) installment of When we fail MrBayes , I’d like to step back and look at how we assess convergence in the first place. I’ve encountered a few datasets now for which convergence problems would be very difficult to diagnose without tools like AWTY . I’ve found two diagnostics in AWTY to be especially useful. First, the slide command shows the posterior probabilities of clades for non-overlapping samples of trees in the sample: basically, a sliding-window of posterior probabilities. If a coarse scale of analysis shows posterior probabilities that vary widely during the course of a single run, this is strong evidence that runs have not converged. I am also a big fan of the compare command, which plots pairwise split frequencies for a series of independent MCMC runs.

As a friend asked yesterday, why aren’t more folks using AWTY to assess convergence? As far as I am aware, this is the only general diagnostic tool out there that is geared towards assessing convergence of topologies, rather than convergence of molecular evolutionary parameters. As such it, is explicitly addressing the one set of parameters that are (usually) of greatest interest to systematists: the tree itself.

With this in mind, I conducted an informal survey of convergence diagnostics in the literature. I looked at all articles published in two recent issues of Molecular Phylogenetics and Evolution using Bayesian inference (mostly MrBayes, but a few using BEAST) and tabulated the convergence diagnostics used. MPE seems like it should be a reasonable gauge of methods currently used by practicing systematists, although it would not surprise me if papers in some journals (Systematic Biology?) use convergence assessment that is, on average, more rigorous. Anyway, in 25 studies:

3 studies reported only that they “examined stationarity of LnL values” or something to this effect. I hesitate to say that this is the worst possible test of convergence, because 5 studies reported no test of convergence whatsoever and another tested for convergence by ‘discarding burn-in’. I think most readers of this blog would agree that these are generally not adequate.

The most frequent class involved some variation of analyzing multiple runs (11 studies); this includes checking the standard deviation of split frequencies for independent runs (6 studies) and comparing posterior probabilities for independent runs (2 studies). I think this is a good general strategy, but the majority of these considered only 2 independent runs. This is not good. Let’s imagine that treespace for your dataset contains two (rather different) topologies of high and equal probability. At convergence, your MCMC sampler should visit both of these topologies in proportion to their posterior probability (say, ~47% of the time for each, as no other topologies are nearly as good).

A major problem arises if it takes many generations to move between these topologies. Even for a highly-optimized MCMC sampler, it does not surprise me at all that it might take many millions of generations to move between “distant” regions of treespace, particularly for large datasets. If this is true, two runs is far from adequate, because there is a 50% chance that two independent runs will find the same high-probability topology first, and – if not run for a sufficient number of generations – it will appear as though the runs have converged, based on both similarity of posterior probabilities, standard dev of split frequencies, etc. This is something of a worst-case scenario, because – as I’ve described it – certain clades would appear to have ~1.00 posterior probability, when in fact the true posterior probability might be closer to 0.5. Unfortunately, there appears to be no good way of determining a ‘sufficient’ number of generations a priori, so the only solution here seems to be ‘lots of runs.’

The next most frequent strategy involved checking convergence of molecular evolutionary parameters, either by estimating effective sizes of parameters (6 studies), or by checking the Gelman-Rubin proportional scale reduction factor (1 study). I am skeptical of these approaches, considered alone, because I’ve found that there is often little correspondence between convergence of molecular evolutionary parameters and convergence of topologies. Perhaps this reflects some particularly troublesome datasets that I have worked with, but it does not leave me feeling encouraged. Note that I am not claiming that monitoring these parameters is unimportant, but that it is fundamentally inadequate with respect to our interest in topologies.

Finally, 3 studies used AWTY , which seems to me rather low given the potential utility of the software in diagnosing convergence failure. On the whole, the results of this survey do not encourage me. Is our research community doing enough to diagnose convergence failure in MCMC analyses? How severe is this problem? Maybe I’m making a mountain out of a molehill here based on my own experience with a few poorly-behaved datasets. But looking at the literature, it is hard to convince myself that most studies are adequately diagnosing convergence problems, and I can’t help but feel a bit unsettled by all of this.

Thursday, April 23, 2009

Dechronization Interviews Joe Felsenstein

This week, I've conducted an interview over email with Joe Felsenstein. Dr. Felsenstein requires no introduction, really. If you're doing something in phylogenetics or comparative methods, chances are, Joe thought of how to do it 20 years ago.

Most of the questions below are from me (LH) but a couple come from Dan Rabosky (DR). Many thanks to Joe for participating.

LH: What are the most exciting recent developments in systematics / comparative methods?

JF: The availability of genome-scale information is certainly one. The arrival of a generation of young researchers who are comfortable with statistical and computational approaches is another. But the most important development is reflected in recent work on coalescent trees of gene copies within trees of species. What this does is tie together between-species molecular evolution and within-species population genetics. Those two lines of work have been developing almost independently since the 1960s. But now, with population samples of sequences at multiple loci in multiple related species, they are coming back together. This is not another Modern Synthesis, but it is a major event that needs a name. How about the "Family Reunion"? Long-estranged relatives who have not been in touch are getting together.

LH: Take us back to the beginnings, back when you were working on phylogenetic and comparative methods for your PhD thesis. Where did you derive your inspiration? Did you anticipate the impact that this work would have on the
field?


JF: I did not anticipate it at all. My original thesis project with Dick Lewontin was a rather grandiose theoretical population genetics macroevolution model -- my idea, not his. It didn't work out and I didn't have any useful results. Meanwhile Lynn Throckmorton and Jack Hubby, whose labs were nearby, needed someone to write a clustering program for protein electrophoresis band data that they had in multiple Drosophila species. I volunteered and was
fascinated by the algorithms. I went on to write parsimony programs for the Camin-Sokal, Dollo, and polymorphism parsimony criteria, and then to work on how to infer trees by likelihood using Anthony Edwards and Luca Cavalli-Sforza's brownian motion approximation to gene frequency drift. Dick finally suggested that I write this up for my thesis, which I did in 1967 (the degree was officially 1968). Through the 1970s I maintained a sideline of work on trees while mostly working in theoretical population genetics. It was really not until about 1978 that I began to see that this was becoming more important, and that it fit in with my interest in evolution beyond the species boundary. So I shifted my work toward trees and dropped out of theoretical population genetics.

DR: A lot of what we do in comparative methods is based on Brownian motion, or models for which BM is a special case (eg OU). As you (Felsenstein) have written, "Brownian motion is a poor model, and so is Ornstein-Uhlenbeck, but just as democracy is the worst method of organizing a society 'except for all the others', so these two models are all we've really got that is tractable. Critics will be admitted to the event, but only if they carry with them another tractable model."

And for discrete traits, we use Markovian models that assume (generally) homogeneous rates through time and among lineages. Undoubtedly, the math for this could get out of hand, but at some point I think we'll have to do something to explore (among other things) more realistic constraint surfaces etc.

Given this, what do you view as "the frontier" for models of continuous and discrete character evolution? New mathematics? Approximate Bayesian approaches that rely on simulation to deal with analytically intractable scenarios?


JF: Hard to see what. I think one framework will be models in which a population "chases" an adaptive peak which is moving. But we need to have some model for how the peak moves, and aside from having a mechanistic and ecological model of the function of the character this is not forthcoming. Nor is it easy to see how adaptive peaks in sister species become different from each other. We're also going to find that the amount of information available to tell different schemes of selection pressure apart will be small. We are going to have to be able to characterize what we can and can't know given the data. Just adding new mathematical tools or lots of simulation will not resolve these dilemmas.

DR: What do you think about the unification of modern (neontological) comparative biology with paleontology? There seems to be a lot of room for progress in this area. Do you have any suggestions for future directions?

JF: Oh thank you thank you thank you for giving me an opportunity to mount the soapbox and hold forth on one of my favorite topics. I've been working on this. See my paper in 2002:

Felsenstein, J. 2002. Quantitative characters, phylogenies, and morphometrics. pp. 27-44 in Morphology, Shape, and Phylogenetics, edited by N. MacLeod. Systematics Association Special Volume Series 64. Taylor and Francis, London.

and watch my Julian Huxley Lecture to the Systematics Association in London in 2008 which is available as a video also with a PDF of my slides.

Basically we can infer the tree of present-day species from molecular data, and then use it for morphological characters (or other measurable continuous or discrete characters) with a Brownian or OU model, to infer phylogenetic covariances of changes of characters. Then we can use these together with the fossil morphology to help place the fossils. (One could also use all this together in a giant likelihood or Bayesian inference but the gain in doing so will be very small as the morphology will add little to the inference of the tree, I think). One can also use bootstrap samples of trees in this, or samples from Bayesian posteriors.

There is lots to be done here and I am rushing to do it, and working with Fred Bookstein on the morphometric angles to this too. I wonder whether statistical frameworks such as this, together with within species quantitative treatment, will not be important in untangling the paleoanthropological mess caused by nonquantitative approaches to hominoid fossils.

LH: What do you think about the current trend in phylogenetics (and, lately, comparative biology) towards Bayesian approaches?

JF: I am a curmudgeon on this, in that Bayesian approaches do not feel right to me. So I have been resisting them. Bayesians were unhappy with the treatment of Bayesian Inference in my book, in that I did not give them four chapters, the last of which ended by declaring victory. I think we're all Bayesians when we come to cross the street, balancing evidence of approaching cars against our priors. But that's where one of the criticisms of Bayesianism comes in -- do we all have the same priors? Is there necessarily a single prior that you can use that will be broadly acceptable to your readership? If not, then maybe the reader of the paper should instead be given the likelihood curve so they can apply their own prior to it. For phylogenies, priors giving equal probability to all topologies (or to all labeled histories) would be noncontroversial. But the part of the prior that puts distributions on branch lengths could be wildly controversial. There is also the issue of whether some things, such as whether the sun will rise tomorrow morning, really should have a prior.

People should be Bayesians if that fits with their philosophy of doing science. But not just because a Bayesian program happens to run faster than a non-Bayesian one. They should also realize that we will continue to have both Bayesians and non-Bayesians. Biologists sometimes think that this controversy emerged in their field and will be settled there -- that one more really good argument and everyone will become a Bayesian. They might not be aware that Bayesian arguments have been around since 1764. There is no new decisive argument that's going to arise in our field.

The issue to contemplate is the priors, not the details of MCMC techniques. We have not yet seen a case where an important conclusion depends strongly on what prior you assume. Perhaps we never will, but if a case like that arises, and causes trouble for Bayesian approaches, people should not be too surprised.

LH: Your work has inspired a generation of comparative biologists. Any
advice for those of us just starting out on our careers?


JF: I have too many opinions on that for this forum. I guess I would urge people to take a long view and to realize that it takes time for methods to be developed, published and used, and to prepare themselves for the new forms of data that are coming. When I submitted my 1985 comparative methods paper, the referees were dubious about it because it required phylogenies, whereas they felt that only classifications were going to be available! A year or two earlier and it might not have been accepted for publication. I would also urge people to become familiar not only with phylogeny methods and statistical techniques, but also with the theoretical side of evolutionary biology. We're entering a period when there is going to be a merger (or Reunion) of between-species phylogenetic inference and within-species population genetics. I'm worried that we are graduating too many people who know what Subtree Pruning and Regrafting is, but who have no idea what Wahlund's Law is, or how mutational load arguments work. Theoretical population genetics is in danger of becoming a lost art, just when it is most needed. Comparative biologists should learn it -- and teach it.

Wednesday, April 15, 2009

Dechronization Interviews Jack Sullivan, Editor-in-Chief of Systematic Biology

I have decided to conduct a series of interviews of prominent evolutionary biologists who work with trees, and post them on this blog. For the first of these, I interviewed Jack Sullivan, Editor-in-Chief of Systematic Biology and a professor in my department at the University of Idaho (photo at left, in his natural habitat). I asked Jack a few questions about the field of systematics and some related issues. It is probably worth noting that I didn’t have a tape recorder or anything like that with me, so Jack’s answers are paraphrased. Thanks to Jack for being my guinea pig, Jack; if there are errors below they are probably mine.

Question: What are the most exciting recent developments in systematics?
I think there are three. First, there are second-order statistical analyses that can now be applied across a sample of trees from a Bayesian posterior distribution. These include biogeography, comparative methods, and macroevolutionary tests. We used to have to rely on a single tree for our analyses; now we can do the same analyses accounting for phylogenetic uncertainty by sampling from the posterior distribution of trees. Second, the explicit accommodation of incongruence in analyses of multilocus data through the use of the coalescent. I think it will be really cool when we can use these approaches to differentiate between incongruence caused by coalescent stochasticity from that caused by nonvertical transmission such as horizontal gene transfer or hybridization. Third, the development of phylogenomics. I remember a symposium debate at the Evolution meetings when I was a graduate student in the early 1990s. The debate was about total evidence approaches versus other methods. During the debate, someone raised the question of, “If we could sequence every single nucleotide in the genome, would we then get the best possible estimate of the phylogeny?” I think that emerging datasets demonstrate that the answer to this question might be, “not necessarily.”

Question: What is the role of Editor-in-Chief of prominent journals?
It really depends on how heavy-handed you want to be. In our journal, Systematic Biology, content is really meant to be driven by the Society for Systematic Biology (SSB). Because of this, I have tried to be less heavy-handed in the journal’s direction. The direction of the journal should be driven by members of the society as reflected by submissions. There are some topics that I wish we had less submissions (for example, phylocode and DNA barcoding). When papers are submitted and go through review with positive results, I am very reluctant to reject them based on the subject matter.

Question: So you view the editors role as more of a service to the society rather than an opportunity to shape the field?
Both. The editor can shape the field by insisting on maintaining the highly rigorous standards for data analysis that Systematic Biology is known for, especially for empirical papers. Particular things that I require as EIC might differ from my predecessors.

Question: What is the difference between a good and a bad review of a paper?
The primary characteristic of an excellent review is that the reviewer has assumed the role of silent partner - this comes from Dick Olmstead when he was the editor. Reviewers do this because it has been done for them at the journal. We have an incredibly valuable tradition of rigorous yet constructive feedback in reviews.

Question: Do you have any advice for the next generation of systematists?
As early as possible, find your niche that differentiates you from all of your peers that are doing great work. You cannot just do “comparative biology of (fill in the blank)” or “molecular phylogeography of (fill in the blank).” Probably the easiest way to think about this is to imagine yourself on an airplane next to an intelligent layperson. Convey to them what is important about what you do in a manner that is unique. This is critical for the job search - it is a rare situation when the audience [of a job talk] is just phylogeneticists or even evolutionary biologists.

Question: OK now I’m going to ask you about two controversial groups. What is your take on the cladists?
The view that statistics are anathema to systematics is dead. All the vitality in the discipline is in statistical approaches.

Question: And how do you think we should respond to the creationists?
Fighting court battles require very different tactics than changing public opinion. To affect public opinion, there are two things we can do:
1. Publicly deconstruct the false dichotomy between macroevolution and microevolution.
2. Engage in a strong public outreach campaign over the importance of evolution in day to day life.

Thursday, April 9, 2009

When We Fail MrBayes…

A recent Dechronization post highlighted the unsuccessful attempts at Bayesian estimation of a large-scale bird phylogeny based on a multi-locus data set by Hackett et al. The apparent failure of MrBayes in this particular case (and under similarly challenging inference scenarios associated with large and/or complex data sets, e.g., Soltis et al., 2007; Moore et al., 2008) appears to raise serious concerns regarding our ability to estimate large-scale phylogeny using Bayesian methods.

However, it is important to carefully consider precisely what such studies have actually demonstrated: that Bayesian estimation of phylogeny appears to be intractable for certain data sets using default settings implemented in a particular program, MrBayes. Unfortunately, these anecdotal observations have led some researchers to a nested series of increasingly dubious and unsubstantiated conclusions. First, that it is impossible to reliably estimate phylogeny for this particular data set under any settings implemented in MrBayes, and more generally, that it is impossible to reliably estimate phylogeny for this particular data set not only using MrBayes but using any Bayesian methods, and finally by following this false premise to its ultimate conclusion, that it is impossible to reliably estimate phylogeny not only for this particular data set, but for any large-scale data set using Bayesian methods.

Although Bayesian estimation of phylogeny appears to succeed for the vast majority of empirical problems, there remain inference problems for which Bayesian estimation is apt to be intransigent, which may be usefully divided into three categories: (1) inference scenarios in which reliable Bayesian (or any other) estimation is likely to be problematic (e.g., whole-genome alignments for extremely large numbers of species); (2) inference scenarios in which rigorous application of existing Bayesian methods are apt to fail; and (3) inference scenarios in which imprudent application of existing Bayesian methods using default settings are apt to fail. I believe that the vast majority of reportedly “impossible” Bayesian phylogeny estimation problems fall within the latter two categories.

Existing implementations, such as MrBayes, approximate the joint posterior probability density of phylogeny and model parameters using some form of MCMC sampling (typically based on the
Metropolis-Hastings algorithm). These methods quietly specify a means of updating the value of each parameter (the proposal mechanisms), the probability of invoking each proposal mechanism (the proposal probability), and the magnitude of the proposed change issued by each proposal mechanism (the tuning parameters). Proposal mechanism design is an art form (there are no hard rules that ensure valid and efficient MCMC sampling for all problems). For this reason, for many (non-phylogenetic) Bayesian inference methods, it is the responsibility of the investigator to explore a range of proposal probabilities and tuning parameterizations that deliver acceptable MCMC performance.

Accordingly, most researchers familiar with Bayesian inference would consider it extremely naïve to expect that any specific MCMC sampling design would perform well for all (or even most) empirical data sets, especially in the very difficult case of phylogeny estimation. Nevertheless, the default settings of existing Bayesian phylogeny estimation programs are so successful that we are actually “
shocked, shocked to find that MrBayes does not solve all of our problems!!”. Without going into detail (as doing so would constitute an entirely separate post), the analyses detailed in the supporting material of the Hackett et al. study reads like a recipe for failure, and I would venture that the putative 'impossibility' of obtaining a reliable estimate with MrBayes in this case falls squarely under the third inference scenario defined above.

What does this mean for our phylogenetic community? First, I would argue that researchers interested in Bayesian estimation of phylogeny need to become much, much more sophisticated about diagnosing MCMC performance, carefully assessing convergence (ensuring that the chain has reached the stationary distribution, which is the joint posterior probability density of interest), mixing (assessing movement of the chain over the stationary distribution in proportion to the posterior probability of the parameter space), and sampling intensity (assessing adequacy of the number of independent samples used to approximate the posterior probability). Second, I believe that developers of Bayesian methods need to encourage and facilitate more vigorous and nuanced exploration of MCMC performance among users of these methods.

Researchers unwilling to develop the requisite knowledge to properly diagnose and troubleshoot MCMC performance should seriously consider alternative strategies, including collaboration with researchers who possess these skills or, of course, pursue alternative inference methods, including ‘fast’ ML approaches. However, it seems that most researchers are equally unclear about the potential deficiencies of the latter methods. Along these lines, and in the spirit of the anecdotal account that inspired this post, I note that I have encountered many data sets for which multiple independent searches using fast ML methods (implemented in GARLI and RAxML) rendered a series of estimates with significantly different MLEs, whereas convergence to a significantly higher mean marginal log likelihood using MrBayes appeared to be unproblematic. Indeed, Hackett et al. note that 80–90% of their fast ML searches converged to solutions with significantly different MLEs!! Moreover, the best of their fast ML searches--apparently based on a partitioned analysis using RAxML--resulted in a phylogeny with a log likelihood of -866,017.07, which is ~5,000–6,500 log likelihood units worse than the ‘unreliable’ plateaus in the time series plots of the marginal log likelihoods estimated with MrBayes!! Clearly, there are no easy solutions to these hard problems...

Friday, April 3, 2009

When MrBayes Fails...

Last summer, Hackett et al. published a widely-read study of phylogenetic relationships among major bird lineages based on 19 independent loci sampled from 169 species (see also Tom Near's previous post). Their study confirmed some patterns suggested by previous phylogenetic studies (e.g., ratites + tinamous as sister to remaining bird species) while also recovering some novel patterns (e.g., passerines sister to parrots [albiet with low support]). One of the more interesting results from their analyses, however, was relegated to the on-line supplement. In this supplement, we learn that all eight of the 10 million generation partitioned analyses they ran in MrBayes apparently failed to reach stationarity (see figure; note that the first 2 million generations are inexplicably trimmed from each analysis as 'burnin-in'). Unpartitioned analyses fared even worse, resulting in immediate crashes "regardless of the memory capacity of the computers used."

Among the partitioned analyses, some continued to shift to new areas of the likelihood surface until relatively late in the analysis. Perhaps even more troubling though was the fact that analyses that did appear to reach a stable plateau sampled significantly different likelihood scores (e.g., -lnL -861,000 v. -lnL 859,500). Is this problem unavoidable in analyses of large datasets?

The most obvious solution would be to simply run the analyses for more than 10 million generations. I've certainly had analyses that required more than 10 million generations to reach stationarity. Perhaps this wasn't done because it took two months on a super computer to run the 10 million generation analyses (anybody know if Hackett et al. or others have implemented longer runs since their paper was published?). Another possibile solution to their problems is to modify the parameters of the MC3 analyses implemented by MrBayes (recall that the MrBayes default is to run two independent MC3 analyses with one cold chain and three heated chains). Hackett et al. explored this possibility by running six analyses with one heated chain and one cold chain (B1-B6) and two analyses with six heated chains and one cold chain (A1-A2). The analyses run with multiple heated chains performed significantly better than those with a single heated chain, perhaps due to the fact that multiple chains are incrementally heated by MrBayes (meaning that the fourth of six heated chains has a flatter likelihood surface than the first). Hackett et al. do not discuss the temperatures used for the heated chains in their analyses, but their results suggest that running multiple heated chains in a single analysis is superior to repeatedly running analyses with only one heated chain.

In any case, Hackett et al.'s ultimate solution was to discard all of their Bayesian analyses and rely instead on parsimony and the fast maximum likelihood methods implemented by GARLI and RAxML. Is this shift away from Bayesian inference in favor of fast maximum likelihood searches for computational reasons a sign of things to come (or has this shift already occurred)? Are the fast maximum likelihood methods ready for prime time, or do people remain uncomfortable with the shortcuts they use to acheive their apparent computational efficiency?

Thursday, March 19, 2009

Bodega Phylogenetics Wiki - Major Revisions

The 10th annual Applied Phylogenetics Workshop took place last week at the Bodega Bay Marine Lab. If you didn't make it to the workshop, but are interested in learning about the latest methods in applied phylogenetics I have good news: we've wikified the workshop! Recent upgrades to the Wiki include new tutorials on how to conduct basic phylogenetic analyses, analyze morphological evolution using the program Brownie, and assesss patterns of community assembly and evolution. We've also expanded our coverage of phylogenetic comparative methods and made progress toward completion of a comprehensive BEAST tutorial. Check it out and don't be shy about signing up as a contributor yourself!