Review: treespace, Statistical Exploration of Landscapes of Phylogenetic Trees¶
Citation
- Jombart, T., Kendall, M., Almagro-Garcia, J., & Colijn, C. (2017). treespace: Statistical exploration of landscapes of phylogenetic trees. Molecular Ecology Resources, 17(6), 1385–1392.
- DOI
Abstract¶
The increasing availability of large genomic data sets as well as the advent of Bayesian phylogenetics facilitates the investigation of phylogenetic incongruence, which can result in the impossibility of representing phylogenetic relationships using a single tree. While sometimes considered as a nuisance, phylogenetic incongruence can also reflect meaningful biological processes as well as relevant statistical uncertainty, both of which can yield valuable insights in evolutionary studies. We introduce a new tool for investigating phylogenetic incongruence through the exploration of phylogenetic tree landscapes. Our approach, implemented in the R package treespace, combines tree metrics and multivariate analysis to provide low-dimensional representations of the topological variability in a set of trees, which can be used for identifying clusters of similar trees and group-specific consensus phylogenies. treespace also provides a user-friendly web interface for interactive data analysis and is integrated alongside existing standards for phylogenetics. It fills a gap in the current phylogenetics toolbox in R and will facilitate the investigation of phylogenetic results.
treespace is the closest contemporary tool to Hifuku, and its discussion names the extension Hifuku implements. It generalizes the tree-space visualization of Hillis et al. (2005) to any tree metric, and treats phylogenetic incongruence as a signal to explore rather than a nuisance to summarize away.
Explore incongruence, do not summarize it¶
Incongruence, the failure of different genes or different analyses to agree on one tree, can carry real biology (horizontal transfer, incomplete lineage sorting, gene loss) or statistical uncertainty. Support values summarize it only when congruence is already high, so Jombart et al. build a workflow to examine it instead. The workflow has four steps: take a set of trees; compute all pairwise distances under a chosen metric; project the distance matrix into a low dimensional Euclidean space by metric multidimensional scaling; and cluster the projection to find "tree islands," each summarized by a representative median tree. The package offers seven metrics, among them the Robinson-Foulds count (Robinson & Foulds (1981)), the branch score of Kuhner and Felsenstein (kuhner1994branch), the Billera-Holmes-Vogtmann geodesic, the path-difference metric, and the Kendall-Colijn metric; non-Euclidean distances are made Euclidean by the Cailliez transformation before scaling. The worked example, seventeen dengue-4 sequences, shows neighbor-joining, maximum-likelihood, and Bayesian trees separating into distinct clusters that disagree on the placement of one clade.
The gap it names, and Hifuku fills¶
The discussion of treespace states the missing ingredient plainly. Its authors note that regions of tree space are defined both by topology and by the trees' "log-likelihood under a specific evolutionary model," and write that although the package does not do so, "it would be interesting to incorporate information on tree log-likelihood as weights in the analysis." Hifuku's map is that map: the elevation is the alignment log-likelihood, computed exactly at every tree the survey visits. The same discussion points to t-SNE as a promising nonlinear projection for tree sets, which is the method Hifuku uses to stitch its per-gene charts into one frame.
What Hifuku takes and changes¶
Hifuku shares treespace's stance, that incongruence is the object of study and a metric embedding is the way to see it, and shares two of its metrics, the branch score (kuhner1994branch) and RF (Robinson & Foulds (1981)), which Hifuku merges into one normalized distance. The difference is in what is placed on the map. treespace projects a fixed set of already-inferred or already-sampled trees and reads off clusters. Hifuku instead illuminates the landscape with a quality-diversity search, keeping the best tree in each region, and fixes a shared three-anchor chart so that many genes occupy one common, reproducible frame rather than a layout that shifts with the sample. Where treespace summarizes each cluster with a median tree, Hifuku reports the elevation everywhere the survey reached.