Skip to content

Review: Analysis and Visualization of Tree Space

Citation

  • Hillis, D. M., Heath, T. A., & St. John, K. (2005). Analysis and visualization of tree space. Systematic Biology, 54(3), 471–482.
  • DOI

Abstract

We explored the use of multidimensional scaling (MDS) of tree-to-tree pairwise distances to visualize the relationships among sets of phylogenetic trees. We found the technique to be useful for exploring "tree islands" (sets of topologically related trees among larger sets of near-optimal trees), for comparing sets of trees obtained from bootstrapping and Bayesian sampling, for comparing trees obtained from the analysis of several different genes, and for comparing multiple Bayesian analyses. The technique was also useful as a teaching aid for illustrating the progress of a Bayesian analysis and as an exploratory tool for examining large sets of phylogenetic trees. We also identified some limitations to the method, including distortions of the multidimensional tree space into two dimensions through the MDS technique, and the definition of the MDS-defined space based on a limited sample of trees. Nonetheless, the technique is a useful approach for the analysis of large sets of phylogenetic trees.


This is the founding paper for visualizing tree space, and the direct ancestor of Hifuku's map. It makes the case Hifuku also makes: a large collection of trees should be looked at, not reduced to one summary.

Consensus discards structure

A phylogenetic analysis usually returns many trees, the equally-parsimonious solutions, a bootstrap set, or a Bayesian posterior sample, and the common practice is to collapse them into a single consensus. Hillis et al. observe that consensus discards most of the information: distinct "tree islands" of solutions that support different biological histories are averaged into one poorly-resolved tree. They propose to visualize the collection instead. Pairwise Robinson-Foulds distances (Robinson & Foulds (1981)), weighted or unweighted, are laid out in two dimensions by multidimensional scaling, which minimizes a stress function between the true and the plotted distances. Points are colored by optimality score or by which analysis produced them, and any point can be expanded to its tree.

The applications they demonstrate map onto Hifuku's uses one for one: a Bayesian sample occupies a tighter region than a bootstrap set; single-gene tree clouds sit offset from the true tree while the concatenation centers on it; two Bayesian runs that sample the same region have converged; and the progress of an MCMC chain, colored by likelihood, moves from a low-scoring start toward a high-scoring basin. The multi-gene layout is the direct predecessor of Hifuku's NCLDV marker map, and the likelihood-colored chain prefigures reading elevation from the log-likelihood.

The distortion they warned about

Hillis et al. are careful about what a two-dimensional picture can and cannot show, and their cautions are the problems Hifuku is built to confront.

  • Projection distorts. Their Figure 10 takes a set of trees that are all exactly one split (RF = 2) from a reference. In the true space these trees lie on a sphere, equidistant from the reference; forced into two dimensions the sphere collapses to a circle, so some trees appear nearer the reference than others although all are equally distant. They advise the reader to "check the primary distance matrix before interpreting too much" about the spatial layout.
  • The space is sample-defined. The MDS layout is recomputed, and redistorted, whenever a new tree is added, so it does not exist independently of the trees drawn. Ideally one would define the space over all possible trees.
  • RF is coarse. Two trees can differ by the placement of a single taxon yet reach the maximum RF distance.

What Hifuku takes and changes

Hifuku keeps the central idea, that a map of tree space is more honest than a summary statistic, and answers the three cautions directly. It illuminates the likelihood landscape, keeping the best tree in each region rather than plotting a fixed sample. It fixes a shared three-anchor chart, so many genes land in one common frame and the space does not redefine each time a tree is added. Its metric combines the branch score (kuhner1994branch) with RF (Robinson & Foulds (1981)) in a form that is squared-Euclidean by construction, so the only obstruction to a flat map is dimensional, never curvature. And rather than ask the reader to check the distance matrix, it draws the residual distortion, as bent scale bars on the assembled map and as the out-of-plane residual at every point.