Skip to content

Review: An approximately unbiased test of phylogenetic tree selection

Citation

  • Shimodaira, H. (2002). An approximately unbiased test of phylogenetic tree selection. Systematic Biology, 51(3), 492-508.
  • DOI

Abstract

An approximately unbiased (AU) test that uses a newly devised multiscale bootstrap technique was developed for general hypothesis testing of regions in an attempt to reduce test bias. It was applied to maximum-likelihood tree selection for obtaining the confidence set of trees. The AU test is based on the theory of Efron et al. (Proc. Natl. Acad. Sci. USA 93:13429-13434, 1996), but the new method provides higher-order accuracy yet simpler implementation. The AU test, like the Shimodaira-Hasegawa (SH) test, adjusts the selection bias overlooked in the standard use of the bootstrap probability test, the Kishino-Hasegawa test. The selection bias comes from comparing many trees at the same time and often leads to overconfidence in the wrong trees. The SH test, though safe to use, may exhibit another type of bias such that it appears overly conservative. Here I show that the AU test is less biased than other methods in typical cases of tree selection. These points are illustrated in a simulation study as well as in the analysis of mammalian mitochondrial protein sequences.


Shimodaira (2002) completes the sequence that runs from the Kishino-Hasegawa test through the Shimodaira-Hasegawa test. It states plainly what each of the three controls, and supplies the log-likelihood decomposition that the earlier two rely on.

The log likelihood as a sum over sites

Equation 1 writes the score of tree \( i \) as

\[Y_i = \sum_{n=1}^{N} X_{i,n}\]

where \( N \) is the sequence length and \( X_{i,n} \) is the site-wise log likelihood of tree \( i \) at site \( n \). The paper notes that \( Y_i \) is a random variable, because it is calculated from the molecular sequences, and that assuming independence of the sites makes its generation equivalent to sampling sites from imaginary sequences of infinite length.

Equation 2 extends this to a sample of length \( N' \) that may differ from \( N \),

\[Y_i^{*} = \frac{N}{N'} \sum_{j=1}^{N'} X_{i,n_j}\]

with the factor \( N/N' \) making the result comparable to \( Y_i \). Varying \( N' \) is the multiscale bootstrap the AU test is built on. The central limit theorem applied to this sum is what justifies the normal approximation of \( Y_i^{*} - Y_j^{*} \) used in the KH test.

What each test controls

The paper is direct about the standing of the three procedures.

The bootstrap probability test is biased, as discussed by several authors. The KH test controls neither type-1 error nor bias, given its selection bias: the likelihood of the maximum-likelihood tree is biased upward, because a maximum over many trees can easily reach a large value by chance, and typically many trees have been compared when the choice is made.

The SH test is excellent in terms of type-1 error but heavily biased, which is why it appears conservative for comparisons of many trees. Strimmer and Rambaut (2001) observed that the number of trees in its confidence set tends to grow with the number of trees compared. The paper attributes this to the SH test being derived under a pessimistic assumption.

The AU test is derived under a moderate assumption and in most cases works better than the SH test, while in special cases it may violate type-1 error. It is not monotone; the paper notes that monotonicity and unbiasedness are not compatible and that the test relaxes the former to reduce the latter.

Relevance to Hifuku

Equation 1 is the statement Hifuku's log-likelihood gate depends on: the score of a tree is a sum over sites, so a margin on that score has to scale with the number of sites. The gate follows the \( \sqrt{n} \) standard error of Kishino and Hasegawa's equation 12, and equation 1 is where the decomposition that makes that scaling meaningful is set out most clearly.

The paper's account of selection bias sets the boundary of what an archive reports. A survey holds the best tree found in each niche and gates on the distance from the best found anywhere, which is a maximum over many topologies and so carries exactly the upward bias described here. The archive is therefore read as a map of the region a survey reached, and the elevation of a niche as the best log likelihood found there, rather than as a confidence statement about any tree in it.

None of the three tests is implemented in Hifuku. Their use here is to fix the shape of the gate and to bound the claims made about the result. CONSEL (Shimodaira and Hasegawa, 2001) implements the AU test for a candidate set that has already been chosen, which is the natural downstream step for a small set of topologies drawn from an archive.

In-document navigation

Section Content
Introduction Bias in the bootstrap probability, KH and SH tests, and what each controls
Methods, Null Hypotheses, equations 1 to 3 The log likelihood as a sum over sites, and the hypothesis regions
Bootstrap Resampling, equation 2 The multiscale bootstrap and the role of \( N' \)
Coverage Probability, equation 4 Type-1 error, unbiasedness, and the confidence set
The Multiscale Bootstrap The steps of the AU test
Simulation and mammalian mt analysis Comparison of the tests in practice
Appendix Technical details and the geometric theory