Skip to content

Review: Evaluation of the maximum likelihood estimate of the evolutionary tree topologies from DNA sequence data

Citation

  • Kishino, H., & Hasegawa, M. (1989). Evaluation of the maximum likelihood estimate of the evolutionary tree topologies from DNA sequence data, and the branching order in Hominoidea. Journal of Molecular Evolution, 29(2), 170-179.
  • DOI

Abstract

A maximum likelihood method for inferring evolutionary trees from DNA sequence data was developed by Felsenstein (1981). In evaluating the extent to which the maximum likelihood tree is a significantly better representation of the true tree, it is important to estimate the variance of the difference between log likelihood of different tree topologies. Bootstrap resampling can be used for this purpose (Hasegawa et al. 1988; Hasegawa and Kishino 1989), but it imposes a great computation burden. To overcome this difficulty, we developed a new method for estimating the variance by expressing it explicitly. The method was applied to DNA sequence data from primates in order to evaluate the maximum likelihood branching order among Hominoidea. It was shown that, although the orangutan is convincingly placed as an outgroup of a human and African apes clade, the branching order among human, chimpanzee, and gorilla cannot be determined confidently from the DNA sequence data presently available when the evolutionary rate constancy is not assumed.


Kishino and Hasegawa (1989) give the scale on which a difference of log likelihoods varies. Hifuku uses that scale to set the log-likelihood gate of the elite-archive survey, which bounds how far from the best tree a survey explores.

Sites as units of sampling

The paper treats homologous sites as the units of sampling, independent of one another and homogeneous along the sequence. For \( n \) sites the log likelihood of topology \( i \) is a sum over them,

\[l_{(i)}(\hat\theta_{(i)}|X) = \sum_{k=1}^{n} \log f_{(i)}(X_k|\hat\theta_{(i)})\]

which the authors note is, for a large sample, in the form of a sum of independently and identically distributed random variables. The central limit theorem then gives its asymptotic normality, and its variance is estimated by the sample variance of the per-site terms (equation 11).

The variance of a difference

The quantity that matters for comparing two topologies is the difference of their log likelihoods, and equation 12 estimates its variance directly:

\[\hat V = \widehat{\mathrm{Var}}\bigl[l_{(2)}(\hat\theta_{(2)}|X) - l_{(1)}(\hat\theta_{(1)}|X)\bigr] = \frac{n}{n-1} \sum_{k=1}^{n} \left\{ \log \frac{f_{(2)}(X_k|\hat\theta_{(2)})}{f_{(1)}(X_k|\hat\theta_{(1)})} - \frac{1}{n} \sum_{k=1}^{n} \log \frac{f_{(2)}(X_k|\hat\theta_{(2)})}{f_{(1)}(X_k|\hat\theta_{(1)})} \right\}^2\]

This is \( n \) times the sample variance of the per-site log-likelihood ratio. Writing that per-site standard deviation as \( s \), the standard error of the difference is \( \sqrt{n}\,s \). Equation 13 builds the 95 percent interval from \( \pm 1.96 \sqrt{\hat V} \).

The method expresses the variance explicitly to avoid the cost of bootstrap resampling, and the paper reports that its estimates agree well with bootstrap values, running slightly smaller. Felsenstein incorporated it into DNAML in PHYLIP 3.1.

What the reported errors show

Table 1 lists log likelihoods and their differences for five hominoid topologies over seven data sets, from 232 to 5198 sites. Taking the column for tree 4, a topology the data reject, the standard error of the difference from tree 1 runs from 5.92 at 232 sites to 22.66 at 5198.

Those errors are far more nearly constant after division by \( \sqrt{n} \) than after division by \( n \): over a 22-fold range of sequence length, the first varies by a factor of 2.4 and the second by 6.6. The measurement is reproduced in docs/theory/algebra/kishino1989logl-margin-scaling.py.

Relevance to Hifuku

The elite-archive survey keeps a tree when its log likelihood is within a margin of the best found. The margin scales as \( \sqrt{n} \) in the number of sites, following the standard error of equation 12, so the gate holds a comparable region of tree space for alignments of different length. The coefficient lives in constants.ARCHIVE_LOGL_MARGIN_COEFF and the gate is applied in driver._floor. A caller who wants a fixed number of nats sets logl_margin, and one who wants an absolute floor sets logl_floor.

The coefficient places the gate about seven times the 95 percent half-width of equation 13, using the per-site standard deviations of Table 1. The gate bounds exploration rather than reporting a confidence set, and a survey crosses the region between competing topologies rather than stopping at the edge of the set it would report. What the paper fixes here is the shape of the margin, and the width is a survey setting.

The paper's own conclusion is a caution worth carrying: the branching order among human, chimpanzee and gorilla could not be resolved from the data then available, because the standard errors of the differences exceeded the differences themselves. A survey reports the scale on which its log likelihoods vary for the same reason.

In-document navigation

Section Content
Method for Estimating the Variance, equations 1 to 9 Setup, the substitution model, and the posterior probability of a topology
Equations 10 and 11 The log likelihood as a sum over sites, and the variance of one topology
Equation 12 The variance of the difference between two topologies
Equation 13 The 95 percent confidence interval on the posterior probability
Table 1 Log likelihoods, differences and standard errors over seven data sets
Appendix A, equation 20 Derivation of the variance estimate
Appendix B When the value of equation 12 is of the order of the number of free parameters