Skip to content

Review: Multiple comparisons of log-likelihoods with applications to phylogenetic inference

Citation

  • Shimodaira, H., & Hasegawa, M. (1999). Multiple comparisons of log-likelihoods with applications to phylogenetic inference. Molecular Biology and Evolution, 16(8), 1114-1116.
  • DOI

Summary

This letter is short and carries no abstract. It modifies the Kishino-Hasegawa test so that comparing many topologies at once accounts for the multiplicity of the comparison, and it defines the confidence set of topologies that the modified test produces.


Kishino and Hasegawa (1989) compare two topologies, one of them fixed in advance. Shimodaira and Hasegawa observe that the test is often used to compare many, and that doing so overlooks the sampling error in the selection of the topology being tested. The result is overconfidence in a wrong tree.

The multiplicity problem

Let \( L_\alpha \) be the maximum log likelihood under topology \( \alpha \), for \( M \) candidate topologies. The KH test compares \( L_\alpha \) against a prespecified \( L_\beta \), and \( (L_\beta - L_\alpha)/\hat\sigma_{\alpha\beta} \) is asymptotically standard normal.

The difficulty arises when \( L_\alpha \) is compared against \( L_{\hat\alpha} = \max\{L_1, \ldots, L_M\} \), because \( \hat\alpha \) is chosen by the same data. The distribution of \( L_{\hat\alpha} - L_\alpha \) has to account for that selection.

The procedure

The letter gives six steps. The test statistic is the difference from the best candidate,

\[T_\alpha = \max\{L_1 - L_\alpha, \ldots, L_M - L_\alpha\} = L_{\hat\alpha} - L_\alpha\]

Bootstrap replicates of the vector \( (L_1, \ldots, L_M) \) are generated and stored as an \( M \times N \) array; the RELL method of Kishino, Miyata and Hasegawa (1990) makes this affordable. Each row is then centered by subtracting its own mean, which the authors describe as regarding the replicates as generated under the least favorable configuration. The statistic is recomputed on each centered column, and the \( p \)-value \( P_\alpha \) is the fraction of replicates whose recomputed statistic exceeds the observed \( T_\alpha \).

The confidence set \( \mathcal{T} \) collects the topologies with \( P_\alpha \geq P^* \). Its coverage satisfies

\[P_C \geq 1 - P^*\]

with equality holding at the least favorable configuration.

The size of the confidence set

Remark 5 qualifies that inequality. Since \( P_C \geq 1 - P^* \) holds in general rather than with equality, the confidence set can be larger than the stated level requires, and the effect grows as more topologies are compared. The authors advise keeping \( M \) as small as possible by eliminating extremely unlikely topologies, and note that all possible topologies need not be included when the question concerns particular biological hypotheses.

Table 1 shows the effect on 15 bifurcating topologies of a mammal data set. At \( P^* = 0.1 \), the bootstrap selection probability and the KH test both admit the best three trees, the multiple-comparison method admits the best seven, and the standardized form admits an eighth. The authors describe the two multiple-comparison methods as conservative, and note that the first two give smaller confidence limits while not being guaranteed to satisfy the coverage inequality.

Relevance to Hifuku

An elite-archive survey holds thousands of topologies at once, one per filled niche of the chart, so a survey compares many topologies by construction.

The log-likelihood gate takes its shape from equation 12 of Kishino and Hasegawa, whose \( \sqrt{n} \) scaling this letter leaves untouched. What the letter governs is how far such a margin can be read as a statement about confidence. It cannot be read that way over an archive: the gate is one margin applied to every topology found, without a centering step and without a correction for the selection of the best. Hifuku treats the gate as a bound on exploration, and the archive as a map of the region a survey reached.

Remark 5 also bears on chart selection. The advice to keep \( M \) small by removing unlikely candidates is the reasoning that leads charts.ChartTable.select to discard triples below the quality and overlap thresholds before any survey runs.

In-document navigation

Section Content
Opening paragraphs Why the KH test overstates confidence when many topologies are compared
Steps 1 to 6 The test statistic, the bootstrap array, the centering, and the p-value
Equation 1 Coverage of the confidence set, \( P_C \geq 1 - P^* \)
Remarks 1 to 6 Conditions, the least favorable configuration, and the size of the candidate set
Table 1 Fifteen mammal topologies under four methods