Review: PhylomeDB v4: Zooming into the Plurality of Evolutionary Histories of a Genome¶
Citation
- Huerta-Cepas, J., Capella-Gutiérrez, S., Pryszcz, L. P., Marcet-Houben, M., & Gabaldón, T. (2014). PhylomeDB v4: zooming into the plurality of evolutionary histories of a genome. Nucleic Acids Research, 42(D1), D897–D902.
- DOI
Abstract¶
PhylomeDB v4 substantially expands the collection of genome-wide phylomes, adds support for browsing multiple evolutionary histories per gene family, and introduces improved orthology and paralogy predictions based on the expanded tree collection. The database now covers over 700,000 trees from more than 100 phylomes.
PhylomeDB v4 is a major expansion of the PhylomeDB resource (see PhylomeDB v3), doubling the number of phylomes and introducing new tools for comparing and summarizing the plurality of evolutionary histories inferred for gene families. The title's "plurality of evolutionary histories" refers to the fact that different genes in a genome tell different stories about organismal history due to horizontal gene transfer, gene duplication, and lineage sorting; PhylomeDB v4 makes it possible to browse all of these histories together.
Key additions in v4 include an improved web interface with phylogenetic tree visualization, programmatic REST API access, and integration of multiple reference phylomes allowing cross-phylome comparisons. The database also includes per-tree model selection information (the substitution model that was selected during inference), which is directly useful for reproducing the likelihood landscape under the correct model.
PhylomeDB is the source of Hifuku's real-data smoke test: three Candida
albicans gene alignments (datasets/phylome205/, genes LZU, N5C, and M28).
Their neighbor-joining trees serve as the anchor triangle, and the map-sanity
test renders and checks the likelihood landscape of each. The v4 release records
per-tree model-selection metadata, which documents the substitution model used
in the original inference.
Future work. The broader v4 collection would support testing across a wider range of taxonomic groups and gene families, from deep to shallow divergences and across rate regimes.