Review: PhylomeDB v3.0: An Expanding Repository of Genome-Wide Collections of Trees, Alignments and Phylogeny-Based Orthology and Paralogy Predictions¶
Citation
- Huerta-Cepas, J., Capella-Gutierrez, S., Pryszcz, L. P., Denisov, I., Kormes, D., Marcet-Houben, M., & Gabaldón, T. (2011). PhylomeDB v3.0: an expanding repository of genome-wide collections of trees, alignments and phylogeny-based orthology and paralogy predictions. Nucleic Acids Research, 39(Database issue), D556–D560.
- DOI
Abstract¶
PhylomeDB is a public database storing genome-wide collections of gene trees (phylomes) with associated multiple sequence alignments and phylogeny-based orthology and paralogy predictions. Version 3 adds new phylomes, improved navigation, and REST API access. The database provides a resource for comparative genomics, gene family evolution, and functional annotation.
A phylome is a complete collection of maximum-likelihood gene trees for every gene in a genome, built under a consistent methodology (alignment, model selection, tree inference). PhylomeDB collects and serves these phylomes for a growing set of model and non-model organisms, providing the broader research community with pre-computed, high-quality phylogenetic trees without needing to re-run the computationally intensive inference step.
Version 3 of PhylomeDB added several new phylomes including human, mouse, and a selection of key model organisms. Each phylome entry includes the multiple sequence alignment, the best-fit substitution model selected by ModelFinder (or its equivalent at the time), and the maximum-likelihood tree topology with branch lengths.
Hifuku uses PhylomeDB as a source of empirical phylogenetic datasets for development and benchmarking. Real gene trees from PhylomeDB provide: biologically realistic tree topologies (with typical rates of duplication, loss, and speciation), empirically calibrated branch length distributions, and alignments in the size range (tens to hundreds of taxa, hundreds to thousands of sites) that Hifuku is designed for. These properties are difficult to reproduce faithfully with purely synthetic datasets.
See also PhylomeDB v4 for the expanded version of the database.