Skip to content

Review: Diversification of Giant and Large Eukaryotic dsDNA Viruses Predated the Origin of Modern Eukaryotes

Citation

  • Guglielmini, J., Woo, A. C., Krupovic, M., Forterre, P., & Gaia, M. (2019). Diversification of giant and large eukaryotic dsDNA viruses predated the origin of modern eukaryotes. Proceedings of the National Academy of Sciences, 116(39), 19585–19592.
  • DOI

Abstract

Giant and large eukaryotic double-stranded DNA viruses from the Nucleo-Cytoplasmic Large DNA Virus (NCLDV) assemblage represent a remarkably diverse and potentially important component of the eukaryotic virome. However, their origin(s), evolution, and potential roles in the emergence of modern eukaryotes remain subjects of intense debate. Here we present robust phylogenetic trees of NCLDVs, based on the 8 most conserved proteins responsible for virion morphogenesis and informational processes. Our results uncover the evolutionary relationships between different NCLDV families and support the existence of 2 superclades of NCLDVs, each encompassing several families. We present evidence strongly suggesting that the NCLDV core genes, which were involved in both informational processes and virion formation, were acquired vertically from a common ancestor. Among them, the largest subunits of the DNA-dependent RNA polymerase were transferred between 2 clades of NCLDVs and proto-eukaryotes, giving rise to 3 eukaryotic DNA-dependent RNA polymerases. Our results strongly suggest that these transfers and the diversification of NCLDVs predated the emergence of modern eukaryotes, emphasizing the major role of these viruses in the evolution of cellular domains.


This paper supplies both the marker genes and the motivating problem for Hifuku's NCLDV survey. It reconstructs the phylogeny of the giant viruses from their most conserved proteins, and its central methodological difficulty, measuring how well the single-gene trees agree, is the difficulty Hifuku's map is built to answer.

The eight core genes

Guglielmini et al. identify the NCLDV core gene set by orthology across a curated collection of genomes (96 collected, 73 retained for the conservation survey). Three proteins are strictly conserved in every genome: the family-B DNA polymerase (DNApol B), the D5-like primase-helicase, and a homolog of the Poxvirus Late Transcription Factor 3 (VLTF3). Three more are added despite scattered loss: transcription elongation factor II-S (TFIIS), the A32-family genome packaging ATPase (pATPase), and the major capsid protein (MCP). The MCP has no homolog in pandoraviruses, and the pATPase is absent from Pithovirus, Cedratvirus, and Orpheovirus. The last two markers are the two largest subunits of the DNA-dependent RNA polymerase (RNAP-a and RNAP-b), present in 92% of the genomes; because these subunits are universal across Archaea, Bacteria, and Eukarya, they also relate the viruses to cellular life. The eight proteins cover both jobs of a virion: building the particle (MCP, pATPase) and running its information processing (the polymerases, the primase, and the transcription factors).

A congruent core, and what it implies

Two independent reconstructions from the eight-gene set recover the same topology: a Bayesian CAT-GTR analysis of the concatenation, and a subtree prune-and-regraft supertree built from the single-gene trees. The authors read this concordance as the group's true vertical history. In that tree the NCLDVs fall into two superclades: MAPI (Marseilleviridae, Ascoviridae, the Pitho-like viruses, and Iridoviridae) and PAM (Phycodnaviridae, Asfarviridae, and the Megavirales). A shared, congruent signal argues that the core genes descend vertically from one common ancestor rather than being assembled piecemeal by later transfer. Positioning the viruses against the cellular RNA polymerases then dates that ancestor: the diversification of the NCLDVs predates the last eukaryotic common ancestor, and two of the three eukaryotic RNA polymerases (RNAP-II, and probably RNAP-I) appear to have been transferred from NCLDV lineages into proto-eukaryotes. The viruses are placed near the root of the eukaryotic informational machinery, not derived from it.

Scoring congruence by hand

The whole argument rests on the single-gene trees carrying the same signal as the concatenation, and the paper measures this directly. Each single-gene maximum-likelihood tree is scored against a set of reference features, the clades read off the Bayesian concatenated tree. Every feature is checked for presence or absence in every tree, features are weighted by how often they recur, and each tree is scored by how many reference features it recovers. The two shortest markers, TFIIS and the VLTF3-like protein, resolve poorly; Poxviridae and Aureococcus anophagefferens form long, unstable branches and are removed to avoid long-branch artifacts. This per-feature counting is careful, but it is manual, and it collapses a whole tree to a tally against one reference. It is the step Hifuku replaces.

The marker set in Hifuku

The ten NCLDV alignments surveyed in this project are this core set. Eight are the core genes above: capsid (MCP), dnapol (DNApol B), primase, vltf3 (VLTF3), tf2s (TFIIS), pATPase, and rnapol1/rnapol2 (the two RNAP subunits). The remaining two, capsid-polinto and pATPase-polinto, are the Polinton-encoded homologs of the capsid and packaging ATPase that Guglielmini et al. use as the outgroup to root the NCLDV tree.

Where the paper scores a gene tree by counting recovered clades, Hifuku places each gene's likelihood landscape on a shared two-dimensional chart under the normalized clade branch score (kuhner1994branch) and Robinson-Foulds (Robinson & Foulds (1981)) metric, so congruence becomes a distance and disagreement becomes a direction. The neighbor-joining tree (saitou1987nj) of each alignment fixes the chart anchors. The survey renders the paper's qualitative reading as geometry: most markers cluster, and the D5-like primase, the gene flagged as topologically divergent in a companion analysis of these same trees, is the one that sits off the shared chart plane.


Related reading: The NCLDV core genes, revisited works through the same congruence question using patristic-distance correlations.