Skip to content

Survey IO

A survey is the self-contained HDF5 output of a run: the taxon namespace, the shared alignment pool, and the per-chart elite archives, bundled in one file. SurveyPlan describes what to run and round-trips through a TOML plan file. SurveyRun prepares the shared context and turns the plan into addressable units of work. Survey holds the result, saves it, and reloads it, and frame/cloud unpack the elites for downstream analysis.

SurveyPlan dataclass

A declarative survey: the alignment pool, the charts, and the parameters.

The plan is the same object as the plan file. :meth:from_toml and :meth:to_toml round-trip it, so a run is reproducible and shareable as one small text file.

The plan validates its own labels at construction. Every anchor and every gene must name an entry of sequence_alignments, and a chart must name three distinct anchors, so a typo raises where the user wrote it.

Attributes:

Name Type Description
sequence_alignments dict

Label to alignment path. The [alignments] table of the plan file.

charts list[ChartSpec]

The charts to survey.

model str

Substitution model name, or "auto".

w float

Topology weight of the tree metric, in [0, 1].

description str

Free text stored with the survey.

params dict

Overrides for :class:~hifuku.params.ArchiveParams, keyed by field name. A field that no layer names keeps its default.

gpu_params dict

Overrides for :class:~hifuku.params.GpuParams. A device value on a host run raises, so a setting cannot be quietly ignored.

Source code in src/hifuku/survey_io.py
@dataclass
class SurveyPlan:
    """A declarative survey: the alignment pool, the charts, and the parameters.

    The plan is the same object as the plan file.  :meth:`from_toml` and
    :meth:`to_toml` round-trip it, so a run is reproducible and shareable as one
    small text file.

    The plan validates its own labels at construction.  Every anchor and every
    gene must name an entry of ``sequence_alignments``, and a chart must name
    three distinct anchors, so a typo raises where the user wrote it.

    Attributes
    ----------
    sequence_alignments : dict
        Label to alignment path.  The ``[alignments]`` table of the plan file.
    charts : list[ChartSpec]
        The charts to survey.
    model : str
        Substitution model name, or ``"auto"``.
    w : float
        Topology weight of the tree metric, in [0, 1].
    description : str
        Free text stored with the survey.
    params : dict
        Overrides for :class:`~hifuku.params.ArchiveParams`, keyed by field
        name.  A field that no layer names keeps its default.
    gpu_params : dict
        Overrides for :class:`~hifuku.params.GpuParams`.  A device value on a
        host run raises, so a setting cannot be quietly ignored.
    """

    sequence_alignments: dict          # label -> alignment path
    charts: list                       # list[ChartSpec]
    model: str = "LG"
    w: float = 0.5
    description: str = ""
    params: dict = field(default_factory=dict)
    gpu_params: dict = field(default_factory=dict)

    def __post_init__(self):
        from hifuku.params import ArchiveParams, GpuParams
        self.charts = [c if isinstance(c, ChartSpec) else ChartSpec(**c)
                       for c in self.charts]
        known = set(self.sequence_alignments)
        for i, chart in enumerate(self.charts):
            for label in chart.anchors:
                if label not in known:
                    raise ValueError(
                        f"chart {i} names anchor {label!r}, which is not in "
                        f"alignments; known labels are {sorted(known)}")
            for label in chart.genes:
                if label not in known:
                    raise ValueError(
                        f"chart {i} names gene {label!r}, which is not in "
                        f"alignments; known labels are {sorted(known)}")
        # Reject an unknown parameter name here, where the user wrote it.
        ArchiveParams.resolve(self.params)
        GpuParams.resolve(self.gpu_params)

    @classmethod
    def from_toml(cls, path) -> "SurveyPlan":
        """Read a plan file.

        The file holds ``description``, ``model`` and ``w`` at the top level, an
        ``[alignments]`` table of label to path, optional ``[params]`` and
        ``[gpu]`` tables, and one ``[[charts]]`` table per chart.
        """
        import tomllib
        with open(path, "rb") as fh:
            raw = tomllib.load(fh)
        return cls(
            sequence_alignments=dict(raw.get("alignments", {})),
            charts=[ChartSpec(anchors=tuple(c["anchors"]), genes=list(c["genes"]))
                    for c in raw.get("charts", [])],
            model=raw.get("model", "LG"),
            w=float(raw.get("w", 0.5)),
            description=raw.get("description", ""),
            params=dict(raw.get("params", {})),
            gpu_params=dict(raw.get("gpu", {})),
        )

    def to_toml(self, path) -> None:
        """Write the plan file that :meth:`from_toml` reads."""
        def fmt(value):
            if isinstance(value, bool):
                return "true" if value else "false"
            if isinstance(value, str):
                return json.dumps(value)
            if value is None:
                raise ValueError(
                    "a plan file cannot store None; omit the key instead")
            return repr(value)

        lines = [f"description = {json.dumps(self.description)}",
                 f"model = {json.dumps(self.model)}",
                 f"w = {self.w!r}", ""]
        if self.params:
            lines.append("[params]")
            lines += [f"{k} = {fmt(v)}" for k, v in sorted(self.params.items())]
            lines.append("")
        if self.gpu_params:
            lines.append("[gpu]")
            lines += [f"{k} = {fmt(v)}" for k, v in sorted(self.gpu_params.items())]
            lines.append("")
        lines.append("[alignments]")
        lines += [f"{k} = {json.dumps(str(v))}"
                  for k, v in self.sequence_alignments.items()]
        for chart in self.charts:
            lines += ["", "[[charts]]",
                      f"anchors = {json.dumps(list(chart.anchors))}",
                      f"genes   = {json.dumps(list(chart.genes))}"]
        Path(path).write_text("\n".join(lines) + "\n")

from_toml classmethod

from_toml(path) -> 'SurveyPlan'

Read a plan file.

The file holds description, model and w at the top level, an [alignments] table of label to path, optional [params] and [gpu] tables, and one [[charts]] table per chart.

Source code in src/hifuku/survey_io.py
@classmethod
def from_toml(cls, path) -> "SurveyPlan":
    """Read a plan file.

    The file holds ``description``, ``model`` and ``w`` at the top level, an
    ``[alignments]`` table of label to path, optional ``[params]`` and
    ``[gpu]`` tables, and one ``[[charts]]`` table per chart.
    """
    import tomllib
    with open(path, "rb") as fh:
        raw = tomllib.load(fh)
    return cls(
        sequence_alignments=dict(raw.get("alignments", {})),
        charts=[ChartSpec(anchors=tuple(c["anchors"]), genes=list(c["genes"]))
                for c in raw.get("charts", [])],
        model=raw.get("model", "LG"),
        w=float(raw.get("w", 0.5)),
        description=raw.get("description", ""),
        params=dict(raw.get("params", {})),
        gpu_params=dict(raw.get("gpu", {})),
    )

to_toml

to_toml(path) -> None

Write the plan file that :meth:from_toml reads.

Source code in src/hifuku/survey_io.py
def to_toml(self, path) -> None:
    """Write the plan file that :meth:`from_toml` reads."""
    def fmt(value):
        if isinstance(value, bool):
            return "true" if value else "false"
        if isinstance(value, str):
            return json.dumps(value)
        if value is None:
            raise ValueError(
                "a plan file cannot store None; omit the key instead")
        return repr(value)

    lines = [f"description = {json.dumps(self.description)}",
             f"model = {json.dumps(self.model)}",
             f"w = {self.w!r}", ""]
    if self.params:
        lines.append("[params]")
        lines += [f"{k} = {fmt(v)}" for k, v in sorted(self.params.items())]
        lines.append("")
    if self.gpu_params:
        lines.append("[gpu]")
        lines += [f"{k} = {fmt(v)}" for k, v in sorted(self.gpu_params.items())]
        lines.append("")
    lines.append("[alignments]")
    lines += [f"{k} = {json.dumps(str(v))}"
              for k, v in self.sequence_alignments.items()]
    for chart in self.charts:
        lines += ["", "[[charts]]",
                  f"anchors = {json.dumps(list(chart.anchors))}",
                  f"genes   = {json.dumps(list(chart.genes))}"]
    Path(path).write_text("\n".join(lines) + "\n")

ChartSpec dataclass

One chart to survey: three anchor labels and the genes to drop onto it.

genes may include labels that are not among anchors: a gene can be surveyed on a chart it is not a vertex of.

Source code in src/hifuku/survey_io.py
@dataclass
class ChartSpec:
    """One chart to survey: three anchor labels and the genes to drop onto it.

    ``genes`` may include labels that are not among ``anchors``: a gene can be
    surveyed on a chart it is not a vertex of.
    """

    anchors: tuple
    genes: list

    def __post_init__(self):
        self.anchors = tuple(self.anchors)
        self.genes = list(self.genes)
        if len(self.anchors) != 3:
            raise ValueError(
                f"a chart needs three anchors, got {len(self.anchors)}: "
                f"{list(self.anchors)}")
        if len(set(self.anchors)) != 3:
            raise ValueError(
                f"a chart needs three distinct anchors, got {list(self.anchors)}")

SurveyRun

One prepared survey: the shared context, the work units, and the results.

:meth:prepare builds the namespace, the NJ tree of every alignment and the AnchorTriangle of every chart before any survey runs, so a bad triangle on the fourth chart raises in seconds.

prepare is a function of the plan. A worker on another host rebuilds the same context from the same paths, which AlignmentInfo records with a SHA-256 digest.

plan = SurveyPlan.from_toml("ncldv.toml")
run  = SurveyRun.prepare(plan, device="cuda")
for task in run.tasks:
    run.execute(task)
run.survey().save("ncldv.h5")
Source code in src/hifuku/survey_io.py
class SurveyRun:
    """One prepared survey: the shared context, the work units, and the results.

    :meth:`prepare` builds the namespace, the NJ tree of every alignment and the
    ``AnchorTriangle`` of every chart before any survey runs, so a bad triangle
    on the fourth chart raises in seconds.

    ``prepare`` is a function of the plan.  A worker on another host rebuilds
    the same context from the same paths, which ``AlignmentInfo`` records with a
    SHA-256 digest.

        plan = SurveyPlan.from_toml("ncldv.toml")
        run  = SurveyRun.prepare(plan, device="cuda")
        for task in run.tasks:
            run.execute(task)
        run.survey().save("ncldv.h5")
    """

    def __init__(self, plan, device, backend, taxon_table, alignments,
                 alns, njs, model, model_label, triangles, params):
        self.plan = plan
        self.device = device
        self.backend = backend
        self.taxon_table = taxon_table
        self.alignments = alignments        # list[AlignmentInfo]
        self.params = params
        self._alns = alns
        self._njs = njs
        self._model = model
        self._model_label = model_label
        self._triangles = triangles         # list[AnchorTriangle], one per chart
        self._records = {}                  # (chart, gene) -> SurveyRecord
        self.tasks = [SurveyTask(chart=ci, gene=gene,
                                 digest=self._digest(ci, gene))
                      for ci, spec in enumerate(plan.charts)
                      for gene in spec.genes]

    @classmethod
    def prepare(cls, plan, *, device="gpu", backend=None) -> "SurveyRun":
        """Build the shared context and validate every chart.

        ``backend`` defaults to the host or device backend that ``device``
        names.  Inject one to run the orchestration against another engine.
        """
        from hifuku.alignment import load_alignment, peek_seq_type
        from hifuku.anchor_triangle import AnchorTriangle
        from hifuku.cli import _resolve_model
        from hifuku.nj import nj_tree
        from hifuku.params import ArchiveParams, GpuParams
        from hifuku.taxa import TaxonTable

        if backend is None:
            backend = _resolve_backend(device, plan.gpu_params)
        elif plan.gpu_params:
            GpuParams.resolve(plan.gpu_params)

        tt = TaxonTable()
        labels = list(plan.sequence_alignments)
        paths = {l: plan.sequence_alignments[l] for l in labels}
        alns = {l: load_alignment(str(paths[l]), tt) for l in labels}
        model, model_label = _resolve_model(
            plan.model, peek_seq_type(str(paths[labels[0]])))
        njs = {l: nj_tree(alns[l], tt) for l in labels}
        alignments = [AlignmentInfo(l, njs[l], str(paths[l]), _sha256(paths[l]))
                      for l in labels]

        # Every triangle before any survey, so a bad chart reports at once.
        triangles = []
        for ci, spec in enumerate(plan.charts):
            r = [njs[a] for a in spec.anchors]
            try:
                triangles.append(
                    AnchorTriangle.from_anchors(r[0], r[1], r[2], w=plan.w))
            except ValueError as exc:
                raise ValueError(
                    f"chart {ci} with anchors {list(spec.anchors)} is not a "
                    f"usable chart:\n  {exc}") from exc

        params = ArchiveParams.resolve(backend.caps.defaults, plan.params)
        return cls(plan, device, backend, tt, alignments, alns, njs, model,
                   model_label, triangles, params)

    def _digest(self, chart: int, gene: str) -> str:
        """SHA-256 over everything that decides the result of one work unit."""
        spec = self.plan.charts[chart]
        by_label = {a.label: a for a in self.alignments}
        payload = {
            "anchors": [[a, by_label[a].alignment_sha256] for a in spec.anchors],
            "gene": [gene, by_label[gene].alignment_sha256],
            "model": self._model_label,
            "w": self.plan.w,
            "params": {k: getattr(self.params, k) for k in
                       sorted(f.name for f in fields(self.params))},
        }
        blob = json.dumps(payload, sort_keys=True, separators=(",", ":"))
        return hashlib.sha256(blob.encode("utf-8")).hexdigest()

    def context(self, task: SurveyTask):
        """The :class:`~hifuku.backend.RunContext` for one work unit."""
        from hifuku.backend import RunContext
        spec = self.plan.charts[task.chart]
        r = [self._njs[a] for a in spec.anchors]
        return RunContext(anchors=(r[0], r[1], r[2]),
                          alignment=self._alns[task.gene], model=self._model,
                          start_trees=[self._njs[task.gene]],
                          triangle=self._triangles[task.chart],
                          taxon_table=self.taxon_table)

    def execute(self, task: SurveyTask, *, monitor=False) -> SurveyRecord:
        """Survey one gene on one chart and keep the record.

        ``monitor`` reaches :func:`~hifuku.driver.run_survey`, which calls it at
        every checkpoint.  A monitor that names ``task`` receives this task.
        """
        from hifuku.driver import run_survey
        result = run_survey(self.backend, self.context(task), self.params,
                            monitor=monitor, task=task)
        record = SurveyRecord(
            gene=task.gene, archive=result.archive, history=result.history,
            converged=result.converged, halt_iter=result.halt_iter,
            model=self._model_label, n_iters=self.params.n_iters,
            seed=self.params.seed, device=self.device)
        self._records[(task.chart, task.gene)] = record
        return record

    def survey(self, *, description="") -> Survey:
        """Assemble the executed records into a :class:`Survey`.

        A chart holds the genes that ran on it, in the order the plan names
        them.  A gene that no call executed is left out.
        """
        charts = []
        for ci, spec in enumerate(self.plan.charts):
            surveys = [self._records[(ci, g)] for g in spec.genes
                       if (ci, g) in self._records]
            charts.append(Chart(triangle=self._triangles[ci],
                                anchors=spec.anchors, surveys=surveys))
        return Survey(taxon_table=self.taxon_table, alignments=self.alignments,
                      charts=charts,
                      description=description or self.plan.description,
                      metric_w=self.plan.w)

prepare classmethod

prepare(plan, *, device='gpu', backend=None) -> 'SurveyRun'

Build the shared context and validate every chart.

backend defaults to the host or device backend that device names. Inject one to run the orchestration against another engine.

Source code in src/hifuku/survey_io.py
@classmethod
def prepare(cls, plan, *, device="gpu", backend=None) -> "SurveyRun":
    """Build the shared context and validate every chart.

    ``backend`` defaults to the host or device backend that ``device``
    names.  Inject one to run the orchestration against another engine.
    """
    from hifuku.alignment import load_alignment, peek_seq_type
    from hifuku.anchor_triangle import AnchorTriangle
    from hifuku.cli import _resolve_model
    from hifuku.nj import nj_tree
    from hifuku.params import ArchiveParams, GpuParams
    from hifuku.taxa import TaxonTable

    if backend is None:
        backend = _resolve_backend(device, plan.gpu_params)
    elif plan.gpu_params:
        GpuParams.resolve(plan.gpu_params)

    tt = TaxonTable()
    labels = list(plan.sequence_alignments)
    paths = {l: plan.sequence_alignments[l] for l in labels}
    alns = {l: load_alignment(str(paths[l]), tt) for l in labels}
    model, model_label = _resolve_model(
        plan.model, peek_seq_type(str(paths[labels[0]])))
    njs = {l: nj_tree(alns[l], tt) for l in labels}
    alignments = [AlignmentInfo(l, njs[l], str(paths[l]), _sha256(paths[l]))
                  for l in labels]

    # Every triangle before any survey, so a bad chart reports at once.
    triangles = []
    for ci, spec in enumerate(plan.charts):
        r = [njs[a] for a in spec.anchors]
        try:
            triangles.append(
                AnchorTriangle.from_anchors(r[0], r[1], r[2], w=plan.w))
        except ValueError as exc:
            raise ValueError(
                f"chart {ci} with anchors {list(spec.anchors)} is not a "
                f"usable chart:\n  {exc}") from exc

    params = ArchiveParams.resolve(backend.caps.defaults, plan.params)
    return cls(plan, device, backend, tt, alignments, alns, njs, model,
               model_label, triangles, params)

context

context(task: SurveyTask)

The :class:~hifuku.backend.RunContext for one work unit.

Source code in src/hifuku/survey_io.py
def context(self, task: SurveyTask):
    """The :class:`~hifuku.backend.RunContext` for one work unit."""
    from hifuku.backend import RunContext
    spec = self.plan.charts[task.chart]
    r = [self._njs[a] for a in spec.anchors]
    return RunContext(anchors=(r[0], r[1], r[2]),
                      alignment=self._alns[task.gene], model=self._model,
                      start_trees=[self._njs[task.gene]],
                      triangle=self._triangles[task.chart],
                      taxon_table=self.taxon_table)

execute

execute(task: SurveyTask, *, monitor=False) -> SurveyRecord

Survey one gene on one chart and keep the record.

monitor reaches :func:~hifuku.driver.run_survey, which calls it at every checkpoint. A monitor that names task receives this task.

Source code in src/hifuku/survey_io.py
def execute(self, task: SurveyTask, *, monitor=False) -> SurveyRecord:
    """Survey one gene on one chart and keep the record.

    ``monitor`` reaches :func:`~hifuku.driver.run_survey`, which calls it at
    every checkpoint.  A monitor that names ``task`` receives this task.
    """
    from hifuku.driver import run_survey
    result = run_survey(self.backend, self.context(task), self.params,
                        monitor=monitor, task=task)
    record = SurveyRecord(
        gene=task.gene, archive=result.archive, history=result.history,
        converged=result.converged, halt_iter=result.halt_iter,
        model=self._model_label, n_iters=self.params.n_iters,
        seed=self.params.seed, device=self.device)
    self._records[(task.chart, task.gene)] = record
    return record

survey

survey(*, description='') -> Survey

Assemble the executed records into a :class:Survey.

A chart holds the genes that ran on it, in the order the plan names them. A gene that no call executed is left out.

Source code in src/hifuku/survey_io.py
def survey(self, *, description="") -> Survey:
    """Assemble the executed records into a :class:`Survey`.

    A chart holds the genes that ran on it, in the order the plan names
    them.  A gene that no call executed is left out.
    """
    charts = []
    for ci, spec in enumerate(self.plan.charts):
        surveys = [self._records[(ci, g)] for g in spec.genes
                   if (ci, g) in self._records]
        charts.append(Chart(triangle=self._triangles[ci],
                            anchors=spec.anchors, surveys=surveys))
    return Survey(taxon_table=self.taxon_table, alignments=self.alignments,
                  charts=charts,
                  description=description or self.plan.description,
                  metric_w=self.plan.w)

SurveyTask dataclass

One gene surveyed on one chart: the unit of work.

The task holds a chart index, a gene label and a digest. Every field is a string or an integer, so the task round-trips through JSON and travels to another host as data.

digest covers the anchor labels, the gene, its alignment SHA-256, the model, the metric weight and every run parameter. Two runs that would produce the same result therefore carry the same key, and a caller resumes with if task.key() in done: continue.

Source code in src/hifuku/survey_io.py
@dataclass(frozen=True)
class SurveyTask:
    """One gene surveyed on one chart: the unit of work.

    The task holds a chart index, a gene label and a digest.  Every field is a
    string or an integer, so the task round-trips through JSON and travels to
    another host as data.

    ``digest`` covers the anchor labels, the gene, its alignment SHA-256, the
    model, the metric weight and every run parameter.  Two runs that would
    produce the same result therefore carry the same key, and a caller resumes
    with ``if task.key() in done: continue``.
    """

    chart: int
    gene: str
    digest: str

    def key(self) -> str:
        """The content digest of this unit of work."""
        return self.digest

    def as_dict(self) -> dict:
        """A plain mapping, ready for JSON."""
        return {"chart": self.chart, "gene": self.gene, "digest": self.digest}

key

key() -> str

The content digest of this unit of work.

Source code in src/hifuku/survey_io.py
def key(self) -> str:
    """The content digest of this unit of work."""
    return self.digest

as_dict

as_dict() -> dict

A plain mapping, ready for JSON.

Source code in src/hifuku/survey_io.py
def as_dict(self) -> dict:
    """A plain mapping, ready for JSON."""
    return {"chart": self.chart, "gene": self.gene, "digest": self.digest}

Survey

A survey: a shared taxon namespace, an alignment pool, and charts.

Construct one directly, or with :meth:run. save and load are a plain, predictable round trip; they always do the same thing and hold no resume or caching logic (that lives in the caller).

Source code in src/hifuku/survey_io.py
class Survey:
    """A survey: a shared taxon namespace, an alignment pool, and charts.

    Construct one directly, or with :meth:`run`.  ``save`` and ``load`` are a
    plain, predictable round trip; they always do the same thing and hold no
    resume or caching logic (that lives in the caller).
    """

    def __init__(self, taxon_table, alignments, charts, *, description="",
                 metric_w=0.5, provenance=None, ordinations=None):
        self.taxon_table = taxon_table
        self.alignments = list(alignments)     # list[AlignmentInfo], the pool
        self.charts = list(charts)             # list[Chart]
        self.description = description
        self.metric_w = metric_w
        self.provenance = provenance or {}
        self.ordinations = list(ordinations or [])   # list[Ordination]
        self._by_label = {a.label: a for a in self.alignments}

    @property
    def n_taxa(self):
        return len(self.taxon_table)

    @classmethod
    def run(cls, plan, *, device="gpu", backend=None, description="",
            monitor=False) -> "Survey":
        """Run every (chart, gene) survey in ``plan`` and return the Survey.

        This is :class:`SurveyRun` with the loop written out.  Call
        ``SurveyRun`` directly to watch a run, to skip finished work, or to send
        a task to another host.
        """
        run = SurveyRun.prepare(plan, device=device, backend=backend)
        for task in run.tasks:
            run.execute(task, monitor=monitor)
        return run.survey(description=description)

    @property
    def metric(self):
        """The cloud-wide tree metric of this survey.

        Built from the NJ tree of every alignment in the pool, so analysis over
        the whole cloud uses one normalization.  A chart carries its own
        ``CBSMetric``, whose ``c_ref`` and ``r_ref`` come from the anchors of
        that chart alone, so this is not the metric of any one chart.
        """
        from hifuku.anchor_triangle import CBSMetric
        return CBSMetric.for_anchors([a.nj_tree for a in self.alignments],
                                     w=self.metric_w)

    def alignment(self, label):
        """AlignmentInfo for a pool label."""
        return self._by_label[label]

    def archive(self, gene, chart=0):
        """EliteArchive for a surveyed gene on a chart (index or Chart)."""
        c = chart if isinstance(chart, Chart) else self.charts[chart]
        return c.archive(gene)

    def frame(self, backend="pandas"):
        """Tidy frame of every elite niche across all charts and surveys.

        Columns: ``gene, chart, i, j, lam1, lam2, logL, z``.
        """
        rows = []
        for ci, chart in enumerate(self.charts):
            for rec in chart.surveys:
                arch = rec.archive
                for key in arch.keys:
                    lam1, lam2 = arch.niche_center(key)
                    e = arch.elites[key]
                    rows.append((rec.gene, ci, key[0], key[1],
                                 lam1, lam2, e.logL, e.z))
        cols = ["gene", "chart", "i", "j", "lam1", "lam2", "logL", "z"]
        if backend == "polars":
            import polars as pl
            return pl.DataFrame(rows, schema=cols, orient="row")
        import pandas as pd
        return pd.DataFrame(rows, columns=cols)

    def cloud(self):
        """Union of every elite tree across all charts, each with ``(gene, chart)``
        provenance.  Returns a list of :class:`CloudPoint`."""
        pts = []
        for ci, chart in enumerate(self.charts):
            for rec in chart.surveys:
                arch = rec.archive
                for key in arch.keys:
                    e = arch.elites[key]
                    pts.append(CloudPoint(tree=arch.tree_at(key), gene=rec.gene,
                                          chart=ci, key=key, logL=e.logL, z=e.z))
        return pts

    def cloud_identity(self) -> list:
        """The ``(gene, chart, i, j)`` of every cloud point, in cloud order.

        This is what an embedding's rows line up with.  :meth:`cloud` builds the
        points in the same order, charts then surveys then niches, so the two
        agree by construction.
        """
        return [(rec.gene, ci, int(key[0]), int(key[1]))
                for ci, chart in enumerate(self.charts)
                for rec in chart.surveys
                for key in rec.archive.keys]

    def cloud_digest(self) -> str:
        """SHA-256 over :meth:`cloud_identity`.

        Two runs of one plan give one digest, because the digest covers which
        niches were filled and not what landed in them.  Adding a chart changes
        it, which is the case that would otherwise misalign a stored embedding.
        """
        blob = json.dumps(self.cloud_identity(), separators=(",", ":"))
        return hashlib.sha256(blob.encode("utf-8")).hexdigest()

    def ordination(self, label: str):
        """The stored embedding with this label."""
        for o in self.ordinations:
            if o.label == label:
                return o
        raise KeyError(
            f"no ordination labelled {label!r}; this survey holds "
            f"{[o.label for o in self.ordinations]}")

    def add_ordination(self, X, *, label: str, method: str, params=None,
                       faithfulness=None):
        """Keep an embedding of this survey's cloud.

        Fills in the point count and the cloud digest, so a later load can tell
        whether the survey still matches the embedding.

        Parameters
        ----------
        X : numpy.ndarray
            ``(n, dim)`` coordinates, one row per cloud point in cloud order.
        label : str
            A name, unique within this survey.
        method : str
            How the embedding was produced.
        params : dict, optional
            What the method was given.
        faithfulness : Faithfulness, optional
            The score of this embedding.

        Returns
        -------
        Ordination
        """
        X = np.asarray(X, dtype=np.float64)
        n = len(self.cloud_identity())
        if X.ndim != 2 or X.shape[0] != n:
            raise ValueError(
                f"the embedding has {X.shape[0]} rows and the cloud holds {n} "
                "points; they line up by position")
        if any(o.label == label for o in self.ordinations):
            raise ValueError(f"this survey already holds an ordination "
                             f"labelled {label!r}")
        record = Ordination(
            label=label, method=method, params=dict(params or {}), X=X,
            faithfulness=faithfulness,
            created=datetime.datetime.now(datetime.timezone.utc).isoformat(),
            n_points=n, cloud_digest=self.cloud_digest())
        self.ordinations.append(record)
        return record

    def save(self, path, *, device="cpu", author=None) -> None:
        """Write the survey to a schema-v2 HDF5 file."""
        from hifuku.anchor_triangle import triangle_quality

        label_to_idx = {a.label: i for i, a in enumerate(self.alignments)}
        with h5py.File(path, "w") as f:
            f.attrs["schema_version"] = _SCHEMA_VERSION
            f.attrs["field_type"] = _FIELD_TYPE
            f.attrs["metric"] = "normalized_branch_score"
            f.attrs["metric_w"] = float(self.metric_w)
            f.attrs["n_alignments"] = len(self.alignments)
            f.attrs["n_charts"] = len(self.charts)
            f.attrs["created"] = datetime.datetime.now(
                datetime.timezone.utc).isoformat()
            f.attrs["description"] = self.description

            f.create_dataset("namespace/taxa",
                             data=np.array(self.taxon_table.labels, dtype=object),
                             dtype=_STR_DTYPE)

            for i, a in enumerate(self.alignments):
                g = f.create_group(f"sequence_alignments/{i}")
                g.attrs["label"] = a.label
                g.attrs["alignment_path"] = str(a.alignment_path or "")
                g.attrs["alignment_sha256"] = a.alignment_sha256 or ""
                g.create_dataset("nj_newick", data=a.nj_tree.as_newick(),
                                 dtype=_STR_DTYPE)

            for ci, chart in enumerate(self.charts):
                tri = chart.triangle
                m = tri.tree_metric
                c = f.create_group(f"charts/{ci}")
                c.attrs["anchor_alignment_indices"] = np.array(
                    [label_to_idx[l] for l in chart.anchors], dtype=np.int32)
                c.attrs["metric_w"] = float(m.w)
                c.attrs["c_ref"] = float(m.c_ref)
                c.attrs["r_ref"] = float(m.r_ref)
                arch0 = chart.surveys[0].archive
                c.attrs["cell"] = float(arch0.cell)
                c.attrs["origin"] = float(arch0.origin)
                Q, _, _ = triangle_quality(tri.P1, tri.P2, tri.P3)
                c.attrs["Q"] = float(Q)
                c.create_dataset("P", data=np.stack(
                    [tri.P1, tri.P2, tri.P3]).astype(np.float64))
                c.create_dataset("dist", data=tri.dist.astype(np.float64))
                c.create_dataset("qual", data=tri.qual.astype(np.float64))

                for si, rec in enumerate(chart.surveys):
                    arch = rec.archive
                    keys = list(arch.keys)
                    s = c.create_group(f"surveys/{si}")
                    s.attrs["alignment_index"] = label_to_idx[rec.gene]
                    s.attrs["gene"] = rec.gene
                    s.attrs["n_filled"] = arch.n_filled
                    s.attrs["best_logL"] = arch.best_logL()
                    s.attrs["converged"] = int(rec.converged)
                    s.attrs["halt_iter"] = int(rec.halt_iter)
                    s.attrs["device"] = rec.device
                    s.attrs["model"] = rec.model
                    s.attrs["n_iters"] = int(rec.n_iters)
                    s.attrs["seed"] = int(rec.seed)
                    niche = np.array(keys, dtype=np.int32).reshape(-1, 2)
                    logL = np.array([arch.elites[k].logL for k in keys],
                                    dtype=np.float64)
                    z = np.array([arch.elites[k].z for k in keys],
                                 dtype=np.float64)
                    newick = np.array([arch.tree_at(k).as_newick() for k in keys],
                                      dtype=object)
                    s.create_dataset("niche_ij", data=niche)
                    s.create_dataset("logL", data=logL)
                    s.create_dataset("z", data=z)
                    s.create_dataset("tree", data=newick, dtype=_STR_DTYPE,
                                     compression="gzip")
                    s.create_dataset("history",
                                     data=rec.history.astype(np.float64))

            group = f.create_group("ordinations")
            for oi, o in enumerate(self.ordinations):
                g = group.create_group(str(oi))
                g.attrs["label"] = o.label
                g.attrs["method"] = o.method
                g.attrs["created"] = o.created
                g.attrs["n_points"] = int(o.n_points)
                g.attrs["cloud_digest"] = o.cloud_digest
                g.attrs["params"] = json.dumps(o.params, sort_keys=True)
                g.create_dataset("X", data=np.asarray(o.X, dtype=np.float64))
                if o.faithfulness is not None:
                    fg = g.create_group("faithfulness")
                    fg.attrs["k"] = int(o.faithfulness.k)
                    fg.attrs["trustworthiness"] = float(o.faithfulness.trustworthiness)
                    fg.attrs["continuity"] = float(o.faithfulness.continuity)
                    fg.attrs["ideal_rank"] = float(o.faithfulness.ideal_rank)
                    fg.attrs["false_neighbor_rate"] = float(
                        o.faithfulness.false_neighbor_rate)
                    fg.create_dataset("neighbor_rank", data=np.asarray(
                        o.faithfulness.neighbor_rank, dtype=np.float64))

            prov = f.create_group("provenance")
            for key, val in _provenance(device, author).items():
                prov.attrs[key] = val

    @classmethod
    def load(cls, path) -> "Survey":
        """Reconstruct a Survey from a schema-v2 file (inverse of :meth:`save`)."""
        from hifuku.taxa import TaxonTable
        from hifuku.tree_io import read_tree
        from hifuku.anchor_triangle import AnchorTriangle, CBSMetric
        from hifuku.archive import Elite, EliteArchive, _set_global_leaf_indices

        with h5py.File(path, "r") as f:
            version = int(f.attrs.get("schema_version", 0))
            if version not in _READABLE_SCHEMA_VERSIONS:
                raise ValueError(
                    f"unsupported schema_version {version} in {path}; this "
                    f"build reads {list(_READABLE_SCHEMA_VERSIONS)}")
            description = _s(f.attrs.get("description", ""))
            metric_w = float(f.attrs["metric_w"])
            provenance = {k: _s(v) for k, v in f["provenance"].attrs.items()}

            tt = TaxonTable()
            for label in f["namespace/taxa"][:]:
                tt.require(_s(label))

            alignments = []
            for i in range(int(f.attrs["n_alignments"])):
                g = f[f"sequence_alignments/{i}"]
                nj = read_tree(_s(g["nj_newick"][()]), tt, from_string=True)
                _set_global_leaf_indices(nj, tt)
                alignments.append(AlignmentInfo(
                    _s(g.attrs["label"]), nj,
                    _s(g.attrs["alignment_path"]),
                    _s(g.attrs["alignment_sha256"])))

            charts = []
            for ci in range(int(f.attrs["n_charts"])):
                c = f[f"charts/{ci}"]
                aidx = np.asarray(c.attrs["anchor_alignment_indices"])
                anchors = tuple(alignments[int(j)].label for j in aidx)
                r = [alignments[int(j)].nj_tree for j in aidx]
                metric = CBSMetric(w=float(c.attrs["metric_w"]),
                                   c_ref=float(c.attrs["c_ref"]),
                                   r_ref=float(c.attrs["r_ref"]))
                P = c["P"][:]
                tri = AnchorTriangle(P1=P[0], P2=P[1], P3=P[2],
                                     r1=r[0], r2=r[1], r3=r[2],
                                     dist=c["dist"][:], qual=c["qual"][:],
                                     metric=metric)
                cell = float(c.attrs["cell"])
                origin = float(c.attrs["origin"])
                surveys = []
                sgroup = c["surveys"]
                for si in range(len(sgroup)):
                    s = sgroup[str(si)]
                    niche = s["niche_ij"][:]
                    logL = s["logL"][:]
                    z = s["z"][:]
                    newick = s["tree"][:]
                    src = {}
                    arch = EliteArchive(cell=cell, origin=origin, anchors=tri,
                                        taxon_table=tt, metric=metric,
                                        tree_source=src)
                    for row in range(niche.shape[0]):
                        key = (int(niche[row, 0]), int(niche[row, 1]))
                        arch.elites[key] = Elite(float(logL[row]), None,
                                                 float(z[row]))
                        arch.keys.append(key)
                        src[key] = _s(newick[row])
                    surveys.append(SurveyRecord(
                        gene=_s(s.attrs["gene"]), archive=arch,
                        history=s["history"][:],
                        converged=bool(s.attrs["converged"]),
                        halt_iter=int(s.attrs["halt_iter"]),
                        model=_s(s.attrs["model"]),
                        n_iters=int(s.attrs["n_iters"]),
                        seed=int(s.attrs["seed"]),
                        device=_s(s.attrs["device"])))
                charts.append(Chart(triangle=tri, anchors=anchors,
                                    surveys=surveys))

        survey = cls(taxon_table=tt, alignments=alignments, charts=charts,
                     description=description, metric_w=metric_w,
                     provenance=provenance)
        survey.ordinations = _read_ordinations(path, survey)
        return survey

metric property

metric

The cloud-wide tree metric of this survey.

Built from the NJ tree of every alignment in the pool, so analysis over the whole cloud uses one normalization. A chart carries its own CBSMetric, whose c_ref and r_ref come from the anchors of that chart alone, so this is not the metric of any one chart.

run classmethod

run(plan, *, device='gpu', backend=None, description='', monitor=False) -> 'Survey'

Run every (chart, gene) survey in plan and return the Survey.

This is :class:SurveyRun with the loop written out. Call SurveyRun directly to watch a run, to skip finished work, or to send a task to another host.

Source code in src/hifuku/survey_io.py
@classmethod
def run(cls, plan, *, device="gpu", backend=None, description="",
        monitor=False) -> "Survey":
    """Run every (chart, gene) survey in ``plan`` and return the Survey.

    This is :class:`SurveyRun` with the loop written out.  Call
    ``SurveyRun`` directly to watch a run, to skip finished work, or to send
    a task to another host.
    """
    run = SurveyRun.prepare(plan, device=device, backend=backend)
    for task in run.tasks:
        run.execute(task, monitor=monitor)
    return run.survey(description=description)

alignment

alignment(label)

AlignmentInfo for a pool label.

Source code in src/hifuku/survey_io.py
def alignment(self, label):
    """AlignmentInfo for a pool label."""
    return self._by_label[label]

archive

archive(gene, chart=0)

EliteArchive for a surveyed gene on a chart (index or Chart).

Source code in src/hifuku/survey_io.py
def archive(self, gene, chart=0):
    """EliteArchive for a surveyed gene on a chart (index or Chart)."""
    c = chart if isinstance(chart, Chart) else self.charts[chart]
    return c.archive(gene)

frame

frame(backend='pandas')

Tidy frame of every elite niche across all charts and surveys.

Columns: gene, chart, i, j, lam1, lam2, logL, z.

Source code in src/hifuku/survey_io.py
def frame(self, backend="pandas"):
    """Tidy frame of every elite niche across all charts and surveys.

    Columns: ``gene, chart, i, j, lam1, lam2, logL, z``.
    """
    rows = []
    for ci, chart in enumerate(self.charts):
        for rec in chart.surveys:
            arch = rec.archive
            for key in arch.keys:
                lam1, lam2 = arch.niche_center(key)
                e = arch.elites[key]
                rows.append((rec.gene, ci, key[0], key[1],
                             lam1, lam2, e.logL, e.z))
    cols = ["gene", "chart", "i", "j", "lam1", "lam2", "logL", "z"]
    if backend == "polars":
        import polars as pl
        return pl.DataFrame(rows, schema=cols, orient="row")
    import pandas as pd
    return pd.DataFrame(rows, columns=cols)

cloud

cloud()

Union of every elite tree across all charts, each with (gene, chart) provenance. Returns a list of :class:CloudPoint.

Source code in src/hifuku/survey_io.py
def cloud(self):
    """Union of every elite tree across all charts, each with ``(gene, chart)``
    provenance.  Returns a list of :class:`CloudPoint`."""
    pts = []
    for ci, chart in enumerate(self.charts):
        for rec in chart.surveys:
            arch = rec.archive
            for key in arch.keys:
                e = arch.elites[key]
                pts.append(CloudPoint(tree=arch.tree_at(key), gene=rec.gene,
                                      chart=ci, key=key, logL=e.logL, z=e.z))
    return pts

cloud_identity

cloud_identity() -> list

The (gene, chart, i, j) of every cloud point, in cloud order.

This is what an embedding's rows line up with. :meth:cloud builds the points in the same order, charts then surveys then niches, so the two agree by construction.

Source code in src/hifuku/survey_io.py
def cloud_identity(self) -> list:
    """The ``(gene, chart, i, j)`` of every cloud point, in cloud order.

    This is what an embedding's rows line up with.  :meth:`cloud` builds the
    points in the same order, charts then surveys then niches, so the two
    agree by construction.
    """
    return [(rec.gene, ci, int(key[0]), int(key[1]))
            for ci, chart in enumerate(self.charts)
            for rec in chart.surveys
            for key in rec.archive.keys]

cloud_digest

cloud_digest() -> str

SHA-256 over :meth:cloud_identity.

Two runs of one plan give one digest, because the digest covers which niches were filled and not what landed in them. Adding a chart changes it, which is the case that would otherwise misalign a stored embedding.

Source code in src/hifuku/survey_io.py
def cloud_digest(self) -> str:
    """SHA-256 over :meth:`cloud_identity`.

    Two runs of one plan give one digest, because the digest covers which
    niches were filled and not what landed in them.  Adding a chart changes
    it, which is the case that would otherwise misalign a stored embedding.
    """
    blob = json.dumps(self.cloud_identity(), separators=(",", ":"))
    return hashlib.sha256(blob.encode("utf-8")).hexdigest()

ordination

ordination(label: str)

The stored embedding with this label.

Source code in src/hifuku/survey_io.py
def ordination(self, label: str):
    """The stored embedding with this label."""
    for o in self.ordinations:
        if o.label == label:
            return o
    raise KeyError(
        f"no ordination labelled {label!r}; this survey holds "
        f"{[o.label for o in self.ordinations]}")

add_ordination

add_ordination(X, *, label: str, method: str, params=None, faithfulness=None)

Keep an embedding of this survey's cloud.

Fills in the point count and the cloud digest, so a later load can tell whether the survey still matches the embedding.

Parameters:

Name Type Description Default
X ndarray

(n, dim) coordinates, one row per cloud point in cloud order.

required
label str

A name, unique within this survey.

required
method str

How the embedding was produced.

required
params dict

What the method was given.

None
faithfulness Faithfulness

The score of this embedding.

None

Returns:

Type Description
Ordination
Source code in src/hifuku/survey_io.py
def add_ordination(self, X, *, label: str, method: str, params=None,
                   faithfulness=None):
    """Keep an embedding of this survey's cloud.

    Fills in the point count and the cloud digest, so a later load can tell
    whether the survey still matches the embedding.

    Parameters
    ----------
    X : numpy.ndarray
        ``(n, dim)`` coordinates, one row per cloud point in cloud order.
    label : str
        A name, unique within this survey.
    method : str
        How the embedding was produced.
    params : dict, optional
        What the method was given.
    faithfulness : Faithfulness, optional
        The score of this embedding.

    Returns
    -------
    Ordination
    """
    X = np.asarray(X, dtype=np.float64)
    n = len(self.cloud_identity())
    if X.ndim != 2 or X.shape[0] != n:
        raise ValueError(
            f"the embedding has {X.shape[0]} rows and the cloud holds {n} "
            "points; they line up by position")
    if any(o.label == label for o in self.ordinations):
        raise ValueError(f"this survey already holds an ordination "
                         f"labelled {label!r}")
    record = Ordination(
        label=label, method=method, params=dict(params or {}), X=X,
        faithfulness=faithfulness,
        created=datetime.datetime.now(datetime.timezone.utc).isoformat(),
        n_points=n, cloud_digest=self.cloud_digest())
    self.ordinations.append(record)
    return record

save

save(path, *, device='cpu', author=None) -> None

Write the survey to a schema-v2 HDF5 file.

Source code in src/hifuku/survey_io.py
def save(self, path, *, device="cpu", author=None) -> None:
    """Write the survey to a schema-v2 HDF5 file."""
    from hifuku.anchor_triangle import triangle_quality

    label_to_idx = {a.label: i for i, a in enumerate(self.alignments)}
    with h5py.File(path, "w") as f:
        f.attrs["schema_version"] = _SCHEMA_VERSION
        f.attrs["field_type"] = _FIELD_TYPE
        f.attrs["metric"] = "normalized_branch_score"
        f.attrs["metric_w"] = float(self.metric_w)
        f.attrs["n_alignments"] = len(self.alignments)
        f.attrs["n_charts"] = len(self.charts)
        f.attrs["created"] = datetime.datetime.now(
            datetime.timezone.utc).isoformat()
        f.attrs["description"] = self.description

        f.create_dataset("namespace/taxa",
                         data=np.array(self.taxon_table.labels, dtype=object),
                         dtype=_STR_DTYPE)

        for i, a in enumerate(self.alignments):
            g = f.create_group(f"sequence_alignments/{i}")
            g.attrs["label"] = a.label
            g.attrs["alignment_path"] = str(a.alignment_path or "")
            g.attrs["alignment_sha256"] = a.alignment_sha256 or ""
            g.create_dataset("nj_newick", data=a.nj_tree.as_newick(),
                             dtype=_STR_DTYPE)

        for ci, chart in enumerate(self.charts):
            tri = chart.triangle
            m = tri.tree_metric
            c = f.create_group(f"charts/{ci}")
            c.attrs["anchor_alignment_indices"] = np.array(
                [label_to_idx[l] for l in chart.anchors], dtype=np.int32)
            c.attrs["metric_w"] = float(m.w)
            c.attrs["c_ref"] = float(m.c_ref)
            c.attrs["r_ref"] = float(m.r_ref)
            arch0 = chart.surveys[0].archive
            c.attrs["cell"] = float(arch0.cell)
            c.attrs["origin"] = float(arch0.origin)
            Q, _, _ = triangle_quality(tri.P1, tri.P2, tri.P3)
            c.attrs["Q"] = float(Q)
            c.create_dataset("P", data=np.stack(
                [tri.P1, tri.P2, tri.P3]).astype(np.float64))
            c.create_dataset("dist", data=tri.dist.astype(np.float64))
            c.create_dataset("qual", data=tri.qual.astype(np.float64))

            for si, rec in enumerate(chart.surveys):
                arch = rec.archive
                keys = list(arch.keys)
                s = c.create_group(f"surveys/{si}")
                s.attrs["alignment_index"] = label_to_idx[rec.gene]
                s.attrs["gene"] = rec.gene
                s.attrs["n_filled"] = arch.n_filled
                s.attrs["best_logL"] = arch.best_logL()
                s.attrs["converged"] = int(rec.converged)
                s.attrs["halt_iter"] = int(rec.halt_iter)
                s.attrs["device"] = rec.device
                s.attrs["model"] = rec.model
                s.attrs["n_iters"] = int(rec.n_iters)
                s.attrs["seed"] = int(rec.seed)
                niche = np.array(keys, dtype=np.int32).reshape(-1, 2)
                logL = np.array([arch.elites[k].logL for k in keys],
                                dtype=np.float64)
                z = np.array([arch.elites[k].z for k in keys],
                             dtype=np.float64)
                newick = np.array([arch.tree_at(k).as_newick() for k in keys],
                                  dtype=object)
                s.create_dataset("niche_ij", data=niche)
                s.create_dataset("logL", data=logL)
                s.create_dataset("z", data=z)
                s.create_dataset("tree", data=newick, dtype=_STR_DTYPE,
                                 compression="gzip")
                s.create_dataset("history",
                                 data=rec.history.astype(np.float64))

        group = f.create_group("ordinations")
        for oi, o in enumerate(self.ordinations):
            g = group.create_group(str(oi))
            g.attrs["label"] = o.label
            g.attrs["method"] = o.method
            g.attrs["created"] = o.created
            g.attrs["n_points"] = int(o.n_points)
            g.attrs["cloud_digest"] = o.cloud_digest
            g.attrs["params"] = json.dumps(o.params, sort_keys=True)
            g.create_dataset("X", data=np.asarray(o.X, dtype=np.float64))
            if o.faithfulness is not None:
                fg = g.create_group("faithfulness")
                fg.attrs["k"] = int(o.faithfulness.k)
                fg.attrs["trustworthiness"] = float(o.faithfulness.trustworthiness)
                fg.attrs["continuity"] = float(o.faithfulness.continuity)
                fg.attrs["ideal_rank"] = float(o.faithfulness.ideal_rank)
                fg.attrs["false_neighbor_rate"] = float(
                    o.faithfulness.false_neighbor_rate)
                fg.create_dataset("neighbor_rank", data=np.asarray(
                    o.faithfulness.neighbor_rank, dtype=np.float64))

        prov = f.create_group("provenance")
        for key, val in _provenance(device, author).items():
            prov.attrs[key] = val

load classmethod

load(path) -> 'Survey'

Reconstruct a Survey from a schema-v2 file (inverse of :meth:save).

Source code in src/hifuku/survey_io.py
@classmethod
def load(cls, path) -> "Survey":
    """Reconstruct a Survey from a schema-v2 file (inverse of :meth:`save`)."""
    from hifuku.taxa import TaxonTable
    from hifuku.tree_io import read_tree
    from hifuku.anchor_triangle import AnchorTriangle, CBSMetric
    from hifuku.archive import Elite, EliteArchive, _set_global_leaf_indices

    with h5py.File(path, "r") as f:
        version = int(f.attrs.get("schema_version", 0))
        if version not in _READABLE_SCHEMA_VERSIONS:
            raise ValueError(
                f"unsupported schema_version {version} in {path}; this "
                f"build reads {list(_READABLE_SCHEMA_VERSIONS)}")
        description = _s(f.attrs.get("description", ""))
        metric_w = float(f.attrs["metric_w"])
        provenance = {k: _s(v) for k, v in f["provenance"].attrs.items()}

        tt = TaxonTable()
        for label in f["namespace/taxa"][:]:
            tt.require(_s(label))

        alignments = []
        for i in range(int(f.attrs["n_alignments"])):
            g = f[f"sequence_alignments/{i}"]
            nj = read_tree(_s(g["nj_newick"][()]), tt, from_string=True)
            _set_global_leaf_indices(nj, tt)
            alignments.append(AlignmentInfo(
                _s(g.attrs["label"]), nj,
                _s(g.attrs["alignment_path"]),
                _s(g.attrs["alignment_sha256"])))

        charts = []
        for ci in range(int(f.attrs["n_charts"])):
            c = f[f"charts/{ci}"]
            aidx = np.asarray(c.attrs["anchor_alignment_indices"])
            anchors = tuple(alignments[int(j)].label for j in aidx)
            r = [alignments[int(j)].nj_tree for j in aidx]
            metric = CBSMetric(w=float(c.attrs["metric_w"]),
                               c_ref=float(c.attrs["c_ref"]),
                               r_ref=float(c.attrs["r_ref"]))
            P = c["P"][:]
            tri = AnchorTriangle(P1=P[0], P2=P[1], P3=P[2],
                                 r1=r[0], r2=r[1], r3=r[2],
                                 dist=c["dist"][:], qual=c["qual"][:],
                                 metric=metric)
            cell = float(c.attrs["cell"])
            origin = float(c.attrs["origin"])
            surveys = []
            sgroup = c["surveys"]
            for si in range(len(sgroup)):
                s = sgroup[str(si)]
                niche = s["niche_ij"][:]
                logL = s["logL"][:]
                z = s["z"][:]
                newick = s["tree"][:]
                src = {}
                arch = EliteArchive(cell=cell, origin=origin, anchors=tri,
                                    taxon_table=tt, metric=metric,
                                    tree_source=src)
                for row in range(niche.shape[0]):
                    key = (int(niche[row, 0]), int(niche[row, 1]))
                    arch.elites[key] = Elite(float(logL[row]), None,
                                             float(z[row]))
                    arch.keys.append(key)
                    src[key] = _s(newick[row])
                surveys.append(SurveyRecord(
                    gene=_s(s.attrs["gene"]), archive=arch,
                    history=s["history"][:],
                    converged=bool(s.attrs["converged"]),
                    halt_iter=int(s.attrs["halt_iter"]),
                    model=_s(s.attrs["model"]),
                    n_iters=int(s.attrs["n_iters"]),
                    seed=int(s.attrs["seed"]),
                    device=_s(s.attrs["device"])))
            charts.append(Chart(triangle=tri, anchors=anchors,
                                surveys=surveys))

    survey = cls(taxon_table=tt, alignments=alignments, charts=charts,
                 description=description, metric_w=metric_w,
                 provenance=provenance)
    survey.ordinations = _read_ordinations(path, survey)
    return survey

Chart

One chart: its AnchorTriangle, the three anchor labels, and the surveys (one per surveyed gene) run on it.

Source code in src/hifuku/survey_io.py
class Chart:
    """One chart: its ``AnchorTriangle``, the three anchor labels, and the surveys
    (one per surveyed gene) run on it."""

    def __init__(self, triangle, anchors, surveys):
        self.triangle = triangle
        self.anchors = tuple(anchors)          # three alignment labels
        self.surveys = list(surveys)           # list[SurveyRecord]
        self._by_gene = {r.gene: r for r in self.surveys}

    @property
    def genes(self):
        """Surveyed gene labels, in survey order."""
        return [r.gene for r in self.surveys]

    def archive(self, gene):
        """EliteArchive for a surveyed gene by label."""
        return self._by_gene[gene].archive

    def record(self, gene):
        """SurveyRecord for a surveyed gene by label."""
        return self._by_gene[gene]

genes property

genes

Surveyed gene labels, in survey order.

archive

archive(gene)

EliteArchive for a surveyed gene by label.

Source code in src/hifuku/survey_io.py
def archive(self, gene):
    """EliteArchive for a surveyed gene by label."""
    return self._by_gene[gene].archive

record

record(gene)

SurveyRecord for a surveyed gene by label.

Source code in src/hifuku/survey_io.py
def record(self, gene):
    """SurveyRecord for a surveyed gene by label."""
    return self._by_gene[gene]