Skip to content

Checkpoint Monitors

A monitor is a plain callable that names the values it wants as its parameters. The driver reads the signature once, at setup, and then calls the monitor by keyword with exactly those values.

def watch(iters, n_filled, best_logL):    # free
    bar.update(iters)

def watch(iters, logL_map):               # one bounded transfer
    queue.put((iters, logL_map))

def watch(iters, trees):                  # state read and rebuild
    popup.update(iters, trees)

A value that costs a device read is read only when a monitor names it. An unknown name raises before any survey work starts. Every snapshot is a copy and can be pickled, so a monitor that sends checkpoints to another process is one line.

bind_monitor

bind_monitor(monitor, *, caps, n_checkpoints: int, gate_window: int = 1) -> BoundMonitor | None

Resolve a monitor against the vocabulary and the backend.

monitor=None selects the default reporter. monitor=False gives None, and the driver then calls nothing. Any other value must be a callable.

Raises:

Type Description
ValueError

An unknown parameter name, or a trees request that the backend cannot serve.

Source code in src/hifuku/monitor.py
def bind_monitor(monitor, *, caps, n_checkpoints: int,
                 gate_window: int = 1) -> BoundMonitor | None:
    """Resolve a monitor against the vocabulary and the backend.

    ``monitor=None`` selects the default reporter.  ``monitor=False`` gives
    None, and the driver then calls nothing.  Any other value must be a
    callable.

    Raises
    ------
    ValueError
        An unknown parameter name, or a ``trees`` request that the backend
        cannot serve.
    """
    if monitor is False:
        return None
    if monitor is None:
        monitor = default_reporter(n_checkpoints, gate_window)
    if not callable(monitor):
        raise TypeError(
            f"monitor must be callable, None or False, got {type(monitor).__name__}")

    names = _requested_names(monitor)
    unknown = names - MONITOR_NAMES
    if unknown:
        raise ValueError(
            f"unknown monitor parameter {sorted(unknown)}; "
            f"use one of {sorted(MONITOR_NAMES)}")

    wants_trees = "trees" in names
    if wants_trees and not caps.can_read_trees:
        raise ValueError(
            f"the {caps.name} backend cannot serve 'trees' to a monitor")
    if wants_trees and not caps.unified_memory:
        warnings.warn(
            f"the {caps.name} backend has no unified memory, so a 'trees' "
            "monitor copies the archive state at every checkpoint.  Lower "
            "checkpoint_every or drop 'trees' if the run slows down.",
            RuntimeWarning, stacklevel=3)

    return BoundMonitor(fn=monitor, names=names, wants_map="logL_map" in names,
                        wants_trees=wants_trees)

Checkpoint dataclass

One checkpoint of a running survey.

logL_map and trees are None unless a monitor named them.

Source code in src/hifuku/monitor.py
@dataclass(frozen=True)
class Checkpoint:
    """One checkpoint of a running survey.

    ``logL_map`` and ``trees`` are None unless a monitor named them.
    """

    task: Any
    iters: int
    n_filled: int
    best_logL: float
    coverage: float
    precision: float
    converged: bool
    final: bool
    elapsed: float
    logL_map: np.ndarray | None = None
    trees: dict | None = None

BoundMonitor dataclass

A monitor with its parameter names resolved.

Attributes:

Name Type Description
fn callable

The monitor.

names frozenset[str]

The values to pass. A **kwargs monitor gets FREE_NAMES.

wants_map bool

True when the driver must read the log-likelihood grid.

wants_trees bool

True when the driver must read the archive state and rebuild the trees.

Source code in src/hifuku/monitor.py
@dataclass(frozen=True)
class BoundMonitor:
    """A monitor with its parameter names resolved.

    Attributes
    ----------
    fn : callable
        The monitor.
    names : frozenset[str]
        The values to pass.  A ``**kwargs`` monitor gets ``FREE_NAMES``.
    wants_map : bool
        True when the driver must read the log-likelihood grid.
    wants_trees : bool
        True when the driver must read the archive state and rebuild the trees.
    """

    fn: Any
    names: frozenset
    wants_map: bool
    wants_trees: bool

    def __call__(self, checkpoint: Checkpoint) -> None:
        self.fn(**{name: getattr(checkpoint, name) for name in self.names})

default_reporter

default_reporter(n_checkpoints: int, gate_window: int = 1, stream=None)

Build the monitor that runs when the caller gives none.

The reporter writes survey statistics to stderr. It is an ordinary monitor that names free values only, so it goes through the same signature check as any other monitor.

The cadence comes from the budget. The reporter writes one line every max(1, n_checkpoints // 8) checkpoints and always writes the last one. A run of any size therefore gives about five to ten lines per gene. Output goes to stderr so that stdout stays clean.

The halt rates need a trailing window of gate_window checkpoints. Until that many exist the rates have no value, and the line reports how much of the window has filled instead. A reader then sees progress toward a halt rather than a column of nan.

Source code in src/hifuku/monitor.py
def default_reporter(n_checkpoints: int, gate_window: int = 1, stream=None):
    """Build the monitor that runs when the caller gives none.

    The reporter writes survey statistics to stderr.  It is an ordinary monitor
    that names free values only, so it goes through the same signature check as
    any other monitor.

    The cadence comes from the budget.  The reporter writes one line every
    ``max(1, n_checkpoints // 8)`` checkpoints and always writes the last one.
    A run of any size therefore gives about five to ten lines per gene.  Output
    goes to stderr so that stdout stays clean.

    The halt rates need a trailing window of ``gate_window`` checkpoints.  Until
    that many exist the rates have no value, and the line reports how much of
    the window has filled instead.  A reader then sees progress toward a halt
    rather than a column of ``nan``.
    """
    every = max(1, n_checkpoints // _DEFAULT_REPORTS)
    seen = [0]

    def report(iters, n_filled, best_logL, coverage, precision, final):
        seen[0] += 1
        if not final and seen[0] % every:
            return
        out = sys.stderr if stream is None else stream
        line = (f"  iters={iters:>10d}  niches={n_filled:>7d}  "
                f"best={best_logL:>14.3f}")
        if math.isnan(coverage) or math.isnan(precision):
            line += f"  halt window {min(seen[0], gate_window)}/{gate_window}"
        else:
            line += f"  coverage={coverage:.3e}  precision={precision:.3e}"
        print(line, file=out)

    return report