Backend Protocol¶
A backend starts the kernels of one device and reads its archive. The driver is written once against this protocol, so a backend that meets it runs correctly.
Two rules make the protocol work. submit_batch returns without waiting, which
lets the driver queue the next chunk before it calls the monitor. A backend
draws its own parent niches, because the host engine draws a fresh parent at
every step and the device engine draws one per walker per pass.
The pattern follows BEAGLE, which solves the same problem in the same field. Ayres et al. (2019).
Backend
¶
Bases: Protocol
The device interface.
A handle is opaque to the driver. A backend returns one from prepare
and gets it back on every later call.
Source code in src/hifuku/backend.py
prepare
¶
seed
¶
Place the anchor trees and the start trees in the archive.
A seed tree is placed without the log-likelihood gate. Raise
ValueError when no seed tree lands on the chart grid.
submit_batch
¶
Queue n_variations variations. Do not synchronize.
Draw the parents from keys, which holds the niche keys the driver
knows about. Apply floor as the log-likelihood gate: a variation
below the floor does not make a niche. Place every candidate before
this call returns control to the driver, or queue the placement on the
same stream.
Source code in src/hifuku/backend.py
sync
¶
read_stats
¶
read_keys
¶
Return the niche keys added at or after index since.
The driver holds the full list and asks only for the new entries, so a
device backend copies a small slice and not the whole array. The key
values are opaque to the driver, which gives them back to
submit_batch.
Source code in src/hifuku/backend.py
read_map
¶
Return the dense log-likelihood grid, NaN where unoccupied.
This is one bounded transfer. The array is a copy, never a live buffer.
read_trees
¶
Return {(i, j): HifukuTree} for every occupied niche.
The trees are copies. The CPU archive changes during a run, so a reference to a live tree would show a later state than the checkpoint.
Capabilities
dataclass
¶
What a backend is and what it can do.
Attributes:
| Name | Type | Description |
|---|---|---|
name |
str
|
Backend name, for example |
unified_memory |
bool
|
True when a state read is pointer arithmetic. On a discrete device this
is False and the driver gives one warning for a |
defaults |
Mapping[str, Any]
|
|
can_read_trees |
bool
|
True when |
Source code in src/hifuku/backend.py
Stats
dataclass
¶
Survey statistics read at one checkpoint.
Every value is cumulative from the start of the run. The driver takes the difference between two checkpoints to get the coverage and precision rates, which is why a backend must not reset these counters.
Attributes:
| Name | Type | Description |
|---|---|---|
n_keys |
int
|
Occupied niches. |
gain |
float
|
Total elite gain, summed over every replacement of an elite. |
best |
float
|
Best log-likelihood found. |
elite_range |
float
|
|
Source code in src/hifuku/backend.py
RunContext
dataclass
¶
The live objects that one survey runs against.
A backend gets this once, in prepare. The context holds trees, an
alignment and a model, so it does not travel between hosts. A
SurveyTask holds labels and indices instead, and a worker rebuilds the
context from the plan.
Attributes:
| Name | Type | Description |
|---|---|---|
anchors |
tuple
|
The three anchor trees that fix the chart. |
alignment, model |
The alignment under survey and its substitution model. |
|
start_trees |
list
|
Extra seed trees, normally the NJ tree of the surveyed gene. |
triangle |
AnchorTriangle
|
Anchor geometry. It carries the chart metric. |
taxon_table |
TaxonTable
|
The shared namespace. |
Source code in src/hifuku/backend.py
CpuBackend
¶
Run the survey on the host.
A state read is a dictionary lookup, so unified_memory is True and a
trees monitor costs nothing beyond building the mapping.
Source code in src/hifuku/cpu_backend.py
74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 | |
prepare
¶
Allocate the archive and the generator. This does no survey work.
Source code in src/hifuku/cpu_backend.py
seed
¶
Place the anchor trees and the start trees without the gate.
Source code in src/hifuku/cpu_backend.py
submit_batch
¶
Run n_variations variations as whole passes.
A host archive is always current, so this backend draws its parents
from its own key list rather than from the keys the driver holds. A
parent therefore comes from the niches that exist at the moment of the
draw, which is the sampling of the sequential survey.
Source code in src/hifuku/cpu_backend.py
sync
¶
read_stats
¶
Read the cumulative survey statistics.
Non-finite elites are skipped, as they are in best_logL and in the
device reduction. max keeps the first item a comparison cannot
beat, so a NaN among the values would set the best and the range by its
position in the dictionary rather than by its magnitude.
Source code in src/hifuku/cpu_backend.py
read_keys
¶
Return the positions of the keys added at or after since.
The values index archive.keys, which only grows. They are opaque to
the driver.
Source code in src/hifuku/cpu_backend.py
read_map
¶
Return the elite log-likelihood grid, NaN where unoccupied.
read_trees
¶
Return {(i, j): HifukuTree} for every occupied niche.
The mapping is a new dictionary. A move returns a new tree and never changes its parent, so the tree objects of a snapshot stay as they were at the checkpoint.
Source code in src/hifuku/cpu_backend.py
CudaBackend
¶
Run the survey on a CUDA device through the shared driver.
Source code in src/hifuku/gpu_backend.py
116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 | |
prepare
¶
Allocate the chains, the device archive and the buffers.
This compiles the walk kernel for the model's state count and lays out the niche allocation. It runs no survey work.
Source code in src/hifuku/gpu_backend.py
128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 | |
seed
¶
Place the anchor trees and the start trees without the gate.
The descriptor and the log-likelihood come from the host, so the seeded archive matches the CPU survey exactly.
Source code in src/hifuku/gpu_backend.py
submit_batch
¶
Queue whole passes of walk and place. Do not synchronize.
The assignments for the whole chunk upload once and the floor holds for the chunk, so no pass waits on the host.
Source code in src/hifuku/gpu_backend.py
sync
¶
read_stats
¶
Read the survey statistics as device scalars.
The elite range is best less the smallest occupied fitness. A rising
relative floor can leave early elites below it, so the smallest elite is
reduced on the device rather than taken as best - floor.
Source code in src/hifuku/gpu_backend.py
read_keys
¶
Copy the niche keys appended at or after since, in niche order.
Only the new slice moves, so a checkpoint does not copy the whole key array. A key equal to the sink is dropped: it marks a candidate that fell outside the allocation.
The slice is sorted before it is returned. A winning thread claims its slot in the key array with an atomic counter, so the device stores the keys of one pass in the order the threads finish. The driver hands this list back for parent selection, so an unsorted list would feed the thread schedule into the search and a survey would differ from run to run. The set of keys a pass adds is fixed, so sorting each slice makes the whole key list a function of the run parameters.
Source code in src/hifuku/gpu_backend.py
read_map
¶
Return the elite log-likelihood grid, NaN where unoccupied.
The grid spans the bounding box of the filled niches, which is the same window the host backend reports.
Source code in src/hifuku/gpu_backend.py
read_trees
¶
Rebuild {(i, j): HifukuTree} from the device archive state.
Source code in src/hifuku/gpu_backend.py
finalize
¶
Rebuild the host archive from the device fitness and state.