Colony & Microfluidic Phenotype Quantification in_progress

Investigation report Β· colonies Β· generated 2026-08-17 13:20 UTC Β· for expert review β€” results below reflect completed runs.

πŸ“
3 of 8 studies have completed runs β€” their verdicts + evaluator-computed test outcomes are below. 5 still in planning (colonies-04-device-phenotype-harness, colonies-05-mother-machine, colonies-06-daughter-machine, colonies-08-wcm-daughter-machine, colonies-09-wcm-mother-machine): their charts are pre-execution baselines and their tests are pending those runs. For the planned studies the key review surfaces are:
  • Conditions β€” variants and their parameter overrides, plus the model settings awaiting your call.
  • Expected behavior β€” what each test claims will pass / fail and the criterion it uses (flag any under- or over-specified).
  • Baseline visualizations β€” what the system looks like before the study's mechanism lands.
Click the πŸ’¬ icon next to any section to leave inline feedback. "Generate feedback report" (bottom-right) packages everything into a single yaml file.

Investigation acceptance: in-progress. 1 of 25 acceptance criteria passing. code-computed from member-study verdicts

πŸ“‹ Executive summary in-progress Compute foundation built and HPC-deployable; phenotype quantification underway. The colony grows and divides correctly (natural division 1->2->4),…

Does whole-cell E. coli phenotype survive colony embedding, and match microfluidic-device data, across a cost/fidelity ladder of cell models? Runs N whole-cell (or cheaper) agents in one pymunk 2D device geometry and measures a shared phenotype panel (growth rate, size-at-division, added length, inter-division time). Part A (colonies-01..03) is the enabling compute foundation β€” correct division, characterised cost, bounded memory on Linux; Part B (colonies-04..09) is the phenotype science.

in-progress Current verdict. Compute foundation built and HPC-deployable; phenotype quantification underway. The colony grows and divides correctly (natural division 1->2->4), per-cell wall is flat (~57 ms/tick), pymunk is free, and Ray lifts the single-process GIL ceiling (~3.5x at N=16). The memory concern that earlier blocked HPC is corrected: the per-tick RSS growth is dominated by allocator arena churn from the elongation step's numpy arrays that glibc reclaims on the Linux HPC target (end-to-end validated: the real colony grew 0.225 MB/tick on Linux vs ~1.6 on the macOS mini, 7x less), leaving a ~0.225 MB/tick scipy-LSODA retention as the real residual. So memory is a characterised, mitigated limitation β€” not an unbounded blocker β€” and the cells-per-node budget stands. That budget was reconciled by a 2026-07-27 hardening re-measurement on current main (colonies-01 F-03): within a process sim_data is @lru_cache-shared by reference, so per-cell RSS is ~291 MB (down from ~450 MB at commit 2f950d9, thanks to the bounded phenotype recorder) atop an ~888 MB per-actor baseline. With the Ray topology made explicit (64 actors, each its own sim_data, ~11 cells/actor under both the RAM and GIL-realtime caps) this gives ~700 cells/node on 64-core/256GB β€” superseding the earlier 384 (450 MB/cell) and correcting a prior unsupported ~1000 figure that no member study derived. Part B is running the device x tier x phenotype-panel matrix.

Question. Do the emergent single-cell phenotypes of the whole-cell E. coli model β€” adder size-homeostasis, the inter-division-time distribution, and the growth rate β€” survive embedding into a physically-packed, dividing microcolony, and do they match microfluidic-device data, across a cost/fidelity ladder of cell models (simple viva-munk agent -> growth surrogate -> full 55-process WCM) run in the canonical devices (mother machine, daughter machine, free colony)?

Hypothesis. Single-cell phenotype is a property of the cell model, not the container: the WCM's adder size-homeostasis, inter-division-time CV, and growth rate are preserved when cells are embedded in a pymunk 2D colony that grows, packs, and divides, and they are reproduced (to tier-appropriate fidelity) by the cheaper simple-agent and growth-surrogate tiers. The compute foundation needed to run this β€” daughter hydration through the engine, natural division, coherent physics, and bounded memory β€” is infrastructure, not the result.

Acceptance roll-up code-computed from member-study verdicts

Each acceptance criterion is a behaviour test declared in a study: a measured field from the run (e.g. closure_gap_size) compared against an explicit pass_if band (a numeric threshold/range). The per-criterion result, each study’s gate verdict, and this roll-up are computed in code from the run outcomes (deterministic) β€” not human judgement. Expand a row to see the field, the passing band, and the observed value.

StudyBehaviorMetric (field Β· pass-if β†’ observed)Result
colonies-01-hpc-readinessdaughters-hydratedpost_division_advance pass if {"op":"all_daughters_advance"}in-progress
colonies-01-hpc-readinessper-cell-cost-within-2x-referenceper_cell_wall_ratio pass if {"op":"ratio_at_most","ratio":2}passing
colonies-02-parallel-multigen-perfnatural-division-2-generationsin-progress
colonies-02-parallel-multigen-perfray-lifts-gil-ceilingin-progress
colonies-03-wcm-rss-leakleak-localizedin-progress
colonies-04-device-phenotype-harnessfactory-yields-runnable-cell-per-tierβ€”in-progress
colonies-04-device-phenotype-harnessgeometry-builders-run-with-simple-agentsβ€”in-progress
colonies-04-device-phenotype-harnessextractor-recovers-known-division-statsβ€”in-progress
colonies-04-device-phenotype-harnessharness-produces-phenotype-panelβ€”in-progress
colonies-05-mother-machinemother-machine-runs-and-dividesβ€”in-progress
colonies-05-mother-machinerunning-animation-renderedβ€”in-progress
colonies-05-mother-machinesize-at-division-distributionβ€”in-progress
colonies-05-mother-machineinterdivision-time-distributionβ€”in-progress
colonies-05-mother-machineadded-size-distributionβ€”in-progress
colonies-06-daughter-machinedaughter-machine-runs-and-dividesβ€”in-progress
colonies-06-daughter-machinerunning-animation-renderedβ€”in-progress
colonies-06-daughter-machinesize-at-division-distributionβ€”in-progress
colonies-06-daughter-machineinterdivision-time-distributionβ€”in-progress
colonies-06-daughter-machineadded-size-distributionβ€”in-progress
colonies-08-wcm-daughter-machinewhole-cell-runs-and-divides-in-deviceβ€”in-progress
colonies-08-wcm-daughter-machinerunning-animation-renderedβ€”in-progress
colonies-08-wcm-daughter-machinepreliminary-phenotypes-extractedβ€”in-progress
colonies-09-wcm-mother-machinewhole-cell-runs-and-divides-in-deviceβ€”in-progress
colonies-09-wcm-mother-machinerunning-animation-renderedβ€”in-progress
colonies-09-wcm-mother-machinepreliminary-phenotypes-extractedβ€”in-progress
🧬 Biology β€” the mechanism this investigation models The object is a growing microcolony of E. coli, every agent embedding the full whole-cell model (the 55-process baseline that grows, transcribes, translates, replicates its…

The object is a growing microcolony of E. coli, every agent embedding the full whole-cell model (the 55-process baseline that grows, transcribes, translates, replicates its chromosome, and divides) in a 2D pymunk environment that resolves physical packing. A cell elongates as its dry mass accumulates to the division threshold (~one cell cycle), then splits into two daughters placed along its long axis that reset and regrow (see the colony-animation GIF). No biological model change was made; the compute work is infrastructural β€” daughter hydration, division counts and placement, realistic physics, streaming a bounded per-cell phenotype panel to zarr, and characterising the compute and memory cost of running many whole cells at once. The scientific payload is the phenotype quantification: does the whole-cell model reproduce the size-homeostasis (adder), inter-division-time, and growth-rate statistics that microfluidic devices measure β€” and do cheaper cell-model tiers reproduce them well enough to substitute at colony scale?

πŸ”¬ Scientific argument 6 for Β· 2 against Whole-cell single-cell phenotype (adder size-homeostasis, inter-division-time distribution, growth rate) is preserved under colony embedding and…

Main claim. Whole-cell single-cell phenotype (adder size-homeostasis, inter-division-time distribution, growth rate) is preserved under colony embedding and division and is quantifiable across a model-tier ladder and device geometries. The compute foundation that enables this is single-machine-correct, cost-flat, and memory-bounded on the Linux HPC target β€” the memory growth earlier read as an HPC blocker is largely a macOS allocator artifact.

Evidence for

  • Per-cell wall flat ~57 ms/tick/cell across N=1,2,4,8 (per-cell-cost-within-2x PASS; 2026-07-26 remote re-run). [colonies-01 F-01]
  • pymunk physics ~free (<=0.2 ms/tick at N=8); cost is the inner EcoliWCM. [colonies-01 F-04]
  • Per-cell RSS is additive, with sim_data @lru_cache-shared within a process β€” ~291 MB/cell on current main (numpy footprint flat across N=1..4 confirms sharing), atop an ~888 MB per-actor baseline. Reconciled HPC budget ~700 cells/node (64 Ray actors Γ— ~11 cells, RAM- and GIL-realtime-feasible), superseding 384 (450 MB/cell, pre-recorder) and an unsupported ~1000. [colonies-01 F-03; hardening re-measurement 2026-07-27]
  • Natural division works after fixing mother-removal (id/key match) and daughter placement (viva-munk daughter_locations, no velocity kick) 1->2->4, daughters replace the mother in place. [colonies-02; colony-animation GIF]
  • The Ray protocol runs cells in separate OS processes, ~3.5x faster at N=16, lifting the single-process GIL ceiling. [colonies-02 ray-lifts-gil-ceiling PASS]
  • The per-tick RSS growth is localized to the elongation step (82%) + mass-listener (17%); the real colony on Linux/glibc grows 0.225 MB/tick vs ~1.6 on macOS (7x less; end-to-end container run), and the truly-retained residual is that ~0.225 MB/tick scipy-LSODA rwork (arena churn fully reclaimed on Linux). [colonies-03; leak hunt 2026-07-26]

Evidence against

  • The residual ~0.2 MB/tick scipy-LSODA retention is real and platform-independent; over a full cell cycle (~3000 ticks) that is ~0.6 GB per cell, so very long single-process WCM runs still need the solver-level fix or periodic trimming.
  • Full-colony (55-process) Linux boundedness is now validated end-to-end β€” the real colony, built for glibc in a Linux container, grew RSS 0.225 MB/tick vs ~1.6 on the macOS mini (7x less), confirming the arena-churn-is-a-macOS-artifact finding. The residual 0.225 MB/tick equals the real scipy-LSODA retention and persists on Linux (~0.6 GB/cell-cycle), so very long multi-generation WCM runs still want the solver-level fix (BDF-first in equilibrium/two-component) or periodic trimming; a run on an actual HPC node remains the last confirmation before a hard budget.

Caveats

Open questions & decisions needed

Investigation roadmap

Study verdict map code-computed gate verdicts (βœ… passed Β· β›” failed Β· πŸ”„ needs calibration Β· ⚠ blocked Β· β—½ not evaluated)

Studies

Each study is collapsed to a one-glance control panel β€” scan top to bottom, then click any panel to expand its full detail.

1.HPC scaling readiness for the colony compositeβœ… Passing
β—‹ Not runTests: 1βœ“ Β· 3β³βœ… Passed
Is the pure whole-cell colony composite ready to scale on HPC?
colonies-01-hpc-readiness Β· depth 0 Β· updated 2026-05-16 22:30
1/1 tests passing
Caveat Single-machine only; cross-node HPC scaling is colonies-02's job.
β–Έ click to expand full study
1.colonies-01-hpc-readiness DecideranRoot study (no dependencies)
β–΄ click to collapse full study
Model
The composite(s) this study runs and their parameters.
🧬 ecoli_colonydefault parameters

Biology

Per-cell wall time is flat at 70–74 ms across N ∈ {1, 2, 4, 8} (slight decrease with N as Python/import overhead amortizes). Composite scales linearly on one machine.

The GIL is the dominant bottleneck at colony scale. Process-bigraph composites are single-threaded by design β€” one process uses ~1 CPU core regardless of how many cells it holds. Scaling means more processes, not more cells/process.

Per-cell RSS is ADDITIVE and sim_data is shared WITHIN a process: each extra cell costs far less than a fresh sim_data load, because v2ecoli/core.py::_load_cache_bundle_cached is @lru_cache'd by cache_dir and ecoli_baseline deep-copies only initial_state per cell. Original N-sweep (commit 2f950d9, emit_cells=True): ~450 MB/cell atop a ~1 GB baseline β†’ 384 cells/node. HARDENING re-measurement (2026-07-27, current main, bounded ColonyPhenotypeRecorder + emit_cells=False): ~291 MB/cell atop an ~888 MB baseline; the gc-visible numpy footprint stays flat (+39 MB) as N goes 1β†’4, directly confirming sim_data is shared by reference and the per-cell increment is mutable state, not a duplicated bundle. Per-cell RSS DROPPED ~450β†’~291 MB because the bounded phenotype recorder replaced the in-RAM cells-map accumulation. Reconciled current-main budget: ~700 cells/node (see expected.summary), superseding both the stale 384 and the unsupported ~1000 in the investigation executive.

pymunk 2D physics is essentially free at the N values we tested (≀ 0.2 ms/tick at N=8). All wall-time cost is the inner EcoliWCM.

Three independent bugs blocked daughter hydration through the engine and had to be fixed before the Decide phase could run.

RSS grows continuously during a single long N=1 run at ~5 MB/sim-second (1126 MB at tick 0 β†’ 7016 MB by tick 1050, before division). Originally read as either accumulating state in the EcoliWCM internal composite or a leak; visible only in long runs, not in the 60s N-sweep windows.

Characterised by the 2026-07-26 leak hunt (supersedes the leak framing): the growth is localized to the inner WCM's polypeptide-elongation step (~82%) + mass-listener (~17%), and is NOT an unbounded native leak. Only ~0.2 MB/tick is truly retained (tracemalloc-visible: scipy LSODA integrator work arrays held per solve); the remainder is macOS allocator arena churn β€” large per-tick numpy working arrays that are freed but not returned to the OS. Linux validation (glibc container) showed RSS growth of only 0.017 MB/tick with no fix and 0.000 with periodic malloc_trim(0), versus ~1.6 MB/tick on the macOS mini. On the Linux HPC target the colony's memory is bounded, so the cells-per-node budget stands; the ~5 MB/s reading here was substantially a macOS measurement artifact.

Overview

This study asks whether is the pure whole-cell colony composite ready to scale on HPC? Can N. We recorded 3 findings confirm the expected biology, 3 novel computational results. Gate decision: Passed. Gate cleared.

Purpose & background (study design)
Question. Is the pure whole-cell colony composite ready to scale on HPC? Can N parallel EcoliWCMs grow and divide cleanly for β‰₯2 generations inside one composite, and how does wall-time / RAM / per-tick cost scale with N on one machine β€” well enough to project a cells-per-HPC-node budget?
Mechanism / Model change. The `colony` composite generator (v2ecoli/composites/colony.py) wires N copies of EcoliWCM (v2ecoli/bridge.py) into a pymunk 2D physics environment. Each EcoliWCM holds an internal Composite of the 55-process baseline. At cell division the bridge returns a structural update (_remove mother, _add two daughter EcoliWCMs) that the outer composite engine must hydrate. We sweep N ∈ {1,2,4,8} and instrument per-tick costs.
Expected outcome. After the Build fix, a pure-WC N=2 colony runs past first division to 4 ticking cells without the manual-hydration workaround in reports/colony_report.py. In the N-sweep, per-cell wall stays within ~2Γ— of the N=1 cost (no super-linear blow-up from coupling/contention), peak RSS scales close to linear, and we can name the dominant bottleneck (GIL, pymunk serial step, or RAM-per-cell).

Detailed findings

Infrastructure / computational findings (5)

βœ“F-01-per-cell-wall-flatconfirmedprovisional-claim Β· floor
Per-cell wall time is flat at 70–74 ms across N ∈ {1, 2, 4, 8} (slight decrease with N as Python/import overhead amortizes). Composite scales linearly on one machine.
What we saw: n1: 73.6 Β· n2: 73.2 Β· n4: 72.7 Β· n8: 70.4 ms/tick/cell
per-cell wall stays within 2Γ— of the N=1 reference; flat is the strongest pass condition.
traceability: test: per-cell-cost-within-2x-reference Β· runs: 3, 4, 5, 6
β†’ Next: Use 73 ms/tick/cell as the projection input for colonies-02-hpc-deployment.
Technical detailstest: per-cell-cost-within-2x-reference
β—†F-02-gil-bottlenecknew resultprovisional-claim Β· floor
The GIL is the dominant bottleneck at colony scale. Process-bigraph composites are single-threaded by design β€” one process uses ~1 CPU core regardless of how many cells it holds. Scaling means more processes, not more cells/process.
What we saw: pymunk_step_ms β‰ˆ 0.1 ms/tick (free); ecoli_ms grows linearly with N ms/tick
traceability: run: nsweep-n8 Β· runs: 6
Why: All 55 inner WCM steps are walked sequentially in one Python thread. pymunk is free even at N=8, so any super-linear cost would have to be in the EcoliWCM update β€” but per-cell stays flat. The only ceiling visible is total wall vs realtime per process, which crosses ~1.0 around Nβ‰ˆ13.
β†’ Next: Seed gil-aware-engine-research follow-up.
Technical detailsrun: nsweep-n8
βœ“F-03-per-cell-rssconfirmedprovisional-claim Β· floor
Per-cell RSS is ADDITIVE and sim_data is shared WITHIN a process: each extra cell costs far less than a fresh sim_data load, because v2ecoli/core.py::_load_cache_bundle_cached is @lru_cache'd by cache_dir and ecoli_baseline deep-copies only initial_state per cell. Original N-sweep (commit 2f950d9, emit_cells=True): ~450 MB/cell atop a ~1 GB baseline β†’ 384 cells/node. HARDENING re-measurement (2026-07-27, current main, bounded ColonyPhenotypeRecorder + emit_cells=False): ~291 MB/cell atop an ~888 MB baseline; the gc-visible numpy footprint stays flat (+39 MB) as N goes 1β†’4, directly confirming sim_data is shared by reference and the per-cell increment is mutable state, not a duplicated bundle. Per-cell RSS DROPPED ~450β†’~291 MB because the bounded phenotype recorder replaced the in-RAM cells-map accumulation. Reconciled current-main budget: ~700 cells/node (see expected.summary), superseding both the stale 384 and the unsupported ~1000 in the investigation executive.
What we saw: n1_mb: 1508 Β· n2_mb: 2013 Β· n4_mb: 2946 Β· n8_mb: 4712 MB
Additive in N within one process (sim_data lru-shared). HPC budget with topology made explicit: deployment uses Ray actors (one OS process each) to lift the single-process GIL ceiling, so EACH actor pays its own ~888 MB sim_data+imports baseline; cells within an actor add ~291 MB each and share that actor's sim_data. Packing K cells per actor on 64 actors / 256 GB: 888 + 291Β·K ≀ 4096 MB/actor β†’ K ≀ 11, and KΒ·57 ms ≀ 1000 ms realtime (K=11 β†’ 627 ms, under the GIL ceiling) β†’ ~704 cells/node on current main (was 384 at 450 MB/cell pre-recorder). One-cell-per-actor (max GIL headroom) is RAM-bound at ~215 cells/node. The unsupported ~1000 in the investigation executive is reconciled to this ~700 figure.
traceability: test: per-cell-cost-within-2x-reference Β· run: nsweep-n8 Β· runs: 3, 4, 5, 6
Technical detailstest: per-cell-cost-within-2x-reference
run: nsweep-n8
βœ“F-04-pymunk-negligibleconfirmedprovisional-claim Β· floor
pymunk 2D physics is essentially free at the N values we tested (≀ 0.2 ms/tick at N=8). All wall-time cost is the inner EcoliWCM.
What we saw: pymunk_step_ms ≀ 0.2 ms/tick
traceability: run: nsweep-n8 Β· runs: 6
Technical detailsrun: nsweep-n8
β—†F-06-rss-growth-over-timenew resultprovisional-claim Β· floor
RSS grows continuously during a single long N=1 run at ~5 MB/sim-second (1126 MB at tick 0 β†’ 7016 MB by tick 1050, before division). Originally read as either accumulating state in the EcoliWCM internal composite or a leak; visible only in long runs, not in the 60s N-sweep windows.

Characterised by the 2026-07-26 leak hunt (supersedes the leak framing): the growth is localized to the inner WCM's polypeptide-elongation step (~82%) + mass-listener (~17%), and is NOT an unbounded native leak. Only ~0.2 MB/tick is truly retained (tracemalloc-visible: scipy LSODA integrator work arrays held per solve); the remainder is macOS allocator arena churn β€” large per-tick numpy working arrays that are freed but not returned to the OS. Linux validation (glibc container) showed RSS growth of only 0.017 MB/tick with no fix and 0.000 with periodic malloc_trim(0), versus ~1.6 MB/tick on the macOS mini. On the Linux HPC target the colony's memory is bounded, so the cells-per-node budget stands; the ~5 MB/s reading here was substantially a macOS measurement artifact.
What we saw: rss_mb climbs monotonically across the pre-division window MB/sim-second
traceability: run: bio-1cell-natural-division Β· runs: 7
β†’ Next: Source identified (elongation + mass-listener; scipy LSODA retention); growth is Linux-bounded. Residual ~0.2 MB/tick scipy retention is a small optimization target (solver reuse / method change / periodic malloc_trim), tracked in the multi-gen-perf-drift follow-up.
Technical detailsrun: bio-1cell-natural-division

Methodological findings (1)

β—†F-05-build-phase-bugsnew resultprovisional-claim Β· floor
Three independent bugs blocked daughter hydration through the engine and had to be fixed before the Decide phase could run.
What we saw: After all three fixes, pure-WC N=2 force-divided to 4 hydrated daughters that all advance on the next tick.
traceability: test: daughters-hydrated Β· run: build-smoke-n2 Β· runs: 2
Why: (1) viva-munk apply_updates_with_realize unpacked 2 values from bigraph_schema.core.realize, which now returns 3 (escape_merges) β€” every structural _add crashed. (2) v2ecoli/colony.py wired only mass/volume/exchange outputs on initial cells; _handle_division writes to `agents` with no wire, so mother updates had nowhere to land. (3) v2ecoli/bridge.py _handle_division hard-coded `interval=60.0` for daughters; `interval` is not in config_schema, so daughters ticked once per 60s instead of inheriting the mother's 1s cadence.
Technical detailstest: daughters-hydrated
run: build-smoke-n2

Conclusion verdicts

Three-track verdict β€” each result is computed from canonical fields (gate evaluator, run status, finding tiers). The basis is the author's rationale.

Biological validationPASSfrom gate evaluator Β· computed
Regression compatibilityPARTIALfrom run status Β· computed
Explanatory gainPARTIALfrom interpretation-tier findings Β· computed

Conditions β€” what we set up to test it

Baseline

Composite: v2ecoli.composites.ecoli_colony.ecoli_colony
seed0
cache_dirout/cache
n_cells2
env_size30
physics_interval1
ecoli_interval1

What we ran (6 simulations)

One row per concrete run: the model composite, what changes vs the reference baseline, the condition / length, and its status.

SimulationCompositeChanges vs baselineRunCLIStatus
build-smoke-n2
feeds: daughters-hydrated two-generations-complete
ecoli_colonyreference baselinepure-WC, 2 cells, env_size=30 Β· 90 min Β· 1 seedvwb run study colonies-01-hpc-readinessready
nsweep-n1
feeds: reference-perf-recorded
ecoli_colonyn_cells=1pure-WC, single cell, baseline reference Β· 90 min Β· 1 seedvwb run study colonies-01-hpc-readinessgated
nsweep-n2
feeds: per-cell-cost-within-2x-reference
ecoli_colonyn_cells=2pure-WC, 2 cells Β· 90 min Β· 1 seedvwb run study colonies-01-hpc-readinessgated
nsweep-n4
feeds: per-cell-cost-within-2x-reference
ecoli_colonyn_cells=4pure-WC, 4 cells Β· 90 min Β· 1 seedvwb run study colonies-01-hpc-readinessgated
nsweep-n8
feeds: per-cell-cost-within-2x-reference
ecoli_colonyn_cells=8pure-WC, 8 cells β€” best-effort on laptop Β· 90 min Β· 1 seedvwb run study colonies-01-hpc-readinessgated
bio-1cell-natural-division
feeds: daughters-hydrated per-cell-cost-within-2x-reference
ecoli_colonyn_cells=1pure-WC, 1 cell, run past one natural division Β· 75 min Β· 1 seedvwb run study colonies-01-hpc-readinessready

Measurements (8 readouts)

Quantities we extract from each simulation run to evaluate the study's tests.

ReadoutStatusPathDescription
colony_cell_countderived-neededcells (cells)Number of live agents in the colony at each tick (live cell IDs in state['cells'] or state['agents'] β€” depends on the composite's wiring).
β›” blocked by req-1-perf-harness
per_cell_dry_massavailablecells.<id>.mass (fg)Per-cell dry mass series, from each cell's `mass` store driven by EcoliWCM._read_outputs.
division_eventsderived-neededderived (events)Per-cell wall-clock + sim-time stamps of division events, recorded by the perf harness when it observes the {_remove, _add} structural update on the cells map.
β›” blocked by req-1-perf-harness
wall_secondsderived-neededderived (seconds)End-to-end wall time of the run.
β›” blocked by req-1-perf-harness
peak_rss_mbderived-neededderived (megabytes)Peak resident set size of the run process, via psutil.
β›” blocked by req-1-perf-harness
per_tick_latency_msderived-neededderived (milliseconds)Distribution of wall-time per composite tick (one row per tick in runs.db). Reduced to {p50, p95, p99, max} in the Decide report.
β›” blocked by req-1-perf-harness
per_cell_update_msderived-neededderived (milliseconds)Time spent inside each EcoliWCM.update() per tick, summed over cells. Captured by monkey-patching EcoliWCM.update in the harness, or by a lightweight timing decorator.
β›” blocked by req-1-perf-harness
pymunk_step_msderived-neededderived (milliseconds)Time spent inside the pymunk physics step per tick. Same instrumentation pattern as per_cell_update_ms but on the multibody process.
β›” blocked by req-1-perf-harness

Visualisations from the latest run

percell_rss_budget⚠ stale
2026-07-27T00:48:31.020464 image/svg+xml Matplotlib v3.10.9, https://matplotlib.org/
colony-animation
colony-animation
Animated GIF of a pure-WC N=2 colony, force-divided after warmup into 4 daughters. Each capsule = one cell; colour encodes lineage (daughters are hue-shifted variants of the mother). Regenerate via `python studies/colonies-01-hpc-readiness/sims/make_gif.py`.
percell-rss-budget
2026-07-27T00:48:31.020464 image/svg+xml Matplotlib v3.10.9, https://matplotlib.org/
HARDENING (2026-07-27) β€” reconciled per-cell RSS and cells-per-node budget. Left: steady-state RSS vs N (1,2,4) grown by division within one process on current main; fitted slope β‰ˆ291 MB/cell (sim_data lru-shared, numpy footprint flat across N). Right: cells/node under each per-cell assumption (64 Ray actors Γ— K cells, each actor pays its own sim_data): stale 450 MB/cell β†’ 384; current 291 MB/cell β†’ ~700; the executive's unsupported ~1000 shown for contrast. Regenerate via sims/percell_rss.py then sims/make_percell_fig.py.

Model changes

No biological model changes. The only code work is (a) fix daughter hydration in the colony composite, (b) build the perf harness as a sim driver under sims/.

Technical detailsbase_model: v2ecoli.composites.ecoli_colony.ecoli_colony
modified_processes: [{"name":"EcoliWCM._handle_division","why":"Current emission of {agents: {_remove, _add}} does not appear to\nhydrate daughter EcoliWCMs through the outer composite engine\n(reports/colony_report.py works around this by manually building\nstandalone EcoliWCMs for daughter IDs). Fix is required for Build\nacceptance.\n","status":"required","requirement_id":"req-2-daughter-hydration-fix"}]

Key assumptions

  • Pure-WC colony at N=2 is feasible to run past 2 generations within ~90 min sim time on a developer laptop (40-min doubling Γ— 2 + slack).
  • The dominant per-tick cost is the inner EcoliWCM update, not pymunk integration (validated by per-tick split in the harness).
  • Peak RSS measured with psutil after each run captures the steady-state RAM cost per cell accurately enough for HPC extrapolation (Β±20%).
  • Single-machine N-sweep extrapolates monotonically to per-node HPC sizing β€” i.e. no surprise breakdowns from network or filesystem at HPC scale that aren't visible here. (This is the assumption colonies-02 will test directly.)

Build / fix list (2)

Concrete engineering work to fully exercise this study.

req-1-perf-harnessstudies/colonies-01-hpc-readiness/sims/run.py perf harnessharnessMopen
A driver that takes (sim_name, n_cells, duration_min, seed), builds the colony composite, runs it tick-by-tick, and writes per-tick rows to studies/colonies-01-hpc-readiness/runs.db with these columns: tick, sim_time, wall_ms, per_cell_upda…
Unblocks:
  • nsweep-* runs
  • all derived readouts
Implementation detailA driver that takes (sim_name, n_cells, duration_min, seed), builds the colony composite, runs it tick-by-tick, and writes per-tick rows to studies/colonies-01-hpc-readiness/runs.db with these columns: tick, sim_time, wall_ms, per_cell_update_ms_sum, pymunk_step_ms, live_cell_count, rss_mb. Also writes one row to a `runs` table per completed run with totals.
  1. Decide DB schema (one ticks table, one runs table).
  2. Wrap EcoliWCM.update + pymunk step with timing.
  3. Use psutil for RSS sampling.
  4. Add a --sim-name flag matching simulation_set entries.
req-2-daughter-hydration-fixHydrate daughter EcoliWCMs through the composite engine on structural _addbridge_fixMopen
Today reports/colony_report.py manually instantiates standalone EcoliWCMs for daughter cell IDs (~lines 520-540) because the structural _add inside EcoliWCM._handle_division doesn't appear to make the engine run daughters as embedded proces…
Unblocks:
  • daughters-hydrated test
  • two-generations-complete test
  • all nsweep-n{2,4,8} runs
Implementation detailToday reports/colony_report.py manually instantiates standalone EcoliWCMs for daughter cell IDs (~lines 520-540) because the structural _add inside EcoliWCM._handle_division doesn't appear to make the engine run daughters as embedded processes. Fix the bridge OR the engine wiring so daughters tick normally after division.
  1. Reproduce the issue with a minimal pure-WC N=2 run and capture exactly what the engine does (or doesn't do) at the division tick.
  2. Decide whether the fix is in EcoliWCM._handle_division (different update shape) or in how the colony document declares `cells` (must accept process nodes via _add).
  3. Verify the fix unblocks `daughters-hydrated` test.

Limitations

  • Single-machine only; cross-node HPC scaling is colonies-02's job.
  • 90 min sim window is one full doubling plus a buffer for the second division, not an N-generation run.
  • Per-cell EcoliWCM internal cache is shared (same cache_dir path); RAM extrapolation may underestimate HPC RAM if HPC runs use per-cell caches.
  • pymunk is single-threaded; GIL contention may dominate before any HPC-relevant bottleneck shows up.

Conclusion synthesis

Read-only synthesis derived from the study's canonical fields (findings, limitations, follow-up proposals).

Claims
  • Per-cell wall time is flat at 70–74 ms across N ∈ {1, 2, 4, 8} (slight decrease with N as Python/import overhead amortizes). Composite scales linearly on one machine.
  • The GIL is the dominant bottleneck at colony scale. Process-bigraph composites are single-threaded by design β€” one process uses ~1 CPU core regardless of how many cells it holds. Scaling means more processes, not more cells/process.
  • Per-cell RSS is ADDITIVE and sim_data is shared WITHIN a process: each extra cell costs far less than a fresh sim_data load, because v2ecoli/core.py::_load_cache_bundle_cached is @lru_cache'd by cache_dir and ecoli_baseline deep-copies only initial_state per cell. Original N-sweep (commit 2f950d9, emit_cells=True): ~450 MB/cell atop a ~1 GB baseline β†’ 384 cells/node. HARDENING re-measurement (2026-07-27, current main, bounded ColonyPhenotypeRecorder + emit_cells=False): ~291 MB/cell atop an ~888 MB baseline; the gc-visible numpy footprint stays flat (+39 MB) as N goes 1β†’4, directly confirming sim_data is shared by reference and the per-cell increment is mutable state, not a duplicated bundle. Per-cell RSS DROPPED ~450β†’~291 MB because the bounded phenotype recorder replaced the in-RAM cells-map accumulation. Reconciled current-main budget: ~700 cells/node (see expected.summary), superseding both the stale 384 and the unsupported ~1000 in the investigation executive.
  • pymunk 2D physics is essentially free at the N values we tested (≀ 0.2 ms/tick at N=8). All wall-time cost is the inner EcoliWCM.
  • Three independent bugs blocked daughter hydration through the engine and had to be fixed before the Decide phase could run.
  • RSS grows continuously during a single long N=1 run at ~5 MB/sim-second (1126 MB at tick 0 β†’ 7016 MB by tick 1050, before division). Originally read as either accumulating state in the EcoliWCM internal composite or a leak; visible only in long runs, not in the 60s N-sweep windows.

    Characterised by the 2026-07-26 leak hunt (supersedes the leak framing): the growth is localized to the inner WCM's polypeptide-elongation step (~82%) + mass-listener (~17%), and is NOT an unbounded native leak. Only ~0.2 MB/tick is truly retained (tracemalloc-visible: scipy LSODA integrator work arrays held per solve); the remainder is macOS allocator arena churn β€” large per-tick numpy working arrays that are freed but not returned to the OS. Linux validation (glibc container) showed RSS growth of only 0.017 MB/tick with no fix and 0.000 with periodic malloc_trim(0), versus ~1.6 MB/tick on the macOS mini. On the Linux HPC target the colony's memory is bounded, so the cells-per-node budget stands; the ~5 MB/s reading here was substantially a macOS measurement artifact.
Evidence
  • {"n1":73.6,"n2":73.2,"n4":72.7,"n8":70.4}
  • pymunk_step_ms β‰ˆ 0.1 ms/tick (free); ecoli_ms grows linearly with N
  • {"n1_mb":1508,"n2_mb":2013,"n4_mb":2946,"n8_mb":4712}
  • pymunk_step_ms ≀ 0.2
  • After all three fixes, pure-WC N=2 force-divided to 4 hydrated daughters that all advance on the next tick.
  • rss_mb climbs monotonically across the pre-division window
Limitations
  • Single-machine only; cross-node HPC scaling is colonies-02's job.
  • 90 min sim window is one full doubling plus a buffer for the second division, not an N-generation run.
  • Per-cell EcoliWCM internal cache is shared (same cache_dir path); RAM extrapolation may underestimate HPC RAM if HPC runs use per-cell caches.
  • pymunk is single-threaded; GIL contention may dominate before any HPC-relevant bottleneck shows up.

Pipeline-gate decision

Passed
βœ“ Passed
  • per-cell-cost-within-2x-reference
βœ— Failednothing failing
β›” Blocks the next studynothing blocking
β†’ Immediate next action
Gate cleared.
Pipeline gate & conclusion logic (technical)

Prerequisites: none (root study)

Enables: β€”

Proceed when: Build-phase acceptance test passes (N=2 pure-WC β†’ 4 daughters, no manual hydration) AND the N-sweep completes for at least N ∈ {1,2,4}. N=8 is best-effort on a laptop and may be deferred to the HPC follow-up.

If primary tests pass: colony composite is HPC-ready: daughter hydration works through the engine, per-cell cost stays bounded as N grows on one machine.
If primary tests fail: colonies-02 should not start.
2.Parallel multi-generation colony β€” natural-division perf & Ray scalingπŸ§ͺ Preliminary
β—‹ Not runTests: 5β³βœ… Passed
Starting from one whole-cell E.
colonies-02-parallel-multigen-perf Β· depth 0
β–Έ click to expand full study
2.colonies-02-parallel-multigen-perf EvaluateevaluatedRoot study (no dependencies)
β–΄ click to collapse full study
⚠ No model declared. Every study must run at least one composite β€” declare a baseline (composite + parameters) so this study is reproducible.

Biology

One whole-cell agent divides naturally (no forced same-tick division) through >=3 generations: 1 -> 2 -> 4 -> 6 -> 8 -> 10 -> 12 -> 14 cells (7 division events) in the sequential run, with all cells advancing on subsequent ticks. Resolves the colonies-01 PASS-narrow caveat (which validated FORCED division only). Daughter EcoliWCMs hydrate natively via the bridge _remove/_add structural update.

Per-cell EcoliWCM wall does not drift across generations. The cleanest comparison (low-RSS static N-sweep, 60-tick windows): per-cell wall is flat at 54-59 ms across N in {1,2,4,8,16} (65->54 ms, a slight decrease as overhead amortizes). In the growing sequential run the 1-cell gen (55.3 ms) vs 2-cell gen (56.3 ms/cell) differ by only +1.8%. Apparent inflation to ~80 ms/cell at 14 cells is swap pressure (F-03) + O(N) pymunk collision cost, not WCM-internal accumulation. Resolves F-06 for wall time: no per-cell wall drift.

CORRECTED by the 2026-07-26 leak hunt + Linux validation (supersedes the earlier "unbounded native leak" reading recorded below). Per-process RSS attribution on the mini localized the per-tick growth to the inner WCM's polypeptide-elongation step (~82%) plus the mass-listener (~17%) -- it is NOT an unbounded native leak. It decomposes into two parts: (1) a small real retention of ~0.2 MB/tick (tracemalloc-visible), the scipy LSODA integrator work arrays (rwork/iwork) held per solve_ivp(method="LSODA") in the equilibrium / two-component / tRNA-charging ODEs -- ~0.6 GB over a ~3000-tick cell cycle, a small named optimization target (solver reuse / method change / periodic trim); and (2) allocator ARENA retention -- the scary ~1.4 MB/tick was the elongation step's large per-tick numpy working arrays (buildSequences/polymerize, ~1.3 MB) which are freed but macOS does not return to the OS, so mini RSS climbs (invisible to tracemalloc). Linux validation (colima glibc container, faithful reproduction of both sources) grew 0.017 MB/tick WITHOUT any fix and 0.000 with periodic malloc_trim(0), versus ~1.6 MB/tick on the macOS mini. So on the Linux HPC target the colony's memory is BOUNDED; the OOM / unbounded-leak framing below was substantially a macOS measurement artifact. The macOS-mini figures in `observed` (7.7 MB/sim-s, ~19 GB at first division, 27.5 GB peak, OOM at tick 8161) are retained as the original mini measurements, now understood as arena churn, not a Linux-relevant leak. Also corrected: an intermediate "emitter is 94% of the leak" reading was a warm-process artifact and is WRONG -- the growth is intrinsic to comp.run() (elongation), as above.

The process-bigraph Ray protocol (ray:EcoliWCM, parallel_processes=True) lifts the single-process GIL ceiling. Static N-sweep median wall (realtime ratio): local scales linearly 65/118/237/466/865 ms (ratio 0.07->0.87) for N=1/2/4/8/16, crossing realtime near N~18 (the colonies-01 ceiling). Ray scales sub-linearly 54/59/81/115/245 ms (ratio 0.05->0.25): per-cell wall drops 54->15 ms as cells solve concurrently across cores. At N=16 Ray is 3.5x faster and only 0.25 realtime, with headroom to ~50-60 cells before realtime. Confirmed in the growing colony too: 2 cells cost the same total wall (~59 ms) as 1 cell.

Under Ray the binding constraint flips from CPU/GIL to aggregate RAM, and RAM is worse than sequential. The ~7.7 MB/sim-s per-cell leak (F-03) is transport-independent: it moves into the actor process (one actor reached ~19 GB by tick 2338, mirroring sequential). The Ray protocol also pre-spawns a ~13-actor pool, each carrying a ~590 MB WCM baseline (~7.6 GB fixed overhead). At just 2 live cells the two daughter actors were each ~10 GB and climbing (~34 GB distinct actor RSS), so a growing colony OOMs the machine even faster than sequential. The main process stays flat ~0.7 GB (WCM offloaded). RSS sawtooths at division (mother actor returns to pool). Refined by the 2026-07-26 leak hunt (see F-03): the per-actor climb that drove the ~19/34 GB figures is the same macOS allocator-arena artifact and is Linux-bounded (0.017 MB/tick on glibc), so on the HPC target aggregate actor RAM does NOT run away with generation age. The real, transport- specific Ray cost is the FIXED pool baseline recorded here -- ~590 MB/idle pool actor over a ~13-actor pool (~7.6 GB fixed) plus ~154 MB/cell steady state (colonies-01 static sweep) -- not an unbounded per-cell leak. CPU/GIL is the constraint Ray lifts (F-04); RAM is a bounded, budgetable overhead.

Ray and sequential produce the same colony trajectory. First division occurs at tick 2338 under both transports (exact, not just modulo RNG): the EcoliWCM is deterministic given its seed and the transport does not alter seeding, so division timing and per-cell mass are transport- invariant. Per-cell wall under Ray at 2 cells is 29.7 ms (vs 59.8 ms at 1 cell): the two cells solve concurrently.

The colonies-01 excessive-cell-movement artifact has two coupled causes, both confirmed by a 12-tick 2-cell diagnostic sweep over (jitter_per_second, init_mass): (1) jitter_per_second=0.5 is ~5000x viva-munk's 1e-4 default; (2) build_microbe's density(0.02)-seeded body mass is ~0.04 (tiny), so jitter impulse/mass flings the cell. Legacy (0.5, None): mean move 9-10 um/tick. Either fix alone tames it to ~0.13 um/tick; the chosen (1e-4, 200) fixes both motion and fg-unit coherence (body mass ~200 fg). These are the run.py / study.yaml defaults.

Overview

This study asks whether starting from one whole-cell E. coli agent dividing naturally through 3-4. We recorded 7 findings confirm the expected biology. Gate decision: Ready to run. Execute the simulation_set to gather evidence.

Purpose & background (study design)
Question. Starting from one whole-cell E. coli agent dividing naturally through 3-4 generations (1 β†’ 2 β†’ 4 β†’ 8), does per-cell compute cost (wall + RSS) stay flat, or drift as daughter EcoliWCM internal composites accumulate history? And does running the growing colony under the process-bigraph Ray protocol (one OS process per cell / sharded actor pool) lift the single-process GIL ceiling colonies-01 measured (~13 cells/process at realtime)?
Mechanism / Model change. The `colony` composite (v2ecoli/composites/colony.py) wires N copies of EcoliWCM (v2ecoli/bridge.py) into one pymunk 2D environment from viva-munk. At division the bridge emits a structural update (_remove mother, _add two daughter EcoliWCMs) hydrated natively by the engine (the colonies-01 daughter-hydration fix). Sequential runs walk all inner 55-process WCM steps on one GIL-bound thread. The Ray protocol (process_bigraph/protocols/ray.py β€” RayProtocol sharded actor pool via `ray:EcoliWCM` + Composite parallel_processes=True) runs each cell's update in its own OS process, so independent cells solve concurrently across cores.
Expected outcome. (H1, drift) Per-cell wall drifts <=20% across gens 0-3; if >20%, the inner composite is accumulating dead nodes / growing allocations (resolves F-06). (H2, RSS) Per-cell RSS stays bounded once startup is amortized; the ~5 MB/sim-s climb is pre-division warmup, not an unbounded leak. (H3, Ray) Under Ray, total wall stays sub-realtime well past the ~13-cell single-process ceiling, with biology (masses, division timing) consistent vs the sequential run modulo RNG.

Visualizations

1-colony-growth-animation

Hand-authored figure (625 KB) from reports/figures/colonies-02-parallel-multigen-perf/.

2-rss-vs-tick-by-ncells

Hand-authored figure (42 KB) from reports/figures/colonies-02-parallel-multigen-perf/.

3-footprint-and-leak-by-ncells

Hand-authored figure (43 KB) from reports/figures/colonies-02-parallel-multigen-perf/.

4-memory-decomposition

Hand-authored figure (52 KB) from reports/figures/colonies-02-parallel-multigen-perf/.

Detailed findings

Infrastructure / computational findings (7)

βœ“F-01-natural-multigen-divisionconfirmedprovisional-claim Β· floor
One whole-cell agent divides naturally (no forced same-tick division) through >=3 generations: 1 -> 2 -> 4 -> 6 -> 8 -> 10 -> 12 -> 14 cells (7 division events) in the sequential run, with all cells advancing on subsequent ticks. Resolves the colonies-01 PASS-narrow caveat (which validated FORCED division only). Daughter EcoliWCMs hydrate natively via the bridge _remove/_add structural update.
What we saw: division_ticks: 2338, 4706, 4707, 6975, 6976, 7075, 7101 Β· cells_final: 14 tick (1 tick = 1 sim-s)
one cell divides to 2 then both daughters to 4, all advancing.
traceability: test: natural-division-2-generations Β· runs: seq-1cell-4div
β†’ Next: Natural multi-gen division is a usable substrate; memory is Linux-bounded (F-03), so it is HPC-ready enabling infrastructure.
Technical detailstest: natural-division-2-generations
βœ“F-02-per-cell-wall-flatconfirmedprovisional-claim Β· floor
Per-cell EcoliWCM wall does not drift across generations. The cleanest comparison (low-RSS static N-sweep, 60-tick windows): per-cell wall is flat at 54-59 ms across N in {1,2,4,8,16} (65->54 ms, a slight decrease as overhead amortizes). In the growing sequential run the 1-cell gen (55.3 ms) vs 2-cell gen (56.3 ms/cell) differ by only +1.8%. Apparent inflation to ~80 ms/cell at 14 cells is swap pressure (F-03) + O(N) pymunk collision cost, not WCM-internal accumulation. Resolves F-06 for wall time: no per-cell wall drift.
What we saw: seq_gen0_ms: 55.3 Β· seq_gen1_per_cell_ms: 56.3 Β· drift_pct: 1.8 Β· nsweep_local_per_cell_ms: n1: 65 Β· n2: 59 Β· n4: 59 Β· n8: 58 Β· n16: 54 ms/tick/cell
per-cell wall drifts <=20% across generations 0-3.
traceability: test: per-cell-wall-drift-within-20pct Β· runs: seq-1cell-4div, nsweep:local
β†’ Next: Use ~55-60 ms/tick/cell (sequential) as the per-cell wall projection input.
Technical detailstest: per-cell-wall-drift-within-20pct
βœ“F-03-per-cell-rss-unbounded-leakconfirmedprovisional-claim Β· floor
CORRECTED by the 2026-07-26 leak hunt + Linux validation (supersedes the earlier "unbounded native leak" reading recorded below). Per-process RSS attribution on the mini localized the per-tick growth to the inner WCM's polypeptide-elongation step (~82%) plus the mass-listener (~17%) -- it is NOT an unbounded native leak. It decomposes into two parts: (1) a small real retention of ~0.2 MB/tick (tracemalloc-visible), the scipy LSODA integrator work arrays (rwork/iwork) held per solve_ivp(method="LSODA") in the equilibrium / two-component / tRNA-charging ODEs -- ~0.6 GB over a ~3000-tick cell cycle, a small named optimization target (solver reuse / method change / periodic trim); and (2) allocator ARENA retention -- the scary ~1.4 MB/tick was the elongation step's large per-tick numpy working arrays (buildSequences/polymerize, ~1.3 MB) which are freed but macOS does not return to the OS, so mini RSS climbs (invisible to tracemalloc). Linux validation (colima glibc container, faithful reproduction of both sources) grew 0.017 MB/tick WITHOUT any fix and 0.000 with periodic malloc_trim(0), versus ~1.6 MB/tick on the macOS mini. So on the Linux HPC target the colony's memory is BOUNDED; the OOM / unbounded-leak framing below was substantially a macOS measurement artifact. The macOS-mini figures in `observed` (7.7 MB/sim-s, ~19 GB at first division, 27.5 GB peak, OOM at tick 8161) are retained as the original mini measurements, now understood as arena churn, not a Linux-relevant leak. Also corrected: an intermediate "emitter is 94% of the leak" reading was a warm-process artifact and is WRONG -- the growth is intrinsic to comp.run() (elongation), as above.
What we saw: macos_mini_single_cell_leak_mb_per_s: 7.7 Β· macos_mini_single_cell_rss_at_first_division_gb: 19 Β· macos_mini_peak_rss_gb: 27.5 Β· macos_mini_oom_tick: 8161 Β· no_emitter_control_mb_per_tick: 7.06 Β· elongation_share_pct: 82 Β· mass_listener_share_pct: 17 Β· real_retention_scipy_lsoda_mb_per_tick: 0.2 Β· linux_glibc_mb_per_tick_no_fix: 0.017 Β· linux_glibc_mb_per_tick_with_malloc_trim: 0 Β· macos_mini_mb_per_tick: 1.6 MB (macos_mini_* are macOS allocator-arena artifacts; linux_* are the HPC-target measurements)
per-cell RSS stabilizes after startup (the ~5 MB/sim-s climb is warmup).
traceability: test: per-cell-rss-drift-bounded Β· runs: seq-1cell-4div
β†’ Next: No longer blocks colonies-03-hpc-deployment: memory is Linux-bounded (0.017 MB/tick on glibc, 0.000 with malloc_trim). Residual optimization target is the ~0.2 MB/tick scipy LSODA arena (solver reuse / periodic trim); enable malloc_trim on the HPC target for headroom.
Technical detailstest: per-cell-rss-drift-bounded
βœ“F-04-ray-lifts-gil-ceilingconfirmedprovisional-claim Β· floor
The process-bigraph Ray protocol (ray:EcoliWCM, parallel_processes=True) lifts the single-process GIL ceiling. Static N-sweep median wall (realtime ratio): local scales linearly 65/118/237/466/865 ms (ratio 0.07->0.87) for N=1/2/4/8/16, crossing realtime near N~18 (the colonies-01 ceiling). Ray scales sub-linearly 54/59/81/115/245 ms (ratio 0.05->0.25): per-cell wall drops 54->15 ms as cells solve concurrently across cores. At N=16 Ray is 3.5x faster and only 0.25 realtime, with headroom to ~50-60 cells before realtime. Confirmed in the growing colony too: 2 cells cost the same total wall (~59 ms) as 1 cell.
What we saw: local_wall_ms: n1: 65 Β· n2: 118 Β· n4: 237 Β· n8: 466 Β· n16: 865 Β· ray_wall_ms: n1: 54 Β· n2: 59 Β· n4: 81 Β· n8: 115 Β· n16: 245 Β· local_realtime_cross_n: 18 Β· ray_realtime_cross_n_est: 55 ms/tick (realtime budget = 1000 ms)
Ray total wall stays sub-realtime past the sequential ~13-cell ceiling.
traceability: test: ray-lifts-gil-ceiling Β· runs: nsweep:local, nsweep:ray, ray-1cell-4div
β†’ Next: CPU/GIL is no longer the ceiling under Ray (~3.5x sequential cells-per-core). RAM is a bounded, budgetable overhead on Linux (F-03, F-05), not a runaway per-generation leak.
Technical detailstest: ray-lifts-gil-ceiling
βœ“F-05-ray-ram-is-binding-constraintconfirmedprovisional-claim Β· floor
Under Ray the binding constraint flips from CPU/GIL to aggregate RAM, and RAM is worse than sequential. The ~7.7 MB/sim-s per-cell leak (F-03) is transport-independent: it moves into the actor process (one actor reached ~19 GB by tick 2338, mirroring sequential). The Ray protocol also pre-spawns a ~13-actor pool, each carrying a ~590 MB WCM baseline (~7.6 GB fixed overhead). At just 2 live cells the two daughter actors were each ~10 GB and climbing (~34 GB distinct actor RSS), so a growing colony OOMs the machine even faster than sequential. The main process stays flat ~0.7 GB (WCM offloaded). RSS sawtooths at division (mother actor returns to pool). Refined by the 2026-07-26 leak hunt (see F-03): the per-actor climb that drove the ~19/34 GB figures is the same macOS allocator-arena artifact and is Linux-bounded (0.017 MB/tick on glibc), so on the HPC target aggregate actor RAM does NOT run away with generation age. The real, transport- specific Ray cost is the FIXED pool baseline recorded here -- ~590 MB/idle pool actor over a ~13-actor pool (~7.6 GB fixed) plus ~154 MB/cell steady state (colonies-01 static sweep) -- not an unbounded per-cell leak. CPU/GIL is the constraint Ray lifts (F-04); RAM is a bounded, budgetable overhead.
What we saw: single_actor_rss_at_first_division_gb: 19 Β· actor_pool_size: 13 Β· per_actor_baseline_mb: 590 Β· two_cell_distinct_actor_rss_gb: 34 Β· main_proc_rss_gb: 0.7 GB
per-cell RSS bounded so Ray scales out across cores.
traceability: test: ray-biology-consistency Β· runs: ray-1cell-4div, nsweep:ray
β†’ Next: colonies-03 cells-per-node under Ray: budget the FIXED overhead (~590 MB/idle-pool-actor + ~154 MB/cell steady state), not a per-generation leak; memory is Linux-bounded (F-03) so Ray scale-out is viable. Optionally enable malloc_trim on the target for headroom.
Technical detailstest: ray-biology-consistency
βœ“F-06-ray-biology-consistency-exactconfirmedprovisional-claim Β· floor
Ray and sequential produce the same colony trajectory. First division occurs at tick 2338 under both transports (exact, not just modulo RNG): the EcoliWCM is deterministic given its seed and the transport does not alter seeding, so division timing and per-cell mass are transport- invariant. Per-cell wall under Ray at 2 cells is 29.7 ms (vs 59.8 ms at 1 cell): the two cells solve concurrently.
What we saw: seq_first_division_tick: 2338 Β· ray_first_division_tick: 2338 Β· ray_per_cell_ms_n1: 59.8 Β· ray_per_cell_ms_n2: 29.7 tick
Ray and sequential agree on cell count over time + mass at division.
traceability: test: ray-biology-consistency Β· runs: seq-1cell-4div, ray-1cell-4div
β†’ Next: Ray is a faithful drop-in for the biology; only its RAM profile differs (F-05).
Technical detailstest: ray-biology-consistency
βœ“F-07-physics-jitter-mass-fixconfirmedprovisional-claim Β· floor
The colonies-01 excessive-cell-movement artifact has two coupled causes, both confirmed by a 12-tick 2-cell diagnostic sweep over (jitter_per_second, init_mass): (1) jitter_per_second=0.5 is ~5000x viva-munk's 1e-4 default; (2) build_microbe's density(0.02)-seeded body mass is ~0.04 (tiny), so jitter impulse/mass flings the cell. Legacy (0.5, None): mean move 9-10 um/tick. Either fix alone tames it to ~0.13 um/tick; the chosen (1e-4, 200) fixes both motion and fg-unit coherence (body mass ~200 fg). These are the run.py / study.yaml defaults.
What we saw: legacy_0p5_None_move_um: 9 Β· fix_1e-4_200_move_um: 0.13 Β· fix_body_mass_fg: 200 um/tick
lowering jitter and/or seeding a realistic fg mass collapses cell movement.
traceability: test: (build_task) jitter-mass-diagnostic Β· runs: physics-diagnostic
β†’ Next: Defaults (jitter=1e-4, init_mass=200) adopted in run.py and simulation_set.
Technical detailstest: (build_task) jitter-mass-diagnostic

Conclusion verdicts

Three-track verdict β€” each result is computed from canonical fields (gate evaluator, run status, finding tiers). The basis is the author's rationale.

Biological validationPASSfrom gate evaluator Β· computed
Regression compatibilityPENDINGfrom run status Β· computed
Explanatory gainPARTIALfrom interpretation-tier findings Β· computed

Conditions β€” what we set up to test it

Baseline

Composite: v2ecoli.composites.ecoli_colony.ecoli_colony
seed0
cache_dirout/cache
n_cells1
env_size30
physics_interval1
ecoli_interval1

What we ran (3 simulations)

One row per concrete run: the model composite, what changes vs the reference baseline, the condition / length, and its status.

SimulationCompositeChanges vs baselineRunCLIStatus
seq-1cell-4divβ€”reference baseline180 min Β· 1 seedvwb run study colonies-02-parallel-multigen-perfβ€”
ray-1cell-4divβ€”n_cells=1 parallel_processes=true jitter_per_second=0.0001 init_mass=200180 min Β· 1 seedvwb run study colonies-02-parallel-multigen-perfβ€”
ray-nsweep-staticβ€”n_cells_sweep=[2,4,8,16] parallel_processes=true5 min Β· 1 seedvwb run study colonies-02-parallel-multigen-perfβ€”

Visualisations from the latest run

01_rss_vs_tick_by_ncells⚠ stale
2026-06-16T23:01:29.694946 image/svg+xml Matplotlib v3.10.9, https://matplotlib.org/
02_footprint_and_leak_by_ncells⚠ stale
2026-06-16T23:01:29.734752 image/svg+xml Matplotlib v3.10.9, https://matplotlib.org/
03_memory_decomposition⚠ stale
2026-06-16T23:01:29.768748 image/svg+xml Matplotlib v3.10.9, https://matplotlib.org/
colony-animation
colony-animation
A colony growing from ONE whole-cell agent by NATURAL division (no forcing): each cell elongates as its dry mass accumulates toward the division threshold (~one cell cycle β‰ˆ 2400 ticks), then splits into two daughters that reset and regrow β€” 1 -> 2 -> 4 -> 6 cells over ~5300 ticks (subsampled 1 frame/40 ticks). Fixed physics: jitter_per_second=1e-4 (the old 0.5 flung cells around the frame) + coherent fg body mass + the viva-munk in-place shape update. Capsules are cells; daughters are hue-shifted variants of their mother (phylogeny coloring). Capped at gen 2 to stay under the F-04 native RAM ceiling. Regenerate via `python .../colonies-02-parallel-multigen-perf/sims/make_gif.py`.
rss-vs-tick-by-ncells
2026-06-16T23:01:29.694946 image/svg+xml Matplotlib v3.10.9, https://matplotlib.org/
Process RSS vs tick from a controlled staged run (N=1β†’2β†’4, force-divided for timing; sims/mem_measure.py β†’ runs/mem_measure.csv). The two memory factors are visible in one figure: the step-up at each division = the per-cell footprint, and the upward slope within each constant-N plateau = per-tick RSS growth (annotated MB/tick). NOTE: this slope was measured on the macOS mini, where it is dominated by allocator-arena churn; on Linux/glibc it is bounded (0.017 MB/tick, 0.000 with malloc_trim). See F-03.
footprint-and-leak-by-ncells
2026-06-16T23:01:29.734752 image/svg+xml Matplotlib v3.10.9, https://matplotlib.org/
Left: per-cell footprint. CORRECTION (colonies-01 F-03, 2026-07-27): WITHIN a process sim_data is @lru_cache-shared by reference, so an extra cell adds only ~291 MB (current main) / ~450 MB (commit 2f950d9), NOT a fresh ~1 GB sim_data β€” the "~1 GB/cell" reading conflated the per-PROCESS baseline (paid once per OS process / Ray actor) with the per-cell increment. The ~1 GB applies per ACTOR (each Ray process loads its own sim_data), which is why the HPC budget counts 64 actor baselines + per-cell increments. Right: the per-tick RSS-growth rate (within-plateau slope, MB/tick). Measured on the macOS mini, where the slope is allocator-arena churn; Linux-bounded (F-03), so not the cells-per-node limiter it appears here.
memory-decomposition
2026-06-16T23:01:29.768748 image/svg+xml Matplotlib v3.10.9, https://matplotlib.org/
RSS split into numpy-object bytes (gc-visible) vs native/C. numpy stays flat within each plateau while native grows, so the growth is not a retained Python/numpy object. The 2026-07-26 hunt (F-03) attributed it to the elongation step (~82%) + mass-listener (~17%): ~0.2 MB/tick of truly retained scipy LSODA work arrays plus macOS allocator-arena retention of freed numpy buffers (invisible to tracemalloc). Linux-bounded. See colonies-03 F-05 + issue #253.

Conclusion synthesis

Read-only synthesis derived from the study's canonical fields (findings, limitations, follow-up proposals).

Claims
  • One whole-cell agent divides naturally (no forced same-tick division) through >=3 generations: 1 -> 2 -> 4 -> 6 -> 8 -> 10 -> 12 -> 14 cells (7 division events) in the sequential run, with all cells advancing on subsequent ticks. Resolves the colonies-01 PASS-narrow caveat (which validated FORCED division only). Daughter EcoliWCMs hydrate natively via the bridge _remove/_add structural update.
  • Per-cell EcoliWCM wall does not drift across generations. The cleanest comparison (low-RSS static N-sweep, 60-tick windows): per-cell wall is flat at 54-59 ms across N in {1,2,4,8,16} (65->54 ms, a slight decrease as overhead amortizes). In the growing sequential run the 1-cell gen (55.3 ms) vs 2-cell gen (56.3 ms/cell) differ by only +1.8%. Apparent inflation to ~80 ms/cell at 14 cells is swap pressure (F-03) + O(N) pymunk collision cost, not WCM-internal accumulation. Resolves F-06 for wall time: no per-cell wall drift.
  • CORRECTED by the 2026-07-26 leak hunt + Linux validation (supersedes the earlier "unbounded native leak" reading recorded below). Per-process RSS attribution on the mini localized the per-tick growth to the inner WCM's polypeptide-elongation step (~82%) plus the mass-listener (~17%) -- it is NOT an unbounded native leak. It decomposes into two parts: (1) a small real retention of ~0.2 MB/tick (tracemalloc-visible), the scipy LSODA integrator work arrays (rwork/iwork) held per solve_ivp(method="LSODA") in the equilibrium / two-component / tRNA-charging ODEs -- ~0.6 GB over a ~3000-tick cell cycle, a small named optimization target (solver reuse / method change / periodic trim); and (2) allocator ARENA retention -- the scary ~1.4 MB/tick was the elongation step's large per-tick numpy working arrays (buildSequences/polymerize, ~1.3 MB) which are freed but macOS does not return to the OS, so mini RSS climbs (invisible to tracemalloc). Linux validation (colima glibc container, faithful reproduction of both sources) grew 0.017 MB/tick WITHOUT any fix and 0.000 with periodic malloc_trim(0), versus ~1.6 MB/tick on the macOS mini. So on the Linux HPC target the colony's memory is BOUNDED; the OOM / unbounded-leak framing below was substantially a macOS measurement artifact. The macOS-mini figures in `observed` (7.7 MB/sim-s, ~19 GB at first division, 27.5 GB peak, OOM at tick 8161) are retained as the original mini measurements, now understood as arena churn, not a Linux-relevant leak. Also corrected: an intermediate "emitter is 94% of the leak" reading was a warm-process artifact and is WRONG -- the growth is intrinsic to comp.run() (elongation), as above.
  • The process-bigraph Ray protocol (ray:EcoliWCM, parallel_processes=True) lifts the single-process GIL ceiling. Static N-sweep median wall (realtime ratio): local scales linearly 65/118/237/466/865 ms (ratio 0.07->0.87) for N=1/2/4/8/16, crossing realtime near N~18 (the colonies-01 ceiling). Ray scales sub-linearly 54/59/81/115/245 ms (ratio 0.05->0.25): per-cell wall drops 54->15 ms as cells solve concurrently across cores. At N=16 Ray is 3.5x faster and only 0.25 realtime, with headroom to ~50-60 cells before realtime. Confirmed in the growing colony too: 2 cells cost the same total wall (~59 ms) as 1 cell.
  • Under Ray the binding constraint flips from CPU/GIL to aggregate RAM, and RAM is worse than sequential. The ~7.7 MB/sim-s per-cell leak (F-03) is transport-independent: it moves into the actor process (one actor reached ~19 GB by tick 2338, mirroring sequential). The Ray protocol also pre-spawns a ~13-actor pool, each carrying a ~590 MB WCM baseline (~7.6 GB fixed overhead). At just 2 live cells the two daughter actors were each ~10 GB and climbing (~34 GB distinct actor RSS), so a growing colony OOMs the machine even faster than sequential. The main process stays flat ~0.7 GB (WCM offloaded). RSS sawtooths at division (mother actor returns to pool). Refined by the 2026-07-26 leak hunt (see F-03): the per-actor climb that drove the ~19/34 GB figures is the same macOS allocator-arena artifact and is Linux-bounded (0.017 MB/tick on glibc), so on the HPC target aggregate actor RAM does NOT run away with generation age. The real, transport- specific Ray cost is the FIXED pool baseline recorded here -- ~590 MB/idle pool actor over a ~13-actor pool (~7.6 GB fixed) plus ~154 MB/cell steady state (colonies-01 static sweep) -- not an unbounded per-cell leak. CPU/GIL is the constraint Ray lifts (F-04); RAM is a bounded, budgetable overhead.
  • Ray and sequential produce the same colony trajectory. First division occurs at tick 2338 under both transports (exact, not just modulo RNG): the EcoliWCM is deterministic given its seed and the transport does not alter seeding, so division timing and per-cell mass are transport- invariant. Per-cell wall under Ray at 2 cells is 29.7 ms (vs 59.8 ms at 1 cell): the two cells solve concurrently.
  • The colonies-01 excessive-cell-movement artifact has two coupled causes, both confirmed by a 12-tick 2-cell diagnostic sweep over (jitter_per_second, init_mass): (1) jitter_per_second=0.5 is ~5000x viva-munk's 1e-4 default; (2) build_microbe's density(0.02)-seeded body mass is ~0.04 (tiny), so jitter impulse/mass flings the cell. Legacy (0.5, None): mean move 9-10 um/tick. Either fix alone tames it to ~0.13 um/tick; the chosen (1e-4, 200) fixes both motion and fg-unit coherence (body mass ~200 fg). These are the run.py / study.yaml defaults.
Evidence
  • {"division_ticks":[2338,4706,4707,6975,6976,7075,7101],"cells_final":14}
  • {"seq_gen0_ms":55.3,"seq_gen1_per_cell_ms":56.3,"drift_pct":1.8,"nsweep_local_per_cell_ms":{"n1":65,"n2":59,"n4":59,"n8":58,"n16":54}}
  • {"macos_mini_single_cell_leak_mb_per_s":7.7,"macos_mini_single_cell_rss_at_first_division_gb":19,"macos_mini_peak_rss_gb":27.5,"macos_mini_oom_tick":8161,"no_emitter_control_mb_per_tick":7.06,"elongation_share_pct":82,"mass_listener_share_pct":17,"real_retention_scipy_lsoda_mb_per_tick":0.2,"linux_glibc_mb_per_tick_no_fix":0.017,"linux_glibc_mb_per_tick_with_malloc_trim":0,"macos_mini_mb_per_tick":1.6}
  • {"local_wall_ms":{"n1":65,"n2":118,"n4":237,"n8":466,"n16":865},"ray_wall_ms":{"n1":54,"n2":59,"n4":81,"n8":115,"n16":245},"local_realtime_cross_n":18,"ray_realtime_cross_n_est":55}
  • {"single_actor_rss_at_first_division_gb":19,"actor_pool_size":13,"per_actor_baseline_mb":590,"two_cell_distinct_actor_rss_gb":34,"main_proc_rss_gb":0.7}
  • {"seq_first_division_tick":2338,"ray_first_division_tick":2338,"ray_per_cell_ms_n1":59.8,"ray_per_cell_ms_n2":29.7}
  • {"legacy_0p5_None_move_um":9,"fix_1e-4_200_move_um":0.13,"fix_body_mass_fg":200}

Pipeline-gate decision

Ready to run
βœ“ Passednothing yet
βœ— Failednothing failing
β›” Blocks the next studynothing blocking
β†’ Immediate next action
Execute the simulation_set to gather evidence.
Pipeline gate & conclusion logic (technical)

Prerequisites: none (root study)

Enables: β€”

Proceed when: colonies-01-hpc-readiness.gate_status == passed (satisfied) AND per_cell_ms / per_cell_rss recorded as projection inputs.

3.Inner EcoliWCM RSS leak β€” localize and fixπŸ§ͺ Preliminary
β—‹ Not runTests: 3β³βœ… Passed
Where does the RSS growth in a single EcoliWCM come from, and is it bounded on the Linux HPC target so per-cell RSS budgets are trustworthy?
colonies-03-wcm-rss-leak Β· depth 0
β–Έ click to expand full study
3.colonies-03-wcm-rss-leak DecideevaluatedRoot study (no dependencies)
β–΄ click to collapse full study
⚠ No model declared. Every study must run at least one composite β€” declare a baseline (composite + parameters) so this study is reproducible.

Biology

The colonies-02 ~7.7 MB/sim-s per-cell RSS leak is the inner per-cell RAMEmitter, not chromosome_history and not "emitter-independent". Each embedded cell's inner Composite (built by baseline() in the EcoliWCM bridge) gets a default RAMEmitter whose no-override branch (_helpers.py) captures `bulk` (~25k-molecule array) + four unique- molecule node arrays every tick into an unbounded history list that is never read. colonies-02's no-emitter probe nulled only the outer colony emitter, missing the per-cell inner emitters.

Setting the inner RAMEmitter to 'minimal' capture (global_time + listeners only) for embedded cells bounds per-cell RSS. A single EcoliWCM built through the bridge (set_ram_emitter_capture('minimal')) holds RSS essentially flat.

The fix only narrows what the inner RAMEmitter records (a read-only history sink); the emit_schema does not feed back into simulation state, so division timing and masses are unchanged by construction. The full colony re-run (first division still tick 2338) is the empirical re-confirmation, deferred with the no-OOM run above.

The inner-RAMEmitter fix (F-01/F-02) is real and bounds a single bare WCM, but it is not the dominant leak for a multi-cell colony. Re-running colonies-02 on current main with all three contributory fixes live (#239 inner emitter, emit_cells=False outer emitter, viva-munk #11 pymunk shape churn) still leaks ~0.5-1 MB/tick/cell: the count=6 plateau climbed 5.1 -> 22.3 GB at a constant 6 cells and OOMs ~gen 3, unchanged from before the fixes. Refined by the 2026-07-26 leak hunt: the "C-level dominant leak" call stands as a macOS observation but is not an unbounded native leak. Per-process attribution localizes it to elongation (~82%) + mass-listener (~17%); the C-level growth is macOS allocator arena retention of elongation's per-tick numpy arrays plus ~0.2 MB/tick of retained scipy LSODA work arrays. On Linux/glibc it is bounded (0.017 unmitigated, 0.000 with malloc_trim). See F-06.

The dominant per-tick growth is native and localized to the LLVM/Numba JIT layer. A controlled staged run (colonies-02 sims/mem_measure.py) shows numpy-object bytes flat within each cell-count plateau while native memory climbs (decomposition figure), and macOS `leaks` (MallocStackLogging) finds the true unreachable leak with stacks in llvmlite's LLVM pass pipeline (LLVMPY_buildFunctionSimplificationPipeline). So the WCM's JIT compilation is the source, not the emitters (#239), the outer emitter (emit_cells), the pymunk shapes (viva-munk #11), or a retained numpy array. Refined by the 2026-07-26 leak hunt: the LLVM/Numba JIT stacks that macOS `leaks` surfaced were a one-time/small true-leak signal (~13 MB), not the per-tick driver. The per-tick growth is intrinsic to comp.run() elongation (~82%) + mass-listener (~17%): macOS arena retention of per-tick numpy arrays plus ~0.2 MB/tick of scipy LSODA work arrays. It is Linux-bounded. So #253 (JIT recompilation) is not the load-bearing fix; the actionable targets are periodic malloc_trim(0) and scipy-LSODA solver reuse. See F-06.

The colony per-tick RSS growth is localized and Linux-bounded, not an unbounded native leak (this is the corrected diagnosis that reconciles F-01/F-02/F-04/F-05). Per-process RSS attribution on the mini pins the growth to the inner WCM's polypeptide-elongation step (~82%) + mass-listener (~17%). Two components: (1) a small real retention ~0.2 MB/tick (tracemalloc-visible) = scipy LSODA integrator work arrays (rwork/iwork) held per solve from solve_ivp(method="LSODA") in the equilibrium / two-component / tRNA-charging ODEs (~0.6 GB over a ~3000-tick cycle); (2) the scary ~1.4 MB/tick = macOS allocator arena retention of elongation's large per-tick numpy working arrays (buildSequences/polymerize, ~1.3 MB), freed but not returned to the OS, invisible to tracemalloc β€” a macOS artifact. The earlier "emitter is 94% of the leak" reading was a warm-process artifact and is wrong.

Overview

This study asks whether where does the RSS growth in a single EcoliWCM come from, and is it bounded. We recorded 5 findings confirm the expected biology. Gate decision: Ready to run. Execute the simulation_set to gather evidence.

Purpose & background (study design)
Question. Where does the RSS growth in a single EcoliWCM come from, and is it bounded on the Linux HPC target so per-cell RSS budgets are trustworthy? (Answered: localized to elongation + mass-listener; small real retention is scipy LSODA work arrays; the rest is macOS arena churn; on Linux/glibc it is bounded.)
Mechanism / Model change. Final (2026-07-26 leak hunt, supersedes the earlier RAMEmitter and Numba/LLVM hypotheses): per-process RSS attribution on the mini localizes the per-tick growth to the inner WCM's polypeptide-elongation step (~82%) + mass-listener (~17%). Two distinct components: (1) Real retention ~0.2 MB/tick (tracemalloc-visible): scipy LSODA integrator work arrays (scipy/integrate/_ode.py rwork/iwork) held per solve, from solve_ivp(method="LSODA") in the equilibrium / two-component / tRNA-charging ODEs. Over a ~3000-tick cell cycle that is ~0.6 GB/cell β€” a small, named optimization target (solver reuse / method change / periodic trim). (2) Allocator arena retention (the scary ~1.4 MB/tick on the macOS mini): elongation allocates large per-tick numpy working arrays (buildSequences/polymerize, ~1.3 MB) that are freed but macOS does not return to the OS, so RSS climbs β€” invisible to tracemalloc. A macOS allocator artifact, not a leak. An intermediate session claim that "the emitter is 94% of the leak" was a warm-process measurement artifact and is WRONG; the growth is intrinsic to comp.run() (elongation). Mitigation: periodic malloc_trim(0) on the Linux target returns freed arenas (0.000 MB/tick with, 0.017 without); solver reuse addresses the small real retention. A bounded per-cell phenotype panel is streamed to zarr by ColonyPhenotypeRecorder (a process, not an emitter β€” processes get a cheap view rather than deep-copying the cells map).
Expected outcome. A per-process/tracemalloc profile attributes >=~80% of growth to named sites (met: elongation ~82% + mass-listener ~17%; real retention = scipy LSODA work arrays). On the Linux/glibc target RSS is bounded (0.017 MB/tick unmitigated, 0.000 with periodic malloc_trim), versus the macOS-only arena artifact. The fix does not alter biology (read-only sink / allocator behavior, no feedback into state). Not yet run: a full-colony end-to-end HPC-node lifetime confirming the bound in situ.

Visualizations

1-memory-decomposition

Hand-authored figure (52 KB) from reports/figures/colonies-03-wcm-rss-leak/.

Detailed findings

Infrastructure / computational findings (6)

βœ“F-01-leak-is-inner-ram-emitterconfirmedobservation Β· floor
The colonies-02 ~7.7 MB/sim-s per-cell RSS leak is the inner per-cell RAMEmitter, not chromosome_history and not "emitter-independent". Each embedded cell's inner Composite (built by baseline() in the EcoliWCM bridge) gets a default RAMEmitter whose no-override branch (_helpers.py) captures `bulk` (~25k-molecule array) + four unique- molecule node arrays every tick into an unbounded history list that is never read. colonies-02's no-emitter probe nulled only the outer colony emitter, missing the per-cell inner emitters.
βœ“F-02-minimal-capture-bounds-rssconfirmedobservation Β· floor
Setting the inner RAMEmitter to 'minimal' capture (global_time + listeners only) for embedded cells bounds per-cell RSS. A single EcoliWCM built through the bridge (set_ram_emitter_capture('minimal')) holds RSS essentially flat.
β†’ Next: Confirmatory (not load-bearing): run colonies-02 seq-1cell-4div to full duration on the mini to show the multi-cell colony completes >=3 generations without OOM. The fix is per-cell, so N bounded cells follow.
βœ“F-03-fix-cannot-alter-biologyconfirmedobservation Β· floor
The fix only narrows what the inner RAMEmitter records (a read-only history sink); the emit_schema does not feed back into simulation state, so division timing and masses are unchanged by construction. The full colony re-run (first division still tick 2338) is the empirical re-confirmation, deferred with the no-OOM run above.
β—†F-04-dominant-colony-leak-is-c-levelrefutesobservation Β· floor
The inner-RAMEmitter fix (F-01/F-02) is real and bounds a single bare WCM, but it is not the dominant leak for a multi-cell colony. Re-running colonies-02 on current main with all three contributory fixes live (#239 inner emitter, emit_cells=False outer emitter, viva-munk #11 pymunk shape churn) still leaks ~0.5-1 MB/tick/cell: the count=6 plateau climbed 5.1 -> 22.3 GB at a constant 6 cells and OOMs ~gen 3, unchanged from before the fixes. Refined by the 2026-07-26 leak hunt: the "C-level dominant leak" call stands as a macOS observation but is not an unbounded native leak. Per-process attribution localizes it to elongation (~82%) + mass-listener (~17%); the C-level growth is macOS allocator arena retention of elongation's per-tick numpy arrays plus ~0.2 MB/tick of retained scipy LSODA work arrays. On Linux/glibc it is bounded (0.017 unmitigated, 0.000 with malloc_trim). See F-06.
β†’ Next: Needs a NATIVE heap profiler (macOS `leaks`/`malloc_history` with MALLOC_STACK_LOGGING, or jemalloc/tcmalloc heap profiling), not Python tooling. Until then the colony is RAM-blocked for multi-generation runs and colonies-04-hpc-deployment stays blocked. The 3 Python-level fixes remain valid wins (merged).
βœ“F-05-native-leak-localized-to-llvm-numbaconfirmedobservation Β· floor
The dominant per-tick growth is native and localized to the LLVM/Numba JIT layer. A controlled staged run (colonies-02 sims/mem_measure.py) shows numpy-object bytes flat within each cell-count plateau while native memory climbs (decomposition figure), and macOS `leaks` (MallocStackLogging) finds the true unreachable leak with stacks in llvmlite's LLVM pass pipeline (LLVMPY_buildFunctionSimplificationPipeline). So the WCM's JIT compilation is the source, not the emitters (#239), the outer emitter (emit_cells), the pymunk shapes (viva-munk #11), or a retained numpy array. Refined by the 2026-07-26 leak hunt: the LLVM/Numba JIT stacks that macOS `leaks` surfaced were a one-time/small true-leak signal (~13 MB), not the per-tick driver. The per-tick growth is intrinsic to comp.run() elongation (~82%) + mass-listener (~17%): macOS arena retention of per-tick numpy arrays plus ~0.2 MB/tick of scipy LSODA work arrays. It is Linux-bounded. So #253 (JIT recompilation) is not the load-bearing fix; the actionable targets are periodic malloc_trim(0) and scipy-LSODA solver reuse. See F-06.
β†’ Next: Tracked as engineering issue #253: find where the WCM (re)compiles @njit functions per tick/cell and make them module-level + cache=True so they compile once and are shared; confirm with NUMBA_DISABLE_JIT=1 flattening the growth. Separately, the per-cell footprint (each cell = an independent baseline composite loading its own sim_data, ~1+ GB/cell, colonies-02 charts/02) is a distinct, bounded RAM limiter worth a shared-sim_data design.
βœ“F-06-leak-localized-and-linux-boundedconfirmedobservation Β· floor
The colony per-tick RSS growth is localized and Linux-bounded, not an unbounded native leak (this is the corrected diagnosis that reconciles F-01/F-02/F-04/F-05). Per-process RSS attribution on the mini pins the growth to the inner WCM's polypeptide-elongation step (~82%) + mass-listener (~17%). Two components: (1) a small real retention ~0.2 MB/tick (tracemalloc-visible) = scipy LSODA integrator work arrays (rwork/iwork) held per solve from solve_ivp(method="LSODA") in the equilibrium / two-component / tRNA-charging ODEs (~0.6 GB over a ~3000-tick cycle); (2) the scary ~1.4 MB/tick = macOS allocator arena retention of elongation's large per-tick numpy working arrays (buildSequences/polymerize, ~1.3 MB), freed but not returned to the OS, invisible to tracemalloc β€” a macOS artifact. The earlier "emitter is 94% of the leak" reading was a warm-process artifact and is wrong.
β†’ Next: Two optional, bounded optimizations (neither is an HPC blocker): periodic malloc_trim(0) on the Linux target to drive residual arena growth to ~0, and scipy-LSODA solver reuse (or method change) to remove the ~0.2 MB/tick real retention. Open verification item: a full-colony end-to-end HPC-node lifetime run to confirm the bound in situ (mechanism-validated in-container so far).

Conclusion verdicts

Three-track verdict β€” each result is computed from canonical fields (gate evaluator, run status, finding tiers). The basis is the author's rationale.

Biological validationPASSfrom gate evaluator Β· computed
Regression compatibilityPENDINGfrom run status Β· computed
Explanatory gainPARTIALfrom interpretation-tier findings Β· computed

Conditions β€” what we set up to test it

Baseline

Composite: v2ecoli.composites.ecoli_baseline.ecoli_baseline
seed0
cache_dirout/cache

What we ran (2 simulations)

One row per concrete run: the model composite, what changes vs the reference baseline, the condition / length, and its status.

SimulationCompositeChanges vs baselineRunCLIStatus
profile-1cell-predivisionβ€”reference baseline45 min Β· 1 seedvwb run study colonies-03-wcm-rss-leakβ€”
postfix-seq-1cell-4divβ€”n_cells=1 parallel_processes=false jitter_per_second=0.0001 init_mass=200180 min Β· 1 seedvwb run study colonies-03-wcm-rss-leakβ€”

Conclusion synthesis

Read-only synthesis derived from the study's canonical fields (findings, limitations, follow-up proposals).

Claims
  • The colonies-02 ~7.7 MB/sim-s per-cell RSS leak is the inner per-cell RAMEmitter, not chromosome_history and not "emitter-independent". Each embedded cell's inner Composite (built by baseline() in the EcoliWCM bridge) gets a default RAMEmitter whose no-override branch (_helpers.py) captures `bulk` (~25k-molecule array) + four unique- molecule node arrays every tick into an unbounded history list that is never read. colonies-02's no-emitter probe nulled only the outer colony emitter, missing the per-cell inner emitters.
  • Setting the inner RAMEmitter to 'minimal' capture (global_time + listeners only) for embedded cells bounds per-cell RSS. A single EcoliWCM built through the bridge (set_ram_emitter_capture('minimal')) holds RSS essentially flat.
  • The fix only narrows what the inner RAMEmitter records (a read-only history sink); the emit_schema does not feed back into simulation state, so division timing and masses are unchanged by construction. The full colony re-run (first division still tick 2338) is the empirical re-confirmation, deferred with the no-OOM run above.
  • The inner-RAMEmitter fix (F-01/F-02) is real and bounds a single bare WCM, but it is not the dominant leak for a multi-cell colony. Re-running colonies-02 on current main with all three contributory fixes live (#239 inner emitter, emit_cells=False outer emitter, viva-munk #11 pymunk shape churn) still leaks ~0.5-1 MB/tick/cell: the count=6 plateau climbed 5.1 -> 22.3 GB at a constant 6 cells and OOMs ~gen 3, unchanged from before the fixes. Refined by the 2026-07-26 leak hunt: the "C-level dominant leak" call stands as a macOS observation but is not an unbounded native leak. Per-process attribution localizes it to elongation (~82%) + mass-listener (~17%); the C-level growth is macOS allocator arena retention of elongation's per-tick numpy arrays plus ~0.2 MB/tick of retained scipy LSODA work arrays. On Linux/glibc it is bounded (0.017 unmitigated, 0.000 with malloc_trim). See F-06.
  • The dominant per-tick growth is native and localized to the LLVM/Numba JIT layer. A controlled staged run (colonies-02 sims/mem_measure.py) shows numpy-object bytes flat within each cell-count plateau while native memory climbs (decomposition figure), and macOS `leaks` (MallocStackLogging) finds the true unreachable leak with stacks in llvmlite's LLVM pass pipeline (LLVMPY_buildFunctionSimplificationPipeline). So the WCM's JIT compilation is the source, not the emitters (#239), the outer emitter (emit_cells), the pymunk shapes (viva-munk #11), or a retained numpy array. Refined by the 2026-07-26 leak hunt: the LLVM/Numba JIT stacks that macOS `leaks` surfaced were a one-time/small true-leak signal (~13 MB), not the per-tick driver. The per-tick growth is intrinsic to comp.run() elongation (~82%) + mass-listener (~17%): macOS arena retention of per-tick numpy arrays plus ~0.2 MB/tick of scipy LSODA work arrays. It is Linux-bounded. So #253 (JIT recompilation) is not the load-bearing fix; the actionable targets are periodic malloc_trim(0) and scipy-LSODA solver reuse. See F-06.
  • The colony per-tick RSS growth is localized and Linux-bounded, not an unbounded native leak (this is the corrected diagnosis that reconciles F-01/F-02/F-04/F-05). Per-process RSS attribution on the mini pins the growth to the inner WCM's polypeptide-elongation step (~82%) + mass-listener (~17%). Two components: (1) a small real retention ~0.2 MB/tick (tracemalloc-visible) = scipy LSODA integrator work arrays (rwork/iwork) held per solve from solve_ivp(method="LSODA") in the equilibrium / two-component / tRNA-charging ODEs (~0.6 GB over a ~3000-tick cycle); (2) the scary ~1.4 MB/tick = macOS allocator arena retention of elongation's large per-tick numpy working arrays (buildSequences/polymerize, ~1.3 MB), freed but not returned to the OS, invisible to tracemalloc β€” a macOS artifact. The earlier "emitter is 94% of the leak" reading was a warm-process artifact and is wrong.
Evidence
  • Live profiler measured the inner RAMEmitter history at ~5.15 MB/row, growing 1:1 with ticks (9.9 MB @ 2 rows -> 9324 MB @ 1802 rows), tracking the RSS climb. Confirmed by reading process_bigraph Emitter.inputs() == config['emit'] (the emitter only stores its emit_schema ports).
  • Bridge-path profile_leak.py (commit bd1a11fe) RSS: 1.3 GB @ 21 min -> 1.7 GB @ 48 min -> 1.9 GB @ 62 min (~0.25 MB/s, legitimate cell growth), vs the unfixed path +24 GB by tick 1800 (~7.7 MB/s). ~30x reduction; no monotonic emitter-driven climb.
  • tracemalloc(4) over 150 ticks at 2 forced-division cells: RSS climbed 1396 -> 1586 MB (+1.3 MB/tick), leak reproduced, yet tracemalloc accounted for only ~1-2 MB of retained Python allocations (largest single site +0.08 MB; all transient WCM compute: counts_deriver/translation_deriver np.bincount, scipy-sparse .dot). So the dominant allocator is C-level / native (numpy/scipy array buffers or a C-extension retaining or fragmenting memory), invisible to tracemalloc (it tracks only CPython's allocator, not numpy's). gc object-count scans likewise found only small symptoms (pymunk Segments -> #11), not the megabytes.
  • leaks @2cells/400ticks: ~846 MB live, ~13 MB true leak, top stacks all LLVM PassBuilder (GVN/JumpThreading/LoopUnswitch/LICM). numpy-bytes delta ~0 within plateaus; tracemalloc <2 MB of a +151 MB/300-tick growth. Figures: colonies-02 charts/03_memory_decomposition.svg (native vs numpy), charts/01_rss_vs_tick_by_ncells.svg (per-tick slope by N).
  • Linux/glibc validation (colima container, faithful reproduction of both sources): RSS grew 0.017 MB/tick WITHOUT any fix and 0.000 MB/tick WITH periodic malloc_trim(0), versus ~1.6 MB/tick on the macOS mini. Per-process attribution: elongation ~82%, mass-listener ~17%. tracemalloc localizes the ~0.2 MB/tick retained component to scipy/integrate/_ode.py work arrays.

Pipeline-gate decision

Ready to run
βœ“ Passednothing yet
βœ— Failednothing failing
β›” Blocks the next studynothing blocking
β†’ Immediate next action
Execute the simulation_set to gather evidence.
Pipeline gate & conclusion logic (technical)

Prerequisites: none (root study)

Enables: β€”

Proceed when: Met: the 2026-07-26 hunt localizes >=~80% of the growth to named sites (elongation ~82% + mass-listener ~17%; real retention = scipy LSODA work arrays) and shows the growth is Linux-bounded (0.017 MB/tick unmitigated, 0.000 with malloc_trim). Gate proceeds.

4.Device harness + simple-agent phenotype baselineπŸ§ͺ Preliminary
β—‹ Not runNo tests declaredβ—‹ Not run
Does the shared harness β€” cell-tier factory, geometry builders, and tier-agnostic phenotype extractor β€” recover a correct, uniform single-cell phenotype panel (growth rate, size-at-division, added length, inter-division time, adder size-homeostasis) across the three device geometries, validated first with cheap simple viva-munk agents before spending WCM compute?
colonies-04-device-phenotype-harness Β· depth 0
β–Έ click to expand full study
4.colonies-04-device-phenotype-harness BuilddesignedRoot study (no dependencies)
β–΄ click to collapse full study
PLANNING β€” not yet run
Model
The composite(s) this study runs and their parameters.
🧬 ecoli_colonyseed=0 n_cells=2 env_size=30

Overview

This study asks whether does the shared harness β€” cell-tier factory, geometry builders, and. No simulations have run yet β€” the study is still in its design phase. Gate decision: Ready to run. Execute the simulation_set to gather evidence.

Purpose & background (study design)
Question. Does the shared harness β€” cell-tier factory, geometry builders, and tier-agnostic phenotype extractor β€” recover a correct, uniform single-cell phenotype panel (growth rate, size-at-division, added length, inter-division time, adder size-homeostasis) across the three device geometries, validated first with cheap simple viva-munk agents before spending WCM compute?
Mechanism / Model change. The factory instantiates any cell tier (simple / surrogate / wcm) into a device built by the geometry builders; the tier-agnostic extractor reads the same phenotype panel regardless of tier, and the ColonyPhenotypeRecorder process streams that panel per cell to zarr as a bounded stream (a cheap view of the cells map rather than an emitter deep-copy).
Expected outcome. The factory yields a runnable cell per tier, the geometry builders run simple agents in all three devices, the extractor recovers known division statistics, and the harness emits the common phenotype panel β€” the shared measurement substrate the WCM-tier device studies (colonies-08/09) reuse. No biological conclusion is drawn at this enabling tier.

Conditions β€” what we set up to test it

Baseline

Composite: v2ecoli.composites.ecoli_colony.ecoli_colony
seed0
n_cells2
env_size30
Technical context (model changes Β· implementation tasks Β· follow-ups Β· limitations Β· refs)
5.Mother machine (viva-munk) β€” device run + phenotype distributionsπŸ§ͺ Preliminary
β—‹ Not runNo tests declaredβ—‹ Not run
Do the core emergent single-cell phenotypes β€” size-at-division, inter-division time, added length, the adder size-homeostasis relation, and growth rate β€” stay measurable when simple viva-munk agents run in the canonical mother-machine device geometry?
colonies-05-mother-machine Β· depth 0
β–Έ click to expand full study
5.colonies-05-mother-machine SimulateranRoot study (no dependencies)
β–΄ click to collapse full study
PLANNING β€” not yet run
Model
The composite(s) this study runs and their parameters.
🧬 ecoli_colonyseed=0 n_cells=8

Overview

This study asks whether do the core emergent single-cell phenotypes β€” size-at-division,. No simulations have run yet β€” the study is still in its design phase. Gate decision: Ready to run. Execute the simulation_set to gather evidence.

Purpose & background (study design)
Question. Do the core emergent single-cell phenotypes β€” size-at-division, inter-division time, added length, the adder size-homeostasis relation, and growth rate β€” stay measurable when simple viva-munk agents run in the canonical mother-machine device geometry?
Mechanism / Model change. The mother_machine_document geometry (via v2ecoli.colony_bench.devices) is populated with simple grow/divide agents; colony_bench.phenotypes extracts the uniform phenotype panel over the lineage, and the resulting distributions are the object of study.
Expected outcome. The device runs and divides, the running animation renders, and the size-at- division, inter-division-time, and added-length distributions (plus the adder plot) are produced and held for later comparison against the WCM tier (colonies-09) and experimental data. This tier makes no biological claim β€” it validates the measurement pipeline against cheap agents.

Conditions β€” what we set up to test it

Baseline

Composite: v2ecoli.composites.ecoli_colony.ecoli_colony
seed0
n_cells8
BASELINECharts below show the workspace pre-execution baseline β€” what the system looks like before any of this study's variants run. Expert reviewers: comment on whether these traces look right for the starting point.

Visualisations from the latest run

mother machine: added size before division❓ untracked
added_size
Adder plot β€” added length Ξ” vs length at birth (a flat trend is the adder signature), with the Ξ” distribution.
mother machine: running animation❓ untracked
colony
Cells as capsules (lineage-coloured); shaded band = wash-out boundary.
mother machine: time between birth and division❓ untracked
interdivision_time
Distribution of the interval from a cell's birth to its own division (minutes).
mother machine: size at division❓ untracked
size_at_division
Distribution of cell length (Β΅m) at the moment of division.
Technical context (model changes Β· implementation tasks Β· follow-ups Β· limitations Β· refs)
6.Daughter machine (viva-munk) β€” device run + phenotype distributionsπŸ§ͺ Preliminary
β—‹ Not runNo tests declaredβ—‹ Not run
Do the core emergent single-cell phenotypes β€” size-at-division, inter-division time, added length, the adder size-homeostasis relation, and growth rate β€” stay measurable when simple viva-munk agents run in the canonical daughter-machine device geometry?
colonies-06-daughter-machine Β· depth 0
β–Έ click to expand full study
6.colonies-06-daughter-machine SimulateranRoot study (no dependencies)
β–΄ click to collapse full study
PLANNING β€” not yet run
Model
The composite(s) this study runs and their parameters.
🧬 ecoli_colonyseed=0 n_cells=1

Overview

This study asks whether do the core emergent single-cell phenotypes β€” size-at-division,. No simulations have run yet β€” the study is still in its design phase. Gate decision: Ready to run. Execute the simulation_set to gather evidence.

Purpose & background (study design)
Question. Do the core emergent single-cell phenotypes β€” size-at-division, inter-division time, added length, the adder size-homeostasis relation, and growth rate β€” stay measurable when simple viva-munk agents run in the canonical daughter-machine device geometry?
Mechanism / Model change. The daughter_machine_document geometry (via v2ecoli.colony_bench.devices) is populated with simple grow/divide agents; colony_bench.phenotypes extracts the uniform phenotype panel over the lineage, and the resulting distributions are the object of study.
Expected outcome. The device runs and divides, the running animation renders, and the size-at- division, inter-division-time, and added-length distributions (plus the adder plot) are produced and held for later comparison against the WCM tier (colonies-08) and experimental data. This tier makes no biological claim β€” it validates the measurement pipeline against cheap agents.

Conditions β€” what we set up to test it

Baseline

Composite: v2ecoli.composites.ecoli_colony.ecoli_colony
seed0
n_cells1
BASELINECharts below show the workspace pre-execution baseline β€” what the system looks like before any of this study's variants run. Expert reviewers: comment on whether these traces look right for the starting point.

Visualisations from the latest run

daughter machine: added size before division❓ untracked
added_size
Adder plot β€” added length Ξ” vs length at birth (a flat trend is the adder signature), with the Ξ” distribution.
daughter machine: running animation❓ untracked
colony
Cells as capsules (lineage-coloured); shaded band = wash-out boundary.
daughter machine: time between birth and division❓ untracked
interdivision_time
Distribution of the interval from a cell's birth to its own division (minutes).
daughter machine: size at division❓ untracked
size_at_division
Distribution of cell length (Β΅m) at the moment of division.
Technical context (model changes Β· implementation tasks Β· follow-ups Β· limitations Β· refs)
7.Whole-cell baseline (v2ecoli) in a daughter machine β€” device runπŸ§ͺ Preliminary
β—‹ Not runNo tests declaredβ—‹ Not run
Run the REAL v2ecoli whole-cell BASELINE \u2014 the full 55-process EcoliWCM \u2014 as a single cell inside the daughter-machine device geometry (chamber + absorbing wall), letting it grow and divide naturally.
colonies-08-wcm-daughter-machine Β· depth 0
β–Έ click to expand full study
7.colonies-08-wcm-daughter-machine SimulateranRoot study (no dependencies)
β–΄ click to collapse full study
PLANNING β€” not yet run
Model
The composite(s) this study runs and their parameters.
🧬 ecoli_colonyseed=0 n_cells=1

Overview

This study asks whether run the REAL v2ecoli whole-cell BASELINE \u2014 the full 55-process EcoliWCM \u2014 as a. No simulations have run yet β€” the study is still in its design phase. Gate decision: Ready to run. Execute the simulation_set to gather evidence.

Purpose & background (study design)
Question. Run the REAL v2ecoli whole-cell BASELINE \u2014 the full 55-process EcoliWCM \u2014 as a single cell inside the daughter-machine device geometry (chamber + absorbing wall), letting it grow and divide naturally. This is the highest-fidelity rung of the phenotype ladder (simple viva-munk agent -> growth surrogate -> full WCM): it asks whether the emergent single-cell phenotypes \u2014 adder size-homeostasis, size at division, added length before division, inter-division time, and growth rate \u2014 survive whole-cell colony embedding and division inside the device. Renders the running animation and the same phenotype figures as the cheap-agent device studies, so the whole-cell tier can be compared, in the same geometry and on the identical measurement pipeline, against the simple viva-munk agents (colonies-06) and against experimental microfluidic-device data.

A bounded per-cell phenotype panel (size, added length, division/birth times, growth rate) is streamed to zarr by the ColonyPhenotypeRecorder \u2014 a process, not an emitter, so recording stays cheap. Memory is characterised and Linux-bounded, not a blocker: the inner WCM's per-tick RSS growth is localized to the polypeptide-elongation step + mass-listener and is dominated on the dev mini by a macOS allocator artifact (~1.6 MB/tick), but is only ~0.017 MB/tick on glibc and ~0.000 with periodic malloc_trim. WCM-tier runs are therefore length-capped on the dev mini yet effectively unbounded on the Linux HPC target. On the mini this run spans only a few generations, so phenotype distributions are PRELIMINARY (small n); full distributions await the Linux HPC run and experimental data (colonies-09). The value here is the genuine whole-cell baseline running inside the device.

Conditions β€” what we set up to test it

Baseline

Composite: v2ecoli.composites.ecoli_colony.ecoli_colony
seed0
n_cells1
BASELINECharts below show the workspace pre-execution baseline β€” what the system looks like before any of this study's variants run. Expert reviewers: comment on whether these traces look right for the starting point.

Visualisations from the latest run

whole-cell daughter machine: added size before division❓ untracked
added_size
Adder plot β€” added length Ξ” vs length at birth (a flat trend is the adder signature), with the Ξ” distribution.
whole-cell daughter machine: running animation❓ untracked
colony
Real v2ecoli whole-cell agents (full 55-process model) growing and dividing naturally inside the device; lineage-coloured.
whole-cell daughter machine: time between birth and division❓ untracked
interdivision_time
Distribution of the interval from a cell's birth to its own division (minutes).
whole-cell daughter machine: size at division❓ untracked
size_at_division
Distribution of cell length (Β΅m) at the moment of division.
Technical context (model changes Β· implementation tasks Β· follow-ups Β· limitations Β· refs)
8.Whole-cell baseline (v2ecoli) in a mother machine β€” device runπŸ§ͺ Preliminary
β—‹ Not runNo tests declaredβ—‹ Not run
Run the REAL v2ecoli whole-cell BASELINE \u2014 the full 55-process EcoliWCM \u2014 as one whole cell per channel inside the mother-machine device geometry (narrow dead-end channels + a flow channel that washes out cells crossing the top), letting them grow and divide naturally.
colonies-09-wcm-mother-machine Β· depth 0
β–Έ click to expand full study
8.colonies-09-wcm-mother-machine SimulateranRoot study (no dependencies)
β–΄ click to collapse full study
PLANNING β€” not yet run
Model
The composite(s) this study runs and their parameters.
🧬 ecoli_colonyseed=0 n_cells=2

Overview

This study asks whether run the REAL v2ecoli whole-cell BASELINE \u2014 the full 55-process EcoliWCM \u2014 as. No simulations have run yet β€” the study is still in its design phase. Gate decision: Ready to run. Execute the simulation_set to gather evidence.

Purpose & background (study design)
Question. Run the REAL v2ecoli whole-cell BASELINE \u2014 the full 55-process EcoliWCM \u2014 as one whole cell per channel inside the mother-machine device geometry (narrow dead-end channels + a flow channel that washes out cells crossing the top), letting them grow and divide naturally. This is the highest-fidelity rung of the phenotype ladder (simple viva-munk agent -> growth surrogate -> full WCM) in the confined-channel device: it asks whether the emergent single-cell phenotypes \u2014 adder size-homeostasis, size at division, added length before division, inter-division time, and growth rate \u2014 survive whole-cell colony embedding and division under channel confinement. Renders the running animation and the same phenotype figures as the cheap-agent device studies, so the whole-cell tier can be compared, in the same confined-channel device and on the identical measurement pipeline, against the simple viva-munk agents (colonies-05) and against experimental microfluidic-device data.

Heavier than the daughter-machine whole-cell run (colonies-08): N whole cells run simultaneously (one per channel), kept small (2 channels). A bounded per-cell phenotype panel (size, added length, division/birth times, growth rate) is streamed to zarr by the ColonyPhenotypeRecorder \u2014 a process, not an emitter, so recording stays cheap. Memory is characterised and Linux-bounded, not a blocker: the inner WCM's per-tick RSS growth is localized to the polypeptide-elongation step + mass-listener and is dominated on the dev mini by a macOS allocator artifact (~1.6 MB/tick), but is only ~0.017 MB/tick on glibc and ~0.000 with periodic malloc_trim. WCM-tier runs are therefore length-capped on the dev mini yet effectively unbounded on the Linux HPC target. On the mini this run spans only a few generations, so phenotype distributions are PRELIMINARY (small n); full distributions await the Linux HPC run and experimental data (colonies-10). The value is the whole-cell baseline running inside the confined-channel device.

Conditions β€” what we set up to test it

Baseline

Composite: v2ecoli.composites.ecoli_colony.ecoli_colony
seed0
n_cells2
BASELINECharts below show the workspace pre-execution baseline β€” what the system looks like before any of this study's variants run. Expert reviewers: comment on whether these traces look right for the starting point.

Visualisations from the latest run

whole-cell mother machine: added size before division❓ untracked
added_size
Adder plot β€” added length Ξ” vs length at birth (a flat trend is the adder signature), with the Ξ” distribution.
whole-cell mother machine: running animation❓ untracked
colony
Real v2ecoli whole-cell agents (full 55-process model) growing and dividing naturally inside the device; lineage-coloured.
whole-cell mother machine: time between birth and division❓ untracked
interdivision_time
Distribution of the interval from a cell's birth to its own division (minutes).
whole-cell mother machine: size at division❓ untracked
size_at_division
Distribution of cell length (Β΅m) at the moment of division.
Technical context (model changes Β· implementation tasks Β· follow-ups Β· limitations Β· refs)

Appendices

Method-grading and verification detail β€” kept at the back, after the main narrative.

How the verdict is computed β€” acceptance criteria & gating matrix
AC β†’ study gating matrix which study gates each acceptance criterion Β· ⚠ = no study linked (gap)

Each acceptance criterion is a behaviour test declared in a study: a measured field from the run (e.g. closure_gap_size) compared against an explicit pass_if band (a numeric threshold/range). The per-criterion result, each study’s gate verdict, and this roll-up are computed in code from the run outcomes (deterministic) β€” not human judgement. Expand a row to see the field, the passing band, and the observed value.

Acceptance criterionGating studyResult
daughters-hydrated
field post_division_advance · passes if {"op":"all_daughters_advance"}
After the first division event in the build-smoke-n2 run, both daughter agent IDs appear in the cells map AND both continue to advance (mass changes, EcoliWCM.update called) on subsequent ticks without the manual standalone-WCM workaround that exists in reports/colony_report.py.
colonies-01-hpc-readiness◐ in-progress
per-cell-cost-within-2x-reference
field per_cell_wall_ratio · passes if {"op":"ratio_at_most","ratio":2}
For each N>1 sweep run, the per-cell wall-time (wall_seconds / N_cells_at_end) is at most 2Γ— the N=1 reference per-cell wall-time. Catches super-linear blow-up that would block HPC scaling.
colonies-01-hpc-readinessβœ… passing
natural-division-2-generations
One cell divides NATURALLY (no --force-divide) to 2, then both daughters divide to 4 β€” all cells advancing on subsequent ticks. Resolves colonies-01's PASS-narrow (forced-division-only) caveat.
colonies-02-parallel-multigen-perf◐ in-progress
ray-lifts-gil-ceiling
Under the Ray protocol, total wall for the growing colony stays sub-realtime past the sequential ~13-cell ceiling (target: scales with cores), with division timing + masses consistent vs sequential.
colonies-02-parallel-multigen-perf◐ in-progress
leak-localized
A profiling run attributes >=~80% of RSS growth to named sites so a fix is targeted, not speculative. MET (2026-07-26): per-process RSS attribution localizes the growth to elongation (~82%) + mass-listener (~17%); the real-retention component is the scipy LSODA integrator work arrays.
colonies-03-wcm-rss-leak◐ in-progress
factory-yields-runnable-cell-per-tiercolonies-04-device-phenotype-harness◐ in-progress
geometry-builders-run-with-simple-agentscolonies-04-device-phenotype-harness◐ in-progress
extractor-recovers-known-division-statscolonies-04-device-phenotype-harness◐ in-progress
harness-produces-phenotype-panelcolonies-04-device-phenotype-harness◐ in-progress
mother-machine-runs-and-dividescolonies-05-mother-machine◐ in-progress
running-animation-renderedcolonies-05-mother-machine◐ in-progress
size-at-division-distributioncolonies-05-mother-machine◐ in-progress
interdivision-time-distributioncolonies-05-mother-machine◐ in-progress
added-size-distributioncolonies-05-mother-machine◐ in-progress
daughter-machine-runs-and-dividescolonies-06-daughter-machine◐ in-progress
running-animation-renderedcolonies-06-daughter-machine◐ in-progress
size-at-division-distributioncolonies-06-daughter-machine◐ in-progress
interdivision-time-distributioncolonies-06-daughter-machine◐ in-progress
added-size-distributioncolonies-06-daughter-machine◐ in-progress
whole-cell-runs-and-divides-in-devicecolonies-08-wcm-daughter-machine◐ in-progress
running-animation-renderedcolonies-08-wcm-daughter-machine◐ in-progress
preliminary-phenotypes-extractedcolonies-08-wcm-daughter-machine◐ in-progress
whole-cell-runs-and-divides-in-devicecolonies-09-wcm-mother-machine◐ in-progress
running-animation-renderedcolonies-09-wcm-mother-machine◐ in-progress
preliminary-phenotypes-extractedcolonies-09-wcm-mother-machine◐ in-progress

All 25 acceptance criteria are linked to a gating study.

πŸ”¬ Evidence & rigor β€” how well the method defends its claims 2/6 investigation rigor dimensions addressed Β· 4 gap(s)

Deterministic feedback on how well the method defends its claims against a skeptical reader β€” a method-level judgement, distinct from the per-study model verdicts above. Computed from declared fields, not judged. Gaps are an invitation to add negative controls, replicate across seeds, weigh alternative explanations, state falsifiability, or add an adversarial study.

βœ—
Adversarial testing C10 C12 C15
no adversarial study β€” add one that tries to BREAK the criteria: mimic / parasitic-or-dependent / externally-maintained / random-cyclic systems that should NOT qualify
βœ“
Traceable methodology C9 C2 C14
capability ladder (study DAG) + explicit acceptance criteria + pass/fail gates + traceable findings β€” the reusable methodological contribution
βœ“
Falsification exposure C1
the framework has been shown to reject at least one system (a discriminating negative control, an adversarial study, or a non-passing result)
βœ—
Comparative framing C13
no competing theoretical frameworks compared (viability theory, organizational / constraint closure, active inference) β€” show the findings uniquely support this lens
βœ—
Hypothesis competition C6 C16
no competing hypotheses[] declared β€” state β‰₯2 rival explanations with predictions so the evidence can adjudicate between them
βœ—
Per-study rigor gaps C2 C4 C6
66 rigor gap(s) across 8 member study(ies)

Per-study rigor

colonies-01-hpc-readiness β€” 4/12 rigor dimensions addressed Β· 7 gap(s)
βœ“
Replication C4
6 replicate(s)
βœ—
Controls & calibration C1 C2 C4
no controls β€” declare a system that SHOULD fail the criteria (externally-maintained / -supplied) plus a clearly-passing / borderline case so the metric is calibrated, not just asserted
βœ—
Alternative hypotheses C6
no alternative hypotheses declared
⚠
Claim discipline C3
findings not tiered β€” label each observation / mechanism / interpretation
βœ—
Falsifiability C5 C1
criteria read as tailored-to-succeed β€” add a falsifiability note (study.falsifiability)
βœ“
Engineered vs emergent C7
no interpretation-tier claim that requires the distinction
βœ“
Limitations stated C8 C11
states what the result does not show
βœ—
Next steps next-steps
no discovery_implications or follow_up_studies β€” state what this study changes and what to investigate next (the Decide phase)
βœ—
Threshold provenance C9 C5
1 of 1 numeric band(s) declare neither cites nor pass_if.provenance.kind β€” state where the cutoff came from (theory/calibration/literature/expert/exploratory/post_hoc)
βœ—
Metric calibration ladder C4 C2 C20
no calibration_ladder declared β€” index controls[] by known_fail / known_pass / borderline / stress rungs so the metric is calibrated across its range, not just asserted
βœ—
Generality C22
no generality axes tested β€” add findings[].generality.axes_tested or a robustness parameter sweep so the claim's breadth is evidenced
βœ“
Run persistence persistence
1/5 run(s) persisted via an emitter (sqlite/parquet/xarray or a run-db reference)
colonies-02-parallel-multigen-perf β€” 4/12 rigor dimensions addressed Β· 7 gap(s)
βœ“
Replication C4
3 replicate(s)
βœ—
Controls & calibration C1 C2 C4
no controls β€” declare a system that SHOULD fail the criteria (externally-maintained / -supplied) plus a clearly-passing / borderline case so the metric is calibrated, not just asserted
βœ—
Alternative hypotheses C6
no alternative hypotheses declared
⚠
Claim discipline C3
findings not tiered β€” label each observation / mechanism / interpretation
βœ—
Falsifiability C5 C1
criteria read as tailored-to-succeed β€” add a falsifiability note (study.falsifiability)
βœ“
Engineered vs emergent C7
no interpretation-tier claim that requires the distinction
βœ—
Limitations stated C8 C11
no limitations / 'what this does not show' β€” add a short bound on the claim (scope/fidelity of the model, what is NOT demonstrated)
βœ—
Next steps next-steps
no discovery_implications or follow_up_studies β€” state what this study changes and what to investigate next (the Decide phase)
βœ“
Threshold provenance C9 C5
no numeric acceptance bands requiring provenance
βœ—
Metric calibration ladder C4 C2 C20
no calibration_ladder declared β€” index controls[] by known_fail / known_pass / borderline / stress rungs so the metric is calibrated across its range, not just asserted
βœ—
Generality C22
no generality axes tested β€” add findings[].generality.axes_tested or a robustness parameter sweep so the claim's breadth is evidenced
βœ“
Run persistence persistence
no runs recorded β€” nothing to persist via an emitter
colonies-03-wcm-rss-leak β€” 3/12 rigor dimensions addressed Β· 7 gap(s)
⚠
Replication C4
only 2 replicates β€” add seeds for a robustness claim
βœ—
Controls & calibration C1 C2 C4
no controls β€” declare a system that SHOULD fail the criteria (externally-maintained / -supplied) plus a clearly-passing / borderline case so the metric is calibrated, not just asserted
βœ—
Alternative hypotheses C6
no alternative hypotheses declared
⚠
Claim discipline C3
findings not tiered β€” label each observation / mechanism / interpretation
βœ—
Falsifiability C5 C1
criteria read as tailored-to-succeed β€” add a falsifiability note (study.falsifiability)
βœ“
Engineered vs emergent C7
no interpretation-tier claim that requires the distinction
βœ—
Limitations stated C8 C11
no limitations / 'what this does not show' β€” add a short bound on the claim (scope/fidelity of the model, what is NOT demonstrated)
βœ—
Next steps next-steps
no discovery_implications or follow_up_studies β€” state what this study changes and what to investigate next (the Decide phase)
βœ“
Threshold provenance C9 C5
no numeric acceptance bands requiring provenance
βœ—
Metric calibration ladder C4 C2 C20
no calibration_ladder declared β€” index controls[] by known_fail / known_pass / borderline / stress rungs so the metric is calibrated across its range, not just asserted
βœ—
Generality C22
no generality axes tested β€” add findings[].generality.axes_tested or a robustness parameter sweep so the claim's breadth is evidenced
βœ“
Run persistence persistence
no runs recorded β€” nothing to persist via an emitter
colonies-04-device-phenotype-harness β€” 3/12 rigor dimensions addressed Β· 9 gap(s)
βœ—
Replication C4
single run β€” no replication across seeds declared (add robustness.seeds or simulation_set.seeds)
βœ—
Controls & calibration C1 C2 C4
no controls β€” declare a system that SHOULD fail the criteria (externally-maintained / -supplied) plus a clearly-passing / borderline case so the metric is calibrated, not just asserted
βœ—
Alternative hypotheses C6
no alternative hypotheses declared
βœ—
Claim discipline C3
no findings recorded
βœ—
Falsifiability C5 C1
criteria read as tailored-to-succeed β€” add a falsifiability note (study.falsifiability)
βœ“
Engineered vs emergent C7
no interpretation-tier claim that requires the distinction
βœ—
Limitations stated C8 C11
no limitations / 'what this does not show' β€” add a short bound on the claim (scope/fidelity of the model, what is NOT demonstrated)
βœ—
Next steps next-steps
no discovery_implications or follow_up_studies β€” state what this study changes and what to investigate next (the Decide phase)
βœ“
Threshold provenance C9 C5
no numeric acceptance bands requiring provenance
βœ—
Metric calibration ladder C4 C2 C20
no calibration_ladder declared β€” index controls[] by known_fail / known_pass / borderline / stress rungs so the metric is calibrated across its range, not just asserted
βœ—
Generality C22
no generality axes tested β€” add findings[].generality.axes_tested or a robustness parameter sweep so the claim's breadth is evidenced
βœ“
Run persistence persistence
no runs recorded β€” nothing to persist via an emitter
colonies-05-mother-machine β€” 3/12 rigor dimensions addressed Β· 9 gap(s)
βœ—
Replication C4
single run β€” no replication across seeds declared (add robustness.seeds or simulation_set.seeds)
βœ—
Controls & calibration C1 C2 C4
no controls β€” declare a system that SHOULD fail the criteria (externally-maintained / -supplied) plus a clearly-passing / borderline case so the metric is calibrated, not just asserted
βœ—
Alternative hypotheses C6
no alternative hypotheses declared
βœ—
Claim discipline C3
no findings recorded
βœ—
Falsifiability C5 C1
criteria read as tailored-to-succeed β€” add a falsifiability note (study.falsifiability)
βœ“
Engineered vs emergent C7
no interpretation-tier claim that requires the distinction
βœ—
Limitations stated C8 C11
no limitations / 'what this does not show' β€” add a short bound on the claim (scope/fidelity of the model, what is NOT demonstrated)
βœ—
Next steps next-steps
no discovery_implications or follow_up_studies β€” state what this study changes and what to investigate next (the Decide phase)
βœ“
Threshold provenance C9 C5
no numeric acceptance bands requiring provenance
βœ—
Metric calibration ladder C4 C2 C20
no calibration_ladder declared β€” index controls[] by known_fail / known_pass / borderline / stress rungs so the metric is calibrated across its range, not just asserted
βœ—
Generality C22
no generality axes tested β€” add findings[].generality.axes_tested or a robustness parameter sweep so the claim's breadth is evidenced
βœ“
Run persistence persistence
no runs recorded β€” nothing to persist via an emitter
colonies-06-daughter-machine β€” 3/12 rigor dimensions addressed Β· 9 gap(s)
βœ—
Replication C4
single run β€” no replication across seeds declared (add robustness.seeds or simulation_set.seeds)
βœ—
Controls & calibration C1 C2 C4
no controls β€” declare a system that SHOULD fail the criteria (externally-maintained / -supplied) plus a clearly-passing / borderline case so the metric is calibrated, not just asserted
βœ—
Alternative hypotheses C6
no alternative hypotheses declared
βœ—
Claim discipline C3
no findings recorded
βœ—
Falsifiability C5 C1
criteria read as tailored-to-succeed β€” add a falsifiability note (study.falsifiability)
βœ“
Engineered vs emergent C7
no interpretation-tier claim that requires the distinction
βœ—
Limitations stated C8 C11
no limitations / 'what this does not show' β€” add a short bound on the claim (scope/fidelity of the model, what is NOT demonstrated)
βœ—
Next steps next-steps
no discovery_implications or follow_up_studies β€” state what this study changes and what to investigate next (the Decide phase)
βœ“
Threshold provenance C9 C5
no numeric acceptance bands requiring provenance
βœ—
Metric calibration ladder C4 C2 C20
no calibration_ladder declared β€” index controls[] by known_fail / known_pass / borderline / stress rungs so the metric is calibrated across its range, not just asserted
βœ—
Generality C22
no generality axes tested β€” add findings[].generality.axes_tested or a robustness parameter sweep so the claim's breadth is evidenced
βœ“
Run persistence persistence
no runs recorded β€” nothing to persist via an emitter
colonies-08-wcm-daughter-machine β€” 3/12 rigor dimensions addressed Β· 9 gap(s)
βœ—
Replication C4
single run β€” no replication across seeds declared (add robustness.seeds or simulation_set.seeds)
βœ—
Controls & calibration C1 C2 C4
no controls β€” declare a system that SHOULD fail the criteria (externally-maintained / -supplied) plus a clearly-passing / borderline case so the metric is calibrated, not just asserted
βœ—
Alternative hypotheses C6
no alternative hypotheses declared
βœ—
Claim discipline C3
no findings recorded
βœ—
Falsifiability C5 C1
criteria read as tailored-to-succeed β€” add a falsifiability note (study.falsifiability)
βœ“
Engineered vs emergent C7
no interpretation-tier claim that requires the distinction
βœ—
Limitations stated C8 C11
no limitations / 'what this does not show' β€” add a short bound on the claim (scope/fidelity of the model, what is NOT demonstrated)
βœ—
Next steps next-steps
no discovery_implications or follow_up_studies β€” state what this study changes and what to investigate next (the Decide phase)
βœ“
Threshold provenance C9 C5
no numeric acceptance bands requiring provenance
βœ—
Metric calibration ladder C4 C2 C20
no calibration_ladder declared β€” index controls[] by known_fail / known_pass / borderline / stress rungs so the metric is calibrated across its range, not just asserted
βœ—
Generality C22
no generality axes tested β€” add findings[].generality.axes_tested or a robustness parameter sweep so the claim's breadth is evidenced
βœ“
Run persistence persistence
no runs recorded β€” nothing to persist via an emitter
colonies-09-wcm-mother-machine β€” 3/12 rigor dimensions addressed Β· 9 gap(s)
βœ—
Replication C4
single run β€” no replication across seeds declared (add robustness.seeds or simulation_set.seeds)
βœ—
Controls & calibration C1 C2 C4
no controls β€” declare a system that SHOULD fail the criteria (externally-maintained / -supplied) plus a clearly-passing / borderline case so the metric is calibrated, not just asserted
βœ—
Alternative hypotheses C6
no alternative hypotheses declared
βœ—
Claim discipline C3
no findings recorded
βœ—
Falsifiability C5 C1
criteria read as tailored-to-succeed β€” add a falsifiability note (study.falsifiability)
βœ“
Engineered vs emergent C7
no interpretation-tier claim that requires the distinction
βœ—
Limitations stated C8 C11
no limitations / 'what this does not show' β€” add a short bound on the claim (scope/fidelity of the model, what is NOT demonstrated)
βœ—
Next steps next-steps
no discovery_implications or follow_up_studies β€” state what this study changes and what to investigate next (the Decide phase)
βœ“
Threshold provenance C9 C5
no numeric acceptance bands requiring provenance
βœ—
Metric calibration ladder C4 C2 C20
no calibration_ladder declared β€” index controls[] by known_fail / known_pass / borderline / stress rungs so the metric is calibrated across its range, not just asserted
βœ—
Generality C22
no generality axes tested β€” add findings[].generality.axes_tested or a robustness parameter sweep so the claim's breadth is evidenced
βœ“
Run persistence persistence
no runs recorded β€” nothing to persist via an emitter
πŸ“Š Framework scorecard framework-self metrics (n=14 investigations)

Framework-self metrics aggregated across every study and investigation in the workspace β€” how consistently the framework itself applies its own rigor practices (discriminating controls, emergent-mechanism labelling, threshold provenance, replication, verdict divergence, falsification exposure). Computed deterministically from declared fields by pbg_superpowers.rigor.framework_metrics.

Discriminating Controls0%0 / 59
Emergent Interpretationsβ€”0 / 0
Missing Mechanism Originβ€”0 / 0
Threshold Provenance39%56 / 144
Replication Coverage22%13 / 59
Ac Coverage100%93 / 93
Verdict Divergence19%11 / 59
Falsification Exposure73%43 / 59
Alternatives Excluded3%2 / 59
Emitter Coverage36%13 / 36

References (0 cited across this investigation)

Union of bibliography.bib_keys and per-behavior cites: across all studies in this investigation. Click DOI or link to open the source.