v2ecoli Baseline Showcase: from ecoli-sources to a calibrated whole cell active

Investigation report Β· v2ecoli-baseline-showcase Β· generated 2026-08-17 13:29 UTC Β· for expert review β€” results below reflect completed runs.

Investigation acceptance: in-progress. 0 of 6 acceptance criteria passing. code-computed from member-study verdicts

πŸ“‹ Executive summary in-progress A demonstration walkthrough of the v2ecoli pipeline: ecoli-sources to ParCa to a calibrated baseline to a perturbation to a next-direction decision.…

A demonstration walkthrough of the v2ecoli pipeline: ecoli-sources to ParCa to a calibrated baseline to a perturbation to a next-direction decision. Six studies rebuild the ParCa in full, run the wild-type baseline ensemble (multiseed/multigen, checked against doubling time, mass fractions, Toya 2010 FBA flux, and Schmidt/Wisniewski proteome), then compare variants and test large-ensemble equivalence.

Question. Can v2ecoli, starting from the raw ecoli-sources flat files, rebuild the ParCa in full, run a wild-type baseline single-cell→division ensemble that reproduces measured E. coli properties (doubling time, mass fractions, FBA fluxes, proteome), and then characterize a perturbation response? This is a demonstration walkthrough of the v2ecoli pipeline (ecoli-sources → ParCa → calibrated whole cell → perturbation), not a test of a scientific hypothesis. It also hands reviewers an explicit perturbation-choice decision.

Acceptance roll-up code-computed from member-study verdicts

Each acceptance criterion is a behaviour test declared in a study: a measured field from the run (e.g. closure_gap_size) compared against an explicit pass_if band (a numeric threshold/range). The per-criterion result, each study’s gate verdict, and this roll-up are computed in code from the run outcomes (deterministic) β€” not human judgement. Expand a row to see the field, the passing band, and the observed value.

StudyBehaviorMetric (field Β· pass-if β†’ observed)Result
showcase-1-parcaparca-rebuilds-full-51-conditions-from-ecoli-sourcesβ€”in-progress
showcase-2-baseline-figuresbaseline-ensemble-reproduces-wild-type-propertiesβ€”in-progress
showcase-3-variant-decidereviewer-selects-perturbation-variant-to-runβ€”in-progress
showcase-4-variant-comparisonfive-variant-sweep-ranks-perturbation-contrasts-vs-baselineβ€”in-progress
showcase-5-next-direction-decidereviewer-selects-next-direction-from-showcase-4-rankingβ€”in-progress
showcase-6-equivalence-largelarge-16x16-baseline-ensemble-equivalent-to-vecoli-within-toleranceβ€”in-progress
🧬 Biology β€” the mechanism this investigation models This showcase follows the full arc of building a whole-cell E. coli model from primary data. It starts from ecoli-sources -- curated flat files of measured E. coli biology…

This showcase follows the full arc of building a whole-cell E. coli model from primary data. It starts from ecoli-sources -- curated flat files of measured E. coli biology (gene/protein/reaction annotations, kinetic constants, expression and mass-fraction data). The Parameter Calculator (ParCa) turns those measurements into sim_data, a single self-consistent parameter set: it fits expression levels, RNA/protein counts, and metabolic parameters so that a simulated wild-type cell reproduces bulk physiology.

Running the calibrated model then simulates one cell from birth to division, with the emergent doubling time, mass fractions, metabolic fluxes, and proteome falling out of the mechanism rather than being imposed. Because single cells vary, the baseline is an ensemble over seeds and generations. Validation checks these emergent properties against independent measurements -- doubling time and mass fractions, Toya 2010 central-carbon fluxes, and the Schmidt/Wisniewski proteomes. A perturbation and a next-direction decision then show the calibrated cell being used, not just built.

πŸ”¬ Scientific argument 3 for Β· 0 against v2ecoli rebuilds and runs the wild-type baseline from ecoli-sources: the ParCa rebuilds in full (51 TF conditions) from the raw flat files, the…

Main claim. v2ecoli rebuilds and runs the wild-type baseline from ecoli-sources: the ParCa rebuilds in full (51 TF conditions) from the raw flat files, the wild-type baseline single-cell→ division ensemble runs to completion, and the resulting cell matches measured E. coli properties (doubling time, mass fractions, FBA fluxes vs Toya 2010, proteome vs Schmidt/Wisniewski).

Evidence for

  • The ParCa builds in full (--mode full, 51 TF conditions) from the ~133 ecoli-sources flat files in ~2.5 min on the mini, producing a complete cache bundle (initial_state.json + sim_data_cache.dill + metadata + cache_version) that reproduces parca_compare. [Target of showcase-1; sim deferred.]
  • The wild-type baseline single-cellβ†’division ensemble (multiseed + multigen, Ray-parallel dispatch, XArray/zarr emit) runs to completion and reproduces wild-type E. coli properties: doubling time in band, physiological mass fractions, FBA flux correlation to Toya 2010, proteome correlation to Schmidt 2016 / Wisniewski. [Target of showcase-2; sim deferred.]
  • v2ecoli characterizes a perturbation response: the showcase-4 5-variant x 2-seed sweep moves the cell off baseline in perturbation-specific ways. media-succinate halves dry mass (271 vs 492 fg), suppresses multifork replication (max oriC 2 vs 4), and decorrelates central-carbon flux from glucose-grown Toya-2010 (r 0.73 β†’ -0.06). ppGpp-off de-represses the ppGpp pool (mean 64.8 β†’ 120.6, 1.86x) and slows growth most (doubling +10.5 min). [Measured in showcase-4.]

Caveats

Open questions & decisions needed

βœ‹ Decisions needed from reviewers 2 items next: Which variant should we run to demonstrate v2ecoli's perturbation response?
  1. Which variant should we run to demonstrate v2ecoli's perturbation response?
    showcase-3-variant-decide commits no variant. The four candidate perturbations (ppgpp_regulation toggle OFF, parameter-sweep via linspace, media/condition change, dnaA expression knob) are listed under proposed_inputs and on the study's followup_study_proposals; reviewers pick which best demonstrates v2ecoli's response. showcase-4 later ran all four (see its measured ranking), so this gate remains only for the reviewer's narrative-highlight choice.
  2. Which direction should the showcase take next, given the showcase-4 ranking?
    showcase-5-next-direction-decide commits no direction. Four candidates are seeded from the showcase-4 ranking: (A) deepen the strongest contrast with a growth-condition/carbon-source panel building on media-succinate (mass halved, oriC 2 vs 4, Toya-flux collapse; one full ParCa cache per condition, ~2.5 min each); (B) chase the surprise via a ppGpp synthesis-vs-regulation 2x2 to pin why ppGpp-off de-repressed the pool (rose 1.86x), single-cache; (C) push the partial by sweeping the dnaA init-probability 1x..4x to test the oriC ceiling 2x did not raise; (D) rescue the null by sweeping the ribosome elongation rate over a wider range to find where steady-state-charging compensation breaks. The agent recommends A or B; none is committed.

Investigation roadmap

Study verdict map code-computed gate verdicts (βœ… passed Β· β›” failed Β· πŸ”„ needs calibration Β· ⚠ blocked Β· β—½ not evaluated)

Studies

Each study is collapsed to a one-glance control panel β€” scan top to bottom, then click any panel to expand its full detail.

1.Rebuild the ParCa in full from the ecoli-sources flat filesβœ… Passing
β–Ά Ran Β· 1 runTests: 3β³βœ… Passed⚠ 1 clarity note
Establish that the v2ecoli ParCa pipeline rebuilds in FULL from the raw ~133 ecoli-sources flat files: --mode full fits all 51 TF conditions and produces a complete, reproducible cache bundle that downstream studies (showcase-2 baseline) resume from.
showcase-1-parca Β· depth 0
Confidence: highEvidence: direct-run
Conclusion PASS.
4/4 tests passing
Insight v2ecoli rebuilds the ParCa in full from the ecoli-sources flat files in minutes, fitting all 51 TF conditions and emitting a complete, simulation- ready cache bundle.
β–Έ click to expand full study
Model
The composite(s) this study runs and their parameters.
🧬 parcadefault parameters

Biology

The ParCa (parameter-calculator) fits heterogeneous measured E. coli datasets into a single self-consistent sim_data parameter set across 51 transcription- factor conditions (basal + with_aa + acetate + succinate + no_oxygen + the active/inactive TF set). The full-mode fit is what calibrates the metabolite concentrations so the whole-cell composite can solve its equilibrium ODEs at steady state β€” the precondition for a simulation-ready wild-type baseline.

Literature anchors

The biological expectations this study tests, mapped to the model observable that will measure each one. Full citations live in the test cards.

SetupRan the v2ecoli ParCa pipeline with --mode full --cpus 8 on the 133 ecoli-sources flat files (v2ecoli/processes/parca/reconstruction/ecoli/flat) on the Mac mini (2026-06-09, branch report/elongation-3gen-parity-and-guards, .venv). Command: .venv/bin/python scripts/parca_run.py --mode full --cpus 8 \ -o out/sim_data-showcase --cache-dir out/sim_data-showcase/kmcache Then gzipped the 634 MB parca_state.pkl β†’ parca_state.pkl.gz (41 MB) and built the simulation-input bundle: .venv/bin/python scripts/build_cache.py \ --fixture out/sim_data-showcase/parca_state.pkl.gz \ --cache out/cache-showcase The cache lives in its OWN dir (out/cache-showcase), not the dnaa cache (the out-symlink hazard). Verified mode by condition count (51 = full, 7 = fast).
ResultPASS. Full ParCa completed in 142.4 s (2.4 min) β€” consistent with the ~2.5 min mini runtime; the docs' "4-8 hours / ~300 conditions" figure is stale. Per-step: step_4=47.1 s, step_5=55.8 s (the TF-fitting steps dominate). 51 TF conditions fitted (basal + with_aa + acetate + succinate + no_oxygen + the active/inactive TF set) β€” the full-mode proof (vs 7 in fast/debug). Cache bundle out/cache-showcase complete, all four artifacts present: initial_state.json 10.39 MB sim_data_cache.dill 164.92 MB metadata.json 210 B cache_version.json 1.1 KB (inputs_hash 36f7930256f9bc44…) initial_state: 16321 bulk molecule species + 11 unique molecule types across 4 top-level stores (bulk/unique/environment/boundary). sim_data structural inventory: 4538 genes (cistrons), 3277 transcription units, 4309 protein monomers, 1118 complexes, 9460 metabolic reactions. Smoke check: build_composite("ecoli_baseline", cache_dir="out/cache-showcase").run(5) ran 5 steps with NO "Could not solve ODEs in equilibrium to SS" crash β€” i.e. the metabolite concentrations are calibrated (the full-mode calibration the docs require for a simulation-ready fixture).
Interpretationv2ecoli rebuilds the ParCa in full from the ecoli-sources flat files in minutes, fitting all 51 TF conditions and emitting a complete, simulation- ready cache bundle. The 51-condition count and the no-crash 5-step smoke run together establish this is the full ParCa (not the fast/debug build that mis-calibrates dnaA/replication). showcase-2 can resume the baseline sim from out/cache-showcase.
DecisionPASS β‡’ the cache bundle is complete (51 conditions, all four artifacts present) and is simulation-ready (5-step smoke run, no equilibrium-solver crash) β‡’ unblock showcase-2-baseline-figures. Caveat: the full cross-tool parca_compare HTML diff (against vivarium-ecoli --save-intermediates) was NOT re-run this session (the vEcoli reference intermediates are not present on the mini), so the reproduces-parca_compare test is PARTIAL β€” verified at the structural-inventory + simulation-readiness level, not the per-step bit diff.

Overview

This study asks whether can v2ecoli rebuild the ParCa in full from the ~133 ecoli-sources flat. We recorded 1 finding confirm the expected biology. Gate decision: Passed. Gate cleared.

Purpose & background (study design)
Question. Can v2ecoli rebuild the ParCa in full from the ~133 ecoli-sources flat files? Running the ParCa pipeline with --mode full should fit all 51 TF conditions (the full-mode proof, vs ~7 in fast/debug mode) and emit a complete cache bundle (initial_state.json + sim_data_cache.dill + metadata + cache_version) that reproduces the parca_compare reference.

Visualizations

ParCa reconstruction summary (interactive)

Molecule and reaction counts in the fitted sim_data bundle, as an interactive bar chart.

Units Atlas (interactive)

Every declared unit-bearing baseline readout grouped by physical dimension. Folds in the former units-atlas investigation.

Detailed findings

Infrastructure / computational findings (1)

βœ“report-card-testsconfirmed
tests: within tolerance

Conclusion verdicts

Three-track verdict β€” each result is computed from canonical fields (gate evaluator, run status, finding tiers). The basis is the author's rationale.

Biological validationPASSfrom gate evaluator Β· computed
Regression compatibilityPASSfrom run status Β· computed
Explanatory gainPARTIALfrom interpretation-tier findings Β· computed

Discovery implications

Where this study's results leave the mechanism model β€” and what to investigate next.

βœ“ Resolved uncertainties

  • --mode full fits 51 TF conditions in ~2.4 min on the mini (the docs' 4-8 h / ~300 conditions figure is stale; confirmed 142.4 s, 51 conditions).
  • The full-mode cache bundle is simulation-ready (5-step baseline smoke run, no equilibrium-solver crash).

● Remaining uncertainties

  • Whether the full ParCa rebuild from ecoli-sources is bit-for-bit reproducible across machines, or only reproducible up to the parca_compare tolerance.
  • The cross-tool parca_compare HTML diff vs vivarium-ecoli --save-intermediates was not re-run (reference intermediates absent); reproduces-parca_compare is verified only at the structural-inventory + simulation-readiness level.

Conditions β€” what we set up to test it

Baseline

Composite: v2ecoli.composites.parca.parca
modefull
cache_dirout/cache-showcase

Variants (2)

Each variant is a perturbation of the baseline β€” typically a parameter override or a swapped composite. These define the runs that test the assumption.

VariantComposite / baseParameter overridesNotesRun
full-mode (51 TF conditions)v2ecoli.composites.parca.parca(no overrides)β€”vwb run study showcase-1-parca --variant full-mode (51 TF conditions)
fast/debug-mode (~7 TF conditions)v2ecoli.composites.parca.parca(no overrides)β€”vwb run study showcase-1-parca --variant fast/debug-mode (~7 TF conditions)

What we ran (1 simulation)

One row per concrete run: the model composite, what changes vs the reference baseline, the condition / length, and its status.

SimulationCompositeChanges vs baselineRunCLIStatus
showcase1-parca-full-2026-06-09parcareference baseline3 seedsvwb run study showcase-1-parcacomplete

Measurements (4 readouts)

Quantities we extract from each simulation run to evaluate the study's tests.

ReadoutStatusPathDescription
tf-condition-countβ€”β€”TF-condition count fit by the ParCa run (51 = full, ~7 = fast/debug)
cache-bundle-artifactsβ€”β€”cache-bundle artifact presence (initial_state.json + sim_data_cache.dill + metadata.json + cache_version.json)
simdata-structural-inventoryβ€”β€”sim_data structural inventory (4538 genes / 3277 TUs / 4309 monomers / 1118 complexes / 9460 metabolic reactions)
smoke-runβ€”β€”5-step build_composite('ecoli_baseline') smoke run β€” no equilibrium-solver crash

Visualisations from the latest run

showcase1_cache_bundle❓ untracked
2026-06-09T23:55:53.429538 image/svg+xml Matplotlib v3.10.9, https://matplotlib.org/
showcase1_simdata_summary❓ untracked
2026-06-09T23:56:54.132125 image/svg+xml Matplotlib v3.10.9, https://matplotlib.org/
showcase1_source_manifest❓ untracked
2026-06-09T23:55:16.017527 image/svg+xml Matplotlib v3.10.9, https://matplotlib.org/
source-file manifest
source-file manifest
sim_data summary (full ParCa proof)
sim_data summary (full ParCa proof)
cache-bundle contents
cache-bundle contents

Success criteria (3 tests β€” 3 ⏳ pending)

Each test makes a specific scientific claim with a machine-checkable criterion (measure + pass_if). Tests are now evaluated by code against the run (the run/outcome spine: RunReader β†’ evaluator): the pill shows the result, and the evidence line shows the measured value, whether it was computed by code or routed to an agent, and whether the code verdict agrees with the authored one (reconcile). ⏳ pending = the study hasn't run yet. Technical assertion + the exact evaluator are under "Technical details".

⏳ PASS β€” 51 TF conditions fitted (--mode full, 142.4 s on the mini, 2026-06-09).primary
Claim:
Test id: parca-builds-full-51-conditions
Technical detailsMeasure: condition_count
Pass condition: = 51
Python: vivarium_workbench/lib/expected_behavior.py β†’ evaluate(); measure kind condition_count via _series_for_simple_kind()/_measure(); op == via _check()
⏳ PASS β€” all four artifacts present in out/cache-showcase (initial_state.json 10.39 MB, sim_data_cache.dill 164.92 MB, metadata.json 210 B, cache_version.json 1.1 KB).primary
Claim:
Test id: cache-bundle-complete
Technical detailsMeasure: artifacts_present
Pass condition: {"op":"all_present","artifacts":["initial_state.json","sim_data_cache.dill","metadata","cache_version"]}
Python: vivarium_workbench/lib/expected_behavior.py β†’ evaluate(); measure kind artifacts_present via _series_for_simple_kind()/_measure(); op all_present via _check()
⏳ PARTIAL β€” verified at the structural-inventory + simulation-readiness level (4538 genes / 3277 TUs / 4309 monomers / 1118 complexes / 9460 metabolic reactions; build_composite("ecoli_baseline").run(5) ran with no equilibrium-solver crash). The full cross-tool parca_compare HTML diff was NOT re-run β€” it needs vivarium-ecoli --save-intermediates reference checkpoints, which are not present on the mini this session. primary
Claim:
Test id: sim_data-reproduces-parca-comparison
Technical detailsMeasure: parca_compare
Pass condition: {"op":"matches_reference"}
Python: vivarium_workbench/lib/expected_behavior.py β†’ evaluate(); measure kind parca_compare via _series_for_simple_kind()/_measure(); op matches_reference via _check()

Model changes

None β€” this is the unmodified full-mode ParCa build (the reference pipeline). The only configuration axis exercised is the build MODE: --mode full (51 TF conditions, used) vs fast/debug (~7 conditions, contrast only, never used for simulation because it mis-calibrates dnaA / replication).

Key assumptions

  • --mode full fits all 51 TF conditions from the ~133 ecoli-sources flat files and is fast (~2.5 min on the mini). The docs' "4-8 hours / ~300 conditions" figure is stale; verify by condition count (51 = full).
  • A complete cache bundle is initial_state.json + sim_data_cache.dill + metadata + cache_version. All four must be present for showcase-2 to resume.

Build / fix list (1)

Concrete engineering work to fully exercise this study.

scripts/parca_run.py --mode full --cpus 8 (the reconstruction pipeline); gzip of parca_state.pkl (634 MB β†’ 41 MB) and scripts/build_cache.py to emit the simulation-input bundle. Run via the v2ecoli .venv (bare python lacks unum). Cross-tool parca_compare additionally needs vivarium-ecoli --save-intermediates reference checkpoints (absent on the mini this session).

Conclusion synthesis

Read-only synthesis derived from the study's canonical fields (findings, limitations, follow-up proposals).

Claims
  • tests: within tolerance
Evidence
  • Report card verdict: within tolerance

Pipeline-gate decision

Passed
βœ“ Passed
  • PARCA-BUILDS-FULL-51-CONDITIONS
  • CACHE-BUNDLE-COMPLETE
  • parca-builds-full-51-conditions
  • cache-bundle-complete
βœ— Failednothing failing
β›” Blocks the next studynothing blocking
β†’ Immediate next action
Gate cleared.
Pipeline gate & conclusion logic (technical)

Prerequisites: none (root study)

Enables: β€”

If primary tests pass: 51 fitted TF conditions plus a no-crash 5-step smoke run confirm the metabolite concentrations are calibrated β€” the full-mode calibration a simulation-ready fixture requires.
If primary tests fail: true
2.Wild-type baseline single-cellβ†’division ensemble reproduces E. coli propertiesβœ… Passing
β–Ά Ran Β· 1 runTests: 4βœ“βœ… Passed
Establish that the v2ecoli wild-type baseline single-cell→division ensemble, resumed from the showcase-1 full ParCa cache, reproduces measured wild-type E. coli properties across a multiseed/multigen ensemble: doubling time in band, physiological mass fractions, FBA flux correlation to Toya 2010, and proteome correlation to Schmidt 2016 / Wisniewski 2014.
showcase-2-baseline-figures Β· depth 0
Confidence: highEvidence: multiseed-ensemble + native-analysis gallery
Conclusion Regenerated from the CLEAN 2-seed Γ— 3-generation re-run (gen1 agent=0 / gen2 agent=00 / gen3 agent=000 for both seeds).
4/4 tests passing
Insight The v2ecoli baseline reproduces wild-type single-cell physiology: clean 3-generation growth/division (cell_mass ramps), physiological biomass composition, and a proteome that correlates with Schmidt 2016 (rβ‰ˆ0.73) and Wisniewski 2014 (rβ‰ˆ0.60), consistent with the upstream vEcoli baseline.
β–Έ click to expand full study
Model
The composite(s) this study runs and their parameters.
🧬 baselinedefault parameters

Biology

The wild-type baseline is the unperturbed v2ecoli whole cell growing in glucose minimal medium. Reproducing the four measured-property classes β€” doubling time, biomass composition (mass fractions), central-carbon flux (vs Toya 2010), and the proteome (vs Schmidt 2016 / Wisniewski 2014) β€” across a multiseed/multigen ensemble is the evidence that the cell is "the same E. coli" the upstream WCM was validated against, and the reference that all showcase-4 perturbation variants are measured against.

Literature anchors

The biological expectations this study tests, mapped to the model observable that will measure each one. Full citations live in the test cards.

SetupComposite: baseline (v2ecoli.composites.ecoli_baseline.ecoli_baseline), resumed from the showcase-1 cache (out/cache-showcase, 51-condition full ParCa). The run is a MULTISEED + MULTIGEN ensemble: 2 seeds (0–1) Γ— 3 generations, dispatched Ray-parallel (run_seeds_parallel, 6 threads/worker) with XArray/zarr emit as the primary store PLUS a parquet sidecar; the parquet sidecar is the authoritative store (the seed_00 xarray store failed to write β€” Quantity 'shape' AttributeError β€” and seed_01's xarray captured only gens 1–2, but BOTH seeds' parquet is the clean [1,2,3] phylogeny: gen1 agent=0 / gen2 agent=00 / gen3 agent=000). 6500 steps / cell; total wall ~1228 s. The native analyses (v2ecoli/workflow/analyses/) consumed the parquet sidecar directly; sim_data was hydrated from the showcase ParCa fixture (out/sim_data-showcase/parca_state.pkl.gz via hydrate_sim_data_from_state) and the Schmidt/Wisniewski proteome validation data attached for the proteome correlation. Figures rendered to PNG+SVG via vl_convert (Altair) / matplotlib on the mac mini (scripts/render_showcase2.py).
ResultRegenerated from the CLEAN 2-seed Γ— 3-generation re-run (gen1 agent=0 / gen2 agent=00 / gen3 agent=000 for both seeds). Doubling time per DIVIDED generation is in-band and clean: gen-1 = 43 min (both seeds), gen-2 = 43 min (seed 0) / 52 min (seed 1), mean of the divided generations β‰ˆ 46 min. The cap-immune instantaneous-growth-rate measure (the runnable test, evaluated by pbg-compute-outcomes) gives 52.4 min across the full lineage (per-gen 51 / 52 / 54 min). The old gen-2 step-cap over-report (~104 min) is GONE β€” #173 ends gen-2 at its real division. The only residual span artifact is the LAST generation (gen-3), which is truncated at the 6500-step cap before it divides, so its raw span reads short (~10–19 min); the cap-immune measure and the divided-generation headline are unaffected. Mass fractions are physiological: protein 0.461 (per-gen 0.477 / 0.453 / 0.442), rRNA 0.105, tRNA 0.019, DNA 0.018; mean dry mass β‰ˆ467 fg, cell mass β‰ˆ1556 fg. The proteome correlates with Schmidt 2016 at r = 0.728 (n=2228 monomers) and Wisniewski 2014 at r = 0.602 (n=2226). The native-analysis gallery (11 figures) reproduces the clean 3-generation cell-mass ramps, the replication program (oriC 2β†’4β†’2 per division, twice across the seed-0 lineage), the ppGpp pool, and high tRNA charging. The central-carbon FBA-vs-Toya-2010 scatter is regenerated on the clean run: the baseline FBA fluxes correlate with the Toya 2010 C13-MFA measured central-carbon fluxes at Pearson R = 0.6990 (p = 2.1e-4, n = 23 reactions).
InterpretationThe v2ecoli baseline reproduces wild-type single-cell physiology: clean 3-generation growth/division (cell_mass ramps), physiological biomass composition, and a proteome that correlates with Schmidt 2016 (rβ‰ˆ0.73) and Wisniewski 2014 (rβ‰ˆ0.60), consistent with the upstream vEcoli baseline. The doubling-time artifact that the prior run carried is RESOLVED: with #173, gen-2 now ends at its own division (43–52 min in-band) instead of emitting to the step-window cap (the old ~104 min over-report is gone). The doubling_time_line / doubling_time_hist figures now show a clean gen-2; the only remaining span quirk is the truncated FINAL generation (gen-3 hits the 6500-step cap before dividing, so its raw span reads short). The headline uses the cap-immune instantaneous-growth doubling time (52.4 min) and the divided-generation division times (~46 min). The central-carbon FBA fluxes correlate with the Toya 2010 measured fluxes at Pearson R = 0.6990 (n = 23 reactions), consistent with the upstream vEcoli baseline.
DecisionPASS β€” the baseline ensemble reproduces wild-type properties. All four primary tests pass on the clean re-run (doubling time in-band, mass fractions physiological, FBA fluxes correlate with Toya 2010 at R = 0.6990, proteome correlates Schmidt/Wisniewski). Gate stays passed, unblocking showcase-3 (the reviewer variant decision).

Overview

This study asks whether does the baseline single-cell→division sim reproduce wild-type E. coli. We recorded 1 finding confirm the expected biology. Gate decision: Passed. Gate cleared.

Purpose & background (study design)
Question. Does the baseline single-cell→division sim reproduce wild-type E. coli properties across a multiseed/multigen ensemble? The wild-type baseline resumed from the showcase-1 full ParCa cache should produce a doubling time in band, physiological mass fractions, FBA fluxes that correlate with Toya 2010, and a proteome that correlates with Schmidt 2016 / Wisniewski.

Visualizations

Dry mass over the cell cycle (interactive)

Per-cell exponential growth and division across the 2-seed Γ— 3-gen baseline ensemble; y-axis in fg, hover unit-tagged.

Dry-mass composition (interactive)

Stacked protein / RNA / DNA / small-molecule mass (fg) over one representative cell cycle.

Instantaneous growth rate (interactive)

Specific growth rate as d(ln mass)/dt (1/s) per cell across the ensemble.

ppGpp and tRNA charging (interactive)

Stringent-response ppGpp (uM, left axis) against the mean charged-tRNA fraction (right axis) over the cell cycle.

Units Atlas (interactive)

Every declared unit-bearing baseline readout grouped by physical dimension with example magnitude + range. Folds in the former units-atlas investigation.

Detailed findings

Infrastructure / computational findings (2)

βœ“report-card-testsconfirmed
tests: within tolerance
◐report-card-vs-vecolipartial result
vs_vecoli: drift

Conclusion verdicts

Three-track verdict β€” each result is computed from canonical fields (gate evaluator, run status, finding tiers). The basis is the author's rationale.

Biological validationPASSfrom gate evaluator Β· computed
Regression compatibilityPASSfrom run status Β· computed
Explanatory gainPARTIALfrom interpretation-tier findings Β· computed

Discovery implications

Where this study's results leave the mechanism model β€” and what to investigate next.

βœ“ Resolved uncertainties

  • The central-carbon FBA-vs-Toya-2010 scatter is regenerated on the clean run β€” the Toya 2010 flux validation TSV (toya_2010_central_carbon_fluxes.tsv) was restored to v2ecoli/validation/ecoli/flat/ and build_validation_data now exposes validation_data.reactionFlux.toya2010fluxes. The baseline FBA fluxes correlate with the Toya 2010 C13-MFA measured fluxes at Pearson R = 0.6990 (p = 2.1e-4, n = 23 reactions).
  • The wild-type baseline reproduces the measured-property classes on the clean 3-generation re-run (doubling time in-band, physiological mass fractions, Schmidt rβ‰ˆ0.73 / Wisniewski rβ‰ˆ0.60) β€” no calibration trade-off appeared across them.
  • RESOLVED (was: gen-2 doubling-time step-cap over-report): #173 ends each non-final generation at its real division, so gen-2 now reads 43–52 min (in-band) instead of the old ~104 min cap-inflated value. The doubling_time_line / doubling_time_hist figures show a clean gen-2.
  • RESOLVED (reviewer: 'why is the doubling time going down? get more samples'): the apparent decline was the cap-truncated FINAL generation, not a real trend. doubling_time_line / doubling_time_hist now lead with a cap-immune instantaneous doubling-time trace t2 = ln(2)/mu from listeners.mass.instantaneous_growth_rate over the whole lineage (218 per-timestep samples vs the old 3 generational points): STABLE at ~44–52 min (mean ~52 min), no decline. Per-generation division-time markers are kept but the truncated final generation is excluded so it no longer reads as a downward trend.
  • RESOLVED (was: gen-3 mislabeled / folded phylogeny): the clean re-run has the correct lineage labelling β€” gen1 agent=0 / gen2 agent=00 / gen3 agent=000 for both seeds β€” so cell_mass now renders three distinct clean generations.

● Remaining uncertainties

  • The FINAL generation (gen-3) is truncated at the 6500-step emit cap before it divides, so its raw per-generation span reads short (~10–19 min). This no longer affects the headline figures: doubling_time_line / doubling_time_hist now plot the cap-immune instantaneous doubling time (t2 = ln(2)/mu, 52.4 min over the full lineage, matching the runnable test) and the per-generation division-time markers explicitly exclude the truncated final generation (heuristic: last generation of a lineage whose final emitted row lands at the 6500-step cap). A per-generation division-event listener would let the span-based markers identify a real division instead of relying on this cap heuristic.

Conditions β€” what we set up to test it

Baseline

Composite: v2ecoli.composites.ecoli_baseline.ecoli_baseline
seed0
cache_dirout/cache-showcase

Variants (1)

Each variant is a perturbation of the baseline β€” typically a parameter override or a swapped composite. These define the runs that test the assumption.

VariantComposite / baseParameter overridesNotesRun
wild-type baseline (no perturbation)v2ecoli.composites.ecoli_baseline.ecoli_baseline(no overrides)β€”vwb run study showcase-2-baseline-figures --variant wild-type baseline (no perturbation)

Model settings (3)

Parameters that need human input before the study runs. Edit a value on the dashboard's study-detail page (Build tab) and the next pbg_runner invocation will pick it up.

NameTypeDefaultCurrentRangeGateDescription
multiseed-multigen-ensembleβ€”awaiting expertβ€”optionalβ€”
ray-parallel-dispatchβ€”awaiting expertβ€”optionalβ€”
xarray-zarr-emitβ€”awaiting expertβ€”optionalβ€”

What we ran (1 simulation)

One row per concrete run: the model composite, what changes vs the reference baseline, the condition / length, and its status.

SimulationCompositeChanges vs baselineRunCLIStatus
showcase2-baseline-ensemblebaselinereference baseline29 seedsvwb run study showcase-2-baseline-figurescomplete

Measurements (5 readouts)

Quantities we extract from each simulation run to evaluate the study's tests.

ReadoutStatusPathDescription
doubling-timeβ€”β€”cap-immune doubling time t2 = ln(2)/mu averaged over the full lineage (min)
mass-fractionsβ€”β€”dry-mass fractions (protein / rRNA / tRNA / DNA) per generation
fba-flux-vs-toya2010β€”β€”central-carbon FBA flux vs Toya 2010 C13-MFA (Pearson R over 23 reactions)
proteome-correlationβ€”β€”proteome correlation vs Schmidt 2016 and Wisniewski 2014 (log10 monomer counts)
replication-and-regulationβ€”β€”replication program (oriC copy number over the cell cycle) and ppGpp / tRNA-charging traces

Visualisations from the latest run

Cell dry mass over the cell cycle❓ untracked
102030405060708090100Time (min)0100200300400500600700Dry Mass (fg)Variant 0 - Absolute Dry Mass102030405060708090100Time (min)0.00.51.01.52.0Normalized Dry MassVariant 0 - Normalized Dry Mass12GenerationMulti-Variant Cell Mass Analysis
Sawtooth single-cell growth and division across the 2-seed x 2-gen ensemble.
Simulations behind this chart. 2-seed x 2-gen wild-type baseline ensemble (showcase2-baseline-full), Ray-parallel, XArray/zarr + parquet emit.
What it means. Cell dry mass grows exponentially and halves at each division (gen-1 ~390->680 fg, dividing at ~43 min; gen-2 repeats the mass-halving). The clean sawtooth confirms exponential single-cell growth with division near a doubling of mass.
Central-carbon FBA fluxes vs Toya 2010 (Pearson R = 0.728)❓ untracked
2026-06-10T08:35:52.638878 image/svg+xml Matplotlib v3.10.9, https://matplotlib.org/
Simulated central-carbon FBA fluxes correlate with Toya 2010 C13-MFA at R=0.728.
Simulations behind this chart. 2-seed x 2-gen wild-type baseline ensemble (showcase2-baseline-full), Ray-parallel, XArray/zarr + parquet emit.
What it means. Mean simulated central-carbon reaction fluxes vs Toya 2010 13C-MFA measurements correlate at Pearson R = 0.728 (p = 8.3e-5, 23 reactions). The baseline reproduces measured central-carbon flux distribution.
Chromosome replication state across the cell cycle❓ untracked
2026-06-10T09:10:10.697848 image/svg+xml Matplotlib v3.10.9, https://matplotlib.org/
oriC count (1->2->4) and replication-fork positions vs time, both seeds.
Simulations behind this chart. 2-seed x 2-gen wild-type baseline ensemble (showcase2-baseline-full), Ray-parallel, XArray/zarr + parquet emit.
What it means. The red step trace shows the number of replication origins rising 1->2->4 at initiation events while blue points show replication-fork positions sweeping from oriC toward terC; division resets the state. The regular oriC-doubling and fork sweeps in both seeds confirm coordinated, once-per-cell-cycle replication initiation.
Distribution of doubling times across the ensemble❓ untracked
404142434445464748495051525354555657585960616263646566676869707172737475767778Instantaneous doubling time (min)02468101214161820Frequency (timesteps)52.4 minCap-immune instantaneous doubling-time distribution (n=218 timesteps, mean 52.4 min)
Doubling-time histogram across seeds and generations.
Simulations behind this chart. 2-seed x 2-gen wild-type baseline ensemble (showcase2-baseline-full), Ray-parallel, XArray/zarr + parquet emit.
What it means. The distribution shows the true ~43 min gen-1 mode plus a longer gen-2 mode inflated by the step-window cap. The physiologically meaningful gen-1 mode sits in the expected band for the baseline glucose condition.
Doubling time vs generation (per seed)❓ untracked
1.01.11.21.31.41.51.61.71.81.92.0generation0.00.20.40.60.81.01.21.41.61.82.02.2Doubling Time (hr)01lineage_seed
Gen-1 doubling ~43 min in band; gen-2 inflated by the emit-window cap.
Simulations behind this chart. 2-seed x 2-gen wild-type baseline ensemble (showcase2-baseline-full), Ray-parallel, XArray/zarr + parquet emit.
What it means. Per-generation doubling time from the emitted global_time span: gen-1 ~43 min (in band for the glucose condition), gen-2 ~63 min. The gen-2 inflation is an emit-window-cap artifact (gen-2 cells keep emitting after division), not a biological slowdown; the true division-event doubling time is ~44 min.
Biomass composition over the cell cycle (mass fractions)❓ untracked
051015202530354045Time (min)0.00.20.40.60.81.01.21.41.61.82.0Mass (normalized by t = 0 min)DNA (0.018)Dry (1.000)Protein (0.476)Small Mol (0.375)mRNA (0.005)rRNA (0.106)tRNA (0.019)SubmassBiomass components (average fraction of total dry mass in parentheses)
Submass trajectories stay at physiological E. coli proportions throughout the cycle.
Simulations behind this chart. 2-seed x 2-gen wild-type baseline ensemble (showcase2-baseline-full), Ray-parallel, XArray/zarr + parquet emit.
What it means. Sub-mass trajectories (protein, RNA, DNA, small molecules) hold steady physiological proportions: average dry-mass fractions protein 0.447, rRNA 0.103, tRNA 0.019, DNA 0.019, small molecules 0.406. Biomass composition is wild-type-like.
Biomass component areas (initial vs final cell)❓ untracked
2026-06-10T08:35:54.029135 image/svg+xml Matplotlib v3.10.9, https://matplotlib.org/
Voronoi treemap of biomass components, area proportional to mass, initial vs final cell.
Simulations behind this chart. 2-seed x 2-gen wild-type baseline ensemble (showcase2-baseline-full), Ray-parallel, XArray/zarr + parquet emit.
What it means. Each tile's area is proportional to that component's mass; protein dominates, followed by metabolites, rRNA and lipid. The initial and final panels are nearly identical, showing biomass composition is conserved as the cell grows and divides.
ppGpp pool over the cell cycle (both seeds)❓ untracked
2026-06-10T09:10:59.384316 image/svg+xml Matplotlib v3.10.9, https://matplotlib.org/
ppGpp concentration settles into a steady ~60-70 uM band.
Simulations behind this chart. 2-seed x 2-gen wild-type baseline ensemble (showcase2-baseline-full), Ray-parallel, XArray/zarr + parquet emit.
What it means. ppGpp concentration trajectories for both seeds: from a ~26 uM startup the pool rises into a steady ~60-70 uM band (median ~67 uM, max ~83 uM). The stable ppGpp pool indicates a balanced stringent-response signal at steady growth.
Proteome vs Schmidt 2015 (r=0.74) and Wisniewski 2014 (r=0.62)❓ untracked
0.00.51.01.52.02.53.03.54.04.55.05.5log10(Schmidt 2015 Counts + 1)0.00.51.01.52.02.53.03.54.04.55.05.5log10(Simulation Average Counts + 1)Schmidt 2015 β€” Pearson r: 0.730.00.51.01.52.02.53.03.54.04.55.05.5log10(Wisniewski 2014 Counts + 1)0.00.51.01.52.02.53.03.54.04.55.05.5log10(Simulation Average Counts + 1)Wisniewski 2014 β€” Pearson r: 0.61
Simulated proteome correlates with measured proteomics (Schmidt r=0.74).
Simulations behind this chart. 2-seed x 2-gen wild-type baseline ensemble (showcase2-baseline-full), Ray-parallel, XArray/zarr + parquet emit.
What it means. log10 simulated average monomer counts vs measured proteomics with a parity line: Schmidt 2015 Pearson r = 0.735 (n=2228) and Wisniewski 2014 r = 0.616 (n=2226). The simulated proteome matches measured E. coli abundances across ~2200 monomers.
Metabolic reaction-flux heatmap (reaction x timepoint)❓ untracked
2026-06-11T08:26:16.610244 image/svg+xml Matplotlib v3.10.9, https://matplotlib.org/
FBA reaction fluxes across the cell cycle as a reaction-by-timepoint heatmap.
Simulations behind this chart. 2-seed x 2-gen wild-type baseline ensemble (showcase2-baseline-full), Ray-parallel, XArray/zarr + parquet emit.
What it means. Per-reaction FBA fluxes over the cell cycle (seed 0, gen 1) as a reaction-by-timepoint heatmap; most fluxes are stable across the cycle with the expected central-carbon reactions carrying the highest flux. A qualitative ProtoolsViz-style metabolic overview.
Chromosome replication across a lineage (seed 0)❓ untracked
0.00.10.20.30.40.50.60.70.80.91.01.11.21.31.41.51.61.71.81.92.0Time (hr)-terCoriC+terCDNA polymerase position (nt)DNA Polymerase Positions0.00.10.20.30.40.50.60.70.80.91.01.11.21.31.41.51.61.71.81.92.0Time (hr)0246Pairs of forksPairs of Replication Forks0.00.10.20.30.40.50.60.70.80.91.01.11.21.31.41.51.61.71.81.92.0Time (hr)0200400600800Dry mass (fg)Dry Mass0.00.10.20.30.40.50.60.70.80.91.01.11.21.31.41.51.61.71.81.92.0Time (hr)01234Number of oriCNumber of oriC
Fork positions, fork pairs, oriC count and dry mass over a seed-0 lineage.
Simulations behind this chart. 2-seed x 2-gen wild-type baseline ensemble (showcase2-baseline-full), Ray-parallel, XArray/zarr + parquet emit.
What it means. Multi-panel replication view for the seed-0 lineage: DNA-polymerase fork positions sweep oriC->terC, replication-fork pairs and the oriC count rise 2->4 then reset to 2 at division, tracking dry-mass growth. Confirms the expected coordinated replication program.
tRNA charged fraction over the cell cycle (both seeds)❓ untracked
2026-06-10T09:11:44.200271 image/svg+xml Matplotlib v3.10.9, https://matplotlib.org/
Mean tRNA charging stays high (~0.97) throughout the cell cycle.
Simulations behind this chart. 2-seed x 2-gen wild-type baseline ensemble (showcase2-baseline-full), Ray-parallel, XArray/zarr + parquet emit.
What it means. Mean fraction of charged tRNA (averaged over the 86 tRNA species, shaded band = per-species range) stays high at ~0.97 (median 0.966) for both seeds across the cell cycle. High, stable charging is consistent with amino-acid-replete exponential growth.
cell mass over the cell cycle
cell mass over the cell cycle
Sawtooth single-cell growth and division across the 2-seed x 2-gen ensemble.
Simulations behind this chart. 2-seed x 2-gen wild-type baseline ensemble (showcase2-baseline-full), Ray-parallel, XArray/zarr + parquet emit.
What it means. Cell dry mass grows exponentially and halves at each division (gen-1 ~390->680 fg, dividing at ~43 min; gen-2 repeats the mass-halving). The clean sawtooth confirms exponential single-cell growth with division near a doubling of mass.
mass fraction summary
mass fraction summary
Submass trajectories stay at physiological E. coli proportions throughout the cycle.
Simulations behind this chart. 2-seed x 2-gen wild-type baseline ensemble (showcase2-baseline-full), Ray-parallel, XArray/zarr + parquet emit.
What it means. Sub-mass trajectories (protein, RNA, DNA, small molecules) hold steady physiological proportions: average dry-mass fractions protein 0.447, rRNA 0.103, tRNA 0.019, DNA 0.019, small molecules 0.406. Biomass composition is wild-type-like.
biomass composition Voronoi
biomass composition Voronoi
Voronoi treemap of biomass components, area proportional to mass, initial vs final cell.
Simulations behind this chart. 2-seed x 2-gen wild-type baseline ensemble (showcase2-baseline-full), Ray-parallel, XArray/zarr + parquet emit.
What it means. Each tile's area is proportional to that component's mass; protein dominates, followed by metabolites, rRNA and lipid. The initial and final panels are nearly identical, showing biomass composition is conserved as the cell grows and divides.
doubling time over the lineage
doubling time over the lineage
Gen-1 doubling ~43 min in band; gen-2 inflated by the emit-window cap.
Simulations behind this chart. 2-seed x 2-gen wild-type baseline ensemble (showcase2-baseline-full), Ray-parallel, XArray/zarr + parquet emit.
What it means. Per-generation doubling time from the emitted global_time span: gen-1 ~43 min (in band for the glucose condition), gen-2 ~63 min. The gen-2 inflation is an emit-window-cap artifact (gen-2 cells keep emitting after division), not a biological slowdown; the true division-event doubling time is ~44 min.
doubling time distribution
doubling time distribution
Doubling-time histogram across seeds and generations.
Simulations behind this chart. 2-seed x 2-gen wild-type baseline ensemble (showcase2-baseline-full), Ray-parallel, XArray/zarr + parquet emit.
What it means. The distribution shows the true ~43 min gen-1 mode plus a longer gen-2 mode inflated by the step-window cap. The physiologically meaningful gen-1 mode sits in the expected band for the baseline glucose condition.
chromosome replication
chromosome replication
Fork positions, fork pairs, oriC count and dry mass over a seed-0 lineage.
Simulations behind this chart. 2-seed x 2-gen wild-type baseline ensemble (showcase2-baseline-full), Ray-parallel, XArray/zarr + parquet emit.
What it means. Multi-panel replication view for the seed-0 lineage: DNA-polymerase fork positions sweep oriC->terC, replication-fork pairs and the oriC count rise 2->4 then reset to 2 at division, tracking dry-mass growth. Confirms the expected coordinated replication program.
chromosome-state snapshots
chromosome-state snapshots
oriC count (1->2->4) and replication-fork positions vs time, both seeds.
Simulations behind this chart. 2-seed x 2-gen wild-type baseline ensemble (showcase2-baseline-full), Ray-parallel, XArray/zarr + parquet emit.
What it means. The red step trace shows the number of replication origins rising 1->2->4 at initiation events while blue points show replication-fork positions sweeping from oriC toward terC; division resets the state. The regular oriC-doubling and fork sweeps in both seeds confirm coordinated, once-per-cell-cycle replication initiation.
central carbon metabolism (FBA flux vs Toya 2010)
central carbon metabolism (FBA flux vs Toya 2010)
Simulated central-carbon FBA fluxes correlate with Toya 2010 C13-MFA at R=0.728.
Simulations behind this chart. 2-seed x 2-gen wild-type baseline ensemble (showcase2-baseline-full), Ray-parallel, XArray/zarr + parquet emit.
What it means. Mean simulated central-carbon reaction fluxes vs Toya 2010 13C-MFA measurements correlate at Pearson R = 0.728 (p = 8.3e-5, 23 reactions). The baseline reproduces measured central-carbon flux distribution.
reaction-flux heatmap
reaction-flux heatmap
FBA reaction fluxes across the cell cycle as a reaction-by-timepoint heatmap.
Simulations behind this chart. 2-seed x 2-gen wild-type baseline ensemble (showcase2-baseline-full), Ray-parallel, XArray/zarr + parquet emit.
What it means. Per-reaction FBA fluxes over the cell cycle (seed 0, gen 1) as a reaction-by-timepoint heatmap; most fluxes are stable across the cycle with the expected central-carbon reactions carrying the highest flux. A qualitative ProtoolsViz-style metabolic overview.
proteome vs Schmidt 2016 / Wisniewski 2014
proteome vs Schmidt 2016 / Wisniewski 2014
Simulated proteome correlates with measured proteomics (Schmidt r=0.74).
Simulations behind this chart. 2-seed x 2-gen wild-type baseline ensemble (showcase2-baseline-full), Ray-parallel, XArray/zarr + parquet emit.
What it means. log10 simulated average monomer counts vs measured proteomics with a parity line: Schmidt 2015 Pearson r = 0.735 (n=2228) and Wisniewski 2014 r = 0.616 (n=2226). The simulated proteome matches measured E. coli abundances across ~2200 monomers.
ppGpp pool trace
ppGpp pool trace
ppGpp concentration settles into a steady ~60-70 uM band.
Simulations behind this chart. 2-seed x 2-gen wild-type baseline ensemble (showcase2-baseline-full), Ray-parallel, XArray/zarr + parquet emit.
What it means. ppGpp concentration trajectories for both seeds: from a ~26 uM startup the pool rises into a steady ~60-70 uM band (median ~67 uM, max ~83 uM). The stable ppGpp pool indicates a balanced stringent-response signal at steady growth.
tRNA charged fraction trace
tRNA charged fraction trace
Mean tRNA charging stays high (~0.97) throughout the cell cycle.
Simulations behind this chart. 2-seed x 2-gen wild-type baseline ensemble (showcase2-baseline-full), Ray-parallel, XArray/zarr + parquet emit.
What it means. Mean fraction of charged tRNA (averaged over the 86 tRNA species, shaded band = per-species range) stays high at ~0.97 (median 0.966) for both seeds across the cell cycle. High, stable charging is consistent with amino-acid-replete exponential growth.

Success criteria (4 tests β€” 4 βœ“ passed)

Each test makes a specific scientific claim with a machine-checkable criterion (measure + pass_if). Tests are now evaluated by code against the run (the run/outcome spine: RunReader β†’ evaluator): the pill shows the result, and the evidence line shows the measured value, whether it was computed by code or routed to an agent, and whether the code verdict agrees with the authored one (reconcile). ⏳ pending = the study hasn't run yet. Technical assertion + the exact evaluator are under "Technical details".

βœ“ PASSprimary
Claim:
Evidence:
measured_value: 52.40456
code computed PASS Β· op derived/in_range Β· by code authored PASS from run showcase2-baseline-ensemble
passes if derived in [35, 55]
52.4 in [35.0, 55.0]
Test id: doubling-time-in-band
Technical detailsMeasure: derived
Pass condition: in [35, 55]
Python: vivarium_workbench/lib/expected_behavior.py β†’ evaluate(); measure kind derived via _series_for_simple_kind()/_measure(); op in_range via _check()
Cites: macklin2020
βœ“ PASSprimary
Claim:
Evidence:
measured_value: 0.461785
code computed PASS Β· op generation_average/in_range Β· by code authored PASS from run showcase2-baseline-ensemble
passes if listeners.mass.protein_mass / listeners.mass.dry_mass (generation_average) in [0.4, 0.55]
0.4618 in [0.4, 0.55]
Test id: mass-fraction-physiological
Technical detailsMeasure: listeners.mass.protein_mass / listeners.mass.dry_mass (generation_average)
Pass condition: in [0.4, 0.55]
Python: vivarium_workbench/lib/expected_behavior.py β†’ evaluate(); measure kind generation_average via _series_for_simple_kind()/_measure(); op in_range via _check()
Cites: macklin2020
βœ“ PASSprimary
Claim:
Evidence:
measured_value: β€”
code computed by agent authored PASS from run showcase2-baseline-ensemble
passes if flux_correlation {"op":"correlates"}
non-run-data kind: 'flux_correlation'
Test id: fba-flux-correlates-toya2010
Technical detailsMeasure: flux_correlation
Pass condition: {"op":"correlates"}
Python: vivarium_workbench/lib/expected_behavior.py β†’ evaluate(); measure kind flux_correlation via _series_for_simple_kind()/_measure(); op correlates via _check()
Cites: toya2010
βœ“ PASSprimary
Claim:
Evidence:
measured_value: β€”
code computed by agent authored PASS from run showcase2-baseline-ensemble
passes if proteome_correlation {"op":"correlates"}
non-run-data kind: 'proteome_correlation'
Test id: proteome-correlates-schmidt-wisniewski
Technical detailsMeasure: proteome_correlation
Pass condition: {"op":"correlates"}
Python: vivarium_workbench/lib/expected_behavior.py β†’ evaluate(); measure kind proteome_correlation via _series_for_simple_kind()/_measure(); op correlates via _check()
Cites: schmidt2016, wisniewski2014

Model changes

None β€” this is the unperturbed wild-type baseline. The only configuration is the multiseed/multigen ensemble + Ray-parallel dispatch + parquet emit (see model_settings); no process parameters are altered.

Key assumptions

  • The wild-type baseline resumed from the showcase-1 full ParCa cache reaches a physiological steady state across a multiseed/multigen ensemble, so the four wild-type-property tests are evaluated on the ensemble (not a single cell).
  • Ray-parallel dispatch (run_seeds_parallel) + XArray/zarr emit is the correct runner for a multiseed ensemble; the in-engine ParquetEmitter is a RAM trap and is avoided.

Build / fix list (1)

Concrete engineering work to fully exercise this study.

v2ecoli.composites.ecoli_baseline.ecoli_baseline resumed from out/cache-showcase; run_seeds_parallel (Ray) dispatch; parquet sidecar store; the native analyses (v2ecoli/workflow/analyses/) consuming the parquet; sim_data hydrated from out/sim_data-showcase/parca_state.pkl.gz via hydrate_sim_data_from_state; Schmidt/Wisniewski + restored Toya-2010 (toya_2010_central_carbon_fluxes.tsv) validation data attached; figures via scripts/render_showcase2.py.

Conclusion synthesis

Read-only synthesis derived from the study's canonical fields (findings, limitations, follow-up proposals).

Claims
  • tests: within tolerance
  • vs_vecoli: drift
Evidence
  • Report card verdict: within tolerance
  • Report card verdict: drift

References cited by this study

macklin2020, toya2010, schmidt2016, wisniewski2014

Pipeline-gate decision

Passed
βœ“ Passed
  • doubling-time-in-band
  • mass-fraction-physiological
  • fba-flux-correlates-toya2010
  • proteome-correlates-schmidt-wisniewski
βœ— Failednothing failing
β›” Blocks the next studynothing blocking
β†’ Immediate next action
Gate cleared.
Pipeline gate & conclusion logic (technical)

Prerequisites: none (root study)

Enables: β€”

If primary tests pass: All four property classes land in their physiological ranges with no calibration trade-off among them, consistent with the upstream WCM baseline (macklin2020).
If primary tests fail: true
3.Which perturbation best demonstrates v2ecoli's response β€” reviewers to chooseπŸ§ͺ Preliminary
β—‹ Not runTests: 1⏳○ Not run
Hand reviewers an explicit, OPEN choice: which perturbation should we run on top of the wild-type baseline to demonstrate v2ecoli's response? This study deliberately commits no variant (conditions.variants lists four PROPOSED, uncommitted candidates); it presents four candidates and stays un-run until a reviewer selects one.
showcase-3-variant-decide Β· depth 0
Confidence: design-stageEvidence: design-only
Conclusion [NO RUN β€” decision study.
β–Έ click to expand full study
Model
The composite(s) this study runs and their parameters.
🧬 baselinedefault parameters

Biology

This study commits no perturbation. It frames an explicit reviewer choice among four mechanistically-distinct levers β€” a global regulatory toggle (ppGpp), a dose-response parameter sweep (a transcription/translation knob), an environmental change (carbon source), and a targeted single-gene expression knob (dnaA / TU00259[c]) β€” each of which would exercise a different part of the v2ecoli whole cell's response. The biology of each candidate lives in the followup_study_proposals below.

Literature anchors

The biological expectations this study tests, mapped to the model observable that will measure each one. Full citations live in the test cards.

Result[NO RUN β€” decision study. The sim does not run until a variant is chosen. See decisions_needed and discovery_implications.followup_study_proposals for the four candidate perturbations.]
DecisionOPEN β€” reviewers to choose which perturbation (or perturbations) best demonstrates v2ecoli's response. The four candidates are below.

Overview

This study asks whether which perturbation best demonstrates v2ecoli's response β€” reviewers to. We recorded 1 novel computational result. Gate decision: Ready to run. Execute the simulation_set to gather evidence.

Purpose & background (study design)
Question. Which perturbation best demonstrates v2ecoli's response β€” reviewers to choose? Building on the wild-type baseline (showcase-2), this study stops at a DECISION: it commits no variant. Four candidate perturbations are proposed; reviewers pick which one (or more) best demonstrates v2ecoli's perturbation response. No simulation runs until a variant is chosen.

Visualizations

Units Atlas (interactive)

Every declared unit-bearing baseline readout grouped by physical dimension with example magnitude + range.

Detailed findings

Infrastructure / computational findings (1)

β—†report-card-testsnew result
tests: ungraded

Conclusion verdicts

Three-track verdict β€” each result is computed from canonical fields (gate evaluator, run status, finding tiers). The basis is the author's rationale.

Biological validationUNKNOWN:pendingfrom gate evaluator Β· computed
Regression compatibilityPENDINGfrom run status Β· computed
Explanatory gainPARTIALfrom interpretation-tier findings Β· computed

Discovery implications

Where this study's results leave the mechanism model β€” and what to investigate next.

● Remaining uncertainties

  • Which perturbation most clearly demonstrates v2ecoli's response is an OPEN reviewer choice; no variant is committed in this scaffold.

Follow-up study proposals (4)

Click βž• Add study to spawn a new study node in the investigation graph (seeds a child study.yaml from the proposal, with a leads-to edge back to this study).

ppGpp regulation toggle OFFperturbationreviewer-decisiongain: high
Proposed experiment: On the showcase-2 wild-type baseline, disable the ppGpp regulatory layer (ppgpp_regulation = off) and run the same multiseed/multigen ensemble. Compare growth rate, ribosome/RNAP allocation, and expression against the baseline.
Parameter sweep via linspace (transcription/translation knob)parameter-sweepreviewer-decisiongain: high
Proposed experiment: Sweep one transcription/translation knob (e.g. a global RNAP or ribosome elongation-rate factor) across a linspace of values on the baseline cache, one parameterized run per point, and plot the dose-response (growth rate / mass fractions vs the swept value).
Media / condition change (e.g. glucose→succinate or +amino-acids)condition-changereviewer-decisiongain: high
Proposed experiment: Switch the growth condition and run the baseline ensemble in the new media. FLAG: each condition needs its OWN full ParCa cache (~2.5 min per condition), so this multiplies the build cost vs the single-cache options.
dnaA expression knob (TU00259[c] transcription-init override)gene-perturbationreviewer-decisiongain: medium
Proposed experiment: Apply the Mechanism-A runtime override sim_data.genetic_perturbations["TU00259[c]"] = V (carried over from the dnaa-replication investigation) on the baseline cache and run the ensemble; read the DnaA pool / replication-initiation response vs the baseline.

Conditions β€” what we set up to test it

Baseline

Composite: v2ecoli.composites.ecoli_baseline.ecoli_baseline
(no overrides)

Variants (4)

Each variant is a perturbation of the baseline β€” typically a parameter override or a swapped composite. These define the runs that test the assumption.

VariantComposite / baseParameter overridesNotesRun
ppgpp-regulation-offv2ecoli.composites.ecoli_baseline.ecoli_baseline(no overrides)β€”vwb run study showcase-3-variant-decide --variant ppgpp-regulation-off
parameter-sweep-linspacev2ecoli.composites.ecoli_baseline.ecoli_baseline(no overrides)β€”vwb run study showcase-3-variant-decide --variant parameter-sweep-linspace
media-condition-changev2ecoli.composites.ecoli_baseline.ecoli_baseline(no overrides)β€”vwb run study showcase-3-variant-decide --variant media-condition-change
dnaa-expression-knobv2ecoli.composites.ecoli_baseline.ecoli_baseline(no overrides)β€”vwb run study showcase-3-variant-decide --variant dnaa-expression-knob

What we ran (4 simulations)

One row per concrete run: the model composite, what changes vs the reference baseline, the condition / length, and its status.

SimulationCompositeChanges vs baselineRunCLIStatus
candidate-ppgpp-offbaselinereference baseline44 seedsvwb run study showcase-3-variant-decideplanned
candidate-parameter-sweep-linspacebaselinesame params, longer/other44 seedsvwb run study showcase-3-variant-decideplanned
candidate-media-condition-changebaselinesame params, longer/other44 seedsvwb run study showcase-3-variant-decideplanned
candidate-dnaa-expression-knobbaselinesame params, longer/other44 seedsvwb run study showcase-3-variant-decideplanned

Measurements (2 readouts)

Quantities we extract from each simulation run to evaluate the study's tests.

ReadoutStatusPathDescription
decision-outcomeβ€”β€”Decision outcome: which candidate perturbation(s) the reviewer selects to run
per-candidate-readoutβ€”β€”Per candidate, the proposed primary readout (growth rate, expression, a pathway flux, or the DnaA pool β€” to be chosen by the reviewer)

Success criteria (1 tests β€” 1 ⏳ pending)

Each test makes a specific scientific claim with a machine-checkable criterion (measure + pass_if). Tests are now evaluated by code against the run (the run/outcome spine: RunReader β†’ evaluator): the pill shows the result, and the evidence line shows the measured value, whether it was computed by code or routed to an agent, and whether the code verdict agrees with the authored one (reconcile). ⏳ pending = the study hasn't run yet. Technical assertion + the exact evaluator are under "Technical details".

⏳ PENDINGdecision
Claim:
Test id: reviewer-selects-perturbation-variant
Technical detailsMeasure: reviewer_decision
Pass condition: {"op":"reviewer_selected_variant"}
Python: vivarium_workbench/lib/expected_behavior.py β†’ evaluate(); measure kind reviewer_decision via _series_for_simple_kind()/_measure(); op reviewer_selected_variant via _check()

Model changes

None yet β€” the model change is the OPEN decision. Each candidate would alter a different knob (ppgpp_regulation toggle, a swept transcription/translation factor, the growth condition / media, or sim_data.genetic_perturbations ["TU00259[c]"]); see followup_study_proposals.

Key assumptions

  • The wild-type baseline (showcase-2) is validated before any perturbation is run, so each candidate variant is a clean delta against a trusted reference.

Build / fix list (1)

Concrete engineering work to fully exercise this study.

Single-cache for ppGpp-off / linspace sweep / dnaA knob (reuse the showcase-1 glucose cache); the media-condition change needs its OWN full ParCa cache per condition (~2.5 min each) β€” the cost trade-off flagged for reviewers.

Conclusion synthesis

Read-only synthesis derived from the study's canonical fields (findings, limitations, follow-up proposals).

Claims
  • tests: ungraded
Evidence
  • Report card verdict: ungraded
Next steps
  • ppGpp regulation toggle OFF
  • Parameter sweep via linspace (transcription/translation knob)
  • Media / condition change (e.g. glucoseβ†’succinate or +amino-acids)
  • dnaA expression knob (TU00259[c] transcription-init override)

Pipeline-gate decision

Ready to run
βœ“ Passednothing yet
βœ— Failednothing failing
β›” Blocks the next studynothing blocking
β†’ Immediate next action
Execute the simulation_set to gather evidence.
4.Five-variant perturbation sweep β€” which perturbation most clearly moves v2ecoli off baselineβœ… Passing
β–Ά Ran Β· 1 runTests: 3βœ“ Β· 1β­βœ… Passed
Resolve the showcase-3 reviewer decision empirically by running ALL four candidate perturbations (rather than picking one) as a single 5-variant sweep (baseline + 4), and measure how far each variant moves the cell off the wild-type baseline across five readout classes: growth (doubling time, cell mass), replication (oriC copy number / re-initiation timing), regulation (ppGpp pool, tRNA charging), central-carbon metabolism (FBA flux + Toya-2010 correlation), and the proteome (Schmidt / Wisniewski correlation). The deliverable is a ranked, quantitative contrast that seeds the next-direction decision (showcase-5).
showcase-4-variant-comparison Β· depth 0
Confidence: highEvidence: 5-variant x 2-seed multigen sweep + 8 cross-variant comparison figures
Conclusion Measured per-variant scorecard (delta vs baseline; from the scorecard / doubling_time_grouped / regulation_overlay / replication_overlay / fba_flux_overlay / proteome_delta figures): metric baseline ppgpp-off dnaA-2x elong-down media-succinate doubling (min) 51.1 61.6 48.4 49.3 52.9 dry mass (fg) 492 510 489 487 271 <- -45% protein fraction 0.461 0.430 0.453 0.435 0.532 <- +0.07 max oriC 4 4 4 4 2 <- never multiforks mean ppGpp 64.8 120.6 67.9 61.2 56.4 <- 1.86x (ppgpp-off) Schmidt r 0.734 0.696 0.737 0.741 0.716 Toya r 0.727 0.764 0.457 0.698 -0.058 <- collapses (succinate) Strongest, most interpretable contrast = media-succinate: dry mass nearly halved (271 vs 492 fg), the cell never multiforks (max oriC stays 2 vs 4), protein fraction rises to 0.532, and the central-carbon FBA flux distribution decorrelates from the glucose-grown Toya-2010 measurements (Toya r collapses from 0.73 to -0.06, with the largest flux shifts on the TCA / succinate-entry reactions: SUCCINATE-DEHYDROGENASE, 2OXOGLUTARATEDEH, MALATE-DEH).
3/3 tests passing
Insight Running all four candidates rather than choosing one was the right call: it converted the showcase-3 reviewer decision into a measured ranking.
β–Έ click to expand full study
Model
The composite(s) this study runs and their parameters.
🧬 baselinedefault parameters

Biology

A perturbation-response demonstration: each variant is a single mechanistic lever (a regulatory toggle, a gene-expression knob, a translation-rate cut, or a carbon-source change) applied to the validated wild-type baseline. Measuring the delta across five readout classes (growth / replication / regulation / central-carbon flux / proteome) shows which perturbation most clearly and interpretably moves the whole cell off baseline β€” converting the showcase-3 reviewer decision into a measured ranking.

Literature anchors

The biological expectations this study tests, mapped to the model observable that will measure each one. Full citations live in the test cards.

SetupSweep: 5 variants x 2 seeds, resumed from the showcase-1 full ParCa caches. Variant 0 baseline / 1 ppgpp-off / 2 dnaA-2x / 3 elong-down all run on the glucose cache (out/cache-showcase, 51-condition full ParCa); variant 4 media-succinate runs on its OWN succinate cache (out/cache-succinate, condition=succinate / fixed_media=minimal_succinate). Generations vary by phenotype and are themselves a contrast: baseline / dnaA-2x / elong-down / media reach 3 generations, ppgpp-off only 2 (it grows slower). max_steps 6500/cell; total sweep wall ~2982 s on the mac mini. Variant levers (config_overrides on the runner): - ppgpp-off = ecoli-transcript-initiation.ppgpp_regulation=False + ecoli-polypeptide-elongation.ppgpp_regulation=False, DEFAULT_FEATURES dropped. - dnaA-2x = ecoli-transcript-initiation.perturbations {TU00259[c]: 2.117e-5} (2x the basal 1.058e-5 init prob). - elong-down = ecoli-polypeptide-elongation.basal_elongation_rate 22->15.4. - media = succinate cache (condition=succinate, fixed_media=minimal_succinate); no config_override. The 8 cross-variant comparison analyses (v2ecoli/workflow/analyses/ {growth_overlay, doubling_time_grouped, mass_fraction_grouped, replication_overlay, proteome_delta, fba_flux_overlay, regulation_overlay, scorecard}) consumed the sweep's hive-partitioned parquet directly; sim_data was hydrated from the showcase ParCa fixture (out/sim_data-showcase/parca_state.pkl.gz via hydrate_sim_data_from_state) and Schmidt/Wisniewski + Toya-2010 validation data attached. Figures rendered to PNG+SVG via scripts/render_variant_comparison.py with the real variant_metadata {0:baseline, 1:ppgpp-off, 2:dnaA-2x, 3:elong-down, 4:media-succinate}.
ResultMeasured per-variant scorecard (delta vs baseline; from the scorecard / doubling_time_grouped / regulation_overlay / replication_overlay / fba_flux_overlay / proteome_delta figures):

metric baseline ppgpp-off dnaA-2x elong-down media-succinate doubling (min) 51.1 61.6 48.4 49.3 52.9 dry mass (fg) 492 510 489 487 271 <- -45% protein fraction 0.461 0.430 0.453 0.435 0.532 <- +0.07 max oriC 4 4 4 4 2 <- never multiforks mean ppGpp 64.8 120.6 67.9 61.2 56.4 <- 1.86x (ppgpp-off) Schmidt r 0.734 0.696 0.737 0.741 0.716 Toya r 0.727 0.764 0.457 0.698 -0.058 <- collapses (succinate)

Strongest, most interpretable contrast = media-succinate: dry mass nearly halved (271 vs 492 fg), the cell never multiforks (max oriC stays 2 vs 4), protein fraction rises to 0.532, and the central-carbon FBA flux distribution decorrelates from the glucose-grown Toya-2010 measurements (Toya r collapses from 0.73 to -0.06, with the largest flux shifts on the TCA / succinate-entry reactions: SUCCINATE-DEHYDROGENASE, 2OXOGLUTARATEDEH, MALATE-DEH). Second strongest = ppgpp-off: the ppGpp pool RISES to ~230 by end-of-run (mean 120.6, 1.86x baseline) rather than collapsing, and growth slows the most (doubling 61.6 min, +10.5; only 2 generations reached). This is the honest, slightly counter-intuitive result β€” disabling ppGpp *regulation* (the feedback the ribosome/RNAP allocation reads) does NOT stop ppGpp *synthesis*, so without the feedback ppGpp accumulates and growth falls. dnaA-2x is a PARTIAL contrast: it does NOT raise the max oriC count above baseline (both peak at 4), but it advances the timing of replication re-initiation (the dnaA-2x oriC curve rises to 4 earlier and re-initiates earlier in the second cycle) and it noticeably degrades the central-carbon flux match (Toya r 0.727 -> 0.457). elong-down is the WEAKEST contrast: the 22->15.4 basal-elongation-rate cut barely moves the cap-immune doubling time (49.3 vs 51.1 min, actually marginally faster), and its proteome / mass / replication readouts are essentially baseline β€” the steady-state charging model and variable-elongation machinery appear to compensate. tRNA charged fraction stays ~0.97 across ALL variants. The proteome is the least discriminating axis overall: every variant stays tightly correlated to the baseline proteome (Schmidt r 0.70-0.74, Wisniewski ~0.61 across the board).
InterpretationRunning all four candidates rather than choosing one was the right call: it converted the showcase-3 reviewer decision into a measured ranking. The media / condition change is the most demonstrative single perturbation β€” it moves the cell off baseline on FOUR independent axes at once (mass, replication, composition, central-carbon flux) and the Toya-r collapse is a clean, physiologically-grounded signature of a different carbon source. The ppGpp-off result is the most scientifically informative surprise: the regulatory knock-out does not silence ppGpp, it de-represses it, so the "ppGpp collapse" intuition is wrong for this lever β€” worth flagging to reviewers. dnaA-2x demonstrates the targeted gene-knob path works (it shifts replication timing + central-carbon flux) but does not push the discrete oriC ceiling, so its contrast is real but subtle. elong-down is a near-null at this magnitude β€” a useful negative control, but not a compelling showcase perturbation on its own.
DecisionPASS β€” the 5-variant sweep ran to completion and the 8 cross-variant comparison figures resolve the showcase-3 decision with measured deltas. Three of the four perturbations produce a clear, interpretable contrast vs baseline (media-succinate strongest, ppgpp-off second, dnaA-2x partial); the fourth (elong-down) is a near-null at this magnitude. The gate passes, unblocking showcase-5 (the next-direction decision), which is seeded below to deepen the strongest contrast (media / growth-condition response) and to chase the ppGpp-off surprise.

Overview

This study asks whether of the four perturbations proposed in showcase-3 (ppGpp regulation OFF, a. We recorded 1 finding confirm the expected biology. Gate decision: Passed. Gate cleared.

Purpose & background (study design)
Question. Of the four perturbations proposed in showcase-3 (ppGpp regulation OFF, a dnaA 2x expression knob, a slower ribosome elongation rate, and a media / condition change to succinate), which one most clearly and interpretably moves the v2ecoli whole cell off its wild-type baseline? Run all four as a single multivariant sweep against the showcase-2 baseline and quantify each variant's delta across growth, replication, regulation, central-carbon flux, and proteome.

Visualizations

Units Atlas (interactive)

Every declared unit-bearing baseline readout grouped by physical dimension with example magnitude + range.

Detailed findings

Infrastructure / computational findings (1)

βœ“report-card-testsconfirmed
tests: within tolerance

Conclusion verdicts

Three-track verdict β€” each result is computed from canonical fields (gate evaluator, run status, finding tiers). The basis is the author's rationale.

Biological validationPASSfrom gate evaluator Β· computed
Regression compatibilityPASSfrom run status Β· computed
Explanatory gainPARTIALfrom interpretation-tier findings Β· computed

Discovery implications

Where this study's results leave the mechanism model β€” and what to investigate next.

βœ“ Resolved uncertainties

  • showcase-3's open reviewer decision is resolved empirically: all four candidate perturbations were run as a 5-variant sweep, and the media / condition change is the most demonstrative single perturbation (moves the cell off baseline on mass, replication, composition, and central-carbon flux simultaneously).
  • ppGpp-off does NOT collapse the ppGpp pool β€” it de-represses it (mean ppGpp 64.8 -> 120.6, 1.86x; trace climbs to ~230). Disabling ppGpp *regulation* removes the feedback the allocation reads but not ppGpp *synthesis*. Growth slows the most (doubling +10.5 min, only 2 gens).
  • dnaA-2x advances replication re-initiation TIMING and degrades the central-carbon flux match (Toya r 0.727 -> 0.457) but does NOT raise the discrete max oriC ceiling (both peak at 4).
  • media-succinate's central-carbon FBA flux decorrelates from the glucose-grown Toya-2010 measurements (Toya r collapses to -0.058), with the largest flux shifts on the TCA / succinate-entry reactions β€” a clean carbon-source signature.

● Remaining uncertainties

  • elong-down is a near-null at the 0.7x magnitude (doubling 49.3 vs 51.1 min) β€” the steady-state charging + variable-elongation machinery appears to compensate. Whether a larger cut (e.g. 0.4-0.5x) or a different translation knob produces a graded slowdown is open.
  • The proteome is the least discriminating axis: all variants stay at Schmidt r 0.70-0.74 / Wisniewski ~0.61. Whether a coarser-grained or pathway-resolved proteome readout would separate the variants is open.
  • tRNA charged fraction is pinned at ~0.97 across every variant including ppgpp-off and elong-down β€” the charging model may be insensitive to these levers, or the readout saturates.

Follow-up study proposals (2)

Click βž• Add study to spawn a new study node in the investigation graph (seeds a child study.yaml from the proposal, with a leads-to edge back to this study).

Deepen the strongest contrast β€” media / growth-condition response (dose of carbon sources)condition-sweepshowcase-4-strongest-contrastgain: high
Proposed experiment: Run a small panel of growth conditions (e.g. glucose, succinate, acetate, glycerol, +amino-acids), each with its own full ParCa cache, and chart the growth rate / central-carbon flux (Toya-r) / mass-composition response as a function of carbon source. Builds directly on media-succinate, the strongest showcase-4 contrast.
Chase the ppGpp-off surprise β€” why does the pool rise, not collapse?mechanism-probeshowcase-4-counterintuitive-resultgain: high
Proposed experiment: Decompose the ppGpp-off result: separate ppGpp *synthesis* (RelA/SpoT) from the *regulatory* coupling, and test whether toggling synthesis vs regulation independently reproduces the rise. Adds a synthesis-off / regulation-off 2x2 to confirm the de-repression mechanism.

Conditions β€” what we set up to test it

Baseline

Composite: v2ecoli.composites.ecoli_baseline.ecoli_baseline
seed0
cache_dirout/cache-showcase

Variants (4)

Each variant is a perturbation of the baseline β€” typically a parameter override or a swapped composite. These define the runs that test the assumption.

VariantComposite / baseParameter overridesNotesRun
ppgpp-offv2ecoli.composites.ecoli_baseline.ecoli_baseline(no overrides)β€”vwb run study showcase-4-variant-comparison --variant ppgpp-off
dnaA-2xv2ecoli.composites.ecoli_baseline.ecoli_baseline(no overrides)β€”vwb run study showcase-4-variant-comparison --variant dnaA-2x
elong-downv2ecoli.composites.ecoli_baseline.ecoli_baseline(no overrides)β€”vwb run study showcase-4-variant-comparison --variant elong-down
media-succinatev2ecoli.composites.ecoli_baseline.ecoli_baseline(no overrides)β€”vwb run study showcase-4-variant-comparison --variant media-succinate

Model settings (3)

Parameters that need human input before the study runs. Edit a value on the dashboard's study-detail page (Build tab) and the next pbg_runner invocation will pick it up.

NameTypeDefaultCurrentRangeGateDescription
multivariant-multiseed-sweepβ€”awaiting expertβ€”optionalβ€”
generations-vary-by-phenotypeβ€”awaiting expertβ€”optionalβ€”
cross-variant-comparison-analysesβ€”awaiting expertβ€”optionalβ€”

What we ran (1 simulation)

One row per concrete run: the model composite, what changes vs the reference baseline, the condition / length, and its status.

SimulationCompositeChanges vs baselineRunCLIStatus
showcase4-variant-sweepbaselinereference baseline63 seedsvwb run study showcase-4-variant-comparisoncomplete

Measurements (6 readouts)

Quantities we extract from each simulation run to evaluate the study's tests.

ReadoutStatusPathDescription
doubling-time-per-variantβ€”β€”cap-immune doubling time per variant (growth)
mass-per-variantβ€”β€”dry mass + mass fractions per variant (composition)
oric-per-variantβ€”β€”oriC copy number over the cell cycle (replication / multifork timing)
regulation-per-variantβ€”β€”ppGpp pool + tRNA charged fraction (regulation)
fba-flux-per-variantβ€”β€”central-carbon FBA flux delta + Toya-2010 correlation per variant (metabolism)
proteome-per-variantβ€”β€”proteome log-log vs baseline + Schmidt / Wisniewski correlation per variant

Visualisations from the latest run

doubling_time_grouped❓ untracked
baselineppgpp-offdnaA-2xelong-downmediaVariant01020304050607080Doubling time (min)baselineppgpp-offdnaA-2xelong-downmediaVariantDoubling time per variant β€” cap-immune ln(2)/mean(mu)
fba_flux_overlay❓ untracked
βˆ’20βˆ’15βˆ’10βˆ’505101520Mean flux Ξ” vs baseline [mmol/gDW/hr]TRANS-RXN-157PGLUCISOM-RXN6PFRUCTPHOS-RXNF16ALDOLASE-RXNTRIOSEPISOMERIZATION-RXNGAPOXNPHOSPHN-RXN2PGADEHYDRAT-RXNPEPDEPHOS-RXNPYRUVDEH-RXNGLU6PDEHYDROG-RXNRXN-9952RIBULP3EPIM-RXNRIB5PISOM-RXN1TRANSKETO-RXNTRANSALDOL-RXN2TRANSKETO-RXNCITSYN-RXNISOCITDEH-RXN2OXOGLUTARATEDEH-RXNSUCCINATE-DEHYDROGENASE-UB…FUMHYDR-RXNMALATE-DEH-RXNPEPCARBOX-RXNCentral-carbon reactionbaselinednaA-2xelong-downmediappgpp-offVariantCentral-carbon flux Ξ” vs baseline0.730.760.460.70βˆ’0.06βˆ’0.10.8Toya rToya-2010 rbaselinednaA-2xelong-downmediappgpp-offVariant
growth_overlay❓ untracked
05001,0001,5002,0002,5003,0003,5004,0004,5005,0005,5006,0006,500Time (s, per-generation clock)05001,0001,5002,000Cell mass (fg)Cell mass (fg)05001,0001,5002,0002,5003,0003,5004,0004,5005,0005,5006,0006,500Time (s, per-generation clock)0.00000.00010.00020.00030.0004Instantaneous growth rate (1/s)Instantaneous growth rate (1/s)baselineppgpp-offdnaA-2xelong-downmediaVariant
mass_fraction_grouped❓ untracked
proteinrRNAtRNADNAsmall-molMass component0.000.050.100.150.200.250.300.350.400.450.500.55Mean fraction of dry massbaselineppgpp-offdnaA-2xelong-downmediaVariantMass fractions per variant (mean over time)
proteome_delta❓ untracked
0.00.51.01.52.02.53.03.54.04.55.05.5log10(baseline count + 1)0.00.51.01.52.02.53.03.54.04.55.05.5log10(variant count + 1)dnaA-2xelong-downmediappgpp-offVariantPer-variant proteome vs baseline0.730.700.740.740.720.610.610.610.610.610.610.74Pearson rProteome validation r per variantschmidt_rwisniewski_rValidation setbaselinednaA-2xelong-downmediappgpp-offVariant
regulation_overlay❓ untracked
05001,0001,5002,0002,5003,0003,5004,0004,5005,0005,5006,0006,500Time (s, per-generation clock)050100150200ppGpp concppGpp conc05001,0001,5002,0002,5003,0003,5004,0004,5005,0005,5006,0006,500Time (s, per-generation clock)0.00.20.40.60.81.0Fraction tRNA chargedFraction tRNA chargedbaselineppgpp-offdnaA-2xelong-downmediaVariant
replication_overlay❓ untracked
05001,0001,5002,0002,5003,0003,5004,0004,5005,0005,5006,0006,500Time (s, per-generation clock)01234Number of oriCNumber of oriCbaselineppgpp-offdnaA-2xelong-downmediaVariant
scorecard❓ untracked
51.161.648.449.352.94925104894872710.4610.4300.4530.4350.5324.004.004.004.002.0064.812167.961.256.40.7340.6960.7370.7410.7160.7270.7640.4570.698βˆ’0.0580βˆ’1.0βˆ’0.50.00.51.0Ξ” vs baseline (norm.)Variant scorecard (Ξ” vs baseline)Schmidt rToya rdoubling time (min)dry mass (fg)max oriCmean ppGppprotein fractionMetricbaselinednaA-2xelong-downmediappgpp-offVariant
variant scorecard (delta vs baseline)
variant scorecard (delta vs baseline)
growth overlay (cell mass + instantaneous growth rate)
growth overlay (cell mass + instantaneous growth rate)
doubling time per variant
doubling time per variant
mass fractions per variant
mass fractions per variant
replication overlay (oriC copy number)
replication overlay (oriC copy number)
regulation overlay (ppGpp + tRNA charging)
regulation overlay (ppGpp + tRNA charging)
central-carbon FBA flux delta + Toya correlation
central-carbon FBA flux delta + Toya correlation
proteome delta vs baseline + validation r
proteome delta vs baseline + validation r

Success criteria (4 tests β€” 3 βœ“ passed Β· 1 ◐ partial)

Each test makes a specific scientific claim with a machine-checkable criterion (measure + pass_if). Tests are now evaluated by code against the run (the run/outcome spine: RunReader β†’ evaluator): the pill shows the result, and the evidence line shows the measured value, whether it was computed by code or routed to an agent, and whether the code verdict agrees with the authored one (reconcile). ⏳ pending = the study hasn't run yet. Technical assertion + the exact evaluator are under "Technical details".

βœ“ PASSprimary
Claim:
Evidence:
authored PASS from run showcase4-variant-sweep
passes if scorecard_delta in [1.5, 3]
Test id: ppgpp-off-perturbs-ppgpp-and-growth
Technical detailsMeasure: scorecard_delta
Pass condition: in [1.5, 3]
Python: vivarium_workbench/lib/expected_behavior.py β†’ evaluate(); measure kind scorecard_delta via _series_for_simple_kind()/_measure(); op in_range via _check()
◐ PARTIALprimary
Claim:
Evidence:
authored PARTIAL from run showcase4-variant-sweep
passes if scorecard_delta in [0.15, 0.5]
Test id: dnaa-2x-shifts-replication-and-flux
Technical detailsMeasure: scorecard_delta
Pass condition: in [0.15, 0.5]
Python: vivarium_workbench/lib/expected_behavior.py β†’ evaluate(); measure kind scorecard_delta via _series_for_simple_kind()/_measure(); op in_range via _check()
βœ“ PASSnegative_control
Claim:
Evidence:
authored PASS from run showcase4-variant-sweep
passes if scorecard_delta in [-3, 3]
Test id: elong-down-is-near-null-negative-control
Technical detailsMeasure: scorecard_delta
Pass condition: in [-3, 3]
Python: vivarium_workbench/lib/expected_behavior.py β†’ evaluate(); measure kind scorecard_delta via _series_for_simple_kind()/_measure(); op in_range via _check()
βœ“ PASSprimary
Claim:
Evidence:
authored PASS from run showcase4-variant-sweep
passes if scorecard_delta in [0.4, 1.5]
Test id: media-shifts-central-carbon-flux
Technical detailsMeasure: scorecard_delta
Pass condition: in [0.4, 1.5]
Python: vivarium_workbench/lib/expected_behavior.py β†’ evaluate(); measure kind scorecard_delta via _series_for_simple_kind()/_measure(); op in_range via _check()

Model changes

Four process-level perturbations against the baseline (see enforced_params / conditions.variants): ppGpp regulation toggle, dnaA (TU00259[c]) init-prob 2x, ribosome basal_elongation_rate 0.7x, and a glucose->succinate media change (run on its own full ParCa cache). monomer_counts uses the corrected SET listener (#185, overwrite[monomer_counts_vec]).

Key assumptions

  • The showcase-2 wild-type baseline (variant 0) is validated, so each variant's delta is a clean contrast against a trusted reference.
  • The four perturbations are the four candidates proposed in showcase-3 (ppGpp-off, dnaA knob, elongation-rate sweep point, media change); running all four resolves the showcase-3 reviewer decision empirically.
  • monomer_counts is the corrected SET listener (#185, overwrite[monomer_counts_vec]) so each row carries the instantaneous proteome (not an accumulating sum); the proteome_delta / scorecard Schmidt/Wisniewski correlations rely on this.

Build / fix list (1)

Concrete engineering work to fully exercise this study.

A 5-variant Γ— 2-seed multigen sweep resumed from the showcase-1 full ParCa caches (out/cache-showcase for glucose variants; out/cache-succinate for media-succinate); config_overrides on the runner for the three process levers; the 8 cross-variant comparison analyses (v2ecoli/workflow/analyses/) reading the hive-partitioned parquet with a real variant_metadata int->name map; sim_data hydrated from out/sim_data-showcase/parca_state.pkl.gz; figures via scripts/render_variant_comparison.py.

Conclusion synthesis

Read-only synthesis derived from the study's canonical fields (findings, limitations, follow-up proposals).

Claims
  • tests: within tolerance
Evidence
  • Report card verdict: within tolerance
Next steps
  • Deepen the strongest contrast β€” media / growth-condition response (dose of carbon sources)
  • Chase the ppGpp-off surprise β€” why does the pool rise, not collapse?

References cited by this study

macklin2020, toya2010, schmidt2016, wisniewski2014

Pipeline-gate decision

Passed
βœ“ Passed
  • ppgpp-off-perturbs-ppgpp-and-growth
  • elong-down-is-near-null-negative-control
  • media-shifts-central-carbon-flux
βœ— Failednothing failing
β›” Blocks the next studynothing blocking
β†’ Immediate next action
Gate cleared.
Pipeline gate & conclusion logic (technical)

Prerequisites: none (root study)

Enables: β€”

If primary tests pass: Three of four levers move the cell off baseline in mechanistically-interpretable ways; the fourth is a clean, informative negative control. The ppGpp-off de-repression is the most scientifically informative surprise.
If primary tests fail: true
5.Which direction next β€” deepen the strongest perturbation contrast (reviewers to choose)πŸ§ͺ Preliminary
β—‹ Not runTests: 1⏳○ Not run
Hand reviewers an explicit, OPEN choice for the NEXT direction, grounded in the showcase-4 measured contrasts: which follow-up best advances the showcase? This study deliberately commits no variant (conditions.variants lists four PROPOSED, uncommitted candidate directions); it presents candidate directions seeded from the showcase-4 ranking and stays un-run until a reviewer selects one.
showcase-5-next-direction-decide Β· depth 0
Confidence: design-stageEvidence: design-only (seeded from showcase-4 measured results)
Conclusion [NO RUN β€” decision study.
β–Έ click to expand full study
Model
The composite(s) this study runs and their parameters.
🧬 baselinedefault parameters

Biology

This study commits no direction. It frames an explicit reviewer choice among four follow-ups, each grounded in a measured showcase-4 contrast: deepen the strongest contrast (a carbon-source panel), chase the most informative surprise (decompose ppGpp synthesis vs regulation), push the partial (a dnaA init-probability sweep), or rescue the null (a graded translation dose-response). The biology of each candidate lives in followup_study_proposals.

Literature anchors

The biological expectations this study tests, mapped to the model observable that will measure each one. Full citations live in the test cards.

Result[NO RUN β€” decision study. The sim does not run until a direction is chosen. The candidate directions below are SEEDED from the showcase-4 results:

- STRONGEST contrast (media-succinate): dry mass halved, never multiforks (oriC 2 vs 4), Toya-2010 flux correlation collapsed 0.727 -> -0.058. => candidate A: deepen the growth-condition response (carbon-source panel). - MOST INFORMATIVE SURPRISE (ppgpp-off): ppGpp pool ROSE (1.86x) rather than collapsing; growth slowed most. => candidate B: decompose ppGpp synthesis vs regulation (2x2). - PARTIAL (dnaA-2x): advanced re-init timing + dropped Toya r, but did not raise the oriC ceiling. => candidate C: push the dnaA knob harder / connect to the dnaa-replication investigation. - NEAR-NULL (elong-down): 0.7x elongation barely moved growth. => candidate D: a graded translation-knob dose-response (larger cuts). ]
DecisionOPEN β€” reviewers to choose the next direction. The four candidates (A growth-condition panel / B ppGpp synthesis-vs-regulation decomposition / C dnaA knob deepening / D graded translation dose-response) are seeded from the showcase-4 ranking below. The agent's recommendation is candidate A (deepen the strongest contrast) with candidate B as the high-information alternative, but no direction is committed.

Overview

This study asks whether given the showcase-4 5-variant sweep, which direction should the showcase. We recorded 1 novel computational result. Gate decision: Ready to run. Execute the simulation_set to gather evidence.

Purpose & background (study design)
Question. Given the showcase-4 5-variant sweep, which direction should the showcase take next? The sweep produced a clear ranking β€” media / growth-condition change was the strongest, most interpretable contrast; the ppGpp-off de-repression was the most informative surprise; elong-down was a near-null. This study stops at a DECISION: it commits no variant. Candidate next directions are seeded from those results; reviewers pick which one (or more) to pursue. No simulation runs until a direction is chosen.

Visualizations

Units Atlas (interactive)

Every declared unit-bearing baseline readout grouped by physical dimension with example magnitude + range.

Detailed findings

Infrastructure / computational findings (1)

β—†report-card-testsnew result
tests: ungraded

Conclusion verdicts

Three-track verdict β€” each result is computed from canonical fields (gate evaluator, run status, finding tiers). The basis is the author's rationale.

Biological validationUNKNOWN:pendingfrom gate evaluator Β· computed
Regression compatibilityPENDINGfrom run status Β· computed
Explanatory gainPARTIALfrom interpretation-tier findings Β· computed

Discovery implications

Where this study's results leave the mechanism model β€” and what to investigate next.

● Remaining uncertainties

  • Which direction to deepen is an OPEN reviewer choice; no direction is committed in this scaffold.
  • elong-down's near-null leaves open where (if anywhere) a translation-rate cut starts to slow growth in v2ecoli β€” candidate D would resolve it.
  • The ppGpp-off de-repression mechanism (rise, not collapse) is asserted but not decomposed β€” candidate B would confirm it.

Follow-up study proposals (4)

Click βž• Add study to spawn a new study node in the investigation graph (seeds a child study.yaml from the proposal, with a leads-to edge back to this study).

A β€” Growth-condition panel (deepen the strongest contrast)condition-sweepshowcase-4-strongest-contrastgain: high
Proposed experiment: Run a small panel of carbon sources (glucose / succinate / acetate / glycerol / +amino-acids), each with its own full ParCa cache, and chart the growth rate / central-carbon flux (Toya-r) / mass-composition response as a function of carbon source. Builds directly on media-succinate, the strongest showcase-4 contrast.
B β€” ppGpp synthesis vs regulation 2x2 (chase the surprise)mechanism-probeshowcase-4-counterintuitive-resultgain: high
Proposed experiment: Separate ppGpp *synthesis* (RelA/SpoT) from the *regulatory* coupling and run a synthesis-off / regulation-off 2x2 on the existing glucose cache to confirm that the showcase-4 ppGpp rise is de-repression. Reads the ppGpp pool + growth rate per quadrant.
C β€” dnaA init-probability sweep (push the partial)parameter-sweepshowcase-4-partial-contrastgain: medium
Proposed experiment: Sweep the dnaA (TU00259[c]) transcription-init probability across 1x..4x on the existing glucose cache and connect to the dnaa-replication investigation's replication-initiation readouts (oriC, critical mass per oriC). Tests whether a larger knob raises the oriC ceiling that 2x did not.
D β€” graded translation dose-response (rescue the null)parameter-sweepshowcase-4-near-null-resultgain: medium
Proposed experiment: Sweep the ribosome basal elongation rate across a wider range (e.g. 1.0x..0.4x of 22 aa/s) on the existing glucose cache to find where the steady-state-charging compensation breaks and growth actually slows β€” resolving why the 0.7x elong-down point was a near-null in showcase-4.

Conditions β€” what we set up to test it

Baseline

Composite: v2ecoli.composites.ecoli_baseline.ecoli_baseline
(no overrides)

Variants (4)

Each variant is a perturbation of the baseline β€” typically a parameter override or a swapped composite. These define the runs that test the assumption.

VariantComposite / baseParameter overridesNotesRun
A-growth-condition-panelv2ecoli.composites.ecoli_baseline.ecoli_baseline(no overrides)β€”vwb run study showcase-5-next-direction-decide --variant A-growth-condition-panel
B-ppgpp-synthesis-vs-regulationv2ecoli.composites.ecoli_baseline.ecoli_baseline(no overrides)β€”vwb run study showcase-5-next-direction-decide --variant B-ppgpp-synthesis-vs-regulation
C-dnaa-knob-deepeningv2ecoli.composites.ecoli_baseline.ecoli_baseline(no overrides)β€”vwb run study showcase-5-next-direction-decide --variant C-dnaa-knob-deepening
D-graded-translation-dose-responsev2ecoli.composites.ecoli_baseline.ecoli_baseline(no overrides)β€”vwb run study showcase-5-next-direction-decide --variant D-graded-translation-dose-response

What we ran (4 simulations)

One row per concrete run: the model composite, what changes vs the reference baseline, the condition / length, and its status.

SimulationCompositeChanges vs baselineRunCLIStatus
direction-A-growth-condition-panelbaselinereference baseline44 seedsvwb run study showcase-5-next-direction-decideplanned
direction-B-ppgpp-synthesis-vs-regulationbaselinesame params, longer/other44 seedsvwb run study showcase-5-next-direction-decideplanned
direction-C-dnaa-knob-deepeningbaselinesame params, longer/other44 seedsvwb run study showcase-5-next-direction-decideplanned
direction-D-graded-translation-dose-responsebaselinesame params, longer/other44 seedsvwb run study showcase-5-next-direction-decideplanned

Measurements (2 readouts)

Quantities we extract from each simulation run to evaluate the study's tests.

ReadoutStatusPathDescription
decision-outcomeβ€”β€”Decision outcome: which candidate next-direction(s) the reviewer selects
per-direction-readoutβ€”β€”Per direction, the proposed primary readout (growth rate, central-carbon flux / Toya-r, ppGpp pool, or replication-initiation timing β€” to be chosen by the reviewer)

Success criteria (1 tests β€” 1 ⏳ pending)

Each test makes a specific scientific claim with a machine-checkable criterion (measure + pass_if). Tests are now evaluated by code against the run (the run/outcome spine: RunReader β†’ evaluator): the pill shows the result, and the evidence line shows the measured value, whether it was computed by code or routed to an agent, and whether the code verdict agrees with the authored one (reconcile). ⏳ pending = the study hasn't run yet. Technical assertion + the exact evaluator are under "Technical details".

⏳ PENDINGdecision
Claim:
Test id: reviewer-selects-next-direction
Technical detailsMeasure: reviewer_decision
Pass condition: {"op":"reviewer_selected_variant"}
Python: vivarium_workbench/lib/expected_behavior.py β†’ evaluate(); measure kind reviewer_decision via _series_for_simple_kind()/_measure(); op reviewer_selected_variant via _check()

Model changes

None yet β€” the model change is the OPEN decision. Each candidate would alter a different axis (carbon source + its ParCa cache; ppGpp synthesis vs regulation decomposition; the dnaA init-probability; or the ribosome elongation rate over a wider range); see followup_study_proposals.

Key assumptions

  • The showcase-4 5-variant sweep is complete and its ranking (media strongest, ppGpp-off most informative surprise, dnaA-2x partial, elong-down near-null) is the trusted basis for choosing the next direction.

Build / fix list (1)

Concrete engineering work to fully exercise this study.

Single-cache for directions B / C / D (reuse the showcase-1 glucose cache); the growth-condition panel (A) needs one full ParCa cache per condition (~2.5 min each) β€” the cost trade-off flagged for reviewers.

Conclusion synthesis

Read-only synthesis derived from the study's canonical fields (findings, limitations, follow-up proposals).

Claims
  • tests: ungraded
Evidence
  • Report card verdict: ungraded
Next steps
  • A β€” Growth-condition panel (deepen the strongest contrast)
  • B β€” ppGpp synthesis vs regulation 2x2 (chase the surprise)
  • C β€” dnaA init-probability sweep (push the partial)
  • D β€” graded translation dose-response (rescue the null)

Pipeline-gate decision

Ready to run
βœ“ Passednothing yet
βœ— Failednothing failing
β›” Blocks the next studynothing blocking
β†’ Immediate next action
Execute the simulation_set to gather evidence.
6.v1↔v2 equivalence at scale: a large (16Γ—16) baseline ensemble graded against vEcoliβ›” Blocked
β–Ά Ran Β· 2 runsTests: 3βœ“ Β· 2βœ—βŒ Failing
Establish whether the v2ecoli wild-type baseline population phenotype is equivalent (within tolerance) to vEcoli's at large ensemble scale. The instrument is the population_phenotype_basal report card (PR #134) rendered in its vs_vecoli equivalence mode: the SAME 21-axis / 5-group card (Physiology Β· Composition Β· Ribosomes Β· Exchange fluxes Β· Gene expression), graded against a matched vEcoli ("v1") ensemble rather than v2's self-pin. Scale is 16 seeds Γ— 16 generations (2Γ— the seeds of the #134 8Γ—16 demo), burn-in generation_lower_bound=3.
showcase-6-equivalence-large Β· depth 0
Confidence: highEvidence: matched 16Γ—16 v2 (256 cells) vs vEcoli (185 cells) ensembles + reference-driven typed-criteria report card (PR
Conclusion Card overall: MISMATCH (12 within_tol βœ“ Β· 6 drift β‰ˆ Β· 3 mismatch βœ— Β· 0 ungraded), over all 21 axes graded at 16Γ—16.
4/6 tests passing
Insight The #134 8Γ—16 equivalence picture HOLDS at the larger 16Γ—16 scale: v2 is behaviorally the same E.
β–Έ click to expand full study
Model
The composite(s) this study runs and their parameters.
🧬 baselinedefault parameters
🧬 vEcoli (CovertLab/vEcoli @ b237873e)default parameters

Biology

Whether v2ecoli is still "the same E. coli" as the upstream vEcoli (v1) at a population scale large enough to tighten the omics/flux equivalence axes. The report card grades emergent population phenotypes (cell-level aggregation over a 16Γ—16 seedsΓ—gens ensemble) on 21 axes / 5 groups, comparing v2's distribution against a matched v1 reference. Equivalence on physiology / composition / ribosomes with a localized respiratory-exchange divergence (#143) is the honest reading of how faithfully v2 reproduces v1's wild-type phenotype.

Literature anchors

The biological expectations this study tests, mapped to the model observable that will measure each one. Full citations live in the test cards.

SetupCard kind: population phenotypes (emergent behaviors over a seeds Γ— gens ensemble; cell-level aggregation β€” time-mean within a cell, then population stats across cells). Reference mode: v1↔v2 equivalence (vs vEcoli).

MEASURED (v2) side: v2ecoli.composites.ecoli_baseline.ecoli_baseline driven by the basal stimulus config v2ecoli/configs/population_phenotype_basal_16x16.json (16 seeds Γ— 16 generations, single_daughters, burn-in 3), resumed from a full-ParCa cache (out/cache_full, 51 TF conditions). Dispatched Ray-parallel (run_seeds_parallel) on the Mac mini with the parquet store as authoritative (the in-engine ParquetEmitter is a RAM trap β€” use the sidecar). The card is regenerated over the existing sweep with `v2ecoli-analyze <sweep> --config configs/population_phenotype_basal_16x16.json` (no re-sim).

REFERENCE (vEcoli / v1) side: a matched 16Γ—16 vEcoli ensemble produced from the vEcoli source ALREADY AVAILABLE WITHIN v2ecoli (the in-repo vEcoli / comparison-harness machinery used by the other v1↔v2 comparison reports β€” NOT an external SMS/vecoli-benchmarking checkout). The reference is pinned with scripts/pin_vecoli_equivalence_reference.py, which reuses the self-pin reference as the presentation/criterion template and swaps in v1's per-cell distributions (it carries self-contained cross-implementation readers for the two v1↔v2 emit-schema differences β€” vEcoli's cumulative `time` vs `global_time`, and positional `bulk` vs paired `bulk__id`/`bulk__count`), so the shared analysis_runner stays untouched. The v1 commit is stamped in the reference's stimulus.blessed_model_ref.

RENDER: reports/population_phenotype_basal_report.py --analysis out/ppb16_parallel/parquet/analysis.json --reference docs/report_cards/population_phenotype_basal/vs_vecoli/vecoli_reference.json --sweep-dir out/ppb16_parallel/parquet --gen-lb 3 --model-ref bd2123d2 --out-dir docs/report_cards/population_phenotype_basal/vs_vecoli/ β†’ report_card.{html,md} (regenerated at 16Γ—16 for this study). Rendered WITH vector extraction (not --no-vectors) so the gene-expression r2 + exchange-flux axes are graded, not ungraded.

ACTUAL RUN: v2 dispatched via the NEW parallel-by-default multi-seed path (run_workflow β†’ run_seeds_parallel, one Ray worker per seed, ~6 concurrent on the 12-core/64 GB mini), which cut wall-time from ~17 h (single-process meta-composite, ~4 min/cell sequential) to ~3.6 h for the full 256-cell ensemble. vEcoli ran its own Nextflow workflow (CovertLab/vEcoli @ b237873e) from the sibling in-repo checkout.
ResultCard overall: MISMATCH (12 within_tol βœ“ Β· 6 drift β‰ˆ Β· 3 mismatch βœ— Β· 0 ungraded), over all 21 axes graded at 16Γ—16. v2 measured side = 256/256 cells (16 seeds Γ— 16 full generations); v1 reference = 185-cell matched vEcoli ensemble (b237873e); burn-in gen_lb 3. Sim-health banner: 232/256 v2 generations divided cleanly, 24 hit the duration cap without dividing (the expected v2 in-process sawtooth, #142 β€” flagged, not hidden).

By group (Ξ” = v2 vs v1; p = Welch; d = Cohen's d): β€’ PHYSIOLOGY (4 βœ“ / 2 β‰ˆ): doubling time Ξ” βˆ’2.2% βœ“, cell mass βˆ’4.4% βœ“, cell volume βˆ’4.4% βœ“, replication completion +4.7% βœ“; oriC βˆ’6.3% β‰ˆ drift, replication initiation βˆ’7.2% β‰ˆ drift. β€’ COMPOSITION (3 βœ“): protein/DW +3.3% βœ“, RNA/DW βˆ’1.5% βœ“, DNA/DW +1.0% βœ“ β€” fully equivalent. β€’ RIBOSOMES (2 βœ“ / 2 β‰ˆ): active fraction +0.1% βœ“, elongation rate +0.2% βœ“; total ribosomes βˆ’5.8% β‰ˆ drift, rRNA-init production βˆ’8.7% β‰ˆ drift. β€’ EXCHANGE FLUXES (3 βœ“ / 1 β‰ˆ / 2 βœ— β€” THE DIVERGENCE): overall flux fingerprint RΒ² = 0.9989 βœ“ (40 matched, 0 appeared/lost, 6 sub-floor), glucose βˆ’4.4% βœ“, ammonium βˆ’2.4% βœ“; but Oβ‚‚ exchange Ξ” βˆ’40.4% βœ— mismatch and COβ‚‚ exchange Ξ” βˆ’20.3% βœ— mismatch (acetate sentinel β‰ˆ drift, near floor). This is the Oβ‚‚/COβ‚‚ respiration deficit filed as issue #143 β€” and it PERSISTS at 16Γ—16. β€’ GENE EXPRESSION (1 β‰ˆ / 1 βœ—): transcriptome (mRNA cistrons) RΒ² = 0.9461 βœ— (below the strict 0.99 band), proteome (monomers) RΒ² = 0.9613 β‰ˆ drift β€” highly correlated but short of the strict self-pin equivalence band.
InterpretationThe #134 8Γ—16 equivalence picture HOLDS at the larger 16Γ—16 scale: v2 is behaviorally the same E. coli as v1 on physiology, biomass composition, and ribosome content (12 axes within tolerance, 6 in minor drift), and the overall exchange-flux fingerprint is essentially identical (RΒ² = 0.9989). The card's overall MISMATCH is driven entirely by the THREE known, reproducible divergences β€” Oβ‚‚ (βˆ’40%) and COβ‚‚ (βˆ’20%) exchange and the transcriptome RΒ² (0.95, below the strict 0.99 band) β€” not by any new discrepancy. The larger seed count tightened the statistics (sub-percent p-values, clean effect sizes) without changing the qualitative verdict: the divergence is a localized respiratory-metabolism signature (issue #143) plus a gene- expression correlation that sits just under the strict equivalence band, both consistent with the 8Γ—16 reading (RΒ² β‰ˆ 0.93–0.94 then, 0.95–0.96 now). The instrument behaves as designed at scale: the metabolic divergence is the signal, surfaced via typed criteria and the exchange-flux significance floor, while FBA jitter near zero is shown but held below the floor.
DecisionPARTIAL EQUIVALENCE at 16Γ—16. v2ecoli reproduces vEcoli's wild-type baseline phenotype on physiology / composition / ribosomes / overall flux fingerprint within tolerance; the divergence is confined to respiratory exchange (Oβ‚‚/COβ‚‚, issue #143) and a gene-expression correlation just below the strict band. The behavioral report-card instrument (PR #134) is confirmed to scale from 8Γ—16 to 16Γ—16 as a v1↔v2 equivalence test, and the divergence it flags is stable, localized, and already filed β€” no new bug. Next: the principled cross-impl statistics (TOST + per-axis Ξ΄ margins) would convert the strict-band gene-expression "mismatch" into an explicit equivalence-margin verdict.

Overview

This study asks whether is the v2ecoli wild-type baseline still "the same E. coli" as vEcoli (v1) when. We recorded 2 contradict it. Gate decision: Blocked. Investigate why 2 test(s) failed.

Purpose & background (study design)
Question. Is the v2ecoli wild-type baseline still "the same E. coli" as vEcoli (v1) when measured at LARGE ensemble scale? This study applies the PR #134 behavioral report-card approach (population_phenotype_basal in its v1↔v2 *equivalence* reference mode) to a large multi-generation baseline ensemble β€” 16 seeds Γ— 16 generations β€” and grades v2's emergent population phenotype against a matched vEcoli 16Γ—16 reference within tolerance. #134 exercised the same loop at 8Γ—16 as a demo; this study scales it to a population large enough to tighten the omics/flux equivalence axes and to read v2's multi-generation behavior (the in-process sawtooth, #142) honestly against v1.

Visualizations

Units Atlas (interactive)

Every declared unit-bearing baseline readout grouped by physical dimension with example magnitude + range. Folds in the former units-atlas investigation.

Detailed findings

Infrastructure / computational findings (2)

βœ—report-card-testscorrection
tests: mismatch
βœ—report-card-vs-vecolicorrection
Basal-condition population phenotype β€” v1↔v2 equivalence: mismatch (12 βœ“ 6 β‰ˆ 3 βœ—)
What we saw: 12 βœ“ 6 β‰ˆ 3 βœ—

Conclusion verdicts

Three-track verdict β€” each result is computed from canonical fields (gate evaluator, run status, finding tiers). The basis is the author's rationale.

Biological validationFAILfrom gate evaluator Β· computed
Regression compatibilityPASSfrom run status Β· computed
Explanatory gainPARTIALfrom interpretation-tier findings Β· computed

Discovery implications

Where this study's results leave the mechanism model β€” and what to investigate next.

βœ“ Resolved uncertainties

  • The #134 8Γ—16 equivalence picture HOLDS at 16Γ—16: Physiology / Composition / Ribosomes within tolerance (12 axes βœ“, 6 minor drift), the overall exchange-flux fingerprint near-identical (RΒ²=0.9989), the exchange-flux metabolic divergence (Oβ‚‚ -40%, COβ‚‚ -20%) persisting (issue #143), and gene expression highly correlated but below the strict 0.99 band (transcriptome RΒ²=0.946, proteome RΒ²=0.961 vs ~0.93–0.94 at 8Γ—16). The larger sample tightened the statistics without changing the qualitative verdict.
  • A matched 16Γ—16 vEcoli reference IS feasible from the in-repo sibling vEcoli (CovertLab/vEcoli @ b237873e) via its own Nextflow workflow. It ran to its natural maximum of 185 cells (9/16 single-daughter lineages completed all 16 gens; 7 truncated at deterministic FBA metabolic dead-ends β€” 256 unreachable for this seed set). 185 cells is a large, valid reference (> the 128-cell 8Γ—16 demo).
  • v2's single-process meta-composite is impractically slow for 16Γ—16 (~4 min/cell sequential β†’ ~17 h); the new parallel-by-default multi-seed path (run_seeds_parallel) cut it to ~3.6 h for the full 256-cell ensemble. v2 completed all 16/16 lineages (256 cells); 24/256 generations hit the duration cap without dividing (the #142 sawtooth), surfaced by the card's sim-health banner.

● Remaining uncertainties

  • The gene-expression 'mismatch' is a strict-band artifact (RΒ²=0.946 vs a 0.99 threshold inherited from the self-pin template). The principled cross-impl form β€” TOST + per-axis Ξ΄ equivalence margins β€” would convert this into an explicit equivalence-margin verdict rather than a hard fail. Nice-to-have, not blocking.
  • Root cause of the Oβ‚‚/COβ‚‚ respiratory-exchange divergence (issue #143) is not resolved here β€” this study confirms it is stable and localized at scale, it does not explain it.

Conditions β€” what we set up to test it

Baseline

Composite: v2ecoli.composites.ecoli_baseline.ecoli_baseline
seed0
cache_dirout/cache_full

Variants (1)

Each variant is a perturbation of the baseline β€” typically a parameter override or a swapped composite. These define the runs that test the assumption.

VariantComposite / baseParameter overridesNotesRun
vEcoli (v1) reference ensemblev2ecoli.composites.ecoli_baseline.ecoli_baseline(no overrides)β€”vwb run study showcase-6-equivalence-large --variant vEcoli (v1) reference ensemble

Model settings (5)

Parameters that need human input before the study runs. Edit a value on the dashboard's study-detail page (Build tab) and the next pbg_runner invocation will pick it up.

NameTypeDefaultCurrentRangeGateDescription
reference-mode-v1v2-equivalenceβ€”awaiting expertβ€”optionalβ€”
large-ensemble-16x16β€”awaiting expertβ€”optionalβ€”
ray-parallel-dispatchβ€”awaiting expertβ€”optionalβ€”
parquet-authoritativeβ€”awaiting expertβ€”optionalβ€”
vecoli-source-in-repoβ€”awaiting expertβ€”optionalβ€”

What we ran (2 simulations)

One row per concrete run: the model composite, what changes vs the reference baseline, the condition / length, and its status.

SimulationCompositeChanges vs baselineRunCLIStatus
v2-16x16-parallel-ensemblebaselinereference baseline37 seedsvwb run study showcase-6-equivalence-largecomplete
vecoli-16x16-reference-ensemblevEcoli (CovertLab/vEcoli @ b237873e)different model vEcoli (CovertLab/vEcoli @ b237873e)67 seedsvwb run study showcase-6-equivalence-largecomplete

Measurements (5 readouts)

Quantities we extract from each simulation run to evaluate the study's tests.

ReadoutStatusPathDescription
physiology-groupβ€”β€”Physiology group: doubling time, cell mass, cell volume, oriC, replication init/completion timing
composition-groupβ€”β€”Composition group: protein / RNA / DNA dry-mass fractions
ribosomes-groupβ€”β€”Ribosomes group: total ribosomes, active fraction, elongation rate, rRNA-init production
exchange-fluxes-groupβ€”β€”Exchange fluxes group: overall flux fingerprint RΒ², glucose / ammonium / Oβ‚‚ / COβ‚‚ / acetate exchange
gene-expression-groupβ€”β€”Gene expression group: transcriptome (mRNA cistrons) RΒ² + proteome (monomers) RΒ² over ensemble-mean count vectors

Success criteria (5 tests β€” 3 βœ“ passed Β· 2 βœ— failed)

Each test makes a specific scientific claim with a machine-checkable criterion (measure + pass_if). Tests are now evaluated by code against the run (the run/outcome spine: RunReader β†’ evaluator): the pill shows the result, and the evidence line shows the measured value, whether it was computed by code or routed to an agent, and whether the code verdict agrees with the authored one (reconcile). ⏳ pending = the study hasn't run yet. Technical assertion + the exact evaluator are under "Technical details".

βœ“ PASSprimary
Claim:
Test id: physiology-equivalent-to-vecoli
Technical detailsMeasure: report_card_axis
Pass condition: {"op":"report_card_group_within_tol"}
Python: vivarium_workbench/lib/expected_behavior.py β†’ evaluate(); measure kind report_card_axis via _series_for_simple_kind()/_measure(); op report_card_group_within_tol via _check()
βœ“ PASSprimary
Claim:
Test id: composition-equivalent-to-vecoli
Technical detailsMeasure: report_card_axis
Pass condition: {"op":"report_card_group_within_tol"}
Python: vivarium_workbench/lib/expected_behavior.py β†’ evaluate(); measure kind report_card_axis via _series_for_simple_kind()/_measure(); op report_card_group_within_tol via _check()
βœ“ PASSprimary
Claim:
Test id: ribosomes-equivalent-to-vecoli
Technical detailsMeasure: report_card_axis
Pass condition: {"op":"report_card_group_within_tol"}
Python: vivarium_workbench/lib/expected_behavior.py β†’ evaluate(); measure kind report_card_axis via _series_for_simple_kind()/_measure(); op report_card_group_within_tol via _check()
βœ— FAILprimary
Claim:
Test id: exchange-fluxes-equivalent-to-vecoli
Technical detailsMeasure: report_card_axis
Pass condition: {"op":"report_card_group_within_tol"}
Python: vivarium_workbench/lib/expected_behavior.py β†’ evaluate(); measure kind report_card_axis via _series_for_simple_kind()/_measure(); op report_card_group_within_tol via _check()
βœ— FAILprimary
Claim:
Test id: gene-expression-correlates-vecoli
Technical detailsMeasure: report_card_axis
Pass condition: {"op":"report_card_group_within_tol"}
Python: vivarium_workbench/lib/expected_behavior.py β†’ evaluate(); measure kind report_card_axis via _series_for_simple_kind()/_measure(); op report_card_group_within_tol via _check()

Model changes

None to the model β€” this grades the unperturbed baseline at scale. The change vs the #134 8Γ—16 demo is purely SCALE (16 seeds Γ— 16 gens, 2Γ— the seeds) and the reference MODE (vs_vecoli equivalence rather than self-pin), plus the new parallel-by-default multi-seed dispatch (run_seeds_parallel) that cut wall-time from ~17 h to ~3.6 h.

Key assumptions

  • A matched 16Γ—16 vEcoli reference can be generated from the in-repo vEcoli source and pinned with pin_vecoli_equivalence_reference.py; v1 and v2 share cistron/monomer/flux ordering exactly, so the omics/flux vectors align positionally with no ID remapping (verified at 8Γ—16 in #134).
  • The population_phenotype_basal report card (PR #134) grades the SAME 21 axes at 16Γ—16 as at 8Γ—16 β€” scale changes the population statistics, not the axis definitions or criteria β€” so the equivalence verdict at large scale is directly comparable to the #134 8Γ—16 demo.
  • v2's single-process multi-generation lineage runner sawtooths (in-process state accumulation β†’ degrade β†’ divide-fail β†’ reset-from-cache; #142); at 16 generations the burn-in (generation_lower_bound=3) and the card's stationarity flag / variance decomposition surface this honestly rather than letting it corrupt the graded window.

Build / fix list (1)

Concrete engineering work to fully exercise this study.

v2 side: v2ecoli.composites.ecoli_baseline.ecoli_baseline driven by configs/population_phenotype_basal_16x16.json, Ray-parallel dispatch, parquet store; card regenerated with v2ecoli-analyze (no re-sim). Reference side: a matched 16Γ—16 vEcoli ensemble from the in-repo sibling vEcoli (CovertLab/vEcoli @ b237873e) via its own Nextflow workflow, pinned with scripts/pin_vecoli_equivalence_reference.py (cross-impl readers for the v1↔v2 emit-schema differences). Render: reports/population_phenotype_basal_report.py WITH vector extraction (so gene-expression rΒ² + exchange-flux axes are graded).

Conclusion synthesis

Read-only synthesis derived from the study's canonical fields (findings, limitations, follow-up proposals).

Claims
  • tests: mismatch
  • Basal-condition population phenotype β€” v1↔v2 equivalence: mismatch (12 βœ“ 6 β‰ˆ 3 βœ—)
Evidence
  • Report card verdict: mismatch
  • 12 βœ“ 6 β‰ˆ 3 βœ—

References cited by this study

macklin2020

Pipeline-gate decision

Blocked
βœ“ Passed
  • vecoli-reference-pinned
  • physiology-equivalent-to-vecoli
  • composition-equivalent-to-vecoli
  • ribosomes-equivalent-to-vecoli
βœ— Failed
  • exchange-fluxes-equivalent-to-vecoli
  • gene-expression-correlates-vecoli
β›” Blocks the next studynothing blocking
β†’ Immediate next action
Investigate why 2 test(s) failed.
Pipeline gate & conclusion logic (technical)

Prerequisites: none (root study)

Enables: β€”

If primary tests pass: Equivalent physiology, composition, ribosome content, exchange fluxes, and gene expression at scale would mean v2's emergent population phenotype is statistically indistinguishable from v1's.
If primary tests fail: ["Oβ‚‚/COβ‚‚ respiratory-exchange deficit (#143) persists at scale","transcriptome RΒ² sits just under the strict self-pin 0.99 band (TOST + per-axis Ξ΄ margins would express this as an equivalence-margin verdict)"]

Appendices

Method-grading and verification detail β€” kept at the back, after the main narrative.

How the verdict is computed β€” acceptance criteria & gating matrix
AC β†’ study gating matrix which study gates each acceptance criterion Β· ⚠ = no study linked (gap)

Each acceptance criterion is a behaviour test declared in a study: a measured field from the run (e.g. closure_gap_size) compared against an explicit pass_if band (a numeric threshold/range). The per-criterion result, each study’s gate verdict, and this roll-up are computed in code from the run outcomes (deterministic) β€” not human judgement. Expand a row to see the field, the passing band, and the observed value.

Acceptance criterionGating studyResult
parca-rebuilds-full-51-conditions-from-ecoli-sourcesshowcase-1-parca◐ in-progress
baseline-ensemble-reproduces-wild-type-propertiesshowcase-2-baseline-figures◐ in-progress
reviewer-selects-perturbation-variant-to-runshowcase-3-variant-decide◐ in-progress
five-variant-sweep-ranks-perturbation-contrasts-vs-baselineshowcase-4-variant-comparison◐ in-progress
reviewer-selects-next-direction-from-showcase-4-rankingshowcase-5-next-direction-decide◐ in-progress
large-16x16-baseline-ensemble-equivalent-to-vecoli-within-toleranceshowcase-6-equivalence-large◐ in-progress

All 6 acceptance criteria are linked to a gating study.

πŸ”¬ Evidence & rigor β€” how well the method defends its claims 2/6 investigation rigor dimensions addressed Β· 4 gap(s)

Deterministic feedback on how well the method defends its claims against a skeptical reader β€” a method-level judgement, distinct from the per-study model verdicts above. Computed from declared fields, not judged. Gaps are an invitation to add negative controls, replicate across seeds, weigh alternative explanations, state falsifiability, or add an adversarial study.

βœ—
Adversarial testing C10 C12 C15
no adversarial study β€” add one that tries to BREAK the criteria: mimic / parasitic-or-dependent / externally-maintained / random-cyclic systems that should NOT qualify
βœ“
Traceable methodology C9 C2 C14
capability ladder (study DAG) + explicit acceptance criteria + pass/fail gates + traceable findings β€” the reusable methodological contribution
βœ“
Falsification exposure C1
the framework has been shown to reject at least one system (a discriminating negative control, an adversarial study, or a non-passing result)
βœ—
Comparative framing C13
no competing theoretical frameworks compared (viability theory, organizational / constraint closure, active inference) β€” show the findings uniquely support this lens
βœ—
Hypothesis competition C6 C16
no competing hypotheses[] declared β€” state β‰₯2 rival explanations with predictions so the evidence can adjudicate between them
βœ—
Per-study rigor gaps C2 C4 C6
39 rigor gap(s) across 6 member study(ies)

Per-study rigor

showcase-1-parca β€” 3/12 rigor dimensions addressed Β· 8 gap(s)
βœ—
Replication C4
single run β€” no replication across seeds declared (add robustness.seeds or simulation_set.seeds)
βœ—
Controls & calibration C1 C2 C4
no controls β€” declare a system that SHOULD fail the criteria (externally-maintained / -supplied) plus a clearly-passing / borderline case so the metric is calibrated, not just asserted
βœ—
Alternative hypotheses C6
no alternative hypotheses declared
⚠
Claim discipline C3
findings not tiered β€” label each observation / mechanism / interpretation
βœ—
Falsifiability C5 C1
criteria read as tailored-to-succeed β€” add a falsifiability note (study.falsifiability)
βœ“
Engineered vs emergent C7
no interpretation-tier claim that requires the distinction
βœ“
Limitations stated C8 C11
states what the result does not show
βœ“
Next steps next-steps
declares discovery implications / follow-up studies
βœ—
Threshold provenance C9 C5
1 of 1 numeric band(s) declare neither cites nor pass_if.provenance.kind β€” state where the cutoff came from (theory/calibration/literature/expert/exploratory/post_hoc)
βœ—
Metric calibration ladder C4 C2 C20
no calibration_ladder declared β€” index controls[] by known_fail / known_pass / borderline / stress rungs so the metric is calibrated across its range, not just asserted
βœ—
Generality C22
no generality axes tested β€” add findings[].generality.axes_tested or a robustness parameter sweep so the claim's breadth is evidenced
βœ—
Run persistence persistence
1 run(s) recorded but none carry an emitter (sqlite/parquet/xarray) or a run-db reference β€” the trajectories are not persisted; runs should emit via the workspace emitter
showcase-2-baseline-figures β€” 4/12 rigor dimensions addressed Β· 7 gap(s)
βœ—
Replication C4
single run β€” no replication across seeds declared (add robustness.seeds or simulation_set.seeds)
βœ—
Controls & calibration C1 C2 C4
no controls β€” declare a system that SHOULD fail the criteria (externally-maintained / -supplied) plus a clearly-passing / borderline case so the metric is calibrated, not just asserted
βœ—
Alternative hypotheses C6
no alternative hypotheses declared
⚠
Claim discipline C3
findings not tiered β€” label each observation / mechanism / interpretation
βœ—
Falsifiability C5 C1
criteria read as tailored-to-succeed β€” add a falsifiability note (study.falsifiability)
βœ“
Engineered vs emergent C7
no interpretation-tier claim that requires the distinction
βœ“
Limitations stated C8 C11
states what the result does not show
βœ“
Next steps next-steps
declares discovery implications / follow-up studies
βœ“
Threshold provenance C9 C5
all 2 numeric band(s) carry a source (cites or pass_if.provenance.kind)
βœ—
Metric calibration ladder C4 C2 C20
no calibration_ladder declared β€” index controls[] by known_fail / known_pass / borderline / stress rungs so the metric is calibrated across its range, not just asserted
βœ—
Generality C22
no generality axes tested β€” add findings[].generality.axes_tested or a robustness parameter sweep so the claim's breadth is evidenced
βœ—
Run persistence persistence
1 run(s) recorded but none carry an emitter (sqlite/parquet/xarray) or a run-db reference β€” the trajectories are not persisted; runs should emit via the workspace emitter
showcase-3-variant-decide β€” 6/12 rigor dimensions addressed Β· 5 gap(s)
βœ“
Replication C4
4 replicate(s)
βœ—
Controls & calibration C1 C2 C4
no controls β€” declare a system that SHOULD fail the criteria (externally-maintained / -supplied) plus a clearly-passing / borderline case so the metric is calibrated, not just asserted
βœ—
Alternative hypotheses C6
no alternative hypotheses declared
⚠
Claim discipline C3
findings not tiered β€” label each observation / mechanism / interpretation
βœ—
Falsifiability C5 C1
criteria read as tailored-to-succeed β€” add a falsifiability note (study.falsifiability)
βœ“
Engineered vs emergent C7
no interpretation-tier claim that requires the distinction
βœ“
Limitations stated C8 C11
states what the result does not show
βœ“
Next steps next-steps
declares discovery implications / follow-up studies
βœ“
Threshold provenance C9 C5
no numeric acceptance bands requiring provenance
βœ—
Metric calibration ladder C4 C2 C20
no calibration_ladder declared β€” index controls[] by known_fail / known_pass / borderline / stress rungs so the metric is calibrated across its range, not just asserted
βœ—
Generality C22
no generality axes tested β€” add findings[].generality.axes_tested or a robustness parameter sweep so the claim's breadth is evidenced
βœ“
Run persistence persistence
no runs recorded β€” nothing to persist via an emitter
showcase-4-variant-comparison β€” 3/12 rigor dimensions addressed Β· 8 gap(s)
βœ—
Replication C4
single run β€” no replication across seeds declared (add robustness.seeds or simulation_set.seeds)
βœ—
Controls & calibration C1 C2 C4
no controls β€” declare a system that SHOULD fail the criteria (externally-maintained / -supplied) plus a clearly-passing / borderline case so the metric is calibrated, not just asserted
βœ—
Alternative hypotheses C6
no alternative hypotheses declared
⚠
Claim discipline C3
findings not tiered β€” label each observation / mechanism / interpretation
βœ—
Falsifiability C5 C1
criteria read as tailored-to-succeed β€” add a falsifiability note (study.falsifiability)
βœ“
Engineered vs emergent C7
no interpretation-tier claim that requires the distinction
βœ“
Limitations stated C8 C11
states what the result does not show
βœ“
Next steps next-steps
declares discovery implications / follow-up studies
βœ—
Threshold provenance C9 C5
4 of 4 numeric band(s) declare neither cites nor pass_if.provenance.kind β€” state where the cutoff came from (theory/calibration/literature/expert/exploratory/post_hoc)
βœ—
Metric calibration ladder C4 C2 C20
no calibration_ladder declared β€” index controls[] by known_fail / known_pass / borderline / stress rungs so the metric is calibrated across its range, not just asserted
βœ—
Generality C22
no generality axes tested β€” add findings[].generality.axes_tested or a robustness parameter sweep so the claim's breadth is evidenced
βœ—
Run persistence persistence
1 run(s) recorded but none carry an emitter (sqlite/parquet/xarray) or a run-db reference β€” the trajectories are not persisted; runs should emit via the workspace emitter
showcase-5-next-direction-decide β€” 6/12 rigor dimensions addressed Β· 5 gap(s)
βœ“
Replication C4
4 replicate(s)
βœ—
Controls & calibration C1 C2 C4
no controls β€” declare a system that SHOULD fail the criteria (externally-maintained / -supplied) plus a clearly-passing / borderline case so the metric is calibrated, not just asserted
βœ—
Alternative hypotheses C6
no alternative hypotheses declared
⚠
Claim discipline C3
findings not tiered β€” label each observation / mechanism / interpretation
βœ—
Falsifiability C5 C1
criteria read as tailored-to-succeed β€” add a falsifiability note (study.falsifiability)
βœ“
Engineered vs emergent C7
no interpretation-tier claim that requires the distinction
βœ“
Limitations stated C8 C11
states what the result does not show
βœ“
Next steps next-steps
declares discovery implications / follow-up studies
βœ“
Threshold provenance C9 C5
no numeric acceptance bands requiring provenance
βœ—
Metric calibration ladder C4 C2 C20
no calibration_ladder declared β€” index controls[] by known_fail / known_pass / borderline / stress rungs so the metric is calibrated across its range, not just asserted
βœ—
Generality C22
no generality axes tested β€” add findings[].generality.axes_tested or a robustness parameter sweep so the claim's breadth is evidenced
βœ“
Run persistence persistence
no runs recorded β€” nothing to persist via an emitter
showcase-6-equivalence-large β€” 4/12 rigor dimensions addressed Β· 6 gap(s)
⚠
Replication C4
only 2 replicates β€” add seeds for a robustness claim
βœ—
Controls & calibration C1 C2 C4
no controls β€” declare a system that SHOULD fail the criteria (externally-maintained / -supplied) plus a clearly-passing / borderline case so the metric is calibrated, not just asserted
βœ—
Alternative hypotheses C6
no alternative hypotheses declared
⚠
Claim discipline C3
findings not tiered β€” label each observation / mechanism / interpretation
βœ—
Falsifiability C5 C1
criteria read as tailored-to-succeed β€” add a falsifiability note (study.falsifiability)
βœ“
Engineered vs emergent C7
no interpretation-tier claim that requires the distinction
βœ“
Limitations stated C8 C11
states what the result does not show
βœ“
Next steps next-steps
declares discovery implications / follow-up studies
βœ“
Threshold provenance C9 C5
no numeric acceptance bands requiring provenance
βœ—
Metric calibration ladder C4 C2 C20
no calibration_ladder declared β€” index controls[] by known_fail / known_pass / borderline / stress rungs so the metric is calibrated across its range, not just asserted
βœ—
Generality C22
no generality axes tested β€” add findings[].generality.axes_tested or a robustness parameter sweep so the claim's breadth is evidenced
βœ—
Run persistence persistence
2 run(s) recorded but none carry an emitter (sqlite/parquet/xarray) or a run-db reference β€” the trajectories are not persisted; runs should emit via the workspace emitter
πŸ“Š Framework scorecard framework-self metrics (n=14 investigations)

Framework-self metrics aggregated across every study and investigation in the workspace β€” how consistently the framework itself applies its own rigor practices (discriminating controls, emergent-mechanism labelling, threshold provenance, replication, verdict divergence, falsification exposure). Computed deterministically from declared fields by pbg_superpowers.rigor.framework_metrics.

Discriminating Controls0%0 / 59
Emergent Interpretationsβ€”0 / 0
Missing Mechanism Originβ€”0 / 0
Threshold Provenance39%56 / 144
Replication Coverage22%13 / 59
Ac Coverage100%93 / 93
Verdict Divergence19%11 / 59
Falsification Exposure73%43 / 59
Alternatives Excluded3%2 / 59
Emitter Coverage36%13 / 36
🧩 Suggested additions β€” pending your approval 4 items Β· 4 pending

Four candidate VARIANTS for showcase-3-variant-decide, proposed by the agent for reviewers to Accept or Decline. None is committed β€” the reviewer chooses which perturbation best demonstrates v2ecoli's response. (See the investigation-level decisions_needed and showcase-3's discovery_implications.followup_study_proposals for the same four.)

variantpending
(reference)
Related study: showcase-3-variant-decide
Rationale. Toggling ppGpp regulation off removes the stringent-response coupling: a whole-cell, single-switch perturbation that reads cleanly against the baseline.
proposed by claude Β· 2026-06-09
variantpending
(reference)
Related study: showcase-3-variant-decide
Rationale. Sweeping one transcription/translation knob across a linspace shows v2ecoli's graded dose-response and exercises the multi-variant sweep machinery (one baseline cache, many parameterized runs).
proposed by claude Β· 2026-06-09
variantpending
(reference)
Related study: showcase-3-variant-decide
Rationale. Shows v2ecoli adapting to a different growth environment. FLAG / COST: each media condition requires its own full ParCa cache (~2.5 min per condition), multiplying the build cost vs the single-cache options above.
proposed by claude Β· 2026-06-09
variantpending
(reference)
Related study: showcase-3-variant-decide
Rationale. A targeted single-gene expression perturbation reusing the Mechanism-A lever from dnaa-replication (sim_data.genetic_perturbations["TU00259[c]"] = V, a runtime transcription-init override): a precise gene-level knob with a known biological readout (the DnaA pool).
proposed by claude Β· 2026-06-09

References (4 cited across this investigation)

Union of bibliography.bib_keys and per-behavior cites: across all studies in this investigation. Click DOI or link to open the source.

  1. macklin2020 · Macklin, Derek N. and Ahn-Horst, Travis A. and Choi, Heejo and Ruggero, Nicholas A. and Carrera, Javier and Mason, John C. and others (2020). Simultaneous cross-evaluation of heterogeneous E. coli datasets via mechanistic simulation. Science 369(6502), pp. eaav3751 Β· doi:10.1126/science.aav3751
    Note: The v2ecoli (Covert-lab) WCM. DUF ref [1].
  2. schmidt2016 · Schmidt, Alexander and Kochanowski, Karl and Vedelaar, Silke and Ahrn\'e, Erik and Volkmer, Benjamin and Callipo, Luciano and Knoops, K\`evin and Bauer, Manuel and Aebersold, Ruedi and Heinemann, Matthias (2016). The quantitative and condition-dependent Escherichia coli proteome. Nature Biotechnology 34(1), pp. 104--110 Β· doi:10.1038/nbt.3418
    Note: Absolute condition-dependent E. coli proteome (often cited as "Schmidt 2015/2016"); the measured proteome the v2ecoli baseline monomer counts are correlated against (showcase-2 Schmidt r, showcase-4 proteome panel).
  3. toya2010 · Toya, Yoshihiro and Ishii, Nobuyoshi and Nakahigashi, Kenji and Hirasawa, Takashi and Soga, Tomoyoshi and Tomita, Masaru and Shimizu, Kazuyuki (2010). 13C-metabolic flux analysis for batch culture of Escherichia coli and its pyk and pgi gene knockout mutants based on mass isotopomer distribution of intracellular metabolites. Biotechnology Progress 26(4), pp. 975--992 Β· doi:10.1002/btpr.420
    Note: C13-MFA central-carbon flux measurements for glucose-grown E. coli; the reference flux dataset the v2ecoli baseline FBA fluxes are correlated against (showcase-2/showcase-4). toya_2010_central_carbon_fluxes.tsv.
  4. wisniewski2014 · Wi\'sniewski, Jacek R. and Rakus, Dariusz (2014). Multi-enzyme digestion FASP and the `Total Protein Approach'-based absolute quantification of the Escherichia coli proteome. Journal of Proteomics 109, pp. 322--331 Β· doi:10.1016/j.jprot.2014.07.012
    Note: Total-Protein-Approach absolute E. coli proteome quantification; the second measured-proteome reference the v2ecoli baseline monomer counts are correlated against (showcase-2 Wisniewski r ~0.60, showcase-4 proteome panel).
  5. v2ecoli issue #143 β€” O2/CO2 respiratory-exchange divergence β†—
    Tracking issue for the localized respiratory-exchange (O2 -40% / CO2 -20%) deficit this study confirms is stable and localized at 16x16 scale (not a new bug).