Whole-Cell Model Comparison planning

Investigation report Β· whole-cell-model-comparison Β· generated 2026-08-17 13:32 UTC Β· for expert review β€” results below reflect completed runs.

πŸ“‹ Executive summary in-progress Does v2ecoli reproduce vEcoli across nutrient conditions when both engines swap FBA Metabolism for MetabolismRedux? Five studies (basal, with_aa,…

Does v2ecoli reproduce vEcoli across nutrient conditions when both engines swap FBA Metabolism for MetabolismRedux? Five studies (basal, with_aa, succinate, no_oxygen, acetate) run both engines from matched ParCa initial states, single generation, 6 seeds requested per condition (5 effective vEcoli seeds; 4 for acetate), and grade v2ecoli against a live vEcoli reference on cell, dry, protein and RNA mass, and growth rate.

HEADLINE: v2ecoli reproduces genuine vEcoli under the MetabolismRedux swap across all five nutrient conditions. PARCA (t=0 initial-state match): within_tol on ALL five conditions -- both engines start from identical states. STATISTICAL (Welch t-test, gen-1): all four masses (cell, dry, protein, rna) are within_tol on ALL five conditions; no condition is graded a mismatch. Per-condition overall statistical verdict is within_tol on every axis for with_aa; basal, succinate, no_oxygen and acetate are graded drift, driven solely by the instantaneous growth-rate axis (active_RNAP/active_ribosome are ungraded -- not in the local reader's observable set).

RESULT (matched gen-1 median |delta|, standard card): basal cell 0.14% dry 0.15% protein 0.09% rna 0.60% growth 10.4% with_aa cell 0.86% dry 0.86% protein 0.13% rna 0.34% growth 4.3% succinate cell 1.49% dry 1.50% protein 0.92% rna 3.50% growth 13.4% no_oxygen cell 0.70% dry 0.70% protein 0.20% rna 1.23% growth 15.3% acetate cell 0.53% dry 0.53% protein 1.09% rna 2.19% growth 12.1%

INTERPRETATION: masses are within tolerance everywhere and initial states are identical (parca within_tol on all five conditions); the only residual is a growth-rate DRIFT (amber, not a mismatch) on 4 of 5 conditions, plausibly the noisy per-tick instantaneous-growth-rate derivative at single-generation shape with ~5 seeds. Multi-generation was deferred: a pbg division bug surfaces with injected processes, so this run is single-generation by design, not by failure.

Question. Does v2ecoli reproduce vEcoli across nutrient conditions and metabolism process swaps?

🧬 Biology β€” the mechanism this investigation models v2ecoli and vEcoli are two implementations of the same idea: a mechanistic whole-cell model of E. coli that simulates metabolism, gene expression, replication, and division…

v2ecoli and vEcoli are two implementations of the same idea: a mechanistic whole-cell model of E. coli that simulates metabolism, gene expression, replication, and division for a single cell over its cycle. vEcoli is the established Covert-lab process-bigraph model; v2ecoli is the reimplementation this workspace develops.

Because both descend from the same underlying biology and parameter pipeline (ParCa / sim_data), a faithful v2ecoli should reproduce vEcoli's growth and mass trajectories when both start from matched initial states and run the same medium -- differences then point to implementation divergences, not biology. This investigation grades v2ecoli against a live vEcoli reference on cell, dry, protein and RNA mass and growth rate across nutrient conditions, treating vEcoli as the standard the newer engine must match.

Open questions & decisions needed

Investigation roadmap

Study verdict map code-computed gate verdicts (βœ… passed Β· β›” failed Β· πŸ”„ needs calibration Β· ⚠ blocked Β· β—½ not evaluated)

Studies

Each study is collapsed to a one-glance control panel β€” scan top to bottom, then click any panel to expand its full detail.

1.v2ecoli reproduces vEcoli on acetateπŸ§ͺ Preliminary
β–Ά Ran Β· 1 runTests: 8⏳⏳ Tests pending⚠ 1 clarity note
Does v2ecoli reproduce vEcoli on the acetate condition?
acetate Β· depth 0
β–Έ click to expand full study
1.acetateevaluatedRoot study (no dependencies)
β–΄ click to collapse full study
Model
The composite(s) this study runs and their parameters.
🧬 ecoli_baselinemedia="minimal_acetate"

Biology

v2ecoli reproduces vEcoli on acetate (4-seed gen-1, Welch t-test): cell, dry, RNA and protein mass within tolerance; growth rate shows a non-significant drift at the seed-noise level.

Overview

This study asks whether does v2ecoli reproduce vEcoli on the acetate condition?. Gate decision: In progress. Continue analysing run outcomes.

Purpose & background (study design)
Question. Does v2ecoli reproduce vEcoli on the acetate condition?

Visualizations

00-acetate-observables

Hand-authored figure (8 KB) from reports/figures/acetate/.

01-fix-headline

Hand-authored figure (9 KB) from reports/figures/acetate/.

02-growth-before-after

Hand-authored figure (8 KB) from reports/figures/acetate/.

03-rna-before-after

Hand-authored figure (8 KB) from reports/figures/acetate/.

04-rnap-rootcause

Hand-authored figure (8 KB) from reports/figures/acetate/.

05-nutrient-gradient

Hand-authored figure (8 KB) from reports/figures/acetate/.

Detailed findings

Infrastructure / computational findings (1)

◐acetate-vs-vecoli-resultpartial resultobservation Β· floor
v2ecoli reproduces vEcoli on acetate (4-seed gen-1, Welch t-test): cell, dry, RNA and protein mass within tolerance; growth rate shows a non-significant drift at the seed-noise level.
What we saw: growth +15.9% (p=0.28, not significant); RNA +3.6%; cell +2.3%; dry +2.3%; protein +1.3% -- overall drift.

Conclusion verdicts

Three-track verdict β€” each result is computed from canonical fields (gate evaluator, run status, finding tiers). The basis is the author's rationale.

Biological validationPENDINGfrom gate evaluator Β· computed
Regression compatibilityPASSfrom run status Β· computed
Explanatory gainPARTIALfrom interpretation-tier findings Β· computed

Conditions β€” what we set up to test it

Baseline

Composite: v2ecoli.composites.ecoli_baseline.ecoli_baseline
mediaminimal_acetate

What we ran (1 simulation)

One row per concrete run: the model composite, what changes vs the reference baseline, the condition / length, and its status.

SimulationCompositeChanges vs baselineRunCLIStatus
acetate-baselineecoli_baselinereference baselineacetatevwb run study acetatecompleted

Success criteria (8 tests β€” 8 ⏳ pending)

Each test makes a specific scientific claim with a machine-checkable criterion (measure + pass_if). Tests are now evaluated by code against the run (the run/outcome spine: RunReader β†’ evaluator): the pill shows the result, and the evidence line shows the measured value, whether it was computed by code or routed to an agent, and whether the code verdict agrees with the authored one (reconcile). ⏳ pending = the study hasn't run yet. Technical assertion + the exact evaluator are under "Technical details".

ungradedsummary report cardsummary-vs-vecoli
report card summary not generated yet β€” run the comparison.
ungradedparca report cardparca-vs-vecoli
report card parca not generated yet β€” run the comparison.
driftstatistical report cardstatistical-vs-vecoli
mismatchstandard report cardstandard-vs-vecoli
ungradedtrajectory report cardtrajectory-vs-vecoli
report card trajectory not generated yet β€” run the comparison.
ungradeddistribution report carddistribution-vs-vecoli
report card distribution not generated yet β€” run the comparison.
ungradedmetabolism report cardmetabolism-vs-vecoli
report card metabolism not generated yet β€” run the comparison.
ungradedcomposition report cardcomposition-vs-vecoli
report card composition not generated yet β€” run the comparison.

Conclusion synthesis

Read-only synthesis derived from the study's canonical fields (findings, limitations, follow-up proposals).

Claims
  • v2ecoli reproduces vEcoli on acetate (4-seed gen-1, Welch t-test): cell, dry, RNA and protein mass within tolerance; growth rate shows a non-significant drift at the seed-noise level.
Evidence
  • growth +15.9% (p=0.28, not significant); RNA +3.6%; cell +2.3%; dry +2.3%; protein +1.3% -- overall drift.

Pipeline-gate decision

In progress
βœ“ Passednothing yet
βœ— Failednothing failing
β›” Blocks the next studynothing blocking
β†’ Immediate next action
Continue analysing run outcomes.
Pipeline gate & conclusion logic (technical)

Prerequisites: none (root study)

Enables: β€”

2.v2ecoli reproduces vEcoli on basalπŸ§ͺ Preliminary
β–Ά Ran Β· 1 runTests: 8⏳⏳ Tests pending⚠ 1 clarity note
Does v2ecoli reproduce vEcoli on the basal condition?
basal Β· depth 0
β–Έ click to expand full study
2.basalevaluatedRoot study (no dependencies)
β–΄ click to collapse full study
Model
The composite(s) this study runs and their parameters.
🧬 ecoli_baselinedefault parameters

Biology

CONTROL -- v2ecoli reproduces vEcoli on basal (the ParCa fit point): all observables within tolerance (4-seed gen-1).

Overview

This study asks whether does v2ecoli reproduce vEcoli on the basal condition?. We recorded 1 finding confirm the expected biology. Gate decision: In progress. Continue analysing run outcomes.

Purpose & background (study design)
Question. Does v2ecoli reproduce vEcoli on the basal condition?

Visualizations

00-basal-observables

Hand-authored figure (8 KB) from reports/figures/basal/.

01-fix-headline

Hand-authored figure (9 KB) from reports/figures/basal/.

02-growth-before-after

Hand-authored figure (8 KB) from reports/figures/basal/.

03-rna-before-after

Hand-authored figure (8 KB) from reports/figures/basal/.

04-rnap-rootcause

Hand-authored figure (8 KB) from reports/figures/basal/.

05-nutrient-gradient

Hand-authored figure (8 KB) from reports/figures/basal/.

Detailed findings

Infrastructure / computational findings (1)

βœ“basal-vs-vecoli-resultconfirmedobservation Β· floor
CONTROL -- v2ecoli reproduces vEcoli on basal (the ParCa fit point): all observables within tolerance (4-seed gen-1).
What we saw: cell/dry +1.x%, protein/RNA <1.5%, growth ~+5% at seed-noise level -- all within_tol.

Conclusion verdicts

Three-track verdict β€” each result is computed from canonical fields (gate evaluator, run status, finding tiers). The basis is the author's rationale.

Biological validationPENDINGfrom gate evaluator Β· computed
Regression compatibilityPASSfrom run status Β· computed
Explanatory gainPARTIALfrom interpretation-tier findings Β· computed

Conditions β€” what we set up to test it

Baseline

Composite: v2ecoli.composites.ecoli_baseline.ecoli_baseline
(no overrides)

What we ran (1 simulation)

One row per concrete run: the model composite, what changes vs the reference baseline, the condition / length, and its status.

SimulationCompositeChanges vs baselineRunCLIStatus
basal-baselineecoli_baselinereference baselinebasalvwb run study basalcompleted

Visualisations from the latest run

v2ecoli vs vEcoli trajectories (basal, seed 0)❓ untracked
2026-07-22T12:10:50.863542 image/svg+xml Matplotlib v3.10.9, https://matplotlib.org/
Cell/dry/protein/RNA mass and growth rate over 4 generations, both engines from a matched ParCa initial state. Biomass overlays within tolerance; growth rate drifts.

Success criteria (8 tests β€” 8 ⏳ pending)

Each test makes a specific scientific claim with a machine-checkable criterion (measure + pass_if). Tests are now evaluated by code against the run (the run/outcome spine: RunReader β†’ evaluator): the pill shows the result, and the evidence line shows the measured value, whether it was computed by code or routed to an agent, and whether the code verdict agrees with the authored one (reconcile). ⏳ pending = the study hasn't run yet. Technical assertion + the exact evaluator are under "Technical details".

ungradedsummary report cardsummary-vs-vecoli
report card summary not generated yet β€” run the comparison.
driftparca report cardparca-vs-vecoli
within tolstatistical report cardstatistical-vs-vecoli
driftstandard report cardstandard-vs-vecoli
ungradedtrajectory report cardtrajectory-vs-vecoli
report card trajectory not generated yet β€” run the comparison.
ungradeddistribution report carddistribution-vs-vecoli
report card distribution not generated yet β€” run the comparison.
ungradedmetabolism report cardmetabolism-vs-vecoli
report card metabolism not generated yet β€” run the comparison.
ungradedcomposition report cardcomposition-vs-vecoli
report card composition not generated yet β€” run the comparison.

Conclusion synthesis

Read-only synthesis derived from the study's canonical fields (findings, limitations, follow-up proposals).

Claims
  • CONTROL -- v2ecoli reproduces vEcoli on basal (the ParCa fit point): all observables within tolerance (4-seed gen-1).
Evidence
  • cell/dry +1.x%, protein/RNA <1.5%, growth ~+5% at seed-noise level -- all within_tol.

Pipeline-gate decision

In progress
βœ“ Passednothing yet
βœ— Failednothing failing
β›” Blocks the next studynothing blocking
β†’ Immediate next action
Continue analysing run outcomes.
Pipeline gate & conclusion logic (technical)

Prerequisites: none (root study)

Enables: β€”

3.v2ecoli reproduces vEcoli with MetabolismRedux (acetate)πŸ§ͺ Preliminary
β–Ά Ran Β· 1 runTests: 8⏳⏳ Tests pending⚠ 1 clarity note
Does v2ecoli reproduce vEcoli on acetate when both engines swap FBA Metabolism for MetabolismRedux?
metabolism_redux_acetate Β· depth 0
β–Έ click to expand full study
3.metabolism_redux_acetateevaluatedRoot study (no dependencies)
β–΄ click to collapse full study
Model
The composite(s) this study runs and their parameters.
🧬 ecoli_baselinecondition="acetate" swap="ecoli-metabolism-redux"

Biology

Measured (6 seeds requested, 4 effective vEcoli, gen-1, matched timepoints): parca (t=0 initial-state match) within_tol. Statistical card (Welch t-test) overall verdict drift -- driven solely by growth-rate -- all four masses within_tol. Matched gen-1 median |delta|: cell 0.53%, dry 0.53%, protein 1.09%, rna 2.19%, growth 12.1%. No mismatch on any axis; active_RNAP/active_ribosome are ungraded (not in the local reader's observable set).

Overview

This study asks whether does v2ecoli reproduce vEcoli on acetate when both engines swap FBA Metabolism for MetabolismRedux?. We recorded 1 finding confirm the expected biology. Gate decision: In progress. Continue analysing run outcomes.

Purpose & background (study design)
Question. Does v2ecoli reproduce vEcoli on acetate when both engines swap FBA Metabolism for MetabolismRedux?

Detailed findings

Infrastructure / computational findings (1)

βœ“metabolism_redux_acetate-vs-vecoli-resultconfirmedobservation Β· floor
Measured (6 seeds requested, 4 effective vEcoli, gen-1, matched timepoints): parca (t=0 initial-state match) within_tol. Statistical card (Welch t-test) overall verdict drift -- driven solely by growth-rate -- all four masses within_tol. Matched gen-1 median |delta|: cell 0.53%, dry 0.53%, protein 1.09%, rna 2.19%, growth 12.1%. No mismatch on any axis; active_RNAP/active_ribosome are ungraded (not in the local reader's observable set).
What we saw: parca within_tol; cell/dry/protein/rna mass all within_tol (median |delta| 0.53%/0.53%/1.09%/2.19%); growth-rate median |delta| 12.1% (sole driver of the drift verdict).

Conclusion verdicts

Three-track verdict β€” each result is computed from canonical fields (gate evaluator, run status, finding tiers). The basis is the author's rationale.

Biological validationPENDINGfrom gate evaluator Β· computed
Regression compatibilityPASSfrom run status Β· computed
Explanatory gainPARTIALfrom interpretation-tier findings Β· computed

Conditions β€” what we set up to test it

Baseline

Composite: v2ecoli.composites.ecoli_baseline.ecoli_baseline
conditionacetate
swapecoli-metabolism-redux

What we ran (1 simulation)

One row per concrete run: the model composite, what changes vs the reference baseline, the condition / length, and its status.

SimulationCompositeChanges vs baselineRunCLIStatus
metabolism_redux_acetate-baselineecoli_baselinereference baselineacetatevwb run study metabolism_redux_acetatecompleted

Success criteria (8 tests β€” 8 ⏳ pending)

Each test makes a specific scientific claim with a machine-checkable criterion (measure + pass_if). Tests are now evaluated by code against the run (the run/outcome spine: RunReader β†’ evaluator): the pill shows the result, and the evidence line shows the measured value, whether it was computed by code or routed to an agent, and whether the code verdict agrees with the authored one (reconcile). ⏳ pending = the study hasn't run yet. Technical assertion + the exact evaluator are under "Technical details".

ungradedsummary report cardsummary-vs-vecoli
report card summary not generated yet β€” run the comparison.
within tolparca report cardparca-vs-vecoli
driftstatistical report cardstatistical-vs-vecoli
mismatchstandard report cardstandard-vs-vecoli
ungradedtrajectory report cardtrajectory-vs-vecoli
driftdistribution report carddistribution-vs-vecoli
driftmetabolism report cardmetabolism-vs-vecoli
mismatchcomposition report cardcomposition-vs-vecoli

Conclusion synthesis

Read-only synthesis derived from the study's canonical fields (findings, limitations, follow-up proposals).

Claims
  • Measured (6 seeds requested, 4 effective vEcoli, gen-1, matched timepoints): parca (t=0 initial-state match) within_tol. Statistical card (Welch t-test) overall verdict drift -- driven solely by growth-rate -- all four masses within_tol. Matched gen-1 median |delta|: cell 0.53%, dry 0.53%, protein 1.09%, rna 2.19%, growth 12.1%. No mismatch on any axis; active_RNAP/active_ribosome are ungraded (not in the local reader's observable set).
Evidence
  • parca within_tol; cell/dry/protein/rna mass all within_tol (median |delta| 0.53%/0.53%/1.09%/2.19%); growth-rate median |delta| 12.1% (sole driver of the drift verdict).

Pipeline-gate decision

In progress
βœ“ Passednothing yet
βœ— Failednothing failing
β›” Blocks the next studynothing blocking
β†’ Immediate next action
Continue analysing run outcomes.
Pipeline gate & conclusion logic (technical)

Prerequisites: none (root study)

Enables: β€”

4.v2ecoli reproduces vEcoli with MetabolismRedux (basal)πŸ§ͺ Preliminary
β–Ά Ran Β· 1 runTests: 8⏳⏳ Tests pending⚠ 1 clarity note
Does v2ecoli reproduce vEcoli on basal when both engines swap FBA Metabolism for MetabolismRedux?
metabolism_redux_basal Β· depth 0
β–Έ click to expand full study
4.metabolism_redux_basalevaluatedRoot study (no dependencies)
β–΄ click to collapse full study
Model
The composite(s) this study runs and their parameters.
🧬 ecoli_baselinecondition="basal" swap="ecoli-metabolism-redux"

Biology

Measured (6 seeds requested, 5 effective vEcoli, gen-1, matched timepoints): parca (t=0 initial-state match) within_tol. Statistical card (Welch t-test) overall verdict drift -- driven solely by growth-rate -- all four masses within_tol. Matched gen-1 median |delta|: cell 0.14%, dry 0.15%, protein 0.09%, rna 0.60%, growth 10.4%. No mismatch on any axis; active_RNAP/active_ribosome are ungraded (not in the local reader's observable set).

Overview

This study asks whether does v2ecoli reproduce vEcoli on basal when both engines swap FBA Metabolism for MetabolismRedux?. We recorded 1 finding confirm the expected biology. Gate decision: In progress. Continue analysing run outcomes.

Purpose & background (study design)
Question. Does v2ecoli reproduce vEcoli on basal when both engines swap FBA Metabolism for MetabolismRedux?

Detailed findings

Infrastructure / computational findings (1)

βœ“metabolism_redux_basal-vs-vecoli-resultconfirmedobservation Β· floor
Measured (6 seeds requested, 5 effective vEcoli, gen-1, matched timepoints): parca (t=0 initial-state match) within_tol. Statistical card (Welch t-test) overall verdict drift -- driven solely by growth-rate -- all four masses within_tol. Matched gen-1 median |delta|: cell 0.14%, dry 0.15%, protein 0.09%, rna 0.60%, growth 10.4%. No mismatch on any axis; active_RNAP/active_ribosome are ungraded (not in the local reader's observable set).
What we saw: parca within_tol; cell/dry/protein/rna mass all within_tol (median |delta| 0.14%/0.15%/0.09%/0.60%); growth-rate median |delta| 10.4% (sole driver of the drift verdict).

Conclusion verdicts

Three-track verdict β€” each result is computed from canonical fields (gate evaluator, run status, finding tiers). The basis is the author's rationale.

Biological validationPENDINGfrom gate evaluator Β· computed
Regression compatibilityPASSfrom run status Β· computed
Explanatory gainPARTIALfrom interpretation-tier findings Β· computed

Conditions β€” what we set up to test it

Baseline

Composite: v2ecoli.composites.ecoli_baseline.ecoli_baseline
conditionbasal
swapecoli-metabolism-redux

What we ran (1 simulation)

One row per concrete run: the model composite, what changes vs the reference baseline, the condition / length, and its status.

SimulationCompositeChanges vs baselineRunCLIStatus
metabolism_redux_basal-baselineecoli_baselinereference baselinebasalvwb run study metabolism_redux_basalcompleted

Success criteria (8 tests β€” 8 ⏳ pending)

Each test makes a specific scientific claim with a machine-checkable criterion (measure + pass_if). Tests are now evaluated by code against the run (the run/outcome spine: RunReader β†’ evaluator): the pill shows the result, and the evidence line shows the measured value, whether it was computed by code or routed to an agent, and whether the code verdict agrees with the authored one (reconcile). ⏳ pending = the study hasn't run yet. Technical assertion + the exact evaluator are under "Technical details".

ungradedsummary report cardsummary-vs-vecoli
report card summary not generated yet β€” run the comparison.
within tolparca report cardparca-vs-vecoli
driftstatistical report cardstatistical-vs-vecoli
mismatchstandard report cardstandard-vs-vecoli
ungradedtrajectory report cardtrajectory-vs-vecoli
mismatchdistribution report carddistribution-vs-vecoli
mismatchmetabolism report cardmetabolism-vs-vecoli
mismatchcomposition report cardcomposition-vs-vecoli

Conclusion synthesis

Read-only synthesis derived from the study's canonical fields (findings, limitations, follow-up proposals).

Claims
  • Measured (6 seeds requested, 5 effective vEcoli, gen-1, matched timepoints): parca (t=0 initial-state match) within_tol. Statistical card (Welch t-test) overall verdict drift -- driven solely by growth-rate -- all four masses within_tol. Matched gen-1 median |delta|: cell 0.14%, dry 0.15%, protein 0.09%, rna 0.60%, growth 10.4%. No mismatch on any axis; active_RNAP/active_ribosome are ungraded (not in the local reader's observable set).
Evidence
  • parca within_tol; cell/dry/protein/rna mass all within_tol (median |delta| 0.14%/0.15%/0.09%/0.60%); growth-rate median |delta| 10.4% (sole driver of the drift verdict).

Pipeline-gate decision

In progress
βœ“ Passednothing yet
βœ— Failednothing failing
β›” Blocks the next studynothing blocking
β†’ Immediate next action
Continue analysing run outcomes.
Pipeline gate & conclusion logic (technical)

Prerequisites: none (root study)

Enables: β€”

5.v2ecoli reproduces vEcoli with MetabolismRedux (no_oxygen)πŸ§ͺ Preliminary
β–Ά Ran Β· 1 runTests: 8⏳⏳ Tests pending⚠ 1 clarity note
Does v2ecoli reproduce vEcoli on no_oxygen when both engines swap FBA Metabolism for MetabolismRedux?
metabolism_redux_no_oxygen Β· depth 0
β–Έ click to expand full study
5.metabolism_redux_no_oxygenevaluatedRoot study (no dependencies)
β–΄ click to collapse full study
Model
The composite(s) this study runs and their parameters.
🧬 ecoli_baselinecondition="no_oxygen" swap="ecoli-metabolism-redux"

Biology

Measured (6 seeds requested, 5 effective vEcoli, gen-1, matched timepoints): parca (t=0 initial-state match) within_tol. Statistical card (Welch t-test) overall verdict drift -- driven solely by growth-rate -- all four masses within_tol (largest growth-rate delta of the five conditions). Matched gen-1 median |delta|: cell 0.70%, dry 0.70%, protein 0.20%, rna 1.23%, growth 15.3%. No mismatch on any axis; active_RNAP/active_ribosome are ungraded (not in the local reader's observable set).

Overview

This study asks whether does v2ecoli reproduce vEcoli on no_oxygen when both engines swap FBA Metabolism for MetabolismRedux?. We recorded 1 finding confirm the expected biology. Gate decision: In progress. Continue analysing run outcomes.

Purpose & background (study design)
Question. Does v2ecoli reproduce vEcoli on no_oxygen when both engines swap FBA Metabolism for MetabolismRedux?

Detailed findings

Infrastructure / computational findings (1)

βœ“metabolism_redux_no_oxygen-vs-vecoli-resultconfirmedobservation Β· floor
Measured (6 seeds requested, 5 effective vEcoli, gen-1, matched timepoints): parca (t=0 initial-state match) within_tol. Statistical card (Welch t-test) overall verdict drift -- driven solely by growth-rate -- all four masses within_tol (largest growth-rate delta of the five conditions). Matched gen-1 median |delta|: cell 0.70%, dry 0.70%, protein 0.20%, rna 1.23%, growth 15.3%. No mismatch on any axis; active_RNAP/active_ribosome are ungraded (not in the local reader's observable set).
What we saw: parca within_tol; cell/dry/protein/rna mass all within_tol (median |delta| 0.70%/0.70%/0.20%/1.23%); growth-rate median |delta| 15.3% (sole driver of the drift verdict).

Conclusion verdicts

Three-track verdict β€” each result is computed from canonical fields (gate evaluator, run status, finding tiers). The basis is the author's rationale.

Biological validationPENDINGfrom gate evaluator Β· computed
Regression compatibilityPASSfrom run status Β· computed
Explanatory gainPARTIALfrom interpretation-tier findings Β· computed

Conditions β€” what we set up to test it

Baseline

Composite: v2ecoli.composites.ecoli_baseline.ecoli_baseline
conditionno_oxygen
swapecoli-metabolism-redux

What we ran (1 simulation)

One row per concrete run: the model composite, what changes vs the reference baseline, the condition / length, and its status.

SimulationCompositeChanges vs baselineRunCLIStatus
metabolism_redux_no_oxygen-baselineecoli_baselinereference baselineno_oxygenvwb run study metabolism_redux_no_oxygencompleted

Success criteria (8 tests β€” 8 ⏳ pending)

Each test makes a specific scientific claim with a machine-checkable criterion (measure + pass_if). Tests are now evaluated by code against the run (the run/outcome spine: RunReader β†’ evaluator): the pill shows the result, and the evidence line shows the measured value, whether it was computed by code or routed to an agent, and whether the code verdict agrees with the authored one (reconcile). ⏳ pending = the study hasn't run yet. Technical assertion + the exact evaluator are under "Technical details".

ungradedsummary report cardsummary-vs-vecoli
report card summary not generated yet β€” run the comparison.
within tolparca report cardparca-vs-vecoli
driftstatistical report cardstatistical-vs-vecoli
mismatchstandard report cardstandard-vs-vecoli
ungradedtrajectory report cardtrajectory-vs-vecoli
mismatchdistribution report carddistribution-vs-vecoli
mismatchmetabolism report cardmetabolism-vs-vecoli
mismatchcomposition report cardcomposition-vs-vecoli

Conclusion synthesis

Read-only synthesis derived from the study's canonical fields (findings, limitations, follow-up proposals).

Claims
  • Measured (6 seeds requested, 5 effective vEcoli, gen-1, matched timepoints): parca (t=0 initial-state match) within_tol. Statistical card (Welch t-test) overall verdict drift -- driven solely by growth-rate -- all four masses within_tol (largest growth-rate delta of the five conditions). Matched gen-1 median |delta|: cell 0.70%, dry 0.70%, protein 0.20%, rna 1.23%, growth 15.3%. No mismatch on any axis; active_RNAP/active_ribosome are ungraded (not in the local reader's observable set).
Evidence
  • parca within_tol; cell/dry/protein/rna mass all within_tol (median |delta| 0.70%/0.70%/0.20%/1.23%); growth-rate median |delta| 15.3% (sole driver of the drift verdict).

Pipeline-gate decision

In progress
βœ“ Passednothing yet
βœ— Failednothing failing
β›” Blocks the next studynothing blocking
β†’ Immediate next action
Continue analysing run outcomes.
Pipeline gate & conclusion logic (technical)

Prerequisites: none (root study)

Enables: β€”

6.v2ecoli reproduces vEcoli with MetabolismRedux (succinate)πŸ§ͺ Preliminary
β–Ά Ran Β· 1 runTests: 8⏳⏳ Tests pending⚠ 1 clarity note
Does v2ecoli reproduce vEcoli on succinate when both engines swap FBA Metabolism for MetabolismRedux?
metabolism_redux_succinate Β· depth 0
β–Έ click to expand full study
6.metabolism_redux_succinateevaluatedRoot study (no dependencies)
β–΄ click to collapse full study
Model
The composite(s) this study runs and their parameters.
🧬 ecoli_baselinecondition="succinate" swap="ecoli-metabolism-redux"

Biology

Measured (6 seeds requested, 5 effective vEcoli, gen-1, matched timepoints): parca (t=0 initial-state match) within_tol. Statistical card (Welch t-test) overall verdict drift -- driven solely by growth-rate -- all four masses within_tol (largest mass deltas of the five conditions, still within tolerance). Matched gen-1 median |delta|: cell 1.49%, dry 1.50%, protein 0.92%, rna 3.50%, growth 13.4%. No mismatch on any axis; active_RNAP/active_ribosome are ungraded (not in the local reader's observable set).

Overview

This study asks whether does v2ecoli reproduce vEcoli on succinate when both engines swap FBA Metabolism for MetabolismRedux?. We recorded 1 finding confirm the expected biology. Gate decision: In progress. Continue analysing run outcomes.

Purpose & background (study design)
Question. Does v2ecoli reproduce vEcoli on succinate when both engines swap FBA Metabolism for MetabolismRedux?

Detailed findings

Infrastructure / computational findings (1)

βœ“metabolism_redux_succinate-vs-vecoli-resultconfirmedobservation Β· floor
Measured (6 seeds requested, 5 effective vEcoli, gen-1, matched timepoints): parca (t=0 initial-state match) within_tol. Statistical card (Welch t-test) overall verdict drift -- driven solely by growth-rate -- all four masses within_tol (largest mass deltas of the five conditions, still within tolerance). Matched gen-1 median |delta|: cell 1.49%, dry 1.50%, protein 0.92%, rna 3.50%, growth 13.4%. No mismatch on any axis; active_RNAP/active_ribosome are ungraded (not in the local reader's observable set).
What we saw: parca within_tol; cell/dry/protein/rna mass all within_tol (median |delta| 1.49%/1.50%/0.92%/3.50%); growth-rate median |delta| 13.4% (sole driver of the drift verdict).

Conclusion verdicts

Three-track verdict β€” each result is computed from canonical fields (gate evaluator, run status, finding tiers). The basis is the author's rationale.

Biological validationPENDINGfrom gate evaluator Β· computed
Regression compatibilityPASSfrom run status Β· computed
Explanatory gainPARTIALfrom interpretation-tier findings Β· computed

Conditions β€” what we set up to test it

Baseline

Composite: v2ecoli.composites.ecoli_baseline.ecoli_baseline
conditionsuccinate
swapecoli-metabolism-redux

What we ran (1 simulation)

One row per concrete run: the model composite, what changes vs the reference baseline, the condition / length, and its status.

SimulationCompositeChanges vs baselineRunCLIStatus
metabolism_redux_succinate-baselineecoli_baselinereference baselinesuccinatevwb run study metabolism_redux_succinatecompleted

Success criteria (8 tests β€” 8 ⏳ pending)

Each test makes a specific scientific claim with a machine-checkable criterion (measure + pass_if). Tests are now evaluated by code against the run (the run/outcome spine: RunReader β†’ evaluator): the pill shows the result, and the evidence line shows the measured value, whether it was computed by code or routed to an agent, and whether the code verdict agrees with the authored one (reconcile). ⏳ pending = the study hasn't run yet. Technical assertion + the exact evaluator are under "Technical details".

ungradedsummary report cardsummary-vs-vecoli
report card summary not generated yet β€” run the comparison.
within tolparca report cardparca-vs-vecoli
driftstatistical report cardstatistical-vs-vecoli
mismatchstandard report cardstandard-vs-vecoli
ungradedtrajectory report cardtrajectory-vs-vecoli
driftdistribution report carddistribution-vs-vecoli
mismatchmetabolism report cardmetabolism-vs-vecoli
mismatchcomposition report cardcomposition-vs-vecoli

Conclusion synthesis

Read-only synthesis derived from the study's canonical fields (findings, limitations, follow-up proposals).

Claims
  • Measured (6 seeds requested, 5 effective vEcoli, gen-1, matched timepoints): parca (t=0 initial-state match) within_tol. Statistical card (Welch t-test) overall verdict drift -- driven solely by growth-rate -- all four masses within_tol (largest mass deltas of the five conditions, still within tolerance). Matched gen-1 median |delta|: cell 1.49%, dry 1.50%, protein 0.92%, rna 3.50%, growth 13.4%. No mismatch on any axis; active_RNAP/active_ribosome are ungraded (not in the local reader's observable set).
Evidence
  • parca within_tol; cell/dry/protein/rna mass all within_tol (median |delta| 1.49%/1.50%/0.92%/3.50%); growth-rate median |delta| 13.4% (sole driver of the drift verdict).

Pipeline-gate decision

In progress
βœ“ Passednothing yet
βœ— Failednothing failing
β›” Blocks the next studynothing blocking
β†’ Immediate next action
Continue analysing run outcomes.
Pipeline gate & conclusion logic (technical)

Prerequisites: none (root study)

Enables: β€”

7.v2ecoli reproduces vEcoli with MetabolismRedux (with_aa)πŸ§ͺ Preliminary
β–Ά Ran Β· 1 runTests: 8⏳⏳ Tests pending⚠ 1 clarity note
Does v2ecoli reproduce vEcoli on with_aa when both engines swap FBA Metabolism for MetabolismRedux?
metabolism_redux_with_aa Β· depth 0
β–Έ click to expand full study
7.metabolism_redux_with_aaevaluatedRoot study (no dependencies)
β–΄ click to collapse full study
Model
The composite(s) this study runs and their parameters.
🧬 ecoli_baselinecondition="with_aa" swap="ecoli-metabolism-redux"

Biology

Measured (6 seeds requested, 5 effective vEcoli, gen-1, matched timepoints): parca (t=0 initial-state match) within_tol. Statistical card (Welch t-test) overall verdict within_tol -- within_tol on ALL axes including growth rate -- the only fully within_tol condition of the five. Matched gen-1 median |delta|: cell 0.86%, dry 0.86%, protein 0.13%, rna 0.34%, growth 4.3%. No mismatch on any axis; active_RNAP/active_ribosome are ungraded (not in the local reader's observable set).

Overview

This study asks whether does v2ecoli reproduce vEcoli on with_aa when both engines swap FBA Metabolism for MetabolismRedux?. We recorded 1 finding confirm the expected biology. Gate decision: In progress. Continue analysing run outcomes.

Purpose & background (study design)
Question. Does v2ecoli reproduce vEcoli on with_aa when both engines swap FBA Metabolism for MetabolismRedux?

Detailed findings

Infrastructure / computational findings (1)

βœ“metabolism_redux_with_aa-vs-vecoli-resultconfirmedobservation Β· floor
Measured (6 seeds requested, 5 effective vEcoli, gen-1, matched timepoints): parca (t=0 initial-state match) within_tol. Statistical card (Welch t-test) overall verdict within_tol -- within_tol on ALL axes including growth rate -- the only fully within_tol condition of the five. Matched gen-1 median |delta|: cell 0.86%, dry 0.86%, protein 0.13%, rna 0.34%, growth 4.3%. No mismatch on any axis; active_RNAP/active_ribosome are ungraded (not in the local reader's observable set).
What we saw: parca within_tol; cell/dry/protein/rna mass all within_tol (median |delta| 0.86%/0.86%/0.13%/0.34%); growth-rate median |delta| 4.3% (within_tol, no drift).

Conclusion verdicts

Three-track verdict β€” each result is computed from canonical fields (gate evaluator, run status, finding tiers). The basis is the author's rationale.

Biological validationPENDINGfrom gate evaluator Β· computed
Regression compatibilityPASSfrom run status Β· computed
Explanatory gainPARTIALfrom interpretation-tier findings Β· computed

Conditions β€” what we set up to test it

Baseline

Composite: v2ecoli.composites.ecoli_baseline.ecoli_baseline
conditionwith_aa
swapecoli-metabolism-redux

What we ran (1 simulation)

One row per concrete run: the model composite, what changes vs the reference baseline, the condition / length, and its status.

SimulationCompositeChanges vs baselineRunCLIStatus
metabolism_redux_with_aa-baselineecoli_baselinereference baselinewith_aavwb run study metabolism_redux_with_aacompleted

Success criteria (8 tests β€” 8 ⏳ pending)

Each test makes a specific scientific claim with a machine-checkable criterion (measure + pass_if). Tests are now evaluated by code against the run (the run/outcome spine: RunReader β†’ evaluator): the pill shows the result, and the evidence line shows the measured value, whether it was computed by code or routed to an agent, and whether the code verdict agrees with the authored one (reconcile). ⏳ pending = the study hasn't run yet. Technical assertion + the exact evaluator are under "Technical details".

ungradedsummary report cardsummary-vs-vecoli
report card summary not generated yet β€” run the comparison.
within tolparca report cardparca-vs-vecoli
within tolstatistical report cardstatistical-vs-vecoli
within tolstandard report cardstandard-vs-vecoli
ungradedtrajectory report cardtrajectory-vs-vecoli
mismatchdistribution report carddistribution-vs-vecoli
within tolmetabolism report cardmetabolism-vs-vecoli
within tolcomposition report cardcomposition-vs-vecoli

Conclusion synthesis

Read-only synthesis derived from the study's canonical fields (findings, limitations, follow-up proposals).

Claims
  • Measured (6 seeds requested, 5 effective vEcoli, gen-1, matched timepoints): parca (t=0 initial-state match) within_tol. Statistical card (Welch t-test) overall verdict within_tol -- within_tol on ALL axes including growth rate -- the only fully within_tol condition of the five. Matched gen-1 median |delta|: cell 0.86%, dry 0.86%, protein 0.13%, rna 0.34%, growth 4.3%. No mismatch on any axis; active_RNAP/active_ribosome are ungraded (not in the local reader's observable set).
Evidence
  • parca within_tol; cell/dry/protein/rna mass all within_tol (median |delta| 0.86%/0.86%/0.13%/0.34%); growth-rate median |delta| 4.3% (within_tol, no drift).

Pipeline-gate decision

In progress
βœ“ Passednothing yet
βœ— Failednothing failing
β›” Blocks the next studynothing blocking
β†’ Immediate next action
Continue analysing run outcomes.
Pipeline gate & conclusion logic (technical)

Prerequisites: none (root study)

Enables: β€”

8.v2ecoli reproduces vEcoli on no_oxygenπŸ§ͺ Preliminary
β–Ά Ran Β· 1 runTests: 8⏳⏳ Tests pending⚠ 1 clarity note
Does v2ecoli reproduce vEcoli on the no_oxygen condition?
no_oxygen Β· depth 0
β–Έ click to expand full study
8.no_oxygenevaluatedRoot study (no dependencies)
β–΄ click to collapse full study
Model
The composite(s) this study runs and their parameters.
🧬 ecoli_baselinemedia="minimal_minus_oxygen"

Biology

v2ecoli reproduces vEcoli on no_oxygen (4-seed gen-1, Welch t-test): cell, dry, protein and RNA mass within tolerance (4/5 within_tol); growth rate shows a non-significant drift. 2 of 4 seeds hit a pre-existing anaerobic-FBA numerical edge case (BIOTIN homeostatic target near zero); graded on the clean seeds.

Overview

This study asks whether does v2ecoli reproduce vEcoli on the no_oxygen condition?. Gate decision: In progress. Continue analysing run outcomes.

Purpose & background (study design)
Question. Does v2ecoli reproduce vEcoli on the no_oxygen condition?

Visualizations

00-no_oxygen-observables

Hand-authored figure (8 KB) from reports/figures/no_oxygen/.

01-fix-headline

Hand-authored figure (9 KB) from reports/figures/no_oxygen/.

02-growth-before-after

Hand-authored figure (8 KB) from reports/figures/no_oxygen/.

03-rna-before-after

Hand-authored figure (8 KB) from reports/figures/no_oxygen/.

04-rnap-rootcause

Hand-authored figure (8 KB) from reports/figures/no_oxygen/.

05-nutrient-gradient

Hand-authored figure (8 KB) from reports/figures/no_oxygen/.

Detailed findings

Infrastructure / computational findings (1)

◐no_oxygen-vs-vecoli-resultpartial resultobservation Β· floor
v2ecoli reproduces vEcoli on no_oxygen (4-seed gen-1, Welch t-test): cell, dry, protein and RNA mass within tolerance (4/5 within_tol); growth rate shows a non-significant drift. 2 of 4 seeds hit a pre-existing anaerobic-FBA numerical edge case (BIOTIN homeostatic target near zero); graded on the clean seeds.
What we saw: growth +6.7% (p=0.50, not significant); RNA +1.0%; cell +1.0%; dry +1.0%; protein +0.5% -- 4/5 within_tol, overall drift.

Conclusion verdicts

Three-track verdict β€” each result is computed from canonical fields (gate evaluator, run status, finding tiers). The basis is the author's rationale.

Biological validationPENDINGfrom gate evaluator Β· computed
Regression compatibilityPASSfrom run status Β· computed
Explanatory gainPARTIALfrom interpretation-tier findings Β· computed

Conditions β€” what we set up to test it

Baseline

Composite: v2ecoli.composites.ecoli_baseline.ecoli_baseline
mediaminimal_minus_oxygen

What we ran (1 simulation)

One row per concrete run: the model composite, what changes vs the reference baseline, the condition / length, and its status.

SimulationCompositeChanges vs baselineRunCLIStatus
no_oxygen-baselineecoli_baselinereference baselineno_oxygenvwb run study no_oxygencompleted

Success criteria (8 tests β€” 8 ⏳ pending)

Each test makes a specific scientific claim with a machine-checkable criterion (measure + pass_if). Tests are now evaluated by code against the run (the run/outcome spine: RunReader β†’ evaluator): the pill shows the result, and the evidence line shows the measured value, whether it was computed by code or routed to an agent, and whether the code verdict agrees with the authored one (reconcile). ⏳ pending = the study hasn't run yet. Technical assertion + the exact evaluator are under "Technical details".

ungradedsummary report cardsummary-vs-vecoli
report card summary not generated yet β€” run the comparison.
ungradedparca report cardparca-vs-vecoli
report card parca not generated yet β€” run the comparison.
driftstatistical report cardstatistical-vs-vecoli
driftstandard report cardstandard-vs-vecoli
ungradedtrajectory report cardtrajectory-vs-vecoli
report card trajectory not generated yet β€” run the comparison.
ungradeddistribution report carddistribution-vs-vecoli
report card distribution not generated yet β€” run the comparison.
ungradedmetabolism report cardmetabolism-vs-vecoli
report card metabolism not generated yet β€” run the comparison.
ungradedcomposition report cardcomposition-vs-vecoli
report card composition not generated yet β€” run the comparison.

Conclusion synthesis

Read-only synthesis derived from the study's canonical fields (findings, limitations, follow-up proposals).

Claims
  • v2ecoli reproduces vEcoli on no_oxygen (4-seed gen-1, Welch t-test): cell, dry, protein and RNA mass within tolerance (4/5 within_tol); growth rate shows a non-significant drift. 2 of 4 seeds hit a pre-existing anaerobic-FBA numerical edge case (BIOTIN homeostatic target near zero); graded on the clean seeds.
Evidence
  • growth +6.7% (p=0.50, not significant); RNA +1.0%; cell +1.0%; dry +1.0%; protein +0.5% -- 4/5 within_tol, overall drift.

Pipeline-gate decision

In progress
βœ“ Passednothing yet
βœ— Failednothing failing
β›” Blocks the next studynothing blocking
β†’ Immediate next action
Continue analysing run outcomes.
Pipeline gate & conclusion logic (technical)

Prerequisites: none (root study)

Enables: β€”

9.ParCa initial-state agreement β€” v2ecoli vs vEcoliβœ… Passing
β–Ά Ran Β· 1 runTests: 1βœ“βœ… Passed
Do v2ecoli and vEcoli start from the same ParCa-fit initial state (basal)?
parca Β· depth 0
1/1 tests passing
β–Έ click to expand full study
9.parcaevaluatedRoot study (no dependencies)
β–΄ click to collapse full study
Model
The composite(s) this study runs and their parameters.
🧬 ecoli_baselinedefault parameters

Biology

v2ecoli and vEcoli start from the same ParCa-fit initial state on basal: all 4 mass observables within tolerance.

Overview

This study asks whether do v2ecoli and vEcoli start from the same ParCa-fit initial state (basal)?. We recorded 1 finding confirm the expected biology. Gate decision: Passed. Gate cleared. No declared downstream studies β€” review pipeline_gate.enables.

Purpose & background (study design)
Question. Do v2ecoli and vEcoli start from the same ParCa-fit initial state (basal)?

Detailed findings

Infrastructure / computational findings (1)

βœ“parca-parcaconfirmedobservation Β· floor
v2ecoli and vEcoli start from the same ParCa-fit initial state on basal: all 4 mass observables within tolerance.
What we saw: 4 within tolerance (cell mass (fg), dry mass (fg), protein mass (fg), RNA mass (fg))

Conclusion verdicts

Three-track verdict β€” each result is computed from canonical fields (gate evaluator, run status, finding tiers). The basis is the author's rationale.

Biological validationPENDINGfrom gate evaluator Β· computed
Regression compatibilityPASSfrom run status Β· computed
Explanatory gainPARTIALfrom interpretation-tier findings Β· computed

Conditions β€” what we set up to test it

Baseline

Composite: v2ecoli.composites.ecoli_baseline.ecoli_baseline
(no overrides)

What we ran (1 simulation)

One row per concrete run: the model composite, what changes vs the reference baseline, the condition / length, and its status.

SimulationCompositeChanges vs baselineRunCLIStatus
parca-baselineecoli_baselinereference baselinebasalvwb run study parcacompleted

Success criteria (1 tests β€” 1 βœ“ passed)

Each test makes a specific scientific claim with a machine-checkable criterion (measure + pass_if). Tests are now evaluated by code against the run (the run/outcome spine: RunReader β†’ evaluator): the pill shows the result, and the evidence line shows the measured value, whether it was computed by code or routed to an agent, and whether the code verdict agrees with the authored one (reconcile). ⏳ pending = the study hasn't run yet. Technical assertion + the exact evaluator are under "Technical details".

within tolparca report cardparca-vs-vecoli

Conclusion synthesis

Read-only synthesis derived from the study's canonical fields (findings, limitations, follow-up proposals).

Claims
  • v2ecoli and vEcoli start from the same ParCa-fit initial state on basal: all 4 mass observables within tolerance.
Evidence
  • 4 within tolerance (cell mass (fg), dry mass (fg), protein mass (fg), RNA mass (fg))

Pipeline-gate decision

Passed
βœ“ Passed
  • parca-vs-vecoli
βœ— Failednothing failing
β›” Blocks the next studynothing blocking
β†’ Immediate next action
Gate cleared. No declared downstream studies β€” review pipeline_gate.enables.
10.Statistical equivalence β€” v2ecoli baseline vs vEcoli (basal, 4 seeds)βœ… Passing
β–Ά Ran Β· 1 runTests: 1βœ“ Β· 1⏳⏳ Tests pending⚠ 1 clarity note
Is v2ecoli statistically equivalent to vEcoli on basal across 4 seeds?
statistical Β· depth 0
1/1 tests passing
β–Έ click to expand full study
10.statisticalevaluatedRoot study (no dependencies)
β–΄ click to collapse full study
Model
The composite(s) this study runs and their parameters.
🧬 ecoli_baselinedefault parameters

Biology

v2ecoli is statistically equivalent to vEcoli on basal across 4 seeds: all 5 observables within tolerance.

Overview

This study asks whether is v2ecoli statistically equivalent to vEcoli on basal across 4 seeds?. We recorded 1 finding confirm the expected biology. Gate decision: Passed. Gate cleared. No declared downstream studies β€” review pipeline_gate.enables.

Purpose & background (study design)
Question. Is v2ecoli statistically equivalent to vEcoli on basal across 4 seeds?

Detailed findings

Infrastructure / computational findings (1)

βœ“statistical-statisticalconfirmedobservation Β· floor
v2ecoli is statistically equivalent to vEcoli on basal across 4 seeds: all 5 observables within tolerance.
What we saw: 5 within tolerance (Cell mass, Growth rate, Dry mass, Protein mass, RNA mass)

Conclusion verdicts

Three-track verdict β€” each result is computed from canonical fields (gate evaluator, run status, finding tiers). The basis is the author's rationale.

Biological validationPENDINGfrom gate evaluator Β· computed
Regression compatibilityPASSfrom run status Β· computed
Explanatory gainPARTIALfrom interpretation-tier findings Β· computed

Conditions β€” what we set up to test it

Baseline

Composite: v2ecoli.composites.ecoli_baseline.ecoli_baseline
(no overrides)

What we ran (1 simulation)

One row per concrete run: the model composite, what changes vs the reference baseline, the condition / length, and its status.

SimulationCompositeChanges vs baselineRunCLIStatus
statistical-baselineecoli_baselinereference baselinebasalvwb run study statisticalcompleted

Visualisations from the latest run

Statistical equivalence (v2ecoli vs vEcoli, basal, 4 seeds)❓ untracked
2026-07-22T10:56:29.530935 image/svg+xml Matplotlib v3.10.9, https://matplotlib.org/
All 5 graded observables within the Β±5% tolerance band and negligible effect size (|d|<0.2) across 4 seeds β€” statistically equivalent. p = two-sample t-test on per-seed means.

Success criteria (2 tests β€” 1 βœ“ passed Β· 1 ⏳ pending)

Each test makes a specific scientific claim with a machine-checkable criterion (measure + pass_if). Tests are now evaluated by code against the run (the run/outcome spine: RunReader β†’ evaluator): the pill shows the result, and the evidence line shows the measured value, whether it was computed by code or routed to an agent, and whether the code verdict agrees with the authored one (reconcile). ⏳ pending = the study hasn't run yet. Technical assertion + the exact evaluator are under "Technical details".

ungradedconfig report cardconfig-vs-vecoli
within tolstatistical report cardstatistical-vs-vecoli

Conclusion synthesis

Read-only synthesis derived from the study's canonical fields (findings, limitations, follow-up proposals).

Claims
  • v2ecoli is statistically equivalent to vEcoli on basal across 4 seeds: all 5 observables within tolerance.
Evidence
  • 5 within tolerance (Cell mass, Growth rate, Dry mass, Protein mass, RNA mass)

Pipeline-gate decision

Passed
βœ“ Passed
  • statistical-vs-vecoli
βœ— Failednothing failing
β›” Blocks the next studynothing blocking
β†’ Immediate next action
Gate cleared. No declared downstream studies β€” review pipeline_gate.enables.
11.v2ecoli reproduces vEcoli on succinateπŸ§ͺ Preliminary
β–Ά Ran Β· 1 runTests: 8⏳⏳ Tests pending⚠ 1 clarity note
Does v2ecoli reproduce vEcoli on the succinate condition?
succinate Β· depth 0
β–Έ click to expand full study
11.succinateevaluatedRoot study (no dependencies)
β–΄ click to collapse full study
Model
The composite(s) this study runs and their parameters.
🧬 ecoli_baselinemedia="minimal_succinate"

Biology

v2ecoli reproduces vEcoli on succinate (4-seed gen-1, Welch t-test): cell, dry, RNA and protein mass within tolerance (4/5 within_tol); growth rate shows a non-significant drift at basal noise level.

Overview

This study asks whether does v2ecoli reproduce vEcoli on the succinate condition?. Gate decision: In progress. Continue analysing run outcomes.

Purpose & background (study design)
Question. Does v2ecoli reproduce vEcoli on the succinate condition?

Visualizations

00-succinate-observables

Hand-authored figure (8 KB) from reports/figures/succinate/.

01-fix-headline

Hand-authored figure (9 KB) from reports/figures/succinate/.

02-growth-before-after

Hand-authored figure (8 KB) from reports/figures/succinate/.

03-rna-before-after

Hand-authored figure (8 KB) from reports/figures/succinate/.

04-rnap-rootcause

Hand-authored figure (8 KB) from reports/figures/succinate/.

05-nutrient-gradient

Hand-authored figure (8 KB) from reports/figures/succinate/.

Detailed findings

Infrastructure / computational findings (1)

◐succinate-vs-vecoli-resultpartial resultobservation Β· floor
v2ecoli reproduces vEcoli on succinate (4-seed gen-1, Welch t-test): cell, dry, RNA and protein mass within tolerance (4/5 within_tol); growth rate shows a non-significant drift at basal noise level.
What we saw: growth -6.6% (p=0.44, not significant); RNA -2.2%; cell -1.3%; dry -1.4%; protein -0.8% -- overall drift.

Conclusion verdicts

Three-track verdict β€” each result is computed from canonical fields (gate evaluator, run status, finding tiers). The basis is the author's rationale.

Biological validationPENDINGfrom gate evaluator Β· computed
Regression compatibilityPASSfrom run status Β· computed
Explanatory gainPARTIALfrom interpretation-tier findings Β· computed

Conditions β€” what we set up to test it

Baseline

Composite: v2ecoli.composites.ecoli_baseline.ecoli_baseline
mediaminimal_succinate

What we ran (1 simulation)

One row per concrete run: the model composite, what changes vs the reference baseline, the condition / length, and its status.

SimulationCompositeChanges vs baselineRunCLIStatus
succinate-baselineecoli_baselinereference baselinesuccinatevwb run study succinatecompleted

Success criteria (8 tests β€” 8 ⏳ pending)

Each test makes a specific scientific claim with a machine-checkable criterion (measure + pass_if). Tests are now evaluated by code against the run (the run/outcome spine: RunReader β†’ evaluator): the pill shows the result, and the evidence line shows the measured value, whether it was computed by code or routed to an agent, and whether the code verdict agrees with the authored one (reconcile). ⏳ pending = the study hasn't run yet. Technical assertion + the exact evaluator are under "Technical details".

ungradedsummary report cardsummary-vs-vecoli
report card summary not generated yet β€” run the comparison.
ungradedparca report cardparca-vs-vecoli
report card parca not generated yet β€” run the comparison.
driftstatistical report cardstatistical-vs-vecoli
mismatchstandard report cardstandard-vs-vecoli
ungradedtrajectory report cardtrajectory-vs-vecoli
report card trajectory not generated yet β€” run the comparison.
ungradeddistribution report carddistribution-vs-vecoli
report card distribution not generated yet β€” run the comparison.
ungradedmetabolism report cardmetabolism-vs-vecoli
report card metabolism not generated yet β€” run the comparison.
ungradedcomposition report cardcomposition-vs-vecoli
report card composition not generated yet β€” run the comparison.

Conclusion synthesis

Read-only synthesis derived from the study's canonical fields (findings, limitations, follow-up proposals).

Claims
  • v2ecoli reproduces vEcoli on succinate (4-seed gen-1, Welch t-test): cell, dry, RNA and protein mass within tolerance (4/5 within_tol); growth rate shows a non-significant drift at basal noise level.
Evidence
  • growth -6.6% (p=0.44, not significant); RNA -2.2%; cell -1.3%; dry -1.4%; protein -0.8% -- overall drift.

Pipeline-gate decision

In progress
βœ“ Passednothing yet
βœ— Failednothing failing
β›” Blocks the next studynothing blocking
β†’ Immediate next action
Continue analysing run outcomes.
Pipeline gate & conclusion logic (technical)

Prerequisites: none (root study)

Enables: β€”

12.v2ecoli reproduces vEcoli on with_aaπŸ§ͺ Preliminary
β–Ά Ran Β· 1 runTests: 8⏳⏳ Tests pending⚠ 1 clarity note
Does v2ecoli reproduce vEcoli on the with_aa condition?
with_aa Β· depth 0
β–Έ click to expand full study
12.with_aaevaluatedRoot study (no dependencies)
β–΄ click to collapse full study
Model
The composite(s) this study runs and their parameters.
🧬 ecoli_baselinemedia="minimal_plus_amino_acids"

Biology

v2ecoli reproduces vEcoli on with_aa (4-seed gen-1, Welch t-test): all five observables (cell, dry, protein, RNA mass, growth rate) are within tolerance.

Overview

This study asks whether does v2ecoli reproduce vEcoli on the with_aa condition?. We recorded 1 finding confirm the expected biology. Gate decision: In progress. Continue analysing run outcomes.

Purpose & background (study design)
Question. Does v2ecoli reproduce vEcoli on the with_aa condition?

Visualizations

00-with_aa-observables

Hand-authored figure (8 KB) from reports/figures/with_aa/.

01-fix-headline

Hand-authored figure (9 KB) from reports/figures/with_aa/.

02-growth-before-after

Hand-authored figure (8 KB) from reports/figures/with_aa/.

03-rna-before-after

Hand-authored figure (8 KB) from reports/figures/with_aa/.

04-rnap-rootcause

Hand-authored figure (8 KB) from reports/figures/with_aa/.

05-nutrient-gradient

Hand-authored figure (8 KB) from reports/figures/with_aa/.

Detailed findings

Infrastructure / computational findings (1)

βœ“with_aa-vs-vecoli-resultconfirmedobservation Β· floor
v2ecoli reproduces vEcoli on with_aa (4-seed gen-1, Welch t-test): all five observables (cell, dry, protein, RNA mass, growth rate) are within tolerance.
What we saw: growth +1.7% (p=0.53); RNA +0.4% (p=0.89); cell +0.9%; dry +0.9%; protein -0.0% -- all within_tol.

Conclusion verdicts

Three-track verdict β€” each result is computed from canonical fields (gate evaluator, run status, finding tiers). The basis is the author's rationale.

Biological validationPENDINGfrom gate evaluator Β· computed
Regression compatibilityPASSfrom run status Β· computed
Explanatory gainPARTIALfrom interpretation-tier findings Β· computed

Conditions β€” what we set up to test it

Baseline

Composite: v2ecoli.composites.ecoli_baseline.ecoli_baseline
mediaminimal_plus_amino_acids

What we ran (1 simulation)

One row per concrete run: the model composite, what changes vs the reference baseline, the condition / length, and its status.

SimulationCompositeChanges vs baselineRunCLIStatus
with_aa-baselineecoli_baselinereference baselinewith_aavwb run study with_aacompleted

Success criteria (8 tests β€” 8 ⏳ pending)

Each test makes a specific scientific claim with a machine-checkable criterion (measure + pass_if). Tests are now evaluated by code against the run (the run/outcome spine: RunReader β†’ evaluator): the pill shows the result, and the evidence line shows the measured value, whether it was computed by code or routed to an agent, and whether the code verdict agrees with the authored one (reconcile). ⏳ pending = the study hasn't run yet. Technical assertion + the exact evaluator are under "Technical details".

ungradedsummary report cardsummary-vs-vecoli
report card summary not generated yet β€” run the comparison.
ungradedparca report cardparca-vs-vecoli
report card parca not generated yet β€” run the comparison.
within tolstatistical report cardstatistical-vs-vecoli
driftstandard report cardstandard-vs-vecoli
ungradedtrajectory report cardtrajectory-vs-vecoli
report card trajectory not generated yet β€” run the comparison.
ungradeddistribution report carddistribution-vs-vecoli
report card distribution not generated yet β€” run the comparison.
ungradedmetabolism report cardmetabolism-vs-vecoli
report card metabolism not generated yet β€” run the comparison.
ungradedcomposition report cardcomposition-vs-vecoli
report card composition not generated yet β€” run the comparison.

Conclusion synthesis

Read-only synthesis derived from the study's canonical fields (findings, limitations, follow-up proposals).

Claims
  • v2ecoli reproduces vEcoli on with_aa (4-seed gen-1, Welch t-test): all five observables (cell, dry, protein, RNA mass, growth rate) are within tolerance.
Evidence
  • growth +1.7% (p=0.53); RNA +0.4% (p=0.89); cell +0.9%; dry +0.9%; protein -0.0% -- all within_tol.

Pipeline-gate decision

In progress
βœ“ Passednothing yet
βœ— Failednothing failing
β›” Blocks the next studynothing blocking
β†’ Immediate next action
Continue analysing run outcomes.
Pipeline gate & conclusion logic (technical)

Prerequisites: none (root study)

Enables: β€”

Appendices

Method-grading and verification detail β€” kept at the back, after the main narrative.

πŸ”¬ Evidence & rigor β€” how well the method defends its claims 1/5 investigation rigor dimensions addressed Β· 4 gap(s)

Deterministic feedback on how well the method defends its claims against a skeptical reader β€” a method-level judgement, distinct from the per-study model verdicts above. Computed from declared fields, not judged. Gaps are an invitation to add negative controls, replicate across seeds, weigh alternative explanations, state falsifiability, or add an adversarial study.

βœ—
Adversarial testing C10 C12 C15
no adversarial study β€” add one that tries to BREAK the criteria: mimic / parasitic-or-dependent / externally-maintained / random-cyclic systems that should NOT qualify
βœ“
Falsification exposure C1
the framework has been shown to reject at least one system (a discriminating negative control, an adversarial study, or a non-passing result)
βœ—
Comparative framing C13
no competing theoretical frameworks compared (viability theory, organizational / constraint closure, active inference) β€” show the findings uniquely support this lens
βœ—
Hypothesis competition C6 C16
no competing hypotheses[] declared β€” state β‰₯2 rival explanations with predictions so the evidence can adjudicate between them
βœ—
Per-study rigor gaps C2 C4 C6
108 rigor gap(s) across 12 member study(ies)

Per-study rigor

parca β€” 2/12 rigor dimensions addressed Β· 9 gap(s)
βœ—
Replication C4
single run β€” no replication across seeds declared (add robustness.seeds or simulation_set.seeds)
βœ—
Controls & calibration C1 C2 C4
no controls β€” declare a system that SHOULD fail the criteria (externally-maintained / -supplied) plus a clearly-passing / borderline case so the metric is calibrated, not just asserted
βœ—
Alternative hypotheses C6
no alternative hypotheses declared
⚠
Claim discipline C3
findings not tiered β€” label each observation / mechanism / interpretation
βœ—
Falsifiability C5 C1
criteria read as tailored-to-succeed β€” add a falsifiability note (study.falsifiability)
βœ“
Engineered vs emergent C7
no interpretation-tier claim that requires the distinction
βœ—
Limitations stated C8 C11
no limitations / 'what this does not show' β€” add a short bound on the claim (scope/fidelity of the model, what is NOT demonstrated)
βœ—
Next steps next-steps
no discovery_implications or follow_up_studies β€” state what this study changes and what to investigate next (the Decide phase)
βœ“
Threshold provenance C9 C5
no numeric acceptance bands requiring provenance
βœ—
Metric calibration ladder C4 C2 C20
no calibration_ladder declared β€” index controls[] by known_fail / known_pass / borderline / stress rungs so the metric is calibrated across its range, not just asserted
βœ—
Generality C22
no generality axes tested β€” add findings[].generality.axes_tested or a robustness parameter sweep so the claim's breadth is evidenced
βœ—
Run persistence persistence
1 run(s) recorded but none carry an emitter (sqlite/parquet/xarray) or a run-db reference β€” the trajectories are not persisted; runs should emit via the workspace emitter
basal β€” 2/12 rigor dimensions addressed Β· 9 gap(s)
βœ—
Replication C4
single run β€” no replication across seeds declared (add robustness.seeds or simulation_set.seeds)
βœ—
Controls & calibration C1 C2 C4
no controls β€” declare a system that SHOULD fail the criteria (externally-maintained / -supplied) plus a clearly-passing / borderline case so the metric is calibrated, not just asserted
βœ—
Alternative hypotheses C6
no alternative hypotheses declared
⚠
Claim discipline C3
findings not tiered β€” label each observation / mechanism / interpretation
βœ—
Falsifiability C5 C1
criteria read as tailored-to-succeed β€” add a falsifiability note (study.falsifiability)
βœ“
Engineered vs emergent C7
no interpretation-tier claim that requires the distinction
βœ—
Limitations stated C8 C11
no limitations / 'what this does not show' β€” add a short bound on the claim (scope/fidelity of the model, what is NOT demonstrated)
βœ—
Next steps next-steps
no discovery_implications or follow_up_studies β€” state what this study changes and what to investigate next (the Decide phase)
βœ“
Threshold provenance C9 C5
no numeric acceptance bands requiring provenance
βœ—
Metric calibration ladder C4 C2 C20
no calibration_ladder declared β€” index controls[] by known_fail / known_pass / borderline / stress rungs so the metric is calibrated across its range, not just asserted
βœ—
Generality C22
no generality axes tested β€” add findings[].generality.axes_tested or a robustness parameter sweep so the claim's breadth is evidenced
βœ—
Run persistence persistence
1 run(s) recorded but none carry an emitter (sqlite/parquet/xarray) or a run-db reference β€” the trajectories are not persisted; runs should emit via the workspace emitter
with_aa β€” 2/12 rigor dimensions addressed Β· 9 gap(s)
βœ—
Replication C4
single run β€” no replication across seeds declared (add robustness.seeds or simulation_set.seeds)
βœ—
Controls & calibration C1 C2 C4
no controls β€” declare a system that SHOULD fail the criteria (externally-maintained / -supplied) plus a clearly-passing / borderline case so the metric is calibrated, not just asserted
βœ—
Alternative hypotheses C6
no alternative hypotheses declared
⚠
Claim discipline C3
findings not tiered β€” label each observation / mechanism / interpretation
βœ—
Falsifiability C5 C1
criteria read as tailored-to-succeed β€” add a falsifiability note (study.falsifiability)
βœ“
Engineered vs emergent C7
no interpretation-tier claim that requires the distinction
βœ—
Limitations stated C8 C11
no limitations / 'what this does not show' β€” add a short bound on the claim (scope/fidelity of the model, what is NOT demonstrated)
βœ—
Next steps next-steps
no discovery_implications or follow_up_studies β€” state what this study changes and what to investigate next (the Decide phase)
βœ“
Threshold provenance C9 C5
no numeric acceptance bands requiring provenance
βœ—
Metric calibration ladder C4 C2 C20
no calibration_ladder declared β€” index controls[] by known_fail / known_pass / borderline / stress rungs so the metric is calibrated across its range, not just asserted
βœ—
Generality C22
no generality axes tested β€” add findings[].generality.axes_tested or a robustness parameter sweep so the claim's breadth is evidenced
βœ—
Run persistence persistence
1 run(s) recorded but none carry an emitter (sqlite/parquet/xarray) or a run-db reference β€” the trajectories are not persisted; runs should emit via the workspace emitter
acetate β€” 2/12 rigor dimensions addressed Β· 9 gap(s)
βœ—
Replication C4
single run β€” no replication across seeds declared (add robustness.seeds or simulation_set.seeds)
βœ—
Controls & calibration C1 C2 C4
no controls β€” declare a system that SHOULD fail the criteria (externally-maintained / -supplied) plus a clearly-passing / borderline case so the metric is calibrated, not just asserted
βœ—
Alternative hypotheses C6
no alternative hypotheses declared
⚠
Claim discipline C3
findings not tiered β€” label each observation / mechanism / interpretation
βœ—
Falsifiability C5 C1
criteria read as tailored-to-succeed β€” add a falsifiability note (study.falsifiability)
βœ“
Engineered vs emergent C7
no interpretation-tier claim that requires the distinction
βœ—
Limitations stated C8 C11
no limitations / 'what this does not show' β€” add a short bound on the claim (scope/fidelity of the model, what is NOT demonstrated)
βœ—
Next steps next-steps
no discovery_implications or follow_up_studies β€” state what this study changes and what to investigate next (the Decide phase)
βœ“
Threshold provenance C9 C5
no numeric acceptance bands requiring provenance
βœ—
Metric calibration ladder C4 C2 C20
no calibration_ladder declared β€” index controls[] by known_fail / known_pass / borderline / stress rungs so the metric is calibrated across its range, not just asserted
βœ—
Generality C22
no generality axes tested β€” add findings[].generality.axes_tested or a robustness parameter sweep so the claim's breadth is evidenced
βœ—
Run persistence persistence
1 run(s) recorded but none carry an emitter (sqlite/parquet/xarray) or a run-db reference β€” the trajectories are not persisted; runs should emit via the workspace emitter
succinate β€” 2/12 rigor dimensions addressed Β· 9 gap(s)
βœ—
Replication C4
single run β€” no replication across seeds declared (add robustness.seeds or simulation_set.seeds)
βœ—
Controls & calibration C1 C2 C4
no controls β€” declare a system that SHOULD fail the criteria (externally-maintained / -supplied) plus a clearly-passing / borderline case so the metric is calibrated, not just asserted
βœ—
Alternative hypotheses C6
no alternative hypotheses declared
⚠
Claim discipline C3
findings not tiered β€” label each observation / mechanism / interpretation
βœ—
Falsifiability C5 C1
criteria read as tailored-to-succeed β€” add a falsifiability note (study.falsifiability)
βœ“
Engineered vs emergent C7
no interpretation-tier claim that requires the distinction
βœ—
Limitations stated C8 C11
no limitations / 'what this does not show' β€” add a short bound on the claim (scope/fidelity of the model, what is NOT demonstrated)
βœ—
Next steps next-steps
no discovery_implications or follow_up_studies β€” state what this study changes and what to investigate next (the Decide phase)
βœ“
Threshold provenance C9 C5
no numeric acceptance bands requiring provenance
βœ—
Metric calibration ladder C4 C2 C20
no calibration_ladder declared β€” index controls[] by known_fail / known_pass / borderline / stress rungs so the metric is calibrated across its range, not just asserted
βœ—
Generality C22
no generality axes tested β€” add findings[].generality.axes_tested or a robustness parameter sweep so the claim's breadth is evidenced
βœ—
Run persistence persistence
1 run(s) recorded but none carry an emitter (sqlite/parquet/xarray) or a run-db reference β€” the trajectories are not persisted; runs should emit via the workspace emitter
no_oxygen β€” 2/12 rigor dimensions addressed Β· 9 gap(s)
βœ—
Replication C4
single run β€” no replication across seeds declared (add robustness.seeds or simulation_set.seeds)
βœ—
Controls & calibration C1 C2 C4
no controls β€” declare a system that SHOULD fail the criteria (externally-maintained / -supplied) plus a clearly-passing / borderline case so the metric is calibrated, not just asserted
βœ—
Alternative hypotheses C6
no alternative hypotheses declared
⚠
Claim discipline C3
findings not tiered β€” label each observation / mechanism / interpretation
βœ—
Falsifiability C5 C1
criteria read as tailored-to-succeed β€” add a falsifiability note (study.falsifiability)
βœ“
Engineered vs emergent C7
no interpretation-tier claim that requires the distinction
βœ—
Limitations stated C8 C11
no limitations / 'what this does not show' β€” add a short bound on the claim (scope/fidelity of the model, what is NOT demonstrated)
βœ—
Next steps next-steps
no discovery_implications or follow_up_studies β€” state what this study changes and what to investigate next (the Decide phase)
βœ“
Threshold provenance C9 C5
no numeric acceptance bands requiring provenance
βœ—
Metric calibration ladder C4 C2 C20
no calibration_ladder declared β€” index controls[] by known_fail / known_pass / borderline / stress rungs so the metric is calibrated across its range, not just asserted
βœ—
Generality C22
no generality axes tested β€” add findings[].generality.axes_tested or a robustness parameter sweep so the claim's breadth is evidenced
βœ—
Run persistence persistence
1 run(s) recorded but none carry an emitter (sqlite/parquet/xarray) or a run-db reference β€” the trajectories are not persisted; runs should emit via the workspace emitter
metabolism_redux_basal β€” 2/12 rigor dimensions addressed Β· 9 gap(s)
βœ—
Replication C4
single run β€” no replication across seeds declared (add robustness.seeds or simulation_set.seeds)
βœ—
Controls & calibration C1 C2 C4
no controls β€” declare a system that SHOULD fail the criteria (externally-maintained / -supplied) plus a clearly-passing / borderline case so the metric is calibrated, not just asserted
βœ—
Alternative hypotheses C6
no alternative hypotheses declared
⚠
Claim discipline C3
findings not tiered β€” label each observation / mechanism / interpretation
βœ—
Falsifiability C5 C1
criteria read as tailored-to-succeed β€” add a falsifiability note (study.falsifiability)
βœ“
Engineered vs emergent C7
no interpretation-tier claim that requires the distinction
βœ—
Limitations stated C8 C11
no limitations / 'what this does not show' β€” add a short bound on the claim (scope/fidelity of the model, what is NOT demonstrated)
βœ—
Next steps next-steps
no discovery_implications or follow_up_studies β€” state what this study changes and what to investigate next (the Decide phase)
βœ“
Threshold provenance C9 C5
no numeric acceptance bands requiring provenance
βœ—
Metric calibration ladder C4 C2 C20
no calibration_ladder declared β€” index controls[] by known_fail / known_pass / borderline / stress rungs so the metric is calibrated across its range, not just asserted
βœ—
Generality C22
no generality axes tested β€” add findings[].generality.axes_tested or a robustness parameter sweep so the claim's breadth is evidenced
βœ—
Run persistence persistence
1 run(s) recorded but none carry an emitter (sqlite/parquet/xarray) or a run-db reference β€” the trajectories are not persisted; runs should emit via the workspace emitter
metabolism_redux_with_aa β€” 2/12 rigor dimensions addressed Β· 9 gap(s)
βœ—
Replication C4
single run β€” no replication across seeds declared (add robustness.seeds or simulation_set.seeds)
βœ—
Controls & calibration C1 C2 C4
no controls β€” declare a system that SHOULD fail the criteria (externally-maintained / -supplied) plus a clearly-passing / borderline case so the metric is calibrated, not just asserted
βœ—
Alternative hypotheses C6
no alternative hypotheses declared
⚠
Claim discipline C3
findings not tiered β€” label each observation / mechanism / interpretation
βœ—
Falsifiability C5 C1
criteria read as tailored-to-succeed β€” add a falsifiability note (study.falsifiability)
βœ“
Engineered vs emergent C7
no interpretation-tier claim that requires the distinction
βœ—
Limitations stated C8 C11
no limitations / 'what this does not show' β€” add a short bound on the claim (scope/fidelity of the model, what is NOT demonstrated)
βœ—
Next steps next-steps
no discovery_implications or follow_up_studies β€” state what this study changes and what to investigate next (the Decide phase)
βœ“
Threshold provenance C9 C5
no numeric acceptance bands requiring provenance
βœ—
Metric calibration ladder C4 C2 C20
no calibration_ladder declared β€” index controls[] by known_fail / known_pass / borderline / stress rungs so the metric is calibrated across its range, not just asserted
βœ—
Generality C22
no generality axes tested β€” add findings[].generality.axes_tested or a robustness parameter sweep so the claim's breadth is evidenced
βœ—
Run persistence persistence
1 run(s) recorded but none carry an emitter (sqlite/parquet/xarray) or a run-db reference β€” the trajectories are not persisted; runs should emit via the workspace emitter
metabolism_redux_acetate β€” 2/12 rigor dimensions addressed Β· 9 gap(s)
βœ—
Replication C4
single run β€” no replication across seeds declared (add robustness.seeds or simulation_set.seeds)
βœ—
Controls & calibration C1 C2 C4
no controls β€” declare a system that SHOULD fail the criteria (externally-maintained / -supplied) plus a clearly-passing / borderline case so the metric is calibrated, not just asserted
βœ—
Alternative hypotheses C6
no alternative hypotheses declared
⚠
Claim discipline C3
findings not tiered β€” label each observation / mechanism / interpretation
βœ—
Falsifiability C5 C1
criteria read as tailored-to-succeed β€” add a falsifiability note (study.falsifiability)
βœ“
Engineered vs emergent C7
no interpretation-tier claim that requires the distinction
βœ—
Limitations stated C8 C11
no limitations / 'what this does not show' β€” add a short bound on the claim (scope/fidelity of the model, what is NOT demonstrated)
βœ—
Next steps next-steps
no discovery_implications or follow_up_studies β€” state what this study changes and what to investigate next (the Decide phase)
βœ“
Threshold provenance C9 C5
no numeric acceptance bands requiring provenance
βœ—
Metric calibration ladder C4 C2 C20
no calibration_ladder declared β€” index controls[] by known_fail / known_pass / borderline / stress rungs so the metric is calibrated across its range, not just asserted
βœ—
Generality C22
no generality axes tested β€” add findings[].generality.axes_tested or a robustness parameter sweep so the claim's breadth is evidenced
βœ—
Run persistence persistence
1 run(s) recorded but none carry an emitter (sqlite/parquet/xarray) or a run-db reference β€” the trajectories are not persisted; runs should emit via the workspace emitter
metabolism_redux_succinate β€” 2/12 rigor dimensions addressed Β· 9 gap(s)
βœ—
Replication C4
single run β€” no replication across seeds declared (add robustness.seeds or simulation_set.seeds)
βœ—
Controls & calibration C1 C2 C4
no controls β€” declare a system that SHOULD fail the criteria (externally-maintained / -supplied) plus a clearly-passing / borderline case so the metric is calibrated, not just asserted
βœ—
Alternative hypotheses C6
no alternative hypotheses declared
⚠
Claim discipline C3
findings not tiered β€” label each observation / mechanism / interpretation
βœ—
Falsifiability C5 C1
criteria read as tailored-to-succeed β€” add a falsifiability note (study.falsifiability)
βœ“
Engineered vs emergent C7
no interpretation-tier claim that requires the distinction
βœ—
Limitations stated C8 C11
no limitations / 'what this does not show' β€” add a short bound on the claim (scope/fidelity of the model, what is NOT demonstrated)
βœ—
Next steps next-steps
no discovery_implications or follow_up_studies β€” state what this study changes and what to investigate next (the Decide phase)
βœ“
Threshold provenance C9 C5
no numeric acceptance bands requiring provenance
βœ—
Metric calibration ladder C4 C2 C20
no calibration_ladder declared β€” index controls[] by known_fail / known_pass / borderline / stress rungs so the metric is calibrated across its range, not just asserted
βœ—
Generality C22
no generality axes tested β€” add findings[].generality.axes_tested or a robustness parameter sweep so the claim's breadth is evidenced
βœ—
Run persistence persistence
1 run(s) recorded but none carry an emitter (sqlite/parquet/xarray) or a run-db reference β€” the trajectories are not persisted; runs should emit via the workspace emitter
metabolism_redux_no_oxygen β€” 2/12 rigor dimensions addressed Β· 9 gap(s)
βœ—
Replication C4
single run β€” no replication across seeds declared (add robustness.seeds or simulation_set.seeds)
βœ—
Controls & calibration C1 C2 C4
no controls β€” declare a system that SHOULD fail the criteria (externally-maintained / -supplied) plus a clearly-passing / borderline case so the metric is calibrated, not just asserted
βœ—
Alternative hypotheses C6
no alternative hypotheses declared
⚠
Claim discipline C3
findings not tiered β€” label each observation / mechanism / interpretation
βœ—
Falsifiability C5 C1
criteria read as tailored-to-succeed β€” add a falsifiability note (study.falsifiability)
βœ“
Engineered vs emergent C7
no interpretation-tier claim that requires the distinction
βœ—
Limitations stated C8 C11
no limitations / 'what this does not show' β€” add a short bound on the claim (scope/fidelity of the model, what is NOT demonstrated)
βœ—
Next steps next-steps
no discovery_implications or follow_up_studies β€” state what this study changes and what to investigate next (the Decide phase)
βœ“
Threshold provenance C9 C5
no numeric acceptance bands requiring provenance
βœ—
Metric calibration ladder C4 C2 C20
no calibration_ladder declared β€” index controls[] by known_fail / known_pass / borderline / stress rungs so the metric is calibrated across its range, not just asserted
βœ—
Generality C22
no generality axes tested β€” add findings[].generality.axes_tested or a robustness parameter sweep so the claim's breadth is evidenced
βœ—
Run persistence persistence
1 run(s) recorded but none carry an emitter (sqlite/parquet/xarray) or a run-db reference β€” the trajectories are not persisted; runs should emit via the workspace emitter
statistical β€” 2/12 rigor dimensions addressed Β· 9 gap(s)
βœ—
Replication C4
single run β€” no replication across seeds declared (add robustness.seeds or simulation_set.seeds)
βœ—
Controls & calibration C1 C2 C4
no controls β€” declare a system that SHOULD fail the criteria (externally-maintained / -supplied) plus a clearly-passing / borderline case so the metric is calibrated, not just asserted
βœ—
Alternative hypotheses C6
no alternative hypotheses declared
⚠
Claim discipline C3
findings not tiered β€” label each observation / mechanism / interpretation
βœ—
Falsifiability C5 C1
criteria read as tailored-to-succeed β€” add a falsifiability note (study.falsifiability)
βœ“
Engineered vs emergent C7
no interpretation-tier claim that requires the distinction
βœ—
Limitations stated C8 C11
no limitations / 'what this does not show' β€” add a short bound on the claim (scope/fidelity of the model, what is NOT demonstrated)
βœ—
Next steps next-steps
no discovery_implications or follow_up_studies β€” state what this study changes and what to investigate next (the Decide phase)
βœ“
Threshold provenance C9 C5
no numeric acceptance bands requiring provenance
βœ—
Metric calibration ladder C4 C2 C20
no calibration_ladder declared β€” index controls[] by known_fail / known_pass / borderline / stress rungs so the metric is calibrated across its range, not just asserted
βœ—
Generality C22
no generality axes tested β€” add findings[].generality.axes_tested or a robustness parameter sweep so the claim's breadth is evidenced
βœ—
Run persistence persistence
1 run(s) recorded but none carry an emitter (sqlite/parquet/xarray) or a run-db reference β€” the trajectories are not persisted; runs should emit via the workspace emitter
πŸ“Š Framework scorecard framework-self metrics (n=14 investigations)

Framework-self metrics aggregated across every study and investigation in the workspace β€” how consistently the framework itself applies its own rigor practices (discriminating controls, emergent-mechanism labelling, threshold provenance, replication, verdict divergence, falsification exposure). Computed deterministically from declared fields by pbg_superpowers.rigor.framework_metrics.

Discriminating Controls0%0 / 59
Emergent Interpretationsβ€”0 / 0
Missing Mechanism Originβ€”0 / 0
Threshold Provenance39%56 / 144
Replication Coverage22%13 / 59
Ac Coverage100%93 / 93
Verdict Divergence19%11 / 59
Falsification Exposure73%43 / 59
Alternatives Excluded3%2 / 59
Emitter Coverage36%13 / 36

References (0 cited across this investigation)

Union of bibliography.bib_keys and per-behavior cites: across all studies in this investigation. Click DOI or link to open the source.