Asked to run 120 frozen candidates through the evaluation funnel, one gate per firing, and decide each one. Sixteen firings repaired around 30 defects in the measuring apparatus, three live gate runs all failed, and zero candidates were decided.
| Quantity | Value | Meaning |
|---|---|---|
| Candidates decided | 0 of 120 | 0 admitted, 0 rejected, 0 parked |
| Firings vs published pages | 16 records, 6 pages on disk | 10 firings unpublished, about 62% understated |
| Control pass probability | 4.07e-19 | the frozen control was arithmetically unpassable |
| Known-positive anchor correlation | 0.9604 with its own target (n=88,040) | an alternative anchor measures 0.3609 |
| Tail size, calibration vs full N | 0.0000 (n=3,713) vs 0.08 (n=88,040) | calibration certified a different test than the one that runs |
| Firing | What it did |
|---|---|
| 1 | The harness could not start โ a path computed one directory short. Fixed; no gate ran |
| 2 | Found the control could never pass, and the halved memory ceiling was never applied to the server |
| 6 | Defect queue emptied; one unmade measurement stood between the harness and real data |
| 7-11 | The gate ran at three thresholds โ every run failed, and the calibration window, not the data, was the limit |
| 16 | Showed calibration certifies a different test than the one that runs; declared not operational |
Three named, unresolved OPERATOR decisions, all still unassigned: widen the calibration window at the finest threshold; replace the known-positive anchor that is 96% correlated with its own target; and decide whether the pre-certification step should exist at all. A derived-minimum artifact was quarantined twice and never committed. Separately, 10 of 16 firings have no published page โ that needs whoever owns the publish step.
The declared deliverable โ 120 verdicts โ was never started, and three named blocking decisions plus an unpublished 62% of the ledger are concretely outstanding on the operator.
'NOT OPERATIONAL (2026-07-31). 0 of 120 decided. The fix list is a symptom.' (manifest.json status)
Lifecycle status is process state, not a judgement of the findings. Results are stated as numbers with their uncertainty.
Campaign regime-invariant-orthogonality ยท judging the 120 harvested liquidity / crowding / fragility candidates ยท append-only spoke ยท โ the harvest that produced them
Things this loop cannot decide for itself. Append-only โ resolved entries are struck through, never removed.
substrate_integrity.py lives only on worktree-fix-substrate-dedup, and every read in this campaign depends on it. The campaign PR #675 is stacked on that branch and must be retargeted to main once #674 squash-merges.โ resolved 2026-07-30: #674 squash-merged; #675 rebased onto main and its diff now carries Axis-2 work only.AXIS2-ENVELOPE.json halves the frozen 5 GiB ceiling so two concurrent queries stay inside the operator's 5 GB cap. The 2026-07-27 feasibility sweep must be re-run at the new ceiling before SEAL′ fires. A Code 241 there is evidence: the campaign serialises to one query at a time. The cap is not raised either way.N_min = 100,000 at ฯ=0.064; 500 dbps has ~180k usable rows and 750 dbps ~86k. Y_tail at 750 dbps spreads ~86k rows over 27 regime cells, leaving ~160 clustered exceedances each. Expect UNVERIFIABLE-PARK(substrate-underpowered) for entire cells โ pre-registered as an outcome, not discovered later.FINAL corrects the read; it does not fix the table. The remedy is min_age_to_force_merge_seconds, a production ALTER and the operator's call. OPTIMIZE โฆ FINAL is not proposed โ it can permanently worsen duplicate accumulation.seal_prime_controls.py:39 computed REPO = EVAL.parents[4], one level short: that is <repo>/findings, not <repo>. sys.path received <repo>/findings/findings/evolution/shared_data and the C4 fixture paths became <repo>/findings/tests/fixtures/โฆ, so the module raised ModuleNotFoundError at line 43 before reaching any of its own logic. Arithmetic, not an outstanding precondition.โ resolved 2026-07-31 (iter 1, cdc39da2): measured on the deploy target, fixed marker-anchored rather than by bumping the index, 3 test legs red beforehand.run_c2 compares the ClopperโPearson upper 95% bound against FWER_MAX = 0.05 over 200 draws. Max passing is 3 hits of 200; for a test correctly sized at ฮฑ=0.05 by construction (E[hits]=10), P(all 9 cells pass) = 4.07e-19. Reproduced independently at the frozen kernel. It also contradicts the sibling gate in the same file โ choose_block_length accepts an L only if its size lands inside binomial_size_band(300) = (0.0267, 0.0767), so the two gates cannot both be satisfied. The harness turned ยง5's "FWER โค 0.05" into "size significantly below ฮฑ". Fix is (a) the size-band test the file already imports or (b) a one-sided exact binomial โ both defensible, not equivalent, and it changes a frozen control, so it is amendment territory. The loop is deliberately not choosing.install_shim.sh installs the base 5 GiB envelope; preflight.py reads SPIKE_ENVELOPE_JSON, which the runner never exports. So a firing is capped at 5 GiB, not 2.5 GiB, and assert_envelope_in_force passes while confirming the untightened envelope โ two concurrent queries could reach 10 GiB, exactly what AXIS2-ENVELOPE.json exists to prevent. Repointing the variable does not help: preflight.py E1 is an equality tripwire, so 2.5 GiB would fail as DRIFT. The file's own _tightening_rule claim that "E1 enforces this direction" is false. Compounds the entry above: the ceiling was not merely unvalidated โ the machinery could not have applied it.run_seal_prime.sh's header listed the 2.5 GiB ceiling among closed preconditions, quoting a cell count and a peak-RSS figure, while the envelope, LEDGER row 0k and attention entry A2 above all said PENDING โ and no artifact carrying those numbers existed in the repo or on bigblack, where the campaign's evidence/ and results/ were both empty.โ resolved 2026-07-31 (iter 1, 44ff0d92): the ceiling is now a runtime gate that asks the envelope and refuses on PENDING. Proven to bite โ run_seal_prime.sh 750 โ REFUSING TO FIRE, exit 1. A2 above therefore now blocks by machinery, not by memory.Nothing outstanding โ the loop is unblocked.
Does a candidate measure something real and useful about execution cost and fragility โ beyond the trivial volume baseline โ and does that hold across market regimes?
Axis-2 candidates are orthogonal to price by construction. The 19 grounded instruments from the realness campaign all ask "does this predict forward return?" โ so pointing them at these 120 would score every one at roughly zero and reject the entire cohort. It would look like a clean, rigorous result.
The target is therefore swapped. A liquidity metric is judged on whether it predicts what it costs to trade, and above all on whether it warns before costs blow up. The warehouse already records exactly that, so this is an existing yardstick rather than one invented for the occasion.
Cheapest filters first, so the most candidates die for the least compute. Default is EXCLUDE โ a candidate advances only by clearing the gate. One gate per loop firing.
| Gate | Asks | Kind |
|---|---|---|
| SEAL | do the controls themselves work? (C1โC4 self-test, before any real candidate) | self-test |
| F0 | computable at all โ spot-legal ยท no order-book requirement ยท a Python path exists | deterministic |
| F1 | non-degenerate and dense enough | inferential |
| F2 | not a duplicate โ vs the shipped set and within the cohort | inferential |
| F3 | beats the raw-volume baseline โ expected to do most of the killing | inferential |
| F4 | real and useful against the cost targets, multiplicity-corrected, power-backed | inferential |
| F5 | holds across market regimes โ the ship bar | inferential |
One firing = one gate = one commit on PR #686 = one page in the ledger below. The laptop orchestrates and never touches ClickHouse; bigblack computes under a shim that caps the server. Read the two lanes as "who does what".
PREREG.md before any candidate was measured.readonly=2; compute confined to 5 cores / 5 GB / no swap. Cap-kills are recorded as evidence โ the cap is never raised.Mostly negative, and that is the anticipated shape rather than a failure. Around 19 candidates park immediately for want of order-book data that does not exist anywhere in this system; a share are rejected on tooling; many will prove to be the same few ideas wearing different hats; and most survivors are expected to fail the volume baseline. The comparable axis-1 campaign ended 16 admitted against 45 killed.
Each firing appends one row here and emits one self-contained page index_iter_<N>_<slug>.html in this folder. Rows and pages are never edited or deleted โ corrections add a new sibling row citing the parent. Mirrors the markdown SSoT axis-2-cascade-free/evaluation/LEDGER.md.
| Iter | Date | Gate | Decided | Running | Page | Note |
|---|---|---|---|---|---|---|
| 6 | 2026-07-31 | SEAL′ @ 750 โ not run | 0 | 0 / 120 | iter 6 โ | DEFECT QUEUE EMPTY. D25 (no cell could reproduce its own numbers), D26 (telemetry excluded the client subprocesses and the server), D27 (seven inlined literals โ and a T4 guard blind to the worst offender). SEALโฒ is now blocked by a missing measurement (A2), not a defect. |
| 5 | 2026-07-31 | SEAL′ @ 750 โ not run | 0 | 0 / 120 | iter 5 โ | D23: one L from one pair over the oldest rows, applied to nine โ and is_tail hardcoded False made the Y_tail exceedance floor unreachable, biasing the cascade target toward over-rejection. Now per-target. 0 blockers; D25/D26/D27 left. |
| 4 | 2026-07-31 | SEAL′ @ 750 โ not run | 0 | 0 / 120 | iter 4 โ | Three controls reporting unearned success: D20 (Y_delta imputed onto excluded rows), D21 (positive control fell back to the target โ a tautology), D22 (1 detection of 18 passed). The D20 fix introduced a defect โ zrank propagates NaN โ caught by running it. |
| 3 | 2026-07-31 | SEAL′ @ 750 โ not run | 0 | 0 / 120 | iter 3 โ | Last blocker closed. D14 (N_min floored by nb! โ derived nbโฅ5; nb=4 measured blind on a perfect signal), D19 (one NaN โ p pinned to floor, false-positive direction), D18 (latent). D17 retired โ did not reproduce. 0 blockers remain. |
| 2 | 2026-07-31 | SEAL′ @ 750 โ not run | 0 | 0 / 120 | iter 2 โ | Rule-0 repairs. D12 (C2 could never pass โ operator ruling โ size band), D13+D24 (N_min binding decorative), D15 (envelope gate skippable), D16 (ceiling neither verified nor enforced โ measured 5 GiB in force, now 2.5 GiB). 5 closed, 20 open. |
| 1 | 2026-07-31 | SEAL′ @ 750 โ not run | 0 | 0 / 120 | iter 1 โ | Rule-0 stop. Two defects in the machinery (D10 harness could not resolve its repo root; D11 runner declared an unevidenced precondition closed), both fixed test-first. No candidate touched. |
Markdown SSoT lives in the audit folder findings/evolution/audits/2026-07-22-regime-invariant-orthogonality-campaign/axis-2-cascade-free/evaluation/. Cohort frozen at catalog commit e6eaa6be. Substrate BTCUSDT ร {100, 250, 500, 750} dbps, 3s cost-realism labels, 2018-01 โ 2026-07.
Compute runs on bigblack inside one dated, fully deletable campaign folder โ /home/nasimubd/lab/2026-07-24-axis2-cascade-free-evaluation/ โ holding the environment, a git worktree of the campaign branch, the harnesses and all artifacts. Nothing this campaign creates lives outside it. Results count only once committed to the campaign PR; a run sitting only in the lab is an unbacked result.