โ€บNavigation
Reconciliation โ€” open ยท HALTED

Asked to run 120 frozen candidates through the evaluation funnel, one gate per firing, and decide each one. Sixteen firings repaired around 30 defects in the measuring apparatus, three live gate runs all failed, and zero candidates were decided.

What it measured
QuantityValueMeaning
Candidates decided0 of 1200 admitted, 0 rejected, 0 parked
Firings vs published pages16 records, 6 pages on disk10 firings unpublished, about 62% understated
Control pass probability4.07e-19the frozen control was arithmetically unpassable
Known-positive anchor correlation0.9604 with its own target (n=88,040)an alternative anchor measures 0.3609
Tail size, calibration vs full N0.0000 (n=3,713) vs 0.08 (n=88,040)calibration certified a different test than the one that runs
How it ran
FiringWhat it did
1The harness could not start โ€” a path computed one directory short. Fixed; no gate ran
2Found the control could never pass, and the halved memory ceiling was never applied to the server
6Defect queue emptied; one unmade measurement stood between the harness and real data
7-11The gate ran at three thresholds โ€” every run failed, and the calibration window, not the data, was the limit
16Showed calibration certifies a different test than the one that runs; declared not operational
What it produced
What is outstanding, and who owns it

Three named, unresolved OPERATOR decisions, all still unassigned: widen the calibration window at the finest threshold; replace the known-positive anchor that is 96% correlated with its own target; and decide whether the pre-certification step should exist at all. A derived-minimum artifact was quarantined twice and never committed. Separately, 10 of 16 firings have no published page โ€” that needs whoever owns the publish step.

Why this status

The declared deliverable โ€” 120 verdicts โ€” was never started, and three named blocking decisions plus an unpublished 62% of the ledger are concretely outstanding on the operator.

'NOT OPERATIONAL (2026-07-31). 0 of 120 decided. The fix list is a symptom.' (manifest.json status)

Lifecycle status is process state, not a judgement of the findings. Results are stated as numbers with their uncertainty.

Axis-2 ยท Cascade-Free โ€” candidate evaluation FIRING 1 SPENT ON RULE 0 ยท HARNESS REPAIRED ยท NO GATE RUN

Campaign regime-invariant-orthogonality ยท judging the 120 harvested liquidity / crowding / fragility candidates ยท append-only spoke ยท โ† the harvest that produced them

120candidates in frozen cohort
0decided
0admitted
SEAL′gate due next

โš  Operator attention

Things this loop cannot decide for itself. Append-only โ€” resolved entries are struck through, never removed.

The question

Does a candidate measure something real and useful about execution cost and fragility โ€” beyond the trivial volume baseline โ€” and does that hold across market regimes?

The trap this campaign is built to avoid

Axis-2 candidates are orthogonal to price by construction. The 19 grounded instruments from the realness campaign all ask "does this predict forward return?" โ€” so pointing them at these 120 would score every one at roughly zero and reject the entire cohort. It would look like a clean, rigorous result.

The target is therefore swapped. A liquidity metric is judged on whether it predicts what it costs to trade, and above all on whether it warns before costs blow up. The warehouse already records exactly that, so this is an existing yardstick rather than one invented for the occasion.

The funnel

Cheapest filters first, so the most candidates die for the least compute. Default is EXCLUDE โ€” a candidate advances only by clearing the gate. One gate per loop firing.

GateAsksKind
SEALdo the controls themselves work? (C1โ€“C4 self-test, before any real candidate)self-test
F0computable at all โ€” spot-legal ยท no order-book requirement ยท a Python path existsdeterministic
F1non-degenerate and dense enoughinferential
F2not a duplicate โ€” vs the shipped set and within the cohortinferential
F3beats the raw-volume baseline  โ† expected to do most of the killinginferential
F4real and useful against the cost targets, multiplicity-corrected, power-backedinferential
F5holds across market regimes โ€” the ship barinferential

How the loop actually works

One firing = one gate = one commit on PR #686 = one page in the ledger below. The laptop orchestrates and never touches ClickHouse; bigblack computes under a shim that caps the server. Read the two lanes as "who does what".

Axis-2 cascade-free evaluation โ€” funnel and firing cycle Top: the SEAL-prime to F5 funnel, 120 candidates entering, default EXCLUDE at every gate. Bottom: the per-firing cycle split into a laptop lane and a bigblack lane, from preflight through publish, with the parked path and the two current blockers marked. โ‘  The funnel โ€” 120 candidates, default EXCLUDE, one gate per firing 120 frozen cohort SEALโ€ฒ C2 ยท C3 ยท POWER-vs-N grades no candidate F0 computable F1 degeneracy F2 redundancy F3 beats volume F4 real & useful F5 regime-invariant SHIP ADMIT default EXCLUDE โ€” a candidate advances only by clearing the gate ยท REJECTED / UNVERIFIABLE-PARK F3 is expected to do most of the killing. F1 and F4 both gate on the N_min that SEALโ€ฒ derives โ€” which is why SEALโ€ฒ runs first and produces something, rather than only self-checking. โ‘ก One firing โ€” 45 min, one gate, one commit, one page ๐Ÿ’ป LAPTOP โ€” orchestrates, never queries ClickHouse ๐Ÿ–ฅ BIGBLACK โ€” computes, readonly=2, capped /loop 45m โ†’ reads LOOP-PROMPT.md the trigger is fixed; the prompt file self-updates. Re-arm every ~7 days. A-1 FIRST FIRING ONLY โ€” bootstrap lab folder ยท worktree ยท venv ยท spike:install (the shim) A0 PREFLIGHT โ€” spike:preflight asks the SERVER what it sees ยท negative + positive control fail โ†’ PARKED page, end firing A1 STOP? all 120 terminal โ†’ CAMPAIGN COMPLETE checked FIRST, every firing. Otherwise: pick the ONE gate due. A2 LEAKAGE GUARD T1โ€“T4 any failure โ†’ HARD STOP. Publish, end firing. A3 RUN THE GATE substrate_integrity.py โ†’ FINAL + total ORDER BY + assert_deduplicated ch_concurrency.py โ†’ โ‰ค 2 queries in flight, waiting not erroring systemd-run scope ยท 5 cores / 5 GB / no swap ยท cap-kill = evidence A4 ATTACK THE RESULT, then A5 LEDGER too-good-to-be-true is a leakage signature ยท append-only row + provenance A6 PUBLISH NEW index_iter_<N>.html ยท append to index + attention[] ยท manifest + telemetry scoped rsync of THIS FOLDER only โ€” no --delete, never dashboard:deploy A7 COMMIT to PR #675 verify OPEN before pushing ยท never self-merge โ€” the operator merges next firing โš  BEFORE THE FIRST FIRING โœ” PR #674 merged โ€” substrate_integrity.py is on main โ˜ Validate the 2.5 GiB per-query ceiling two concurrent queries ร— 2.5 GiB = the 5 GiB cap, exactly re-run the 07-27 sweep at the new ceiling on the worst cell Code 241 โ†’ serialise to 1 query. The cap is never raised. โ˜ seal_prime_controls.py โ€” the SEALโ€ฒ harness itself

Rules that cannot be bent

What to expect

Mostly negative, and that is the anticipated shape rather than a failure. Around 19 candidates park immediately for want of order-book data that does not exist anywhere in this system; a share are rejected on tooling; many will prove to be the same few ideas wearing different hats; and most survivors are expected to fail the volume baseline. The comparable axis-1 campaign ended 16 admitted against 45 killed.

Evaluation ledger โ€” append-only ยท newest first ยท one page per firing

Each firing appends one row here and emits one self-contained page index_iter_<N>_<slug>.html in this folder. Rows and pages are never edited or deleted โ€” corrections add a new sibling row citing the parent. Mirrors the markdown SSoT axis-2-cascade-free/evaluation/LEDGER.md.

IterDateGateDecidedRunningPageNote
62026-07-31SEAL′ @ 750 โ€” not run00 / 120 iter 6 โ†’ DEFECT QUEUE EMPTY. D25 (no cell could reproduce its own numbers), D26 (telemetry excluded the client subprocesses and the server), D27 (seven inlined literals โ€” and a T4 guard blind to the worst offender). SEALโ€ฒ is now blocked by a missing measurement (A2), not a defect.
52026-07-31SEAL′ @ 750 โ€” not run00 / 120 iter 5 โ†’ D23: one L from one pair over the oldest rows, applied to nine โ€” and is_tail hardcoded False made the Y_tail exceedance floor unreachable, biasing the cascade target toward over-rejection. Now per-target. 0 blockers; D25/D26/D27 left.
42026-07-31SEAL′ @ 750 โ€” not run00 / 120 iter 4 โ†’ Three controls reporting unearned success: D20 (Y_delta imputed onto excluded rows), D21 (positive control fell back to the target โ€” a tautology), D22 (1 detection of 18 passed). The D20 fix introduced a defect โ€” zrank propagates NaN โ€” caught by running it.
32026-07-31SEAL′ @ 750 โ€” not run00 / 120 iter 3 โ†’ Last blocker closed. D14 (N_min floored by nb! โ€” derived nbโ‰ฅ5; nb=4 measured blind on a perfect signal), D19 (one NaN โ†’ p pinned to floor, false-positive direction), D18 (latent). D17 retired โ€” did not reproduce. 0 blockers remain.
22026-07-31SEAL′ @ 750 โ€” not run00 / 120 iter 2 โ†’ Rule-0 repairs. D12 (C2 could never pass โ€” operator ruling โ†’ size band), D13+D24 (N_min binding decorative), D15 (envelope gate skippable), D16 (ceiling neither verified nor enforced โ€” measured 5 GiB in force, now 2.5 GiB). 5 closed, 20 open.
12026-07-31SEAL′ @ 750 โ€” not run00 / 120 iter 1 โ†’ Rule-0 stop. Two defects in the machinery (D10 harness could not resolve its repo root; D11 runner declared an unevidenced precondition closed), both fixed test-first. No candidate touched.

Provenance

Markdown SSoT lives in the audit folder findings/evolution/audits/2026-07-22-regime-invariant-orthogonality-campaign/axis-2-cascade-free/evaluation/. Cohort frozen at catalog commit e6eaa6be. Substrate BTCUSDT ร— {100, 250, 500, 750} dbps, 3s cost-realism labels, 2018-01 โ†’ 2026-07.

Compute runs on bigblack inside one dated, fully deletable campaign folder โ€” /home/nasimubd/lab/2026-07-24-axis2-cascade-free-evaluation/ โ€” holding the environment, a git worktree of the campaign branch, the harnesses and all artifacts. Nothing this campaign creates lives outside it. Results count only once committed to the campaign PR; a run sitting only in the lab is an unbacked result.