Iteration 22 โ The Entrance Exam Falsified Itself: STABLE Is an ฮต-Frontier EXAM v1 KILLEDCANDIDATE A ALIVE
In plain words: the new e-value instrument sat the entrance exam โ identify the one relationship we "knew" was perfectly regime-invariant (the committed duplicate pair) โ and rejected it, loudly (e=43). Before blaming the instrument, the loop measured the ground truth itself: the "perfect" duplicate is actually 99.9999% linear, not 100%. The remaining dust โ two millionths of the variance โ is itself regime-structured, and with twenty thousand data points any honest test must see it. So the exam's premise was wrong, not the instrument: no valid test can pass an exactness exam on real data at our sample sizes. The deep consequence lands on the campaign's one open goal: "perfectly stable" is not a yes/no question on real markets โ it is a tolerance question ("stable up to how much dust?"), and the tolerance must be written down before testing. Even the old instrument's behavior gets partially rehabilitated in hindsight โ it too may have been seeing real dust โ though its rejection stands, because it has no tolerance dial at all. The new instrument does everything right (its calibration and its zero-guard both verified live) and proceeds to the re-specified, tolerance-anchored exam next firing.
Preflight (A0): 2026-07-08 09:19 UTC โ load1 5.30/32c (โค24) ยท 22 GiB avail ยท CH + sidecar active โ ALL PASS. Rebased on moved main (5 commits). Two capped runs (0.21 min exam + measurement probe), readonly loaders, permutations of real rows only.
EXAM v1 (exactness form): identify {ofi} for target = turnover_imbalance
e({ofi}) = 43.4 โ REJECTED, no degenerate-zero flag
โ MEASURE THE GROUND TRUTH: pooled Pearson 0.999999 ยท per-env Spearman
0.999967โ0.999999 โ ฮต-exact, NOT exact; residual dust ~2e-6 of variance,
env-structured โ at nโ20k any calibrated test MUST reject
โ KILLED-QUESTION (row 29): exact-invariance identification is unsatisfiable
as an exam or a goal at our N โ STABLE is formally an ฮต-FRONTIER
โ row 30 corrects iter-12's mechanism attribution (verdict unchanged:
causalicp has no tolerance dial โ blind for the ฮต-question)
โ CANDIDATE A ALIVE: calibrated e-values โ ยท degenerate guard verified โ
NEXT: ฮต-relaxed exam โ tolerance ฮตโ anchored to the committed duplicate's own
residual heterogeneity (real-data-derived, no fiat); identified = subsets
indistinguishable from duplicate-level dust.
โถ Next iteration
Next firing: pre-register the ฮต-relaxed entrance exam (ฮตโ = the duplicate pair's own measured residual-heterogeneity scale โ committed-anchored, tagged DERIVED) and run it: candidate A must accept {ofi} at ฮตโ and reject non-ofi subsets at the same ฮตโ. Pass โ ladder walk. Standing for the operator: the 73 cycle-1 draft statuses and the Frontier #4 precondition review.
Iteration 22 ยท 2026-07-08 ยท exam v1 killed by measurement (rows 29โ30) ยท candidate A alive ยท capped ยท zero synthetic data ยท append-only