โ€บNavigation

Iteration 22 โ€” The Entrance Exam Falsified Itself: STABLE Is an ฮต-Frontier EXAM v1 KILLED CANDIDATE A ALIVE

In plain words: the new e-value instrument sat the entrance exam โ€” identify the one relationship we "knew" was perfectly regime-invariant (the committed duplicate pair) โ€” and rejected it, loudly (e=43). Before blaming the instrument, the loop measured the ground truth itself: the "perfect" duplicate is actually 99.9999% linear, not 100%. The remaining dust โ€” two millionths of the variance โ€” is itself regime-structured, and with twenty thousand data points any honest test must see it. So the exam's premise was wrong, not the instrument: no valid test can pass an exactness exam on real data at our sample sizes. The deep consequence lands on the campaign's one open goal: "perfectly stable" is not a yes/no question on real markets โ€” it is a tolerance question ("stable up to how much dust?"), and the tolerance must be written down before testing. Even the old instrument's behavior gets partially rehabilitated in hindsight โ€” it too may have been seeing real dust โ€” though its rejection stands, because it has no tolerance dial at all. The new instrument does everything right (its calibration and its zero-guard both verified live) and proceeds to the re-specified, tolerance-anchored exam next firing.

Preflight (A0): 2026-07-08 09:19 UTC โ€” load1 5.30/32c (โ‰ค24) ยท 22 GiB avail ยท CH + sidecar active โ†’ ALL PASS. Rebased on moved main (5 commits). Two capped runs (0.21 min exam + measurement probe), readonly loaders, permutations of real rows only.

EXAM v1 (exactness form): identify {ofi} for target = turnover_imbalance
  e({ofi}) = 43.4  โ†’ REJECTED, no degenerate-zero flag
  โ†’ MEASURE THE GROUND TRUTH: pooled Pearson 0.999999 ยท per-env Spearman
    0.999967โ€“0.999999 โ†’ ฮต-exact, NOT exact; residual dust ~2e-6 of variance,
    env-structured โ†’ at nโ‰ˆ20k any calibrated test MUST reject
  โ†’ KILLED-QUESTION (row 29): exact-invariance identification is unsatisfiable
    as an exam or a goal at our N โ€” STABLE is formally an ฮต-FRONTIER
  โ†’ row 30 corrects iter-12's mechanism attribution (verdict unchanged:
    causalicp has no tolerance dial โ†’ blind for the ฮต-question)
  โ†’ CANDIDATE A ALIVE: calibrated e-values โœ“ ยท degenerate guard verified โœ“
NEXT: ฮต-relaxed exam โ€” tolerance ฮตโ‚€ anchored to the committed duplicate's own
residual heterogeneity (real-data-derived, no fiat); identified = subsets
indistinguishable from duplicate-level dust.

โ–ถ Next iteration

Iteration 22 ยท 2026-07-08 ยท exam v1 killed by measurement (rows 29โ€“30) ยท candidate A alive ยท capped ยท zero synthetic data ยท append-only