โ€บNavigation

Iteration 62 โ€” H-064 Calibration Gap Passes the Entrance Exam EXAM PASS ยท IPP-SHADOW CHECK DECISIVE NEXT

In plain words: the second-to-last candidate of the batch reads a genuinely different quantity from the fallen loss-spread family: not how BIG a model's errors are per regime, but whether the model's confidence statements stay honest in each regime โ€” a published theorem ties "calibrated in every regime" to the very invariance our certificates are about. Its exam was the cleanest of the diagnostics so far: the perfect-by-construction model reads zero divergence (through the pre-registered handling of its peculiar edge case), and the deliberately wrong model's calibration divergence beats everything label-shuffling of the same real values can produce by a factor of six โ€” at every choice of the binning knob, which was swept on schedule. But its decisive fight is named and it's a familiar one: our certified IPP instrument already reads distributional honesty per feature โ€” if the calibration lens turns out to be IPP's shadow, the redundancy gate ends it. The field test decides.

Preflight (A0): 2026-07-10 11:06 โ€” load1 3.52/32c (โ‰ค24) ยท no heavy jobs ยท 19 GiB avail, si/so=0 ยท CH + sidecar + kintsugi active โ†’ ALL PASS. Sync: origin/main unchanged. One capped run, 0.51 min, readonly=2 loader, real rows only (permutations of real PIT values per ยง0).

Exam results (pre-registered form, calibration_gap_exam.py)

LegReadingRuling
E1 โ€” perfect model reads zero divergenceExact-zero residuals โ†’ point-mass predictive โ†’ pre-registered guard: perfectly calibrated everywhere, divergence = 0 (the F014-trap handling, clean)PASS
E2 โ€” wrong model's divergence real (every B)Divergence beats the 97.5th percentile of K=199 env-label permutations of the REAL PIT values at 5.8โ€“6.0ร— โ€” the strongest null margin among the B-03 diagnosticsPASS
E3 โ€” binning dial (dossier flag)B swept {10, 20, 50}: identical verdicts at every binning; the census-scale dial check remains decisive per the standing lawPASS
Named census adjacency (decisive)IPP (ADMITTED row 43) reads per-feature distributional invariance via CRPS โ€” the calibration lens may be its shadow; census M3 vs ipp_n_acc + the standard stack + the loss-spread corpse-affinity check decideIPP-SHADOW FIGHT NEXT
Verdict (row 73)EXAM PASS โ€” checkpoint; ladder walk unlocked

โ–ถ Next iteration

Iteration 62 ยท 2026-07-10 ยท CHECKPOINT (row 73) โ€” exam PASS ยท capped (0.51 min) ยท readonly=2 ยท zero generated values ยท append-only ยท evidence: calibration_gap_exam.py ยท calibration_gap_exam_results.json