โ€บNavigation

Iteration 57 โ€” the DRO Cousins Both Pass EXAM PASS ร—2 ยท GROUPED CENSUS NEXT

In plain words: after the map-drawing school went zero-for-three, the queue turned to a completely different question: not "what drives what," but "how much worse is the worst regime than the average?" Two cousin instruments ask it two ways โ€” one over our named market regimes (the worst-regime gap), one over the worst FRACTION of moments regardless of labels (the tail-risk excess, with a tunable tail size). Their joint exam was spotless. Handed the model that is perfect by construction, both read exactly zero โ€” no worst regime, no bad tail, precisely as the ground truth demands. Handed a deliberately wrong model, both lit up: the worst regime was 2.4 loss-units worse than average, and the worst 10% of moments carried nearly 8 units of excess pain. The tail-size dial passed its built-in mathematics check (deeper tail must read worse โ€” it did, 7.9 โ†’ 2.7 โ†’ 0.9). Next: the field test as a pair, where the named question is whether the two cousins are secretly one instrument.

Preflight (A0): 2026-07-10 06:05 โ€” load1 2.50/32c (โ‰ค24) ยท no heavy jobs ยท 37 GiB avail, si trickle/so=0 ยท CH + sidecar + kintsugi active โ†’ ALL PASS. Sync: origin/main unchanged. One capped run, 0.45 min, readonly=2 loader, deterministic algebra, real rows only.

Paired exam results (pre-registered form, dro_cousins_exam.py)

LegReadingRuling
H-060 E1 โ€” true model reads uniformgap = 0 exactly (every per-env loss < 10โปยนยฒ, degenerate guard)PASS
H-060 E2 โ€” wrong model exposedPer-env losses spread 0.108 โ†’ 3.338; worst-group gap decisively positive (โ‰ˆ 2.4 loss-units above average)PASS
H-061 E1 โ€” true model reads zero tail riskCVaR excess = 0 at every ฮท (all per-observation losses exactly zero)PASS
H-061 E2 โ€” wrong model's tail exposedExcess = 7.92 / 2.69 / 0.92 at ฮท = 0.1 / 0.25 / 0.5 โ€” strictly positive at every tail massPASS
H-061 S1 โ€” the dial's mathematics tripwireCVaR must be nonincreasing in ฮท: 7.92 โ‰ฅ 2.69 โ‰ฅ 0.92 โœ“ (an implementation-bug catch, clean)PASS
F014 trapExact-zero losses handled by the guards โ€” no crashes, no NaNNOT SPRUNG
Verdicts (row 68)EXAM PASS ร—2 โ€” checkpoint; the pair enters the ladder together; ฮท's census-scale dial check remains owed (the standing law)

โ–ถ Next iteration

Iteration 57 ยท 2026-07-10 ยท CHECKPOINT (row 68) โ€” paired exam PASS ร—2 ยท capped (0.45 min) ยท readonly=2 ยท zero generated values ยท append-only ยท evidence: dro_cousins_exam.py ยท dro_cousins_exam_results.json