โ€บNavigation

Iteration 37 โ€” H-017 StabReg Passes the Entrance Exam EXAM PASS ยท ฮบ KNOB EXPOSED

In plain words: the next candidate judges teams of features rather than individuals: a team survives only if it is both steady across regimes AND actually good at predicting. On the identification exam it was flawless โ€” the true team survived with a perfect score, every surviving team contained the right member, the empty team was thrown out emphatically. The subtler test โ€” two byte-identical twins must always be picked together โ€” hit an honest snag twice: on both real targets tried, no team survived at all (nothing was simultaneously steady and predictive โ€” a true fact about the data, not a malfunction). But the run still settled the question: the twins received bit-for-bit identical scores on every team they appeared in, and since selection is decided purely by scores, identical scores mean they can never be separated. Pass. One genuine worry surfaced, though: this candidate carries a tuning dial (how much prediction quality it will trade for stability), and the second run showed the dial genuinely changes answers. Campaign law says: no certification while an unexplained dial is load-bearing. Resolving that dial is the next fight.

Preflight (A0): 2026-07-13 โ€” load1 3.74/32c (โ‰ค24) ยท 16 GiB avail ยท CH + sidecar active โ†’ ALL PASS. Sync note: main moved (forex b5-10 campaign landed); a rebase hit conflicts in 331 generated dashboard pages โ†’ aborted, merged main instead, regenerated nav via the builders, pushed clean (PR pickup verified). Two capped exam runs (0.5 min each), permutations of real rows only.

Exam results (pre-registered form)

LegReadingRuling
Part A identification{duration_us}: e=1.0 (degenerate) + MSE=0.0 โ†’ survives ยท all 4 survivors contain duration ยท weight(duration)=1.0 ยท โˆ… rejected (e=1.4M)PASS โ€” perfect separation
Part B co-selection (burstiness target)ZERO surviving teams โ†’ weight gap 0=0 over an empty familyVACUOUS โ€” recorded, not accepted
Part Bโ€ฒ co-selection (intensity_q1000 target, twin pair among predictors)Family again empty (stable teams not best-predictive; best-predictive teams unstable โ€” a truth about the data). But swap-symmetric teams scored bit-identically through both screens ({duration_us} โ‰ก {intra_duration_us}: e=11.89, MSE=0.953768; supersets 43.53/0.906603)Co-selection ESTABLISHED (measurement + proof: identical scores โ‡’ inseparable in any family)
Verdict (row 46)EXAM PASS โ€” checkpoint; ladder next with two named fights

The two ladder fights, named now: (1) the ฮบ knob โ€” the predictiveness slack (ฮบ=0.01 exam fixture) demonstrably changes which teams survive; per campaign law (ยง0: magic numbers must be derived or shown insensitive) it must be resolved before any ADMIT; (2) the shared-oracle question โ€” the stability screen rides the admitted v5 oracle, so the IAS lesson applies: the M3 attack must show the SET-level readout adds information beyond the per-feature instruments.

โ–ถ Next iteration

Iteration 37 ยท 2026-07-13 ยท CHECKPOINT (row 46) โ€” exam PASS, knob in red ink ยท capped ยท zero synthetic data ยท append-only ยท evidence: stabreg_exam.py ยท stabreg_exam_b2.py ยท result JSONs