Navigation
DashboardProbesRealness Loop › Iter 12 · #10

iteration 12 · 2026-07-23 · P2 · laptop-drives-bigblack

🎯 #10 Harvey–Liu–Zhu t≥3 (M_eff) GROUNDED

The multiple-testing haircut: once you account for everything ever tried, how large a t-stat does a feature really need? All six §7 gates pass, robust across 5 seeds. This closes the realness axis (#2, #3, #6, #8, #9, #10 all grounded).

6 / 6
§7 gates pass
0.043
HAC FWER (perm) ≤ .05
0.82
naïve-OLS FWER (≈19× worse)
5 seeds
all GROUNDED
Preflight (resource-only): load1 0.13 · 42 GiB available · si/so ~0 · ClickHouse active, readonly=2. 5c/5G/no-swap capped, single-thread BLAS.

The six gates

Gate (§7 row 10)ResultTarget
Power1.00 (real look-ahead clears threshold)≥ .8PASS
M_eff recovery5.00 vs true 5 (rel-err 0.02%)±20%PASS
HAC FWER (permutation)0.043 held-out≤ .05PASS
Naïve-OLS FWER0.82 (iid SE, no MTC)≥ .25PASS
Dup M_eff1.00 (feature + exact clone)∈ [1,1.2]PASS
N_min(H)detectable at all tested N ≥ 2000reportedPASS
Harden — no collapse-to-1collapse FWER 0.085 > frozen 0.043 & > .05inflatesPASS

Seed-robustness: 5 seeds all GROUNDED 6/6 — permutation FWER ∈ [.030,.048], naïve ∈ [.79,1.00], collapse ∈ [.085,.182], M_eff-recovery 5.00, dup 1.00, power 1.0.

The load-bearing necessity finding

The §7 analytic rule |t| ≥ max(t*(M_eff), 3.0) assumes the null t-stat is a standard normal. Attacking that assumption: on real autocorrelated data the HAC-t null has heavier tails, so the fixed analytic threshold (3.13) under-controls FWER — 0.058 → 0.075 → 0.083 → 0.060 → 0.130 across the five seeds (one seed blew past α by 2.6×).

The fix, and the operating threshold, is permutation-calibrated (empirical null max-|t|, real values only): it controls FWER by construction whatever the tail, and it adapts — the threshold rises 3.19 → 3.58 exactly on the heavy-tail seed where the fixed analytic threshold fails. Same permutation-refinement pattern as #3 (HAC) and #8 (permutation SR0).

M_eff must be the frozen global universe. Collapsing it to 1 (only test the winner, ignore everything tried) drops back to the bare 3.0 floor and inflates FWER to 0.085. The permutation threshold rides above the floor and adapts to the real tail; the fixed floor cannot. The PREREG denominator (PREREG.md:66; integer stamped at first feature-panel run) is load-bearing — 3.0 is a last-resort floor, not the operating value.
★ Realness axis complete. With #10 grounded, all six realness instruments (#2 block-perm null, #3 effective-n+FDR, #6 SFI, #8 DSR, #9 PBO, #10 HLZ) carry terminal GROUNDED verdicts. Next: the usefulness axis — #4 Rank IC.