Navigation
DashboardProbesRealness Loop › Iter 10 · #9

iteration 10 · 2026-07-23 · P2 · laptop-drives-bigblack

🎲 #9 PBO / CSCV CHECKPOINT · in-progress

Probability of Backtest Overfitting: if you pick the in-sample winner, does it stay a winner out-of-sample — or is it a coin flip? The mechanism works (it separates skill from noise), but the precise §7 pass-marks are marginal on real data, so #9 does not graduate yet.

0.11 / 0.48
PBO: skill vs noise
AUC 0.89
separates skill from overfit
null 0.43
exchangeable ≈ chance ✓
1 / 7
strict §7 gates pass
Preflight (resource-only): load1 0.26 · 42 GiB available · si/so ~0 · ClickHouse active, readonly=2. Compute wall ~4 s.
Standing context: this is exactly the PBO-null calibration the operator descoped from #0 to #9 (2026-07-22). It is genuinely hard — and this firing confirms why.

State — mechanism works, thresholds marginal

GateResultTarget
Block-permuted null ≈ .5median PBO 0.43 (structureless per-cell null)[.42,.58]PASS
Overfit-noise highmedian 0.48 ✓ but fire-rate 0.69fire ≥ .95OPEN
Genuine-skill lowmedian 0.109≤ .10OPEN
AUC (skill vs noise)0.888≥ .90OPEN
Known-duplicate\|ΔPBO\| 0.056 (tie-robust; was 0.34 broken)< .05OPEN
Purge lifts PBOlift −0.002 (no cross-boundary leak)≥ .25OPEN
Harden — magnitude gateflipped under the rank changeOPEN

Why a checkpoint, not more tweaking

PBO discriminates real skill (median 0.11) from overfit noise (median 0.48) with AUC 0.89, and the exchangeable null sits at chance (0.43) — the mechanism is sound. But the constructions trade off: fixing the null calibration shifted skill (0.05→0.11) and AUC (0.93→0.888). Forcing all seven frozen thresholds green by further construction changes would be construction-shopping, not grounding — so #9 is checkpointed for a careful, pre-registered pass instead.

Resolution (next firing, pre-registered)

  1. Stronger dominant skill signal (shorter look-ahead) → skill PBO well below .10 and AUC ≥ .90 with margin.
  2. More CSCV blocks (S=14 → 3432 splits) → tighter PBO distributions → overfit fire-rate ≥ .95 and duplicate \|ΔPBO\| < .05.
  3. A real cross-IS/OOS-block leak for purge-lift (3s-label overlap), with an embargo that removes it → lift ≥ .25.
  4. Re-derive the magnitude-gate on a rank-real sub-cost edge vs the round-trip cost.

If the principled improvements still don't ground it cleanly, escalate to the operator — as with the #0 PBO descope. No threshold p-hacking.