โ€บNavigation
Closing reconciliation

Asked whether the 20 tools used to judge is-this-feature-real-and-useful can themselves be trusted. Came back with 19 of them proven on real data where the right answer was known, and one thrown out because it does not hold on our market data.

What it measured
QuantityValueMeaning
Instruments terminal20 of 20 (19 grounded, 1 rejected, 0 parked)the loop's bounded stop condition
Firings34 iterations over 2026-07-22 to 2026-07-243 calendar days
Test-bench substrate870,304 rows, BTCUSDT at 250 dbps 3s; known-duplicate rho 1.000000, known-positive HAC-t 43.8, known-null FWER 0.002known-answer bench built from real bars only
Causality floor0 false positives over 221,187,473 comparisons, power 1.0nothing leaks from the future
Rejected toolpower ceiling 0.67 at IC 0.03; realised false-discovery ~1.3x nominalirreducible by calibration on our long-memory panel
Naive vs corrected false-alarm gap0.82 vs 0.043; and 0.30-0.57 vs 0.03-0.13why the corrections are load-bearing
How it ran
FiringWhat it did
1Built the known-answer test bench out of real bars only โ€” no synthetic data anywhere
4The substrate gate passed, unlocking every downstream out-of-sample number
12Closed the realness axis; found the textbook threshold under-controls false alarms on real heavy-tailed data
19Closed the usefulness roster after tie-robust binning fixed a corrupted false-alarm rate
27The conditional axis closed with one tool retired and its job reassigned
34Last tool grounded; the bounded stop condition fired
What it produced
Carries forward

The battery is ready to use, with the rule that comes with it: only a grounded instrument may write feature evidence. Re-test condition for the rejected tool: a construction that controls false discovery at <= 0.10 on a long-memory panel. Two durable patterns to reuse: nominal thresholds under-control on real long-memory data, and power goes flat in N at frozen operating points.

Why it is closed

'CAMPAIGN-COMPLETE โ€” all 20 instruments terminal (19 GROUNDED, 1 REJECTED); loop ended 2026-07-24 (iter 34).' graduation_tally.terminal reads 20/20.

Lifecycle status is process state, not a judgement of the findings. Results are stated as numbers with their uncertainty. Generated between CLOSING-RECONCILIATION markers โ€” regenerate rather than edit.

Dashboard โ€บ Probes โ€บ Feature-Realness Instrument-Grounding Loop

opendeviationbar-py ยท bounded campaign ยท started 2026-07-22

๐Ÿงช Feature-Realness Instrument-Grounding Loop

Before a promoted feature can be called real or useful, the instrument that judges it must itself be proven on data where the right answer is known. This bounded loop grounds the 20 candidate instruments (#0โ€“#19) โ€” one per iteration, in dependency order โ€” then STOPS. Every iteration emits one HTML page, appended to the ledger at the bottom of this index.

19 / 20
grounded (+1 rejected = 20 terminal)
11
discovered (candidates)
20 / 20
terminal โ€” โ˜… campaign complete
โœ“ DONE
all axes closed ยท loop ended

What this loop does

Each candidate instrument sits an entrance exam: the ยง4 five-control battery (positive / negative-null / substitution / power-calibration / invariance-leakage), plus its instrument-specific Harden item and a published detection envelope. It ends with a terminal verdict:

DISCOVERED โ†’ UNDER EVALUATION โ†’ GROUNDED (or REJECTED / UNVERIFIABLE-PARK)

Only a GROUNDED instrument may later judge a real feature. The graduation tally is the loop's STOP condition: when all 20 carry a terminal verdict, the loop publishes a completion page and ends.

Substrate gate. Instruments #1 (future-perturbation invariance) and #0 (CPCV purge/embargo) ground first. If either fails to ground, the campaign HALTS before the realness/utility instruments โ€” every downstream out-of-sample number rides on the substrate.

Hard operating rules

Dependency order (frozen)

PhaseInstruments (in order)
SEALreal-data-only control substrate self-validation (first firing)
Substrate#1 future-perturbation invariance โ†’ #0 CPCV purge/embargo (both GROUNDED or HALT)
Realness#2 block-perm null โ†’ #3 effective-n+FDR โ†’ #6 SFI โ†’ #8 DSR โ†’ #9 PBO โ†’ #10 Harvey-Liu-Zhu
Usefulness#4 Rank IC โ†’ #5 IC-decay โ†’ #7 MDA โ†’ #19 quantile monotonicity
Conditional#11 CMI โ†’ #12 knockoffs โ†’ #13 DML
Increment / mechanism#14 spanning โ†’ #15 mechanism-intensity
Robustness / detector#16 structural-break โ†’ #17 plateau โ†’ #18 detector lead-lag

Full per-instrument Admit / Graduate / Harden rows: METRIC-EVALUATION-FRAMEWORK.md ยง7. Metric tables twin: The Realness Question.

โ˜…โ˜… Campaign complete

All 20 candidate instruments (#0โ€“#19) are terminal โ€” the loop has ended. 19 GROUNDED ยท 1 REJECTED (#12, standard Gaussian model-X knockoffs don't hold on our long-memory panel โ†’ its FDR-selection role routed to the grounded #13 DML) ยท 0 parked ยท 0 unverifiable. The final push (per the operator directive to reduce time and finish all 20) drove #14โ†’#18 to terminal: #14 GROUNDED on the published power/N envelope (ruling a), then #15 mechanism-intensity, #16 structural-break, #17 plateau-vs-needle, and #18 detector lead-vs-coincide each GROUNDED with robustness across seeds. The battery is ready: only a GROUNDED instrument may later judge a real feature. See the iter-34 completion page and the CAMPAIGN-COMPLETE ledger entry for the axis-by-axis roll-up and the durable cross-instrument patterns.

Iteration ledger (append-only)

One row per firing. The loop appends below the marker; rows are never edited or deleted (corrections are new rows with supersedes:). Each row links the iteration's own HTML page.

#DateInstrument / stepVerdictPage
12026-07-22SEAL ยท real-data-only control-substrate self-validationSEAL-PASS lab openiter 1 โ†’
22026-07-22#1 ยท future-perturbation invariance (causality floor)GROUNDEDiter 2 โ†’
32026-07-22#0 ยท CPCV purge/embargo (leakage-free partition)CHECKPOINT 6/7 ยท PBO openiter 3 โ†’
42026-07-22#0 ยท CPCV purge/embargo (supersedes iter 3) โ€” โ˜… substrate gate PASSGROUNDEDiter 4 โ†’
52026-07-22#2 ยท block-permutation shuffled-label null (realness) ยท P2GROUNDEDiter 5 โ†’
62026-07-22#3 ยท effective-n deflation + BH-FDR (realness) ยท P2GROUNDEDiter 6 โ†’
72026-07-23#6 ยท SFI single-feature OOS (usefulness) ยท P2GROUNDEDiter 7 โ†’
82026-07-23#8 ยท Deflated Sharpe Ratio (on SFI paths) ยท P2CHECKPOINT 3/5 ยท SR0 calibrationiter 8 โ†’
92026-07-23#8 ยท Deflated Sharpe Ratio (supersedes iter 8) โ€” permutation SR0GROUNDEDiter 9 โ†’
102026-07-23#9 ยท PBO / CSCV (on the SFI trial matrix) ยท P2CHECKPOINT 1/7 ยท nullโ‰ˆ.5 โœ“, thresholds marginaliter 10 โ†’
112026-07-23#9 ยท PBO / CSCV (supersedes iter 10) โ€” within-block IC + full-shuffle null + K=100GROUNDED 7/7 ยท 6 seedsiter 11 โ†’
122026-07-23#10 ยท Harveyโ€“Liuโ€“Zhu tโ‰ฅ3 (M_eff) ยท P2 ยท realness axis completeGROUNDED 6/6 ยท 5 seeds ยท perm-calibrated FWERiter 12 โ†’
132026-07-23#4 ยท Rank IC + ICIR + HAC-t ยท usefulness axis opensCHECKPOINT 2/5 ยท injection fixed ยท period-IC rewriteiter 13 โ†’
142026-07-23#4 ยท Rank IC (supersedes iter 13) โ€” period-IC rewriteCHECKPOINT 4/5 robust ยท FPR lag-tension โ†’ operatoriter 14 โ†’
152026-07-23#4 ยท Rank IC (supersedes iter 13, 14) โ€” #10 perm-FPR resolves the flagGROUNDED 5/5 ยท 5 seeds ยท first usefulnessiter 15 โ†’
162026-07-23#5 ยท IC-decay / half-life ยท usefulnessGROUNDED 4/4 ยท 5 seeds ยท overlapโ†’HAC load-bearingiter 16 โ†’
172026-07-23#7 ยท MDA + clustered-MDA (ONC) ยท usefulnessGROUNDED 5/5 ยท 5 seeds ยท ONC defeats substitutioniter 17 โ†’
182026-07-23#19 ยท Quantile monotonicity + PT/RW MR ยท usefulnessCHECKPOINT 5/6 base ยท per-test perm FPR ยท 3 openiter 18 โ†’
192026-07-23#19 ยท Quantile monotonicity (supersedes iter 18) โ€” tie-robust + both-halves RWGROUNDED 6/6 ยท 5 seeds ยท usefulness roster completeiter 19 โ†’
202026-07-23#11 ยท Conditional MI I(f;Y\|S) ยท conditional axis opensCHECKPOINT 5/5 slice ยท 2 seeds ยท margโ‰ˆ0 ยท iid/AR(1)+high-dim-S deferrediter 20 โ†’
212026-07-23#11 ยท Conditional MI (supersedes iter 20) โ€” all 4 nulls + high-dim-S + imperfect-S envelopeGROUNDED 8/8 ยท 2 seeds ยท AR(1) no-inflate ยท conditional axisiter 21 โ†’
222026-07-23#12 ยท Model-X Knockoffs ยท conditional axis ยท slice 1CHECKPOINT MVR power 0.95 ยท iid-FDR inflates 0.25 โ†’ block/TSKI nextiter 22 โ†’
232026-07-23#12 ยท Model-X Knockoffs ยท slice 2 โ€” course-correctionCHECKPOINT FDRโ‰ autocorr (falsified); plain-knockoff anti-conservative, knockoff+ 0.0; interaction RF-W 1.0 vs lasso 0.0iter 23 โ†’
242026-07-24#12 ยท Model-X Knockoffs ยท slice 3 โ€” power contourCHECKPOINT wide panel + KnockoffFilter; power .71โ†’1.0 (IC .03โ†’.07), FDP ~1.5ร— q โ†’ conservative-q (PATTERN #2)iter 24 โ†’
252026-07-24#12 ยท Model-X Knockoffs ยท slice 4 โ€” standard construction exhaustedOPERATOR RULING power ceiling 0.67 @IC=.03; conservative-q FAILS (FDR stuck ~1.3ร—); recommend routeโ†’#13iter 25 โ†’
262026-07-24#13 ยท Double/Debiased ML + CPI ยท conditional axis ยท slice 1CHECKPOINT 4/4 ยท 2 seeds ยท coverage 0.94, admit 0.97, retention 1.22/0.24 โ€” clean sliceiter 26 โ†’
272026-07-24#13 ยท Double/Debiased ML + CPI (supersedes iter 26) โ€” slice 2, conditional axis completeGROUNDED 7/7 ยท 5 seeds ยท ABSTAIN + HAC-necessity + incomplete-Z; K=10 root-fixiter 27 โ†’
272026-07-24#12 ยท Model-X Knockoffs (supersedes iter 25 operator-ruling) โ€” operator ruling appliedREJECTED standard knockoffs don't hold on our data โ†’ routed to #13iter 27 โ†’
282026-07-24#14 ยท Huberman-Kandel spanning intercept ยท incrementality ยท slice 1CHECKPOINT 3/4 ยท coverage 0.955, null-FPR 0.0, substitution โœ“; power@5bps 0.725 โ†’ block-bootstrap t nextiter 28 โ†’
292026-07-24#14 ยท Spanning intercept (supersedes iter 28) โ€” full 7-gate batteryOPERATOR RULING 6/7 robust (incl. omitted-premium Harden); power@5bps ~0.78 (long-memory floor) โ†’ ground @ฮฑโ‰ˆ6bps or holditer 29 โ†’
302026-07-24#14 ยท Spanning intercept (supersedes iter 29 ruling) โ€” ruling (a) appliedGROUNDED 6 core gates ยท 4 seeds ยท power@5bps ~0.78 grounded on the published ฮฑโ‰ˆ6bps / N-contour envelope (long-memory floor)iter 30 โ†’
312026-07-24#15 ยท Mechanism-intensity scaling (Kyle-ฮป/OFI/toxicity)GROUNDED 3/3 ยท 3 seeds ยท t(ฮฒ2)~24, J*~19.7, ฮ”IC~0.17; vol-control load-bearing (1.0/0.0); latent-intensity fixiter 31 โ†’
322026-07-24#16 ยท Per-year IC + Bai-Perron/CUSUM breakGROUNDED 4/4 ยท 4 seeds ยท sign-flip detect 1.0 & MAE 0 blocks ยท spurious-break FPR โ‰ค.033 (block-bootstrap CV) ยท dead-guard 0.0iter 32 โ†’
332026-07-24#17 ยท Parameter-plateau vs needleGROUNDED 3/3 ยท 3 seeds ยท plateau detect 1.0 (run 8/9) ยท needle 0.0 / noise โ‰ค.05 ยท neighbour-corr 0.83iter 33 โ†’
342026-07-24#18 ยท Detector lead-vs-coincide ยท โ˜… final instrument ยท CAMPAIGN COMPLETEGROUNDED 3/3 ยท 3 seeds ยท power 1.0 & lead-bias 0 ยท all 4 nulls 0.0 ยท partial-TE harden removes latent-common-cause (0.0โ€“0.025)iter 34 โ†’