Asked whether the 20 tools used to judge is-this-feature-real-and-useful can themselves be trusted. Came back with 19 of them proven on real data where the right answer was known, and one thrown out because it does not hold on our market data.
| Quantity | Value | Meaning |
|---|---|---|
| Instruments terminal | 20 of 20 (19 grounded, 1 rejected, 0 parked) | the loop's bounded stop condition |
| Firings | 34 iterations over 2026-07-22 to 2026-07-24 | 3 calendar days |
| Test-bench substrate | 870,304 rows, BTCUSDT at 250 dbps 3s; known-duplicate rho 1.000000, known-positive HAC-t 43.8, known-null FWER 0.002 | known-answer bench built from real bars only |
| Causality floor | 0 false positives over 221,187,473 comparisons, power 1.0 | nothing leaks from the future |
| Rejected tool | power ceiling 0.67 at IC 0.03; realised false-discovery ~1.3x nominal | irreducible by calibration on our long-memory panel |
| Naive vs corrected false-alarm gap | 0.82 vs 0.043; and 0.30-0.57 vs 0.03-0.13 | why the corrections are load-bearing |
| Firing | What it did |
|---|---|
| 1 | Built the known-answer test bench out of real bars only โ no synthetic data anywhere |
| 4 | The substrate gate passed, unlocking every downstream out-of-sample number |
| 12 | Closed the realness axis; found the textbook threshold under-controls false alarms on real heavy-tailed data |
| 19 | Closed the usefulness roster after tie-robust binning fixed a corrupted false-alarm rate |
| 27 | The conditional axis closed with one tool retired and its job reassigned |
| 34 | Last tool grounded; the bounded stop condition fired |
The battery is ready to use, with the rule that comes with it: only a grounded instrument may write feature evidence. Re-test condition for the rejected tool: a construction that controls false discovery at <= 0.10 on a long-memory panel. Two durable patterns to reuse: nominal thresholds under-control on real long-memory data, and power goes flat in N at frozen operating points.
'CAMPAIGN-COMPLETE โ all 20 instruments terminal (19 GROUNDED, 1 REJECTED); loop ended 2026-07-24 (iter 34).' graduation_tally.terminal reads 20/20.
Lifecycle status is process state, not a judgement of the findings. Results are stated as numbers with their uncertainty. Generated between CLOSING-RECONCILIATION markers โ regenerate rather than edit.
opendeviationbar-py ยท bounded campaign ยท started 2026-07-22
Before a promoted feature can be called real or useful, the instrument that judges it must itself be proven on data where the right answer is known. This bounded loop grounds the 20 candidate instruments (#0โ#19) โ one per iteration, in dependency order โ then STOPS. Every iteration emits one HTML page, appended to the ledger at the bottom of this index.
Each candidate instrument sits an entrance exam: the ยง4 five-control battery (positive / negative-null / substitution / power-calibration / invariance-leakage), plus its instrument-specific Harden item and a published detection envelope. It ends with a terminal verdict:
DISCOVERED โ UNDER EVALUATION โ GROUNDED (or REJECTED / UNVERIFIABLE-PARK)
Only a GROUNDED instrument may later judge a real feature. The graduation tally is the loop's STOP condition: when all 20 carry a terminal verdict, the loop publishes a completion page and ends.
readonly=2, SELECT only, both substrates.systemd-run --user --scope cap on bigblack. Cap-kills are evidence; the cap is never raised./loop orchestrates on the laptop and dispatches every read + compute step to bigblack over SSH. Published to /home/nasimubd/sites.| Phase | Instruments (in order) |
|---|---|
| SEAL | real-data-only control substrate self-validation (first firing) |
| Substrate | #1 future-perturbation invariance โ #0 CPCV purge/embargo (both GROUNDED or HALT) |
| Realness | #2 block-perm null โ #3 effective-n+FDR โ #6 SFI โ #8 DSR โ #9 PBO โ #10 Harvey-Liu-Zhu |
| Usefulness | #4 Rank IC โ #5 IC-decay โ #7 MDA โ #19 quantile monotonicity |
| Conditional | #11 CMI โ #12 knockoffs โ #13 DML |
| Increment / mechanism | #14 spanning โ #15 mechanism-intensity |
| Robustness / detector | #16 structural-break โ #17 plateau โ #18 detector lead-lag |
Full per-instrument Admit / Graduate / Harden rows: METRIC-EVALUATION-FRAMEWORK.md ยง7. Metric tables twin: The Realness Question.
One row per firing. The loop appends below the marker; rows are never edited or deleted (corrections are new rows with supersedes:). Each row links the iteration's own HTML page.
| # | Date | Instrument / step | Verdict | Page |
|---|---|---|---|---|
| 1 | 2026-07-22 | SEAL ยท real-data-only control-substrate self-validation | SEAL-PASS lab open | iter 1 โ |
| 2 | 2026-07-22 | #1 ยท future-perturbation invariance (causality floor) | GROUNDED | iter 2 โ |
| 3 | 2026-07-22 | #0 ยท CPCV purge/embargo (leakage-free partition) | CHECKPOINT 6/7 ยท PBO open | iter 3 โ |
| 4 | 2026-07-22 | #0 ยท CPCV purge/embargo (supersedes iter 3) โ โ substrate gate PASS | GROUNDED | iter 4 โ |
| 5 | 2026-07-22 | #2 ยท block-permutation shuffled-label null (realness) ยท P2 | GROUNDED | iter 5 โ |
| 6 | 2026-07-22 | #3 ยท effective-n deflation + BH-FDR (realness) ยท P2 | GROUNDED | iter 6 โ |
| 7 | 2026-07-23 | #6 ยท SFI single-feature OOS (usefulness) ยท P2 | GROUNDED | iter 7 โ |
| 8 | 2026-07-23 | #8 ยท Deflated Sharpe Ratio (on SFI paths) ยท P2 | CHECKPOINT 3/5 ยท SR0 calibration | iter 8 โ |
| 9 | 2026-07-23 | #8 ยท Deflated Sharpe Ratio (supersedes iter 8) โ permutation SR0 | GROUNDED | iter 9 โ |
| 10 | 2026-07-23 | #9 ยท PBO / CSCV (on the SFI trial matrix) ยท P2 | CHECKPOINT 1/7 ยท nullโ.5 โ, thresholds marginal | iter 10 โ |
| 11 | 2026-07-23 | #9 ยท PBO / CSCV (supersedes iter 10) โ within-block IC + full-shuffle null + K=100 | GROUNDED 7/7 ยท 6 seeds | iter 11 โ |
| 12 | 2026-07-23 | #10 ยท HarveyโLiuโZhu tโฅ3 (M_eff) ยท P2 ยท realness axis complete | GROUNDED 6/6 ยท 5 seeds ยท perm-calibrated FWER | iter 12 โ |
| 13 | 2026-07-23 | #4 ยท Rank IC + ICIR + HAC-t ยท usefulness axis opens | CHECKPOINT 2/5 ยท injection fixed ยท period-IC rewrite | iter 13 โ |
| 14 | 2026-07-23 | #4 ยท Rank IC (supersedes iter 13) โ period-IC rewrite | CHECKPOINT 4/5 robust ยท FPR lag-tension โ operator | iter 14 โ |
| 15 | 2026-07-23 | #4 ยท Rank IC (supersedes iter 13, 14) โ #10 perm-FPR resolves the flag | GROUNDED 5/5 ยท 5 seeds ยท first usefulness | iter 15 โ |
| 16 | 2026-07-23 | #5 ยท IC-decay / half-life ยท usefulness | GROUNDED 4/4 ยท 5 seeds ยท overlapโHAC load-bearing | iter 16 โ |
| 17 | 2026-07-23 | #7 ยท MDA + clustered-MDA (ONC) ยท usefulness | GROUNDED 5/5 ยท 5 seeds ยท ONC defeats substitution | iter 17 โ |
| 18 | 2026-07-23 | #19 ยท Quantile monotonicity + PT/RW MR ยท usefulness | CHECKPOINT 5/6 base ยท per-test perm FPR ยท 3 open | iter 18 โ |
| 19 | 2026-07-23 | #19 ยท Quantile monotonicity (supersedes iter 18) โ tie-robust + both-halves RW | GROUNDED 6/6 ยท 5 seeds ยท usefulness roster complete | iter 19 โ |
| 20 | 2026-07-23 | #11 ยท Conditional MI I(f;Y\|S) ยท conditional axis opens | CHECKPOINT 5/5 slice ยท 2 seeds ยท margโ0 ยท iid/AR(1)+high-dim-S deferred | iter 20 โ |
| 21 | 2026-07-23 | #11 ยท Conditional MI (supersedes iter 20) โ all 4 nulls + high-dim-S + imperfect-S envelope | GROUNDED 8/8 ยท 2 seeds ยท AR(1) no-inflate ยท conditional axis | iter 21 โ |
| 22 | 2026-07-23 | #12 ยท Model-X Knockoffs ยท conditional axis ยท slice 1 | CHECKPOINT MVR power 0.95 ยท iid-FDR inflates 0.25 โ block/TSKI next | iter 22 โ |
| 23 | 2026-07-23 | #12 ยท Model-X Knockoffs ยท slice 2 โ course-correction | CHECKPOINT FDRโ autocorr (falsified); plain-knockoff anti-conservative, knockoff+ 0.0; interaction RF-W 1.0 vs lasso 0.0 | iter 23 โ |
| 24 | 2026-07-24 | #12 ยท Model-X Knockoffs ยท slice 3 โ power contour | CHECKPOINT wide panel + KnockoffFilter; power .71โ1.0 (IC .03โ.07), FDP ~1.5ร q โ conservative-q (PATTERN #2) | iter 24 โ |
| 25 | 2026-07-24 | #12 ยท Model-X Knockoffs ยท slice 4 โ standard construction exhausted | OPERATOR RULING power ceiling 0.67 @IC=.03; conservative-q FAILS (FDR stuck ~1.3ร); recommend routeโ#13 | iter 25 โ |
| 26 | 2026-07-24 | #13 ยท Double/Debiased ML + CPI ยท conditional axis ยท slice 1 | CHECKPOINT 4/4 ยท 2 seeds ยท coverage 0.94, admit 0.97, retention 1.22/0.24 โ clean slice | iter 26 โ |
| 27 | 2026-07-24 | #13 ยท Double/Debiased ML + CPI (supersedes iter 26) โ slice 2, conditional axis complete | GROUNDED 7/7 ยท 5 seeds ยท ABSTAIN + HAC-necessity + incomplete-Z; K=10 root-fix | iter 27 โ |
| 27 | 2026-07-24 | #12 ยท Model-X Knockoffs (supersedes iter 25 operator-ruling) โ operator ruling applied | REJECTED standard knockoffs don't hold on our data โ routed to #13 | iter 27 โ |
| 28 | 2026-07-24 | #14 ยท Huberman-Kandel spanning intercept ยท incrementality ยท slice 1 | CHECKPOINT 3/4 ยท coverage 0.955, null-FPR 0.0, substitution โ; power@5bps 0.725 โ block-bootstrap t next | iter 28 โ |
| 29 | 2026-07-24 | #14 ยท Spanning intercept (supersedes iter 28) โ full 7-gate battery | OPERATOR RULING 6/7 robust (incl. omitted-premium Harden); power@5bps ~0.78 (long-memory floor) โ ground @ฮฑโ6bps or hold | iter 29 โ |
| 30 | 2026-07-24 | #14 ยท Spanning intercept (supersedes iter 29 ruling) โ ruling (a) applied | GROUNDED 6 core gates ยท 4 seeds ยท power@5bps ~0.78 grounded on the published ฮฑโ6bps / N-contour envelope (long-memory floor) | iter 30 โ |
| 31 | 2026-07-24 | #15 ยท Mechanism-intensity scaling (Kyle-ฮป/OFI/toxicity) | GROUNDED 3/3 ยท 3 seeds ยท t(ฮฒ2)~24, J*~19.7, ฮIC~0.17; vol-control load-bearing (1.0/0.0); latent-intensity fix | iter 31 โ |
| 32 | 2026-07-24 | #16 ยท Per-year IC + Bai-Perron/CUSUM break | GROUNDED 4/4 ยท 4 seeds ยท sign-flip detect 1.0 & MAE 0 blocks ยท spurious-break FPR โค.033 (block-bootstrap CV) ยท dead-guard 0.0 | iter 32 โ |
| 33 | 2026-07-24 | #17 ยท Parameter-plateau vs needle | GROUNDED 3/3 ยท 3 seeds ยท plateau detect 1.0 (run 8/9) ยท needle 0.0 / noise โค.05 ยท neighbour-corr 0.83 | iter 33 โ |
| 34 | 2026-07-24 | #18 ยท Detector lead-vs-coincide ยท โ final instrument ยท CAMPAIGN COMPLETE | GROUNDED 3/3 ยท 3 seeds ยท power 1.0 & lead-bias 0 ยท all 4 nulls 0.0 ยท partial-TE harden removes latent-common-cause (0.0โ0.025) | iter 34 โ |