Forward-Orthogonality ยท feature-declaration expiry ยท robustness review
Does a feature declared orthogonal / conditional today have a trustworthy way to know when that verdict expires in live trading? This reviews the design we already have (the G12 live monitor + the feature-status lifecycle) against what we need now โ robustness as research.
2026-07-15 Existing design: EVALUATING โ not operational Benchmark: campaign laws + SOTA
A feature is graded per regime through an early-exit gate ladder; which gate first disqualifies it sets the grade. Promoted features (CONDITIONAL/STABLE) ship with a regime scope and a live expiry monitor. When the monitor fires, the verdict expires and the feature re-runs the sealed ladder.
feature ร regime โโโบ GATE LADDER (G0 orthogonal-today ยท G3 regime-invariance ยท
G2 forward-lift ยท G10 shuffled-null ยท G11 e-BH)
โ first gate to disqualify = the grade
โโโโโโโโโโโโโฌโโโโโโโโโโโดโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโ
โผ โผ โผ โผ
FAIL FRAGILE CONDITIONAL STABLE-ORTHOGONAL
(=BAN) (=WATCH) (=WATCH) (=ORTHOGONAL)
โโโโโโโโโโโ promoted, scoped to regimes R โโโโโโโ
โ
โผ
G12 / E6 LIVE MONITOR (post-promotion, streaming)
CTM/SKIT wealth martingale + BOCPD alarm, e-BH-merged,
Ville bound: P(โt: W_t โฅ 1/ฮฑ) โค ฮฑ โ "verdict EXPIRED"
โ fires (or: next completed regime slice)
โผ
EXPIRE โ feature reverts to DRAFT โ re-run sealed ladder
โ re-promote (new scope) OR demote
Grounded grading evidence (strong): persistence 94โ95% into next regime ยท flip-ranking AUC 0.95/0.88 (crypto/forex) ยท ACI coverage 89.9%/79.9% ยท FDR-declared sets at q=0.10, PBO 0.00/0.20, deflation passed ยท ICP 0/92 โ every declaration is regime-conditional ยท leakage-gated, twice-replicated. Honest ceiling: ~8/10 structural โ unconditional invariance is provably impossible.
Read: we can grade a feature and we picked the right kind of alarm โ but the alarm itself, the part that says "this verdict has expired," is the weak link, and the automatic re-grounding depends on a detector we haven't built.
CTM + SKIT + e-BH) that the later matrix-admission campaign put under attack โ and it proved that any monitor whose "how big is chance?" null assumes independence cries wolf at 11โ57% instead of 5% on market bars. The design (June 2026) predates that discovery (July 2026), so its default null is the one the campaign banned. Good news: the campaign also certified the fix (the SKIT betting e-process + circular de-alignment), so the hole is closable โ but nothing gates until it passes the certificate attack.
| # | Gap | Breaks | Severity |
|---|---|---|---|
| F1 | Generic CTM fed raw non-conformity scores from autocorrelated bars โ 11โ57% false-alarm (killed 8 instruments). | L-IID (dependence-aware null) | DECISIVE |
| F2 | E6 replay never run โ no positive/negative control; sensitivity + specificity unmeasured. | M5 blind-gauge | DECISIVE |
| F3 | No attribution โ feature-drift vs instrument/substrate-drift indistinguishable. | alarm fatigue / spurious re-cert | HIGH |
| F4 | One ฮฑ-monitor per (feature ร regime) ร hundreds โ family-wide false-alarm explosion; e-BH not wired at the monitor layer. | L-IID / L-MAGIC (aggregation) | HIGH |
| F5 | Interim "re-ground after next completed regime slice" needs a live slice-completed detector โ unbuilt (B-05 gated). | auto-expiry inert | HIGH |
| F6 | WATCH covariate weights estimated on ~2K windows; rolling CTM can self-heal toward the new regime. | estimation-error / drift-to-accept | MED |
| F7 | The dependence-aware null's own inputs (block-length / tail-index) can drift; nothing re-checks them. | live L-IID | MED |
| F8 | ~2K bars/regime โ effective sample size can be too thin; no under-power veto. | false negatives | MED |
| F9 | wold_R crypto inversion (0.278) ยท forex PBO watch. (forex ICP wiring defect โ later REPAIRED by the admission loop.) | grading hygiene | LOW |
Priority-ordered; each closes a numbered gap above.
feature verdict; an instrument verdict suppresses expiry.Keep: the anytime-valid guarantee, the e-BH aggregation intent, the honest regime-scoped grammar, and the whole well-grounded grading side โ these are SOTA-correct.
Fix before it can be trusted live, in order: (1) a dependence-aware null that passes the certificate attack, (2) the E6 positive/negative-control validation, (3) an attribution race so a regime turn doesn't spuriously expire still-valid features. Everything else is structural hardening on top.
The single sentence: our expiry design is right, but its default statistics were written before we learned that market memory breaks naรฏve tests โ so the expiry monitor must be re-nulled for dependence and power-proven (E6) before any feature's verdict is allowed to expire on its say-so.
Research assessment ยท no code, no production change, nothing ratified ยท append-only.
SSoT twin: findings/evolution/audits/2026-05-26-forward-orthogonality-prediction/FEATURE-EXPIRY-ROBUSTNESS-ASSESSMENT-2026-07-15.md.
Sources: PHASE-3-PROBE-DESIGN.md (G0โG12) ยท STATUS-GROUNDING-AND-THRESHOLDS.md ยง2 ยท SOTA-RESEARCH-2026-06-23.md ยง2 ยท CAMPAIGN-GROUNDED-SUMMARY.md ยท probes/forward.html ยท matrix-admission LEDGER rows 99โ117 (the two laws) ยท robust-expiry SOTA workflow wf_88195813-778.
Citation-verification: pre-2026 backbone established; 2026-dated SOTA pending the standing arXiv re-verification gate before wiring.