Iteration 8 โ Permutation k-of-N Passes Its Own Null ยท First FRAGILEโCONDITIONAL Split Measured INSTRUMENT INCLUDE-IFFIRST k DISTRIBUTION
What happened in this iteration, in plain words: version 2 of the per-regime counting instrument โ the one where every feature's regime-break score is judged against that same feature's own shuffled-regime distribution โ passed the exact safety test that killed version 1: with regime structure deliberately destroyed, it now correctly reads "holds almost everywhere" (9.35 of 10, where a calibrated truth-teller should read ~9.5, and version 1 read a broken 1.91). With a trustworthy instrument in hand, the campaign gets its first-ever fine-grained answer to "who survives regime changes, and how often": out of 26 features with real historical data, a clean split appears โ a six-member club (volume, buy volume, sell volume, trade counts, bar duration โ the market's activity dial) holds its behavior in 9 of 10 regimes, while 15 features hold in 2 or fewer. FRAGILE vs CONDITIONAL is measurable for the first time. Two honesty notes shipped alongside: (1) 75 of the ~105 columns turn out to be genuinely empty in the historical windows (so iteration 7's "100 features" included ~74 phantom gradings โ its per-feature numbers are hereby corrected on the record); (2) iteration 7's kill stands, but its mechanism story was part artifact โ a NaN-handling defect in the inherited code was a co-cause, and the clean recheck shows the textbook test still fails calibration (7.81, needs โฅ9) while version 2 passes on identical data.
Preflight (A0)
2026-07-03 11:19 UTC โ load1 3.94 / 32 cores (โค24) ยท no competing nasimubd jobs ยท 21 GiB available, si/soโ0 ยท clickhouse-server active ยท sidecar active + /health healthy โ ALL PASS. Three capped read-only runs totaling <2 min (main 0.35 min), readonly=2, 5c/5G scope, pre-authorized (03b), zero generated values (03c โ permutations of real rows only).
The instrument and its acceptance
DESIGN (successor to the killed parametric variant):
per (feature, regime): Chow statistic (env vs rest) judged against the
permutation distribution of the SAME statistic under 1,000 random
re-assignments of REAL rows (sizes preserved) โ matched null BY CONSTRUCTION,
no distributional assumptions. Global per-feature standardization (preserves
regime scale breaks; difference from the old per-env wiring recorded).
Vectorized via sufficient statistics: full run 0.35 min inside the cap.
ACCEPTANCE (must pass before the measurement is trusted):
A1 destroyed-env null mean k = 9.35 of 10 need โฅ 9.0 โ PASS
(version 1 read 1.91 here โ the exact failure that killed it)
A2 real < null 3.92 < 9.35 โ PASS
THE FIRST k DISTRIBUTION (26 real-data features, crypto BTCUSDT@250):
k = 1: โโโโโโโโโโ 10 features โ FRAGILE-material
k = 2: โโโโโ 5
k = 4: โ 1
k = 5: โโ 2
k = 7: โโ 2
k = 9: โโโโโโ 6 โ CONDITIONAL-material:
volume ยท buy_volume ยท sell_volume ยท individual_trade_count ยท
agg_record_count ยท duration_us (the market's ACTIVITY/SIZE family)
Honesty notes (corrections on the record)
#
Note
1
Feature universe corrected: 75 of ~105 numeric columns are ALL-NULL in the 2018โ2024 window subsamples (microstructure/plugin columns unpopulated there โ consistent with the committed "24 crypto effective columns" note). Iteration 7's "100 features graded" therefore included ~74 NaN-phantom gradings; its per-feature k values are unreliable (LEDGER row 12).
2
Iteration 7 mechanism refined, verdict unchanged: the inherited v3-style fill leaves all-NaN columns as NaN โ NaN test statistics counted as rejections (a third silent-failure mode in the v3 lineage, now on record). Clean-pool recheck: the parametric null improves 1.91 โ 7.81 but STILL fails the โฅ9 bar while the permutation successor reads 9.35 on identical data/seed โ residual parametric size distortion under regime mixture is real. Kill stands; mechanism = NaN defect (major co-cause) + mixture distortion (residual, confirmed). LEDGER row 12 supersedes row 9's mechanism sentence only.
(i) M4 incremental-value + M5โฒ real-data power legs before Loop-2 consumes k for statuses ยท (ii) block-permutation refinement (within-env autocorrelation) ยท (iii) Pโฅ10k perms if BH-gated declarations needed (current min p โ 0.001)
Conventions
ฮฑ=0.05 (G3) ยท BH q=0.10 (G11) ยท P=1000 (resolution recorded) ยท seeds STRUCTURAL ยท zero new magic numbers
Numbers
pool 20,604 rows ร 26 features ยท A1 9.35 (pass โฅ9) ยท A2 3.92<9.35 ยท mean k_BH 3.88 ยท wall 0.35 min main + recheck
three sub-2-min capped runs ยท ~30 readonly=2 SELECTs ยท zero writes outside audit folder + dashboard
โถ Next iteration
Iteration 9 = Frontier #2, admission legs: M4 (does k add value?) + M5โฒ (real-data power) for the permutation k-of-N. In plain words: the counting instrument works and its first measurement is in โ but by the lab's own rules its numbers may not ground status declarations until two more gates pass: (M4) adding k to the certified panel must measurably improve held-out forecasts of who stays orthogonal (same machinery as the seal's C2โฒ/C3โฒ), and (M5โฒ) it must demonstrate power on a real-data known answer โ the leave-one-out pattern applies: the activity/size family's k=9 vs the k=1 group gives a committed-anchored contrast to verify against the walk-forward persistence label. Both legs are hermetic-or-capped, pre-authorized. If both pass โ the instrument is fully ADMITTED and the first FRAGILE/CONDITIONAL status declarations become possible on crypto. Steering alternatives: full forex grid (fix-plan step 6), forex k-of-N, or F014 (Frontier #3, identifiability-sharpened).
Iteration 8 ยท 2026-07-03 ยท Frontier #2 successor: instrument INCLUDE-IF'd, first k distribution, two corrections on record ยท capped read-only ยท zero synthetic data ยท append-only