Parsed programmatically from BONEYARD.md at build time โ four append-only sections: 7 pre-campaign seeds from the metric-addition gate, 7 grounded negatives, 1 killed question (impossibility result), and the campaign kills. Kills are wins; several carry re-entry clauses. RECONCILIATION FLAG (honest): the hub counter reads 47 โ 50 physical rows exist; the delta is discovery/killed-question entries counted differently, queued for counter reconciliation next firing.
| Section | Candidate | Kill gate | Recorded reason |
| Seeded (metric-addition gate) | RCIT / FastKCI | M3 DUPLICATE-GAUGE | redundancy-beyond-set already delivered by G2 conditional AUC; borrows G10/G11's null |
| Seeded (metric-addition gate) | Anchor Regression / DRIG | M3 DUPLICATE-GAUGE (weaker sibling of G3) | duality identity, no finite-sample null; "bounds unseen shifts" documented-false (arXiv 2502.02710) |
| Seeded (metric-addition gate) | CPSS (stability selection) | M3 DUPLICATE-GAUGE | duplicates G11 within-window FDR; biased against correlated-useful predictors |
| Seeded (metric-addition gate) | CPCV (purged combinatorial CV) | M1 POLICY (path-spread un-nullable) + M3 | PBO/deflated-z already in G11 over the same null |
| Seeded (metric-addition gate) | Deflated Sharpe (proper, F048) | M3 DUPLICATE-GAUGE / inert | needs a per-feature return-Sharpe object the probe never produces |
| Seeded (metric-addition gate) | Failing-Loudly | M3 DUPLICATE-GAUGE | domain-classifier mode = G5 adv_auc verbatim; DR+KS = excluded univariate-drift family |
| Seeded (metric-addition gate) | Permutation / MarchenkoโPastur threshold | M2 TEXTBOOK-ONLY | MP edge assumes IID โ broken by crypto autocorrelation; significance-null flags everything at nโ78K |
| Seeded (grounded negatives) | `bootci_w`, `stab_freq` | rank-derived, NULL-INVARIANT โ INPUT-ONLY, never gates (M1 POLICY) | โ |
| Seeded (grounded negatives) | EnbPI-on-return | Rยฒโ0 on return target (M4 DEAD-WEIGHT) | โ |
| Seeded (grounded negatives) | Model-X knockoffs | model-synthetic; operator-demoted from real-data core | โ |
| Seeded (grounded negatives) | CODEC / FOCI | operator-excluded; no incremental value over Pearson (M4) | โ |
| Seeded (grounded negatives) | univariate PSI/JS/KS | blind to redundancy flips (M0/M3) | โ |
| Seeded (grounded negatives) | GAM-DVQR | model-synthetic | โ |
| Seeded (grounded negatives) | partial dcor | does not characterize conditional independence (M3) | โ |
| Killed questions | "Unconditional (regime-free) invariance is achievable" | KILLED (pending F014 sharpening) | Rosenfeld 2020; ICP 0/91โ0/92 replicated; NP-hardness reframe (arXiv 2501.17354) โ ฮต-frontier |
| Campaign kills | `bootci_w` (evidence update, 2026-07-03, C2โฒ / LEDGER rows 4+6) | M4 DEAD-WEIGHT (reconfirmed on the real campaign panel) | Held-out ฮAUC +0.004, CI [โ0.0026,+0.0115] โ 0 โ no admissible improvement. **Discovery note:** carries a real sub-bar whisker of persistence-label information (pairing-permutation p=0.025, 250 refits) โ small but statistically detectable; remains INPUT-ONLY/excluded. Do-not-re-propose unless a use case makes a +0.004 AUC increment decision-relevant AND a native null is found (original M1 POLICY kill stands). |
| Campaign kills | Candidate | Kill gate | Recorded reason |
| Campaign kills | Per-environment k-of-N, **parametric Chow e-vs-rest variant** (2026-07-03, LEDGER row 9) | M2 TEXTBOOK-ONLY | Own matched null failed to collapse (mean k under env-destroyed real-row permutation 1.91 vs expected โ10) while calibrated on single-env exchangeable data (4.57% @ ฮฑ=5%) โ the e-vs-REST design's rest side is a 9-regime mixture; mixture heteroskedasticity breaks the F-test's iid-error assumption โ over-rejection by construction. Report-only per ยง0. **Revisit path:** permutation-calibrated per-feature k-of-N (feature's own env-shuffle permutation distribution as the null โ matched by construction); enters the frontier as the Frontier #2 successor slice. |
| Campaign kills | Candidate | Kill gate | Recorded reason |
| Campaign kills | Permutation k-of-N (**regression-invariance flavor**) as a FRAGILEโCONDITIONAL status gauge (2026-07-03, LEDGER rows 13โ14) | M4 DEAD-WEIGHT (gate role) + KILLED-QUESTION (status mapping) | ฮAUC CI [โ0.0033,+0.0145] โ 0; and the naive high-kโCONDITIONAL mapping is falsified decisively: ฯ(k, persistence) = โ0.291 CI [โ0.356,โ0.198] vs null ยฑ0.12 โ k measures anchor-relationship stability, which anti-predicts orthogonality persistence (stably-anchored = stably predictable = non-orthogonal). Instrument retains report-only measurement value; its calibration design (own-permutation null) is sound and carries to successors. **Revisit paths:** (a) verdict-stability k-of-N (per-env orthogonality-verdict counting โ semantically exact for the statuses), (b) inverted-signal candidate ("anchor-coupling stability" as a NON-orthogonality predictor) through the full ladder. |
| Campaign kills | Candidate | Kill gate | Recorded reason |
| Campaign kills | **F014 nonlinear/relaxed ICP** (causalicp wiring, linear + rank variants) as the STABLE-ceiling falsifier (2026-07-03, LEDGER row 17) | M5โฒ BLIND-GAUGE | Failed real-data ground-truth calibration: with the committed duplicate `turnover_imbalance` โก `ofi` as target, linear rejects ALL subsets (n_accepted=0, zero-residual degeneracy) and rank accepts 8 with empty intersection โ cannot identify a PERFECT invariant. All its negatives (incl. committed 0/91โ0-identified) are blindness, not evidence. The 2026-06-16 INCLUDE-IF one-shot is resolved. **Revisit path:** an identification-capable invariance framework (handles exact/degenerate relationships; reports confidence sets, not bare intersections) from a fresh SOTA sweep (Tier-3 refill); meanwhile the pragmatic STABLE evidence path = verdict-stability k_v=10 flags + expiry re-grounding. |
| Campaign kills | **H-017 StabilizedRegression** (Pfister et al. 2021 dual-screen set selection; B-01 feed row) (2026-07-10, LEDGER row 51) | M3 REDUNDANT | The ฮบ dial resolved IN THE CANDIDATE'S FAVOR first โ derived ฮบ_D = 0.0104 (median estimation-noise SE of best-set fold MSEs; the 0.01 exam fixture was, by luck, almost exactly the derived value), readout on a plateau across the whole [0, 0.5] sweep (adjacent Spearman 0.948โ1.0), median ฮn_surv = 0 at ฮบ_D/2..2ฮบ_D, non-degenerate (17.7% non-trivial families at ฮบ_D) โ and then M3 killed it: per-feature survivor count n_surv reads Spearman **0.977** vs the reigning identifier's accepted-set count (e-ICP-fp n_acc) over the 130-subject grid, โฅ the 0.95 bar. Mechanism (the iter-37 shared-oracle worry, confirmed terminal): the stability screen IS the admitted v5 fold-product oracle, and the predictiveness screen adds too little independent variation at census scale; the alternate readout w_sum also peaks against the same instrument (0.804). The set-level lattice readout adds nothing the per-feature instruments do not already say. Do not re-propose as-is; a variant is admissible only with a stability screen NOT built on the admitted v5 oracle. |
| Campaign kills | **H-021 EILLS** (Fan-Fang-Gu-Zhang 2024 Ann. Stat., penalized-invariance selection; B-01 feed row) (2026-07-10, LEDGER row 53) | M3 (ฮณ-fragility census rule) | Entrance exam PASSED cleanly (row 52), but the census killed it: within the pre-registered ฮณ band [1, 1000], adjacent-ฮณ selection agreement broke the 0.90 bar at the 1โ10 edge (0.815) โ and the deeper finding makes the kill safe under ANY band redraw: for ฮณ โฅ 10 the selection collapses to the EMPTY set for ALL 130 subjects (Q(empty)=1 beats every penalized support once ฮณยทpen dominates the โค2.6%-median MSE savings) โ a constant readout, the C2 dead-weight failure mode. The only informative readout left, poi (price-of-invariance), is a variance-explained gauge (worst ยท ฯ ยท =0.663 vs stack โ not redundant, but not an invariance instrument either). Independent corroboration of the iter-22 ฮต-frontier finding (third mechanism: e-ICP exams, StabReg empty families, now EILLS empty collapse). Re-proposal admissible only as an ฮต-relaxed / robust variant with a derived ฮณ (the H-026 invariance-guided-relaxation lead is the natural successor). |
| Campaign kills | **H-004 PCMCI+** (Runge 2020, tigramite 5.2.10.1 GPL-3.0; B-01 feed row) (2026-07-10, LEDGER row 58) | M3 (census dial rule) | Entrance exam PASSED emphatically (row 57: 12/12 clean, the F014 byte-twin trap did not spring, exam-scale dial-invariant) โ but the census killed it on its own role readout: struct_instab (mean pairwise Jaccard distance of per-regime discovered link-sets โ the dossier's 'structural-invariance falsifier') correlates only **0.625** across pc_alpha 0.01โ0.05 (bar 0.90) over 127 subjects on identical rows โ the significance dial reshuffles the very ranking the instrument exists to produce, and no derivation for ฮฑ exists. link_rate was dial-stable (0.955) but is not the role readout; not degenerate (struct_instab median 0.41, spread); not redundant (worst ยท ฯ ยท = 0.493 incl. vs the new CD d_min). Exam-scale dial invariance did NOT generalize to census scale. Re-proposal path recorded: tigramite's pc_alpha=None (internal model selection) or an ฮฑ-integrated/FDR-based link readout would remove the dial โ admissible as a NEW candidate with that derivation. |
| Campaign kills | **H-053 Universal Inference** (Wasserman-Ramdas-Balakrishnan PNAS 2020, split LRT; B-03 feed row 1) (2026-07-10, LEDGER row 60) | M3 (census split-stability rule) | Entrance exam PASSED (row 59: E1/E2/E3 clean, zero new dials โ threshold 20 = certified ETHR = 1/ฮฑ@0.05) โ but the census killed it on the mechanism's one free choice: the per-feature accepted-support count correlates only **0.891** (bar 0.90) between the two split phases, and the fragility is real, not knife-edge โ 15/130 features change verdicts with the halves swapped, several catastrophically (churn_run_count_q1000 and effective_short_entry_bps flip 8-accepted โ 0-accepted). The split-LRT's known 'split hedge' instability, now measured on this substrate. NOTABLY: the named oracle-shadow attack did NOT land (worst ยท ฯ ยท = 0.732 vs eicp_log_min_e โ likelihood-ratio machinery is genuinely distinct from the permutation oracle) and readouts were not degenerate โ the kill is split-fragility alone. Re-proposal path recorded: crossfit/subsample-AVERAGED split-LRT (a mean of e-values is an e-value โ the split choice integrates out) = a NEW admissible candidate. |
| Campaign kills | **H-065 SarganโHansen J** (Hansen 1982 overidentification; B-03 feed row 3, anchor-triad member) (2026-07-10, LEDGER row 64) | M4+M5โฒ (frozen disjunction, both legs fail) | Exam v2 PASSED (row 62) and M3 CLEAR (row 63: genuinely distinct from the certified HSIC-X kernel guard, ฯ=0.484) โ but the finale showed the distinctness carries no verdict value: held-out M4 ฮAUC = โ0.00063, CI95 [โ0.00108, โ0.00028] entirely โค 0 (a small NEGATIVE lift, inside the pairing-null band) and M5โฒb ฯ = โ0.079 (fails the one-sided leg). The parametric validity reading differs from the kernel guard's, but where they differ, the difference does not improve orthogonality-survival verdicts โ HSIC-X remains the shelf's validity guard. Its triad siblings H-066 (strength) and H-067 (weak-robust signal) were ADMITTED the same firing (row 64) โ the kill is role-specific, not family-wide. Re-proposal: only with a demonstrably different validity readout (e.g., per-anchor J decomposition) AND a named verdict channel the kernel guard misses. |
| Campaign kills | **H-027 LPCMCI** (Gerhardus-Runge 2020, latent-confounder PAG discovery; B-03 feed row 6) (2026-07-10, LEDGER row 66) | M3 (decisive census ฮฑ rule) | Entrance exam PASSED 12/12 (row 65, F014 trap clean) โ but the pre-registered decisive check killed it exactly as its sibling: struct_instab (the structural-invariance falsifier, THE role readout) correlates only **0.393** across pc_alpha 0.01โ0.05 on identical rows (bar 0.90; the sibling PCMCI+ read 0.625, row 58) โ the PAG orientation phase makes the latent variant MORE threshold-brittle, not less. link_rate dial-stable (0.957) but not the role readout; not degenerate; not redundant (worst ยท ฯ ยท = 0.695 vs HSIC-X min_p). **Family-level lesson, now measured twice**: constraint-based causal-discovery rankings on this substrate are ฮฑ-brittle at census scale โ exam-scale dial invariance certifies nothing (both siblings aced exams). Re-proposal: only with the dial derived away (pc_alpha=None model selection where the API offers it, or ฮฑ-integrated link readouts). |
| Campaign kills | **H-030 DYNOTEARS** (Pamfil et al. 2020 SVAR structure learning; causalnex 0.12.1 Apache-2.0 on py3.10 โ numpy<1.24 pin noted; B-03 feed row 7) (2026-07-10, LEDGER row 67) | ENTRANCE EXAM (ground-truth duplicate form) | Killed at the door, both legs: **E1 12/18 FAIL** โ the reading of the one CERTAIN quantity in the system (the twin's true contemporaneous coefficient, exactly 1.0 after identical standardization) is ฮป-load-bearing: 0.97 at ฮป=0.05 โ 0.556 at ฮป=0.2 (3%โ44% L1-shrinkage bias, underived); **E2 9/18 FAIL** โ up to 13 PHANTOM edges involving the byte-twin with weights to 0.29, where the constraint-based siblings drew exactly one edge 12/12: the continuous optimizer splits an exact duplicate's signal across spurious lagged structure (the known L1-collinearity pathology, now measured on this substrate). No crashes (F014-clean) โ it fails by being WRONG, not blind. Re-proposal: only with debiased/adaptive-penalty estimation AND a collinearity-safe formulation, plus the still-unwired block-bootstrap trust anchor from its dossier. |
| Campaign kills | **H-061 CVaR/ฯยฒ-DRO bounds** (Levy et al. NeurIPS 2020; B-03 feed row 9, DRO cousin) (2026-07-10, LEDGER row 69) | M3 (census dial rule) | Paired exam PASSED (row 68, exact zeros + S1 monotonicity clean) โ but the standing dial law fired at census scale: adjacent-ฮท Spearman 0.25โ0.5 = **0.897 < 0.90** (knife-edge; 0.1โ0.25 = 0.977, a clean plateau โ the instability appears only toward ฮท=0.5 where CVaR approaches the upper-half mean). RECORDED HONESTLY: this kill would not survive a band redraw, but redrawing the swept band after seeing results is precisely what the frozen discipline forbids (the EILLS symmetry). Costs little: the dial-free cousin H-060 SURVIVED the same census (mutual ฯ = 0.77) and carries the worst-regime role. Not degenerate, not stack-redundant (worst 0.535). Re-proposal: derive ฮท (e.g., from regime-duration fractions) or pre-register a narrower plateau band โ a NEW candidate, cheap to re-run. |
| Campaign kills | **H-060 GroupDRO worst-group gap** (Sagawa et al.; B-03 feed row 8, DRO cousin) (2026-07-10, LEDGER row 70) | M4+M5โฒ (frozen disjunction, both legs fail) | Exam PASSED (row 68) and M3 CLEAR (row 69: dial-free, worst stack ยท ฯ ยท =0.63) โ but the finale was decisive AGAINST: held-out M4 ฮAUC = **โ0.0048**, CI95 [โ0.0080, โ0.0022] entirely negative AND below the pairing-null band's lower edge โ the gap column actively HARMS verdicts. Because the M4 logistic is sign-agnostic, this is not a fixable sign flip: the column's relationship to orthogonality survival is temporally UNSTABLE between calibration and test epochs. M5โฒb ฯ = โ0.116 (CI entirely negative โ worst-regime-divergent features KEEP orthogonality more, the inverted-mechanism family again, cf. rows 9/64). Genuinely novel (max ยท ฯ ยท vs base gates 0.27) and genuinely harmful. Re-proposal: only as a pre-registered regime-lens covariate (not a verdict gate) or with a demonstrated temporally-stable transformation. Both DRO cousins now dead (H-061 row 69). |
| Campaign kills | **H-063 V-REx risk-variance** (Krueger et al. ICML 2021; B-03 feed row 10) (2026-07-10, LEDGER row 72) | M3 (pre-registered fate-transfer rule) | Exam PASSED with its required env-permutation null wired (row 71) โ but the census affinity check was decisive: Spearman(verex_var, the executed H-060 worst-group gap) = **0.982** over 130 subjects on the identical census model โ variance and gap are twin spread-summaries of the SAME per-env losses, both dominated by the worst env. The twin inherits H-060's M4 verdict (actively harmful, temporally unstable direction, row 70) without a finale spend โ the cheap-kill route the row-71 front-loading was built for. Not otherwise stack-redundant (worst 0.623) nor degenerate. The per-env-loss-spread FAMILY is now closed 0-for-3 (H-060 gap ยท H-061 CVaR ยท H-063 var). Re-proposal: any loss-spread readout needs a demonstrated temporally-stable transform first. |
| Campaign kills | **H-028 Convergent Cross Mapping** (Sugihara et al. 2012 Science; pyEDM 2.5.0, license metadata nonstandard โ provenance caution; B-03 feed row 13) (2026-07-10, LEDGER row 77) | ENTRANCE EXAM (ground-truth duplicate form) | The substrate itself killed it: two BYTE-IDENTICAL series cross-map at only **0.87โ0.996 depending on the regime** (covid_crash 0.87โ0.89; 6/9 envรE cells below the derived 0.99 bar). Mechanism: CCM reads coupling through attractor self-predictability โ on stochastic bars the ground-truth reading CONFOUNDS coupling strength with per-regime stochasticity, so the one certain answer in the system is read inconsistently across regimes. The dossier's flagged weakness ('deterministic coupling; stochastic noisy series are its known weakness') confirmed at the cheapest gate. E2 clean (wrong pair โค 0.8 everywhere); convergence present but plateauing below ceiling; F014-clean (0 NaN/crash) โ wrong-on-substrate, not blind. Re-proposal: only a stochasticity-corrected skill normalization (skill relative to the series' own self-prediction ceiling) โ a NEW construction; the 2024 causalized-CCM leakage fix alone does not address this. |
| Campaign kills | H-056 two-sample by betting (ShekharโRamdas, kernel-witness e-process; clean-room build โ reference repo has no license file) | M3 (redundancy census โ MMD-shadow decisive check) | Echo of certificate #13: ฯ=0.982 vs the MMD readout on the same grid, full correlate-profile mirror; the kernel witness makes the bettor a sequential MMD estimator โ a different inference wrapper around the same witness is the same instrument (H-053 precedent, second sighting). Anytime-validity is real but buys nothing at M3. Do not re-propose kernel-witness betting variants; a betting form with a genuinely different witness would need to name its non-MMD delta at M0. (Row 82) |
| Campaign kills | H-074 LGC regime-switching dependency test (Stรธve-lineage; clean-room localized-moments build โ lg is GPL-3 R, regime layer unpublished) | Entrance exam (E3 boundary-role leg) | BLIND GAUGE: on the canonical anchor pair no regime pair rejects its own block-bootstrap null (covidโmidcycle 0.119 vs 0.257) while the certified MMD meter reads the same pairs at 3.5โ50ร its null โ tail-local correlation variance makes the max-switch null too heavy; the WHERE resolution costs more than the signal. F014 kill mode. Do not re-propose grid-max local-dependence switch tests without a variance-controlled statistic; note H-071 (the measure) carries a named risk but is distinct. (Row 83) |
| Campaign kills | H-038 energy distance (Szekely-Rizzo; clean-room V-statistic, parameter-free; exam PASSED with derived bit-exactness, row 84) | M3 (grouped distance-cousins census, sibling rule) | NOT an MMD echo (shadow 0.921 < 0.95 โ the geometries genuinely differ from RKHS) but redundant to its transport sibling: ฯ(energy, W1) = 0.978; the pre-registered grouped rule advances the lower-MMD-shadow sibling (W1 at 0.908). An honorable kill โ the readout is sound, the stack just needs at most one distance geometry. Re-proposal only if the W1 line dies AND a non-transport delta is named. (Row 85) |
| Campaign kills | H-071 LGC measure (Tjostheim-Hufthammer local Gaussian correlation; clean-room localized-moments, center-tail contrast readout) | Entrance exam (E3 between-regime leg) | Beat its dead sibling's variance trap WITHIN regimes (contrast CIs resolve in covid + gold) but the BETWEEN-regime contrast difference (0.210) drowns in the block re-split null (0.405) โ the boundary role needs exactly that reading. LGC family 0-for-2: region-resolved dependence cannot certify regime differences above dependent-data noise at this scale, in test OR measure form. Do not re-propose local-dependence boundary instruments without a variance-reduction design (e.g., paired-window contrasts); the within-regime map itself is real and could serve DESCRIPTIVE (non-gate) roles. (Row 87) |
| Campaign kills | H-049 CRQA cross-recurrence DET (Marwan-Kurths; clean-room Chebyshev embedding, fixed-RR 5%) | Entrance exam (E2 coupling-reality surrogate leg) | Autocorrelation gauge in coupling clothing: de-aligning the pair via circular shift (marginals + autocorr preserved) moves DET by ~0.001 in every env โ the reading carries no alignment information; its between-regime differences are marginal-structure differences already owned by certificates #13/#14. Do not re-propose recurrence-based COUPLING instruments on this substrate without demonstrating surrogate separation at the door; recurrence stats as UNIVARIATE structure readouts would be a different M0 claim. (Row 88) |
| Campaign kills | H-068 selection-stability (stabm/Nogueira per-feature decomposition; clean-room) | Entrance exam (E3 door law) | Between-regime instability-profile difference (0.169) drowns in the block re-split null (0.192) over 124 features โ the boundary reading is pseudo-regime noise. Third family felled by the rows-83/87 law. Re-proposal bar: a selection-stability boundary instrument must first show its regime difference clears an honest block null. (Row 89) |
| Campaign kills | H-069 similarity-adjusted stability (Bommert-Rahnenfuhrer; cluster-union adjustment at the campaign 0.95 bar) | Entrance exam (E3 door law) | The adjustment WORKS (union math verified; swap clusters show ยฑ0.3 instability shifts โ and is NOT uniformly stabilizing, recorded) but the adjusted regime difference (0.173) still drowns in its null (0.206). The delta is real at the mechanism level and irrelevant at the boundary. (Row 89) |
| Campaign kills | H-070 RBO top-weighted rank persistence (Webber-Moffat-Zobel; clean-room per-feature decomposition; exam + M3 flawless โ k_v shadow 0.019, the campaign's most orthogonal survivor) | M4/M5' finale | NOVELTY WITHOUT VALUE, the taxonomy's newest kill mechanism: M4 ฮAUC entirely NEGATIVE (โ0.00167, actively harmful, era-stable harm) and M5' direction significantly INVERTED (โ0.184): top-rank persistence predicts orthogonality LOSS โ mechanism datum #7, corroborating iter-9 from an independent channel. Do not re-propose rank-persistence instruments under the one-sided rule; if the operator ever pre-registers an inverted-direction admission path, this readout is the first candidate to revisit. (Row 92) |
| Campaign kills | H-044 Tucker congruence phi (Lorenzo-Seva/ten Berge; clean-room PCA loadings + greedy alignment, per-feature loading-row cosine) | Entrance exam (E3 within-vs-between ordering leg) | Cleared the block re-split door law (0.373 vs 0.276) but the clearance was noise: within-regime split-half congruence (0.587/0.552) is WORSE than between-regime (0.627) โ loading-estimation variance dominates at panel scale (125 features, ~1500 rows, Kaiser r=28); 1โphi is a noise gauge. Re-proposal bar: a factor-drift boundary instrument needs a variance-controlled loading estimator (fewer factors, regularized loadings, or much longer windows) and must pass the ordering leg first. (Row 93) |
| Campaign kills | H-072 MRP minimum regime performance (Alexander-Fabozzi; trivial clean-room min-over-windows of selector strengths) | M3 (NSUB dial law) | Won the k_v fight (shadow 0.092 โ the floor is genuinely not the flip counter) but the min is an EXTREME statistic: at different subsample sizes different windows become the argmin, reshuffling the ranking (ฯ 0.735 across NSUB 1000/1500). Third sighting of extreme-statistic noise inheritance (max-switch H-074, min-floor here). Re-proposal bar: a variance-controlled floor (e.g., soft-min / lower-quantile aggregation) must pass the dial law first โ the k_v-orthogonality makes a successor genuinely worth building. (Row 95) |
| Campaign kills | H-093 Giacomini-White CPA (clean-room Wald-on-HAC, q=8 regime instruments; exam flawless incl. the zero-differential guard) | M3 (NSUB dial law) | The campaign's closest dial-law miss: 0.868 vs the 0.90 bar โ expanding-forecast paths + HAC-inverse sensitivity at moderate n reshuffle ~1/10 of the ranking between subsample sizes. Clean sheet otherwise (stack worst 0.372 โ best of the campaign; conditioning genuinely non-decorative at 0.736 vs plain DM). Re-proposal bar: a variance-controlled CPA readout (p-value scale, stabilized covariance, or longer windows) is the batch's most promising revisit โ the boundary-role information is demonstrably there. (Row 97) |
| Campaign kills | H-092 Diebold-Mariano + HLN per-regime spread (clean-room; row-64 HAC) | Entrance exam (E3 door law) | The t-normalization cancels the regime signal: per-window DM all individually significant [3.9/4.1/2.5] but the cross-regime spread (0.72) is BELOW the pseudo-regime null (1.69) โ mean and variance both scale with regime intensity and the ratio forgets. Family bar for the declaration layer: per-regime readouts must compare MOMENTS, not self-normalized statistics (contrast GW-CPA's joint moment test, real at p=6.7e-8, killed only by the dial law). (Row 98) |
| Campaign kills | H-047 Benjamini-Yekutieli on HAC-t p-values (clean-room; implementation verified exactly correct at E1) | Entrance exam (E2 control claim, all-true-nulls design) | INPUT INVALIDITY, not procedure defect: with every null true by construction, BY-on-HAC-t rejected in 43.4% of replicates (bound 9.4%; BH 61.6%) โ the HAC-asymptotic marginal p-values understate long-run variance on this persistent substrate and no FDR procedure launders invalid inputs. AUTO-RE-ENTRY: BY returns the moment a certified-valid p-source exists (resampling-based โ the H-046 wild bootstrap is the named candidate). Operator flag: HAC p-values at moderate alpha are anti-conservative here; 0.999-threshold certificates have margin, not immunity. (Row 99) |
| Campaign kills | H-046 Romano-Wolf adjusted-p AS A GATE INSTRUMENT (shift-based max-T; procedure exam flawless row 100) | M3 (NSUB dial law) | The per-feature adjusted-p scalar carries only 30 distinct values (K=99 granularity + monotonicity) โ coarse discrete readouts are inherently rank-unstable (dial 0.803). CRITICAL SCOPE NOTE: this kills the INSTRUMENT form only; RW's FWER control validity (0/49 where HAC hit 43%, row 100) stands, and a pipeline-adoption proposal for RW as the declaration layer's inference backbone is filed with the operator (row 101). Do not re-propose adjusted-p ranking instruments at finite K without a continuous-margin readout. (Row 101) |
| Campaign kills | H-094 Hansen SPA (clean-room stationary-bootstrap, raw-moment best-vs-average) | Entrance exam (E3 boundary leg) | Marginal blindness: snooping-corrected p = 0.07 > 0.05 on the canonical pair where the joint moment test certifies regime structure at 6.7e-8 โ cannot clear its own bar where structure is known present. K=99 granularity noted (~0.01 steps); a K=999 successor might clear, but H-095 MCS passed the same door and covers the set-valued scoping need โ SPA revisit is low-priority. (Row 102) |
| Campaign kills | H-095 Model Confidence Set as a per-feature instrument (raw-range elimination, stationary bootstrap; exam coverage verified row 102) | M3 (dial law 0.523; near-degeneracy) | The scoper keeps all 8 regimes for 116/123 features โ power starvation at panel scale makes ยท MCS ยท nearly constant, and its sparse remainder reshuffles under subsampling. Coarse discrete readouts fail the dial law structurally (third sighting: RW adjusted-p, MRP argmin, now ยท MCS ยท ). Natural home is the declaration PIPELINE (set-valued scoping of already-declared features) โ folded into the row-101 RW proposal's scope. Instrument-form revival unlikely. (Row 103) |
| Campaign kills | H-090 exceedance correlation (Ang-Chen; clean-room, derived orientation for negative pairs) | Entrance exam (E3 door law) | The tail asymmetry is REAL (co-crash exceeds co-rally by ~0.25, resolvable within regimes โ the LGC-fate leg passed) but REGIME-INVARIANT: between-regime difference 0.038 vs null 0.214. A substrate-wide structural constant carries no FRAGILE-CONDITIONAL information. Note for future STABLE-side work: a regime-invariant tail-structure reading is exactly the kind of object the STABLE boundary values โ if a per-feature form exists, it would enter under a different M0 claim. (Row 104) |
Published 2026-07-11 ยท operator-directed artifact (LEDGER NOTE row 105) ยท append-only ยท sources: LEDGER.md rows 0a–104, BONEYARD.md (50 rows), CHATTERJEE-THRESHOLD-DECLARATION.md, the 21-slot certified bucket