The loop's finish line โ "every instrument certified, every certificate current, nothing new left to test" โ can never be reached: markets keep producing new regimes (so certificates keep expiring) and researchers keep publishing new instruments (so candidates keep arriving). That impossibility is deliberate: it is the fuel that keeps the loop running 24/7. Yet no turn is wasted โ every single iteration deposits one durable, evidence-backed verdict in the ledger. A finish line stops a runner; a horizon keeps the runner walking forever.
Every M4/M5 run is wrapped in the T1โT4 leakage guard (blocking preflight). Full spec: audit twin โ GATE-LADDER-M0-M5.md.
| # | Report | In plain words | Status |
| 1 | Scaffold: rulebook frozen + forex defect diagnosed | Built the lab, froze its rulebook, seeded the reject-list, and found exactly why forex was invisible (a five-word list of crypto column names) | FROZEN |
| 2 | Self-test seal 1/4: planted duplicate caught (C1) | Fed the lab a trusted instrument photocopied under a fancy new name plus a whisper of noise; the lab threw it out at exactly the right gate (correlation 1.000) while leaving genuinely different things alone | C1 PASS |
| 3 | Self-test seal 2/4: noise gauge rejected as dead weight (C2) | Submitted a fake instrument emitting pure random numbers five times over; the lab measured zero forecasting improvement every time and threw it out โ while a secretly-perfect instrument fed through the same machinery WAS recognized as valuable. Ran on the real 2,992-transition campaign panel. | C2 PASS (superseded by iter 4 re-run) |
| 4 | Seal re-validation (no-synthetic): C1โฒ + C3โฒ PASS, C2โฒ awaiting ruling | All tricks re-run with real data only: duplicate detector passed (exact copy reads 1.000); power test passed spectacularly (hidden strongest instrument re-admitted, 0.31 โ 0.98); dead-weight test surprised โ the gate correctly refused the reject-listed instrument, but it turns out to carry a microscopic real whisker of information (p=0.025), so the loop hard-stopped and asks the operator to rule | C1โฒ C3โฒ PASS ยท C2โฒ RULING |
| 5 | Ruling A recorded ยท C4 replay 9/9 ยท SEAL COMPLETE | The operator ruled the dead-weight trick passed (gate refused the candidate exactly as pre-registered; the whisker is logged as a discovery). The final trick โ re-judging all nine historical candidates through the frozen ladder โ reproduced nine of nine verdicts (7 rejections, 2 scoped admissions; one candidate dies at a cheaper gate, reason stated). All four controls green: the lab is open. | SEAL COMPLETE |
| 6 | Forex ICP repaired (Tier-A): twin-column root cause, 163 features graded | The forex blackout ran deeper than the casebook: two anchor columns are the same column under two names (byte-identical over 701,424 rows), crashing the invariance math for every feature โ silently. Proved on record, twin removed, errors made loud, re-ran capped read-only: 163/167 forex features graded (was zero), crypto control identical on all comparable features. Nothing on forex is regime-invariant either โ now a measurement, not blindness. | REPAIR LANDED |
| 7 | k-of-N: parametric variant killed by its own null ยท "nothing is invariant" reinterpreted | Built the per-regime counting instrument โ and the lab's safety net caught it crying "regime break" on deliberately structure-destroyed data (calibrated on homogeneous data, broken by regime mixtures; mechanism proven, killed at the assumptions gate, successor designed). Bigger: the campaign's "0 of 92 regime-invariant" is the wrong reading โ invariance was never rejected for ANY of 100 features; the engine just couldn't identify WHICH anchor set is invariant. (Mechanism refined + per-feature numbers corrected in iter 8.) | KILL #1 ยท MAJOR FINDING |
| 8 | Permutation k-of-N passes its own null ยท first FRAGILEโCONDITIONAL split measured | Version 2 passed the safety test that killed version 1 (destroyed structure reads 9.35/10; v1 read 1.91). First fine-grained regime-survival answer: the activity dial (volumes, trade counts, duration) holds in 9 of 10 regimes; 15 of 26 features hold in โค2. Plus two corrections on record: 75 columns genuinely empty in historical windows (iter 7 graded ~74 phantoms); iter 7's kill-mechanism was part NaN artifact โ the kill stands. (Axis semantics corrected in iter 9.) | INCLUDE-IF ยท FIRST SPLIT |
| 9 | k-of-N gate excluded ยท inverted-signal discovery | Certification produced the campaign's most interesting result: panel improvement indistinguishable from zero, and โ decisively backwards โ features rated most regime-stable are the MOST likely to lose orthogonality (ฯ=โ0.29, far outside the null band). Mechanism: stable coupling to the core signals = stable predictability = not orthogonal. Gate rejected, mapping question formally killed, discovery booked; successor counts per-regime orthogonality verdicts themselves. | GATE EXCLUDED ยท QUESTION KILLED ยท DISCOVERY |
| 10 | Verdict-stability k-of-N: THE CAMPAIGN'S FIRST ADMISSION | The counter that asks the literal per-regime question passed the full six-gate ladder: the committed duplicate reads exactly 0-of-10 (ground truth), four engineered features read a perfect 10-of-10, provably not a repackaging (ฯ=โ0.75 vs the 0.95 bar), and its history strongly predicts who keeps orthogonal status (+0.59 vs ~0 null). First certificate: ADMIT with scope, role, expiry. FRAGILE vs CONDITIONAL now assignable on crypto โ and iterations 9 and 10 confirm each other's mechanism from opposite sides. | ADMIT #1 |
| 11 | PROPOSAL: the campaign's first status declarations (4 CONDITIONAL drafts) | The campaign did what it was built for: with both decision-rule instruments certified and unexpired, the first feature statuses were drafted under a rule fixed from committed numbers BEFORE assignment (bar: โฅ6 of 10 regimes, panel-corrected). Clean bimodal result: 4 features clear CONDITIONAL at perfect 10-of-10 (also STABLE-candidate flagged), 0 FRAGILE, 12 FAIL-today. Awaiting operator review. | PROPOSAL |
| 12 | F014 falsified as BLIND-GAUGE โ the STABLE ceiling re-opens as an instrument gap | The nonlinear invariance instrument was handed a case with a certain answer โ a real byte-identical duplicate, perfectly invariant everywhere. Both versions failed (linear rejects even the true answer; nonlinear identifies nothing). So every "nothing is invariant" ever produced, including the famous 0-of-91, was blindness. Falsifier falsified; what is missing is a better instrument, not more data. | KILL #3 ยท ONE-SHOT RESOLVED |
| 13 | Forex k_v ADMITTED (ADMIT #2) ยท the first FRAGILE declarations | The regime counter earned its forex certificate on its own merits โ passing even the incremental-value test crypto narrowly missed (+0.032, entirely positive, 2,642 held-out rows). 132 forex features drafted: 60 CONDITIONAL (57 perfect 8-of-8), 9 FRAGILE โ the campaign's first โ 63 FAIL-today. Both substrate proposals await the operator. | ADMIT #2 ยท PROPOSAL #2 |
| 14 | SOTA sweep blocked ยท goal-reconciliation consolidated | The literature sweep for a better invariance instrument could not run (search unavailable; no unverified citations admitted). Instead the scoreboard caught up with 13 iterations of change: 9 certified slots (was 7), 19 boneyard entries (was 15), 73 draft statuses pending review, 6 of 8 cells live (was 2), STABLE gap precisely characterized. | PARKED (sweep) ยท CONSOLIDATED |
| 15 | Full-grid k_v: the pending proposals corroborated grid-wide | Counting extended to all 35 market combinations as a self-attack on the pending drafts: the four crypto draft features keep their records everywhere (petrosian_fd perfect on all 20 crypto combos); the counter reads consistently where it should; 70 FRAGILE sightings concentrate on forex. Sweep still blocked, nothing unverified admitted. | SUPPORTING EVIDENCE |
| 16 | Forex k_v hardened: the M3 concern defused | The flagged near-bar correlation (-0.90 with mean redundancy) got the tough test: the panel received the simple aggregate explicitly, and the counter still improved forecasts on top (+0.016, entirely positive CI, null band 10x smaller). The forex certificate stands on firmer ground; follow-up closed. Sweep blocked a 4th time; nothing unverified admitted. | ADMIT HARDENED |
| 17 | PARKED: awaiting operator reviews and search availability | The loop ran out of work it may do alone: draft statuses and the expiry monitor await the operator; the sweep awaits a working search backend (5 errored attempts); expiry re-grounding awaits the calendar. It parks and states exactly what it waits for. | PARKED |
| 18โ19 | Sweep un-parked via firecrawl: two refill candidates for the STABLE gap | Iter 18 diagnosed the blockage precisely (topic-filter, not infrastructure) and queued the self-hosted scraper path; iter 19 ran the sweep through it: two serious candidates for the open instrument gap โ e-value invariance tests (immune to the perfect-fit breakdown that killed causalicp) and invariance-guided regularization โ both gated on the iteration-12 perfect-duplicate entrance exam. | SWEEP EXECUTED ยท 2 CANDIDATES |
| 20 | Candidate A extraction checkpoint: e-values fit the gap better than hoped | The paper's tests reshuffle real data (mandate-compliant by construction), never divide by residual variance (the old instrument's fatal flaw cannot occur), and yield always-valid sequential monitors serving the expiry boundary too. On-paper reasoning suggests it passes the perfect-duplicate entrance exam โ to be proven on real data next. | CHECKPOINT |
| 21 | Cycle 2 resume: harvest feed imported (68 candidates, 6 batches) | Cycle 1 merged (PR #578); fresh branch + campaign PR #591 opened; the operator directive recorded verbatim; 68 pre-screened candidates imported into the frontier โ STABLE batches first, cheapest kills first, e-value entries folded into the pending candidate-A entrance exam. | CYCLE 2 ยท IMPORTED |
| 22 | The entrance exam falsified itself: STABLE is an ฮต-frontier | The e-value instrument rejected the "perfectly invariant" duplicate โ and measurement showed the exam was wrong, not the instrument: the duplicate is 99.9999% linear, its dust is regime-structured, and any honest test at 20k points must see it. Exact stability is unsatisfiable on real data; STABLE becomes a written-down-tolerance question. Candidate A verified calibrated; tolerance-anchored exam next. | EXAM v1 KILLED ยท CANDIDATE A ALIVE |
| 23 | Exams v2+v3: true anchor found, a power gap remains | The tolerance idea failed backwards (crypto anchor retired โ its dust is regime-structured grit); the byte-identical forex duration twins are the true anchor. There the instrument never rejects the truth (validity โ) but wrong answers slip under the bar โ a power shortfall of the simple e-formula; the paper's optimal constructions are the next slice. | CHECKPOINT |
| 24 | ENTRANCE EXAM PASSED: exact identification on the true anchor | After five exam forms: on real data with one genuinely perfect invariant relationship, the instrument accepted every truth-containing answer (five folds, five exact zeros), rejected every wrong one (evidence up to 960), and identified exactly the right variable. Last obstacles: a power gap (fold-product fix) and a numerics gremlin (raw microsecond scales; standardization silenced it). Candidate A enters the formal ladder. | EXAM PASS |
| 25 | ADMIT #3: e-ICP-fp fills the STABLE slot (forex) ยท first identified sets on crypto | The exam survivor walked the ladder: near-zero correlation with every existing instrument (0.03 vs the regime counter), power credential = the passed exam. Third certificate issued โ the vacant STABLE identification slot filled (forex; crypto conditional pending one leg). First-ever non-empty identified invariant sets on three ordinary crypto features. | ADMIT #3 |
| 26 | ADMIT #4: e-ICP-fp crypto leg โ admitted on both substrates | The crypto conditional resolved by the sealed disjunction: no panel lift (ฮAUC CI โ 0 โ the same near-ceiling wall the regime counter hit) but the direction test passed (ฯ=+0.123, CI [+0.06,+0.17] > 0). Fourth certificate โ the STABLE identifier now covers BOTH substrates; the three identified crypto sets graduate to declaration raw material. Thin null margin flagged; expiry re-tests it automatically. | ADMIT #4 |
| 27 | H-016 seqICP passes the entrance exam on the first form | The label-free invariance test (slices time into blocks by itself โ no regime labels) passed the ground-truth exam in one attempt: truth accepted, wrong subsets rejected, exactly the right variable identified. Caveats flagged (shuffle null vs serial dependence; one fixed block layout). Cheapest kill next: redundancy vs the just-admitted e-ICP-fp. | EXAM PASS ยท LADDER PENDING |
| 28 | seqICP M3: block-null correction load-bearing โ clear, with convergent evidence | Run 1 was a stuck gauge (all-reject, zero information) โ the loop attacked its own measurement, confirmed the dossier-predicted serial-dependence hazard, and re-ran with a neighbourhood-preserving shuffle. Gate cleared non-vacuously (worst 0.25 vs the 0.95 bar); five invariance-rich features, two matching the new identifier's picks โ first cross-instrument corroboration on the STABLE frontier. | M3 CLEAR |
| 29 | seqICP crypto persistence role EXCLUDED: sparse support โ candidate alive for forex | The invariance flag lights up mostly for features outside the certified panel's universe โ on the held-out set it never lit at all, and a constant reading can't predict anything. M4 narrow miss, M5'b unmeasurable โ scoped kill by the pre-registered rule. Bonus: ehlers increment asymmetry flagged in every late-era regime โ third consecutive corroboration of the new identifier's pick. | EXCLUDE (crypto role) |
| 30 | ADMIT #5: seqICP โ the env-free corroborator (forex) | Head-to-head over 130 forex features: graded readout, correlation 0.65 vs the reigning identifier โ related but far from the 0.95 kill bar; the disagreements are the value. Fifth certificate, forex-scoped, in a new role: a label-free second opinion at the STABLE boundary โ insurance for when the regime definitions themselves are wrong. Crypto exclusion stands. | ADMIT #5 |
| 31 | H-023 IAS passes the entrance exam โ redundancy warning recorded | The ancestry readout (smallest steady sets โ union) aced the ground-truth exam: empty answer rejected, exactly one minimal set, exactly the right variable. Red-ink caveat recorded first: it runs on the same underlying test as the reigning identifier โ two dials on one sensor. The redundancy gate decides next. | EXAM PASS ยท M3 NEXT |
| 32 | ADMIT #6: IAS ancestry readout โ M3 cleared razor-thin, output scoped | One measurement set, both combination rules, 130 features: the set-count dial is a near-copy of the identifier (0.937 โ in red, NOT certified); the namesake ancestry dial measures something new (evidence on 4 features where the identifier is silent, 0.82 overall). Sixth certificate, deliberately narrow. Bonus: 43/130 features read regime-steady by themselves โ a new STABLE lens. | ADMIT #6 (scoped) |
| 33 | H-019 IPP passes the entrance exam (CRPS form) | The subtler question: does the whole SHAPE of the prediction hold across regimes โ tails included? Exam passed: perfect answers accepted (CRPS survives exact-zero errors where log-scores blow up), wrong answers rejected with the strongest separation yet, right variable identified. Decisive next: the census of features whose shape shifts while the average holds โ this candidate's entire reason to exist. | EXAM PASS ยท M3 NEXT |
| 34 | ADMIT #7: IPP distributional invariance โ 27 shape-shifters found | Paired lockstep tests (same data, same shuffles) found 27 features whose average holds across all ten regimes while their predictive shape visibly shifts (mean-test ~3 vs shape-test ~65,000 on bid_ask_imbalance) โ exactly what a mean-based STABLE certificate would wrongly bless. Redundancy comfortable (0.76 worst, ~0 vs ancestry). The stack can now see the tail-risk axis of stability. | ADMIT #7 |
| 35 | H-025 HSIC-X passes the entrance exam โ always-alarm risk flagged | The guard-for-the-guards candidate (does an anchor secretly entangle with model errors?) passed both exam legs: silent on the correct model, caught the omitted driver on the broken one. Red-ink worry: it flagged everything in the broken test โ specificity untested. Next: the 130-feature flag-rate census; ring-always = dead weight. | EXAM PASS ยท CENSUS NEXT |
| 36 | ADMIT #8: HSIC-X, the anchor-validity guard | Not an always-alarm: graded census across 130 features โ 37 fully clean, 19 partial, 74 with every anchor residual-entangled (standing caveat on their invariance evidence). Worst correlation 0.44 โ the most independent reading of any certified instrument. The shelf now has a guard under the stability meters, checking the ground they stand on. | ADMIT #8 |
| 37 | H-017 StabReg passes the entrance exam โ the ฮบ knob exposed | The team-judging candidate aced identification (true team survives, weight 1.0, empty team rejected at e=1.4M). Twins-must-co-select hit an honest snag โ no team survived on either real target โ but the twins scored bit-identically everywhere: inseparable in any family, established. Red ink: a tuning dial (prediction-vs-stability slack) demonstrably changes answers โ must be derived or shown insensitive before any certificate. | EXAM PASS ยท ฮบ FIGHT NEXT |
| 38 | PARKED: sidecar down after host reboot ยท cycle 3 opened | Pre-work vitals failed: the production data-collector and its self-healer never came back after the machine restarted, and the loop may not touch production services. Bookkeeping only โ cycle 2 was merged by the operator, so a fresh cycle-3 branch and pull request were opened โ plus this honest status note. The ฮบ-dial fight stays queued for the moment the pipeline is green. | PARKED ยท CYCLE 3 |
| 39โ40 | H-017 StabReg EXCLUDED: ฮบ resolved in its favor, then killed as an oracle shadow | Firing 39: one-line parked checkpoint (services still down). Firing 40: the operator restored the pipeline, and the team-judging candidate got the fairest trial โ the tuning-dial worry dissolved (derived value โ the exam's guess; answers flat across the whole dial range), then the killer question landed: its central reading moves in 97.7% lockstep with an already-certified instrument, because its steadiness test is built on that instrument's own engine. Zero new information: rejected, reason on record. | EXCLUDE (M3, ฯ=0.977) ยท ฮบ RESOLVED |
| 41 | H-021 EILLS passes the entrance exam โ ฮณ-invariant selection | A genuinely different engine: "predict well" and "stay steady" folded into one price, cheapest answer wins. Ground-truth exam flawless โ exactly the right predictor named (price exactly zero), every wrong answer thrown out with a price that grows as the steadiness weight rises, and the answer never budged across three orders of magnitude of that weight. Unlike the shadow just killed, it does not ride our certified engine โ the "do you add anything new?" fight is genuinely open, and it is next. | EXAM PASS ยท M3 NEXT |
| 42 | H-021 EILLS EXCLUDED: the empty-set collapse โ a third witness for the ฮต-frontier | The exam ace failed the field test: on all 130 real features it answers "nothing" the moment the steadiness weight matters โ no feature has a perfectly steady anchor relationship, so the steadiness charge always outweighs the tiny prediction reward (median 2.6%). Same answer for everything = measures nothing (the dead-weight failure our first planted control was built to catch). Silver lining: a third independent engine confirms perfect steadiness does not exist here โ steadiness is a matter of tolerance, and the tolerance-based successor (H-026 lead) is already in the catalog. | EXCLUDE (ฮณ-fragile ยท empty collapse) ยท ฮต-WITNESS #3 |
| 43 | H-018 Causal Dantzig passes the entrance exam โ no dials anywhere | A candidate that hunts steadiness in the differences between regimes: structural relationships cancel perfectly when one regime is subtracted from another. Exam emphatic โ perfect cancellation (fifteen decimal places) for the true answer, enormous leftovers for every wrong one: the widest margin on record. No tuning dials at all, and it tolerates hidden common causes โ an angle absent from the whole shelf. Field test across 130 features next. | EXAM PASS ยท ZERO DIALS ยท M3 NEXT |
| 44 | Causal Dantzig M3 CLEAR โ informative, not redundant | The regime-subtraction candidate survived the field test that killed its two predecessors: it did not go mute (readings spread across the whole range) and it is not an echo of anything we own (strongest resemblance 76%, far below the 95% bar). Two statistical hurdles remain โ held-out verdict improvement and detection power โ both with the serial-dependence-aware designs this market demands. | M3 CLEAR ยท M4/M5โฒ NEXT |
| 45 | ADMIT #9: Causal Dantzig, the hidden-confounder-tolerant invariance meter | The regime-subtraction candidate finished the strong way: its readings measurably improve held-out verdicts (entire uncertainty band positive, far outside lucky pairing) AND its closeness-to-regime-proof score predicts which features keep their edge. The rule required either; it won both. Ninth certificate โ a steadiness meter honest even under hidden common causes, zero dials end to end. Curiosity logged: regime-proof-closer forex features KEEP their edge more often โ the mirror image of the old crypto finding. | ADMIT #9 |
| 46 | H-004 PCMCI+ passes the entrance exam โ the F014 trap did not spring | The batch's last candidate, a full causal-map drawer built for autocorrelated data. Its exam hid a trap โ conditional tests on a byte-identical twin leave zero wiggle room, the exact situation that blinded F014 โ and it walked through cleanly twelve times of twelve: exactly one connection drawn (the correct same-instant edge), no crashes, no undefined numbers, dials never changed the answer. Batch 1 fully examined: six certificates, two honorable deaths, one in ladder. | EXAM PASS ยท 12/12 ยท M3 NEXT |
| 47 | H-004 PCMCI+ EXCLUDED: the ฮฑ dial reshuffles its own role readout ยท B-01 CLOSED | The exam ace failed the field test in the subtlest way yet: its regime-wiring-change ranking โ its entire reason to exist here โ survives a small turn of its sensitivity dial only 62.5% intact (bar 90%), on identical data. Verdicts that hinge on an arbitrary knob certify nothing; out, repair path on record (the library's dial-free mode). Batch 1 closes: ten in, six certificates (the STABLE slot now six instruments deep), three honorable kills, all evidence on file. | EXCLUDE (ฮฑ-fragile) ยท B-01 CLOSED |
| 48 | B-03 opens: H-053 Universal Inference passes the entrance exam | Batch two's cheapest kill attempt: a method honest under ANY conditions โ one data half proposes "differs per regime," the other judges, evidence guaranteed fair with no fine print. Exam flawless, zero tuning dials, its rejection bar coinciding exactly with our certified threshold. The decisive fight is named in advance: it answers the same question as our reigning oracle with a different engine โ and echoes die at the redundancy gate. | EXAM PASS ยท ORACLE-SHADOW NEXT |
| 49 | H-053 Universal Inference EXCLUDED โ the split hedge, measured | Died on its one hidden hinge: the method cuts data in half (one proposes, one judges), and swapping the halves should not matter โ it does (89.1% agreement vs the 90% bar, with some features flipping totally: eight answers accepted one way, zero the other). A judge whose rulings reverse when the jury boxes swap certifies nothing. The echo-of-our-oracle danger named in advance turned out FALSE (73% resemblance), and the textbook repair โ average over many cuts โ is a new admissible candidate. | EXCLUDE (split-fragile) ยท SHADOW MISSED |
| 50 | H-058 consumed by coordination โ the theory root of an instrument we already certified | The queue entry turned out to be the parent theory of certified candidate A. Re-examining the parent would be double work with no good outcome โ a pass duplicates an existing certificate, a fail teaches nothing โ so the no-double-tracking rule closes it by written ruling. Nothing discarded: its one distinct capability (evidence honest under indefinite watching) is exactly the deferred certificate-expiry monitor, so its ready-made package is filed there awaiting the operator's go. | CONSUMED (coordination) ยท LEAD โ #4 |
| 51 | Anchor triad passes exam form v2 โ form v1's margin self-falsified | Three econometric guards examined together: anchor honesty, anchor strength, weak-proof verdicts. All three aced every certain-answer question, yet form v1 said FAIL โ its 100ร cushion turned out to be doubled protection (the exam already corrects for dependence directly). Per the iteration-22 precedent the exam was fixed on the record; form v2 passed cleanly. Both forms on file; the triad enters the field test as a unit. | EXAM PASS (v2) ยท CENSUS NEXT |
| 52 | Anchor triad M3 CLEAR โ parametric โ kernel | All three guards passed the field test, and the pre-named fight gave the best result: the honesty-checker is NOT our certified kernel guard in different clothes โ their readings agree less than half the time, so the classical moment check and the kernel check see different failure surfaces. Both on the shelf = stronger. Triad members resemble each other moderately (same family) but below the bar, and the strength-meter's spread settles its discrimination question. All three march to the finale as a unit. | M3 CLEAR ร3 ยท FINALE NEXT |
| 53 | Triad finale: two admits, one kill โ the anchor-failure surface covered | The final hurdle split the triad: the strength meter and the weak-proof verdict meter delivered real held-out improvements โ certificates ten and eleven โ while the honesty checker's admittedly-different readings improved nothing (slightly the wrong way, in fact). Different โ useful: excluded, kernel guard keeps the honesty job. Discovery: all three meters agree strongly-anchored features are MORE likely to lose their edge โ the iteration-9 inverted mechanism through a new lens. Anchor surface now fully covered. | ADMIT #10 + #11 ยท H-065 OUT |
| 54 | H-027 LPCMCI passes the entrance exam โ the latent-variant F014 check comes back clean | The map-drawer's careful sibling โ honest even when a hidden puppeteer pulls several strings, the blind spot every map-drawer shares. Twelve for twelve: exactly one connection drawn, no crashes, half a minute despite the heaviest machinery yet. But the family history hangs over it: its sibling aced this exam then died when its sensitivity dial reshuffled its own rankings. Same dial, no self-tuning โ the field test, where the dial is decisive, is next. | EXAM PASS ยท 12/12 ยท ฮฑ DECISIVE NEXT |
| 55 | H-027 LPCMCI EXCLUDED โ the family lesson, measured twice | The careful sibling failed the pre-registered decisive check harder than its brother: its regime-wiring ranking survives a dial turn only 39% intact (brother 62.5%, bar 90%) โ the hidden-puppeteer machinery amplifies the dial rather than taming it. Family law, measured twice: this school's rankings are threshold-brittle at scale, and exam success predicts nothing about it. Next: a different school โ a deterministic optimizer with no dial at all. | EXCLUDE (ฮฑ 0.393) ยท FAMILY ร2 |
| 56 | H-030 DYNOTEARS EXCLUDED at the door โ biased on the truth, phantoms everywhere | The dial-free hope died fastest, at the entrance exam: its reading of the one certain answer (exactly 1.0 by construction) ranged 0.97โ0.56 with its penalty knob โ up to 44% underived bias โ and it drew up to thirteen phantom connections for a variable with exactly one true relationship. Never crashed: confidently wrong, not blind โ what the exam exists to catch, at the cheapest price. Structure-learning: zero for three here, three different failure modes, all on record. | EXCLUDE (exam kill) ยท 0/3 SCHOOL |
| 57 | The DRO cousins both pass the paired entrance exam | A different question after the map-drawers' 0-for-3: how much worse is the worst regime than average? Two cousins, two framings โ named regimes vs the worst fraction of moments. Joint exam spotless: exactly zero on the perfect model, decisively lit on the wrong one (worst regime +2.4 units; worst 10% of moments ~8 units of excess), the tail dial passing its built-in mathematics check. Field test as a pair next โ are they secretly one instrument? | EXAM PASS ร2 ยท CENSUS NEXT |
| 58 | The cousins split: the gap survives, CVaR falls on a knife edge | The dial-free worst-regime gap meter came through clean (strongest certified resemblance 63%); the tail-risk cousin fell by the campaign's narrowest margin โ 89.7% vs the 90% bar โ with the honesty on record that a friendlier range would have acquitted it, and that redrawing ranges after seeing results is exactly what the campaign forbids. Cheap kill (the survivor covers the job at 77% overlap), cheap repair. The survivor marches to the finale. | H-060 CLEAR ยท H-061 OUT (0.897) |
| 59 | H-060 EXCLUDED: actively harmful โ the DRO family closes 0-for-2 | The surviving cousin failed the last hurdle in the most damning way: its readings actively WORSEN held-out verdicts, beyond random-pairing noise โ and since the machinery may use readings upside-down, this isn't a fixable sign flip: the arrow points a different way each era. Deposit: for the third time through a third lens, stability-adjacent "worse" features keep their edge MORE โ the iteration-9 inverted mechanism, corroborated three ways. Both cousins honorably dead. | EXCLUDE (harmful) ยท DRO 0/2 |
| 60 | H-063 V-REx passes the entrance exam โ its required null wired in | The regime-variance candidate (how UNEVEN are the regimes together?) passed cleanly: exactly zero on the perfect model, and on the wrong one a reading 2ร anything label-shuffling of the same real losses produces โ its paperwork-required shuffle test wired in on schedule, honest margin on record. A shadow hangs over the ladder: unevenness and the executed worst-regime gap summarize the SAME numbers โ near-identical rankings would transfer the gap's fate, killing this one cheaply at the census. | EXAM PASS ยท AFFINITY CHECK NEXT |
| 61 | H-063 V-REx EXCLUDED by fate transfer โ the loss-spread family closes 0-for-3 | The named shadow materialized precisely: 98.2% rank agreement with the executed worst-regime gap โ variance and worst-minus-average are two summaries of the same ten numbers. The expensive finale lesson (actively harmful, era-flipping direction) transfers to the twin without a second doomed spend โ exactly why the check was front-loaded. A third school closes zero-for-N; two candidates remain in the batch. | EXCLUDE (0.982) ยท SPREAD 0/3 |
| 62 | H-064 calibration gap passes the entrance exam โ the strongest null margin yet | A genuinely different object from the fallen family: not error size per regime, but whether the model's confidence statements stay honest in each regime โ theorem-linked to invariance. Cleanest diagnostic exam so far: zero on the perfect model, six-fold beyond the shuffle null on the wrong one, at every binning setting. Decisive fight named: certified IPP already reads distributional honesty โ a shadow verdict ends it. | EXAM PASS ยท IPP-SHADOW NEXT |
| 63 | H-064 M3 CLEAR โ the IPP shadow missed | Three checks, three passes: the binning knob harmless at scale (96โ98% stable โ the dial that killed two map-drawers), the feared IPP echo absent (67.5% โ related but distinct), and no disguise of the executed loss-spread family (42โ44%). One candidate stands where three schools fell โ only the final hurdles remain, era-stability watching. | M3 CLEAR ร3 ยท FINALE NEXT |
| 64 | ADMIT #12: the calibration gap meter โ the campaign's strongest lift | The survivor of three fallen schools finished with the campaign's best result: confidence-honesty readings improving held-out verdicts by nearly TWICE the previous record, the whole uncertainty band positive. The era-stability question answered itself. Twelfth certificate, for a channel nothing else reads โ and the inverted mechanism is now corroborated four ways, begging for its own investigation. One pair remains in the batch. | ADMIT #12 ยท RECORD LIFT |
| 65 | H-062 W-DRO PARKED with a PROPOSAL โ the certificate is void on dependent data | The candidate's selling point is a guaranteed number, but the guarantee assumes independent samples โ void on market bars, making it exactly the confident-but-unbacked reading this campaign keeps off the shelf. The guarantee-less version would just recreate the executed loss-spread family. Parked with a written repair spec (dependence-aware radius + coverage check on our eight real shifts) and a re-harvest lead. Zero compute spent on a question arithmetic already answers. | PARKED ยท PROPOSAL |
| 66 | H-028 CCM EXCLUDED: the substrate speaks โ B-03 CLOSED | The final batch candidate was killed by the market itself: handed two byte-identical series, its coupling reading ranged 87%โ99.6% by regime โ it reads coupling through self-predicting dynamics, which works for plankton but not noisy markets. A gauge that reads the same truth differently per regime certifies nothing; the dossier's flagged weakness, confirmed at the cheapest gate. B-03 closes 13-for-13: three certificates (incl. the campaign's best), eight honorable kills, one coordination-close, one park. | EXCLUDE (substrate) ยท B-03 CLOSED |
| 67 | B-04 OPENS: H-073 MMD passes exam v2 bit-exactly (v1โs zero point self-falsified) | The queue turned to the boundary in worst shape โ the FRAGILE-vs-CONDITIONAL split rests on a single certified instrument. The new batchโs cheapest candidate, a kernel distance asking whether a featureโs distribution differs between regimes, passed its exam bit-exactly after one instructive stumble: the first exam form demanded a series-vs-byte-twin distance of zero, but the textbookโs unbiased estimator deliberately reads slightly negative on identical samples โ the exam was wrong, not the candidate; the corrected form passed with differences of exactly zero while real regime shifts rang the bell at up to fifty times the shuffle bar. | EXAM PASS (v2) ยท B-04 OPENED |
| 68 | H-073 MMD M3 CLEAR: the incumbent shadow missed by a wide margin | The kernel-distance candidate faced the field test that has killed six candidates: across all 130 features, does its reading repeat something a certified instrument already says? It does not โ against the incumbent regime counter the correlation is a weak NEGATIVE: the features whose distributions shift most between regimes are not the ones the counter flags as verdict-unstable, and two instruments that disagree about which features are regime-sensitive is exactly what a second boundary instrument is for. Recomputed two independent ways, matched to the last bit. One test remains: held-out forecasting value. | M3 CLEAR ยท M4/M5′ NEXT |
| 69 | ADMIT #13: the MMD heterogeneity meter โ thinnest margin, era-attacked before certification | The kernel-distance candidate passed its final test โ barely, and the record says so out loud: held-out improvement with an entirely-positive confidence interval satisfies the frozen rule (certificate #13; the FRAGILE-vs-CONDITIONAL boundary now rests on two instruments), but the lift sits inside the random-re-pairing band every prior admit cleared by 6–18×. The loop attacked its own result โ era split: holds in both halves, no sign flip, most lift in the earlier half. Cautions on the certificate; whether the bar itself should rise is flagged for the operator. Bonus: fifth data point for the both-extremes-lose mechanism. | ADMIT #13 ยท THIN MARGIN โ CAUTIONS SCOPED |
| 70 | H-056 betting passes the entrance exam on the first form: the wealth identity is estimator-clean | The next candidate answers the same question as certificate #13 but as a gambler: it bets on the difference between two samples; honest odds guarantee wealth cannot grow beyond luck if there is none โ accumulated wealth IS the evidence, valid at every moment. Handed a byte-twin, it finds nothing to bet on and wealth stays at exactly 1.0 โ no estimator quirk, no second form needed (unlike MMD’s). On real regime pairs it got rich fast: wealth 405 to 43 million against a bar of 20, calling the difference as early as 110 pairs in. The shadow fight is named: echoes die at the redundancy gate. Reference code has no license โ rebuilt from the paper, zero reuse. | EXAM PASS ยท MMD-SHADOW NEXT |
| 71 | H-056 betting EXCLUDED: the named check fired โ certificate #13’s echo | The gambler was killed by exactly the shadow named before the fight: across all 130 features its wealth ranking replays the certified kernel distance’s ranking at 98.2%, and the mirror is total โ against every reference instrument its readings track the distance meter’s almost number for number. It bets USING the same kernel witness the meter reads: a sequential retelling of the same measurement. The real anytime advantage buys nothing where rankings are the currency. Family lesson, second sighting: a different inference wrapper around the same witness is the same instrument. | EXCLUDE (ECHO ฯ=0.982) ยท KILL #16 |
| 72 | H-074 LGC-switch EXCLUDED at the door: a blind gauge โ the WHERE costs more than the signal | The candidate promised WHERE regimes differ (center vs tails), not just whether. The exam killed it at its own game: asked to detect covid-crash vs mid-cycle โ a difference the certified kernel meter reads at up to fifty times its noise floor โ it could not tell the real split from re-shuffled blocks of the same data, on any pair. Tail dependence is measured from the few points living there; the noise floor is enormous and the answer drowns (the iteration-12 blind-gauge mode). Side lessons: two bit-exact certainty claims failed on floating-point technicalities while the ε-tolerance passed; the sibling measure-candidate stays queued with the risk pinned. | EXCLUDE (BLIND GAUGE) ยท KILL #17 |
| 73 | H-038 energy distance passes the exam: the first derived-then-observed exactness | The second distance cousin โ built from plain absolute differences, no tuning knobs at all โ passed on the first form, with a methodological first: after two exams whose exactness claims broke on floating-point technicalities, this one’s was DERIVED before running, and the machine read exactly 0.0 as proven. Real regime pairs at 3–31× its shuffle bar. Now the fight it was warned about at the door: energy distance is mathematically an MMD in disguise, and the gambler died for exactly that kinship โ one census will judge both remaining cousins against certificate #13’s shadow. | EXAM PASS ยท GROUPED M3 NEXT |
| 74 | The cousins split: energy killed by its sibling, Wasserstein advances | The predicted double execution did not happen โ the miss is the interesting part. Neither cousin echoes certificate #13 (92.1% / 90.8%, under the bar): different geometries genuinely rank features differently, whereas a different wrapper around the same kernel did not. But the cousins echo EACH OTHER (97.8%), and the pre-registered tie-break advances the sibling less similar to the existing certificate. Energy โ exam flawless โ dies an honorable duplicate’s death; transport goes to the final held-out test. One instrument per geometry is what the stack buys. The survivor’s math: exact agreement with an independent library. | H-038 EXCLUDE (#18) ยท H-075 M3 CLEAR |
| 75 | ADMIT #14: the Wasserstein meter โ clean margins where #13 was thin | The transport-geometry meter โ which weights how far probability mass must move between regimes, so tails cost more โ passed the final held-out test with the clean margins its kernel cousin lacked: improvement 4× beyond the random re-pairing band (where #13’s sat inside it) and 3.2× #13’s size; era check inline โ holds in both halves, no flip. Third distributional meter saying shift-prone features lose their edge (six data points now). Pointed operator note: under the stricter bar being considered, #13 would have failed exactly where this one passes. | ADMIT #14 ยท BEYOND-NULL 4.0ร |
| 76 | H-071 LGC measure EXCLUDED: within-regime resolves, between-regime drowns โ the family closes 0-for-2 | The blind gauge’s sibling beat the exact trap that killed its kin โ its center-vs-tails reading proved stable inside a regime โ only to die one door later: a boundary instrument must say the map DIFFERS between regimes, and there its reading (0.21) drowned under an honest noise floor of 0.41. Family verdict zero-for-two: region-resolved dependence describes regime interiors but cannot certify between-regime differences on this data. A real observation fell out: in calm regimes dependence concentrates in the center; in crashes it spreads into the tails. | EXCLUDE (BOUNDARY-BLIND) ยท KILL #19 |
| 77 | H-049 CRQA EXCLUDED: an autocorrelation gauge in coupling clothing | The candidate promised to read whether two series’ DYNAMICS move in aligned patterns. The exam fed it versions with one series rotated in time โ same values, same internal rhythms, all alignment destroyed. If it truly reads alignment, rotated data must score much lower; in all three regimes the de-aligned score came within a thousandth of the real one. The coupling reading reflects each series’ internal rhythm โ marginal structure the certified meters already capture; its regime differences are repackaged information. The surrogate leg is where pretenders die. | EXCLUDE (NO ALIGNMENT INFO) ยท KILL #20 |
| 78 | The stability pair: DOUBLE KILL at the door โ the door law fells its third family | Two candidates entered together: a selection-consistency counter and its smarter sibling that stops counting swaps among near-identical features as instability. Controls passed โ the byte-twin was treated identically everywhere; the adjustment was caught working on camera (honest surprise: it isn’t always calming). Both died at the door that has now killed three families: their regime-to-regime readings differ no more than randomly re-split blocks produce. The meta-pattern discriminates: resampling-style regime readings drown in block noise while real boundary signal cleared the same bar at 3–50×. | DOUBLE EXCLUDE ยท KILLS #21/#22 |
| 79 | H-070 RBO passes the exam: the dial is innocent, the rankings are real | The candidate tracks each feature’s rank across the eight walk-forward windows, top-weighted โ a continuous cousin of the certified flip counter. Three doors, three passes: byte-twin substitution reproduced every number to the bit; its one knob barely moved the ordering (99.7% โ unlike two earlier dial deaths); real rankings proved four times more consistent than shuffled ones. Now the fight it was born for: does the continuous reading say anything the flip counter doesn’t? The incumbent’s shadow is the decisive check next firing. | EXAM PASS ยท k_v-SHADOW NEXT |
| 80 | H-070 RBO M3 CLEAR: the campaign’s most orthogonal survivor (k_v shadow 0.019) | The fight it was born for ended in the most lopsided result of the campaign โ in its favor: expected to overlap heavily with the certified flip counter, its reading shares essentially nothing with it (0.019, where every previous candidate measured at least ±0.35). Worst overlap against the entire 21-instrument stack: 0.63 โ every previous survivor’s worst was at least 0.84. The most genuinely new reading to survive the redundancy gate all campaign. Caveat stated out loud: novelty is not value โ the final held-out test decides. | M3 CLEAR ยท FINALE NEXT |
| 81 | H-070 RBO EXCLUDED: novelty without value โ and mechanism datum #7 | The campaign’s most original survivor died at the final gate and taught two lessons: being genuinely NEW is not being USEFUL โ its reading overlapped with nothing, yet adding it actively hurt held-out forecasts in both halves of the test period. More valuable than a certificate: its direction is significantly backwards โ features that persistently hold top positions are LESS likely to keep their edge. Iteration 9’s discovery, reproduced from an instrument sharing nothing with the one that found it. Seven data points, three families: the inverted-mechanism investigation is proposal-ready. | EXCLUDE (ACTIVELY HARMFUL) ยท KILL #23 |
| 82 | H-044 congruence EXCLUDED: it cleared the door law โ and the ordering leg proved the clearance was noise | The candidate reads whether the market’s hidden factor structure rearranges between regimes โ and became the first since the distance cousins to clear the door law. Then a subtler control exposed the clearance: the structure measured on two halves of the SAME regime must agree better than structures from DIFFERENT regimes, or the drift is measurement wobble. It did not โ same-regime halves disagreed more (0.55–0.59) than different regimes did (0.63): with 125 features and 28 factors on ~1500 rows, estimation noise dominates. The boneyard prescribes the successor’s fix. | EXCLUDE (NOISE GAUGE) ยท KILL #24 |
| 83 | H-072 MRP passes the exam: B-04’s exam phase complete โ the k_v fight decides its close | The batch’s final candidate is its simplest: for each feature, keep the WORST regime reading โ the floor. Three doors, three passes: the byte-twin scored exactly 1.0 in all eight regimes as arithmetic demands; the door law that killed four candidates was cleared honestly (1.44×, with no noisy estimation layer to fake it); and the floor is not merely the average in disguise (84%, under the bar, noted). Whether B-04 closes at 2 certificates and 10 kills or gains a late third rides on the incumbent fight: floor aggregation vs the certified flip counter. | EXAM PASS ยท k_v FIGHT DECIDES |
| 84 | H-072 MRP EXCLUDED: it won the k_v fight and lost to the dial law โ B-04 CLOSES: 2 admits · 10 kills | The batch’s last candidate won the fight everyone was watching (9% overlap with the certified flip counter โ genuinely different) and lost to a technicality that isn’t one: on a slightly smaller sample its ranking reshuffled wholesale (74%). A MINIMUM is set by whichever regime scores lowest, and noisy estimates crown different worst-regimes on different samples. The boneyard prescribes the successor (a soft minimum โ the orthogonality is real). B-04 closes 12-for-12: two certificates (the boundary that rested on one instrument now has three) and ten kills, each by a different named mechanism. | EXCLUDE (DIAL 0.735) ยท B-04 CLOSED |
| 85 | B-06 opens: GW-CPA passes the exam โ every boundary has now fed | The queue turned to the deepest remaining deficit: the CONDITIONAL boundary’s declaration layer โ the formal machinery for saying “this feature works in THESE regimes” with controlled error rates; sixty drafts await it. With this batch open, every boundary has received candidates โ a campaign first. Its cheapest candidate is its most on-point: the econometrics-native test of regime-conditional predictive advantage. Exam passed first form โ byte-twin forecasters produced a loss difference of zero to the last bit, the identical-forecasts guard fired properly, and the real question was decisive at odds of fifteen million to one. | EXAM PASS ยท B-06 OPENED |
| 86 | H-093 GW-CPA EXCLUDED: the closest dial-law miss on record, with the cleanest sheet otherwise | The declaration layer’s opening candidate died the campaign’s narrowest death: on a slightly smaller sample its ranking agreed 86.8% with itself against a 90% bar โ the rule fires mechanically. Everything else was the cleanest sheet any candidate has shown (worst stack overlap 37%; named fights won outright). Two durable goods: the regime-conditioned reading genuinely differs from the plain test (74%) โ so Diebold–Mariano enters next on its own merits โ and a precise successor prescription sits on the boneyard: the information is there, only the estimator wobbles. | EXCLUDE (DIAL 0.868) ยท KILL #26 |
| 87 | H-092 Diebold–Mariano EXCLUDED: the normalization cancels the regime signal | The classic forecast-comparison test died at the door with the sharpest mechanistic story yet: its per-regime readings were individually excellent, but its job is the DIFFERENCES between regimes โ and there its spread came in at less than half what shuffled pseudo-regimes produce. Crisis regimes have proportionally bigger averages AND bigger noise; dividing one by the other comes out the same everywhere โ the division destroys the signal. New family bar: per-regime declaration instruments must compare raw moments, not self-normalized statistics. | EXCLUDE (SIGNAL CANCELLED) ยท KILL #27 |
| 88 | H-047 BY EXCLUDED by input invalidity: the HAC p-values are broken on this substrate | The false-discovery backbone died at the door โ not by its own fault, and the autopsy produced the headline: with every discovery false by construction, a 5% error budget produced false discoveries in 43% of replicates (uncorrected: 62%). The standard time-series p-values understate uncertainty on data this persistent; no correction launders invalid inputs. Resurrection clause built in: BY re-enters once a valid p-source is certified โ and the next candidate is precisely a resampling inference that might be it. Operator flag: 99.9%-threshold certificates have margin, not immunity. | EXCLUDE (INPUT INVALIDITY) ยท KILL #28 |
| 89 | H-046 Romano–Wolf passes: FWER zero where HAC broke at 43% โ the p-source found | Yesterday’s disease met today’s cure: the stepdown procedure’s own inference took the identical torture test โ forty-nine worlds where every discovery is false by construction โ and produced zero false discoveries in all forty-nine. Instead of trusting a textbook formula for uncertainty, it measures it empirically from time-rotations of the data itself. On the real panel: fifteen genuine discoveries (the broken pipeline claimed sixty-one). The wild bootstrap’s sign-flip gray zone is flagged for the operator; the clean rotation-based version was used. The declaration layer has its foundation candidate. | EXAM PASS ยท FWER 0/49 |
| 90 | H-046 RW EXCLUDED as an instrument โ the backbone stands, a pipeline proposal goes to the operator | A kill that must not be misread: as a feature-RANKING instrument the stepdown’s score died at the dial law (thirty distinct values โ a coarse steppy score reshuffles under small data changes, 80.3% vs the 90% bar), but the landmark stands โ as an ERROR-CONTROL engine it was flawless where the standard method failed 43% of the time. The resolution is the contribution: the declaration layer needs a trustworthy inference backbone, not another ranking instrument โ a PIPELINE decision, not a stack admission. The formal proposal sits with the operator. | EXCLUDE AS INSTRUMENT ยท PROPOSAL FILED |
| 91 | The SPA–MCS pair splits: MCS passes with its guarantee verified, SPA dies on a marginal blindness | Two set-minded candidates split: the Model Confidence Set โ “in WHICH regimes is a feature’s advantage indistinguishable from its best?” โ passed all three doors (twin identity; coverage verified on twenty-five all-equal worlds; resolved real structure by scoping out one of eight regimes). Set-valued CONDITIONAL scoping working at the door. Its sibling missed its own bar (0.07 vs 0.05) where structure is certified at fifteen-million-to-one โ the blindness rule fires; marginal, noted, low-priority revisit. Next: MCS’s reading is a COUNT of regimes โ and the certified incumbent k_v is also a count. Same grammar; the census decides. | H-095 PASS ยท H-094 KILL #30 |
| 92 | H-095 MCS EXCLUDED: the set-valued scoper almost never fires at grid scale | The scoper that shone at its exam went nearly silent at scale: asked to scope 123 features, it kept ALL eight regimes for 116 โ the exam’s demonstration was one of only seven features where the meter fires. At this data volume the elimination test lacks power for 94% of the panel; a reading that is almost always “everything qualifies” certifies nothing, and its sparse remainder reshuffles under the slightest perturbation. The coarse-readout pattern is confirmed three times over. Constructive ending: its natural home is the declaration PIPELINE, folded into the pending operator proposal. | EXCLUDE (NEAR-DEGENERATE) ยท KILL #31 |
| 93 | H-090 exceedance correlation EXCLUDED: real tail asymmetry, wrong boundary | The most philosophical death of the campaign: it measured something TRUE that lives at the wrong boundary. After an honest form fix (the canonical pair moves in opposite directions; v2 orients it, no knob), the discovery: crash-vs-rally co-movement asymmetry is real โ about 0.25 extra correlation in the crash tail, cleanly resolved inside each regime, exactly where the local-dependence family drowned. The kill: that asymmetry is nearly identical across regimes (0.04 vs a 0.21 floor). A reading that never changes carries zero boundary information โ STABLE-flavored structure, re-entry note filed. Plus a bookkeeping correction: five queued was wrong; six were, five remain. | EXCLUDE (REGIME-INVARIANT) ยท KILL #32 |
| 94 | H-086 extremogram EXCLUDED: the tail-invariance pattern is named | The second tail-structure candidate died the exact same death as the first โ and two identical deaths make a discovery. Impeccable credentials: fed shuffled data, the real clustering reading beat the shuffled one in every regime (up to double) โ genuine temporal clustering, precisely where the recurrence candidate failed. Between regimes: 0.11 against a 0.65 floor. The pattern is named: TAIL structure is the substrate’s constant; BULK structure is what regimes change. Prediction registered: the next pair are also tail meters โ if they die the same way, the characterization layer closes with a genuine finding and zero instruments. An honest outcome, not a failure. | EXCLUDE (REGIME-INVARIANT) ยท KILL #33 |
| 95 | The tail pair: double kill, prediction confirmed โ the pattern stands at four sightings | Last iteration the loop bet against its own candidates, registering in advance that both tail meters would measure something real and regime-flat. Both landed โ the second at the sharpest margin of the campaign: tail dependence reads a massive 0.84 (17× independence), and it is 0.8378 in the covid crash versus 0.8400 in the calm mid-cycle โ identical to two decimals against a noise floor a hundred times wider. Four tail meters, four clean measurements, four regime-flat readings: proposal-grade science. The tails have one law that never changes with the weather; the body of the distribution is where the weather lives. | DOUBLE EXCLUDE ยท KILLS #34/#35 |
| 96 | H-005 Regime-MCI EXCLUDED: the closest door-law miss on record โ under-powered, not invariant | The different-family candidate died the narrowest door-law death โ itself information. It reads whether intensity CAUSES the next bar’s duration (own pasts accounted for; the significance knob that killed both family predecessors removed by design). The link is real with the sensible negative sign, and โ unlike the four tail meters โ DOES look regime-varying: fifty percent stronger in the crash. The difference fell just short of the noise floor (88% vs the tail meters’ 1–18%). No same-breath rerun (threshold-hacking); instead the boneyard records the batch’s strongest re-proposal case: the same form on full windows. | EXCLUDE (0.88×) ยท KILL #36 |
| 97 | H-089 Hรผsler–Reiss EXCLUDED: the fifth sighting, independently witnessed โ B-06 CLOSES 12/12 | The batch’s last candidate ran under a standing bet โ four tail meters in a row had measured something real that never changes between regimes โ and died exactly as predicted, at 0.53 against its noise floor. One extra check made it the strongest possible close: a wrapper test shows this meter is NOT a re-labelling of the previous kill (correlation 0.23 where a re-labelling scores above 0.95). Five sightings from provably different instruments, all real, all regime-flat: the tails have one law that never changes with the weather; regimes live in the body of the distribution. Batch product: a pipeline proposal + a five-sighting pattern awaiting the operator. | EXCLUDE (0.53×) ยท KILL #37 ยท B-06 CLOSED |
| 98 | B-02 OPENS; H-036 Reduced TE EXCLUDED: the IID-certificate class claims a second member | The new batch’s first candidate pitched speed with rigor: a transfer-entropy meter with a textbook noise formula, no resampling needed. The exam attacked the pitch โ textbook noise formulas assume independent samples, and bars are anything but. Fed provably-null data, the formula cried signal three times too often: 16% false alarms where it promised 5%. Same disease as the HAC kill โ a two-member failure class, proposed as a standing rule. The twist: the reading itself is real (strongest in the CALM regime) and not a duplicate of the incumbent โ the mechanism re-enters later under its surrogate-null variant. Only the certificate is dead. | EXCLUDE (16.2% FA) ยท KILL #38 ยท B-02 OPENED |
| 99 | Four-cousins quadruple kill: every null the family brought is invalid โ the certificate class reaches six members | Four textbook dependence detectors, examined together, all SEE the dependence (4–70× their noise floors) โ and all four died the same death: their noise floors assume independent samples. Hoeffding’s “exact” null false-alarms 57% (worst on record, and it duplicates the incumbent gate anyway); even distance correlation’s block-shuffled repair fails at 24% โ short blocks can’t fake long memory. Six broken certificates now โ law-grade. The exam’s self-check also caught a float-arithmetic bug in our τ* (fixed, verdict unchanged), and refuted the feed’s interchangeability claim: dcor is a genuinely different witness โ its pairing with the certified shift-based floor is the family’s best re-proposal. | EXCLUDE ×4 ยท KILLS #39–#42 |
| 100 | H-055 SKIT: null machinery CERTIFIED, instrument EXCLUDED โ inverted at the admission bar | The batch’s first survivor โ and the ladder did its job. SKIT’s anytime safety guarantee HELD on the dependence-preserving nulls that broke six other certificates (0.7% false alarms) โ banked as the campaign’s second trustworthy noise floor, exactly the part the killed cousins lacked. The census showed a channel nothing in the stack reads (worst overlap 0.59). Then the admission bar killed it: actively harmful in the forward-prediction harness, direction significantly BACKWARDS โ intensity co-movement with duration predicts orthogonality LOSS (inverted-mechanism sighting #8). Kill the instrument, keep the engine, log the physics. | EXCLUDE ยท KILL #43 ยท NULL CERTIFIED |
| 101 | ETE + copula entropy double kill: the shuffle certificates fail too โ eight broken certificates | The transfer-entropy meter came back for its second chance with the resampling noise floor its literature recommends. Better (13% false alarms instead of 16%) but still broken โ shuffling destroys the memory structure de-aligned market series keep. Copula entropy failed far worse (39%) while measuring something spectacularly real that no other instrument reads. Eight broken certificates; and the intensity→duration information flow is now real, twice-confirmed, twice-orphaned โ its way forward is the matched-pair proposal on the operator’s desk. | EXCLUDE ×2 ยท KILLS #44/#45 |
| 102 | PID double ADMIT (#15, #16): B-02 closes 10/10 โ the FAIL boundary’s first certified instruments beyond G0 | The batch that killed eight in a row ends with two admissions. The pair asks what the dead candidates didn’t: of the information two features carry about the NEXT bar, how much is shared, how much unique, how much only appears combined? Airtight entrance anchors (a duplicated feature reads exactly zero unique share), census-clear novel channels (overlaps 0.70/0.57, and only 0.88 with each other), and both cleared the admission bar cleanly โ entire uncertainty interval above zero on held-out forward prediction. Both point backwards (mechanism sightings #9/#10); a June WATCH flag travels with the second. Every actionable feed candidate is now evaluated โ the handover executes next. | ADMIT #15 + #16 ยท B-02 CLOSED |
| 103 | THE HANDOVER: every metric in plain words ยท the live-expiry ADR ยท the flowchart โ loop PARKED | The operator’s directive executed in full: five batches closed (57 of 68; the last 11 wait behind the operator’s own review). Every admitted instrument and every kill explained in plain words, grouped by the seven ways instruments die; a proposed LIVE-EXPIRY architecture (90-day clock, drift tripwire from each instrument’s own admission band, regime-turn sentinel reserved for the gated batch); and the flowchart that turns an expiry into a re-certification instead of a silently stale certificate. Parked until further direction. | HANDOVER ยท ADR PROPOSED ยท PARKED |
Campaign: 2026-07-02-matrix-admission-paradox-loop ยท created 2026-07-03 ยท read-only ยท append-only ยท this page is the human-readable twin; the audit folder is the machine-readable twin