A fourteen-round loop stress-tested the two feature-screening probes and found that most of what had been blamed on the data is actually a property of the measuring tool and its pass rule — including a rule that refused 72% of candidates it should have accepted — but the loop hit its own round cap with three items untouched and every proposed dial change left unsigned for the operator.
Lifecycle, not result. This says where the audit sits in its process — never whether what it found was good.
[The folder contains no CLAUDE.md and no verdict.md, so it carries no terminal declaration of its own. The strongest in-folder status statements are these two verbatim lines from PRE-REGISTRATION-DRAFT-xi-stability-guard.md:] "**Status: DRAFT. Not in force. No dial has been changed.**" ... "**Requires operator sign-off before any candidate is re-scored.**"
2026-08-17. Derived from folder evidence, then adversarially challenged; the challenge pass was upheld.
This audit has no verdict document. Nothing in its folder records a conclusion under any of the names this generator looks for. That is stated rather than papered over — the reconciled status and the sourced claims below are what the folder actually supports.
Every row pairs a claim with the file it came from and the verbatim text in that file. The sources sit above the deploy root, so the quote is embedded and the path is printed as text rather than linked — a link would resolve on a laptop and 404 here.
| Claim | Evidence |
|---|---|
| The loop stopped because it hit its own iteration cap, not because it finished: 14 of 17 attack-list items were addressed and the cron was cancelled. CONFIRMED 14 iterations run; 14 of 17 attack-list items addressed; 3 items (L-3, L-4, L-5) left OPEN | > **STATUS: CAP REACHED 2026-07-30.** 14 iterations run, 14 of 17 items addressed, loop stopped and findings/dashboard/campaigns/2026-07-29-probe-hardening-loop/LOOP_PROMPT.md |
| The campaign's machine-readable lifecycle record marks it HALTED and blocked on the operator since 2026-07-30. CONFIRMED iteration 14; blocked since 2026-07-30 | "status": "HALTED", findings/dashboard/campaigns/2026-07-29-probe-hardening-loop/LOOP_LEDGER.json |
| The repo's ξ kernel matches Chatterjee's own CRAN XICOR 0.4.1 to far inside the 1e-9 tolerance every other kernel in the repo is pinned at, closing the only unpinned statistic in the cascade. MEASURED leg 1: 15 tied-Y cases, all pass, max |repo − XICOR| = 2.44e-15 at tol 1e-9, n=2,000; leg 2: 6 X-tie cases vs 2,000 random XICOR tie-breaks each, all pass at Bonferroni α=0.001667 | "tol": 1e-09, findings/evolution/audits/2026-07-29-probe-hardening-loop/harness/xi_oracle_parity_evidence.json |
| The previously reported 2e-3 disagreement between the repo kernel and "canonical FOSS" was against the third-party xicorpy package, not against canonical XICOR — the audit refutes its own earlier diagnosis. REFUTED reported gap 2e-3; a local simulation put the true simplified-vs-general formula gap at ≈2.9e-2 at a 30.9% Y-tie rate and ≈1e-3 at ≈10% — mutually inconsistent with the 2e-3 figure | ### 1.12 The in-house "2e-3 formula discrepancy vs canonical FOSS" — **REFUTED as diagnosed; requires re-diagnosis before pinning any oracle gate.** findings/evolution/audits/2026-07-29-probe-hardening-loop/SOTA-RESEARCH-CASCADE-DECIDABILITY.md |
| ξ is a randomized statistic when the x-values contain ties, and its run-to-run spread dwarfs the margin that decided a candidate's verdict. MEASURED 200 XICOR calls at n=2,000 with 30.9% x point mass: sd 0.007965, spread 0.0432 = 54× the 0.0008 decision margin; with continuous x the spread was exactly 0 | > "real xicor() over 200 calls: min=0.210049 max=0.253249 sd=0.007965 SPREAD=0.043200 -> spread is 54x the reported 0.0008 margin; 1 sd = 0.0080 = 10.0 margins" findings/evolution/audits/2026-07-29-probe-hardening-loop/SOTA-RESEARCH-CASCADE-DECIDABILITY.md |
| That tie-break jitter shrinks roughly as 1/√n, so at the cascade's real sample size it sits below the 0.006 pass moat but is still several times the 0.0008 margin that decided one candidate. MEASURED rate-valued kernel at n=3,990, 400 reps: sd 0.004358 = 0.73× the 0.006 moat and 5.45× the 0.0008 margin; fitted scaling exponents −0.465 / −0.493 / −0.521 at tie fractions 0.1 / 0.309 / 0.5 (r² 0.984–0.997) | "sd_over_moat": 0.7262990943157506, findings/evolution/audits/2026-07-29-probe-hardening-loop/harness/xi_tie_jitter_scaling_evidence.json |
| Averaging ξ over R independent tie-rearrangements removes the randomness at a measurable compute price, giving a budget rather than a guess. MEASURED sd falls from 0.004840 (R=1) to 0.000407 (R=128), ratio to the 1/√R prediction 0.90–0.99 over 200 outer reps at n=3,990; R=37 to resolve the 0.0008 margin at 1 SE, R=147 for 2× headroom | "meaning": "resolve the 0.0008 margin at 1 SE", findings/evolution/audits/2026-07-29-probe-hardening-loop/harness/xi_derandomisation_evidence.json |
| The zero-tolerance unanimity rule refuses candidates for being near the line rather than for being unstable: at a true ξ of 0.45 it rejected 72% of candidates whose worst-cell reading was below the 0.50 line. MEASURED 60 candidates × 62 cells × n=3,990 per level: false-veto 0% at true ξ ≤ 0.40, 71.7% (43 of 60) at ξ=0.45 with worst-cell mean 0.4740 and 1.62 unstable cells; required tolerance k rises 0 → ≈2 → ≈38 of 62 as ξ goes 0.40 → 0.45 → 0.48 | "false_veto_rate": 0.7166666666666667, findings/evolution/audits/2026-07-29-probe-hardening-loop/harness/unanimity_false_veto_evidence.json |
| A worst-of-N statistic drifts upward as more cells are added, so adding coverage silently tightens the gate — quadrupling the grid consumes the whole pass moat. MEASURED 30 reps, cell n=2,000: displacement at 62 cells 0.0307–0.0347; going 62→248 cells adds +0.0053 to +0.0098 against a 0.006 moat; fitted slope vs σ·√(2 ln N) 0.964–1.079 (r² 0.966–0.997); 0.9 cell overlap damps but does not remove it | "delta_62_248": 0.009797351138853938 findings/evolution/audits/2026-07-29-probe-hardening-loop/harness/worst_of_n_displacement_evidence.json |
| The legacy probe keeps price levels in its comparison panel and the rotation probe drops them — and neither choice is right: a pure function of the closing price is banned by one and passed by the other. MEASURED n=490 rows, 10 candidates: rolling_mean_close_LEVEL reads |ρ|=1.0 (BAN, worst partner vwap) on the legacy panel and 0.1716 (PASS) on the A1 panel; an independent-but-trending series reads 0.7764 legacy vs 0.3571 A1, an inflation of +0.42; 1 of 10 verdicts flips | "candidate": "rolling_mean_close_LEVEL", findings/evolution/audits/2026-07-29-probe-hardening-loop/harness/legacy_price_level_panel_evidence.json |
| Trade-ID columns enter the head-to-head only through the rotation probe's wider column universe, and their presence flips one candidate's verdict — it is a declared treatment, not a bug, but the legacy arm ran a superset panel. MEASURED 6 columns present in legal_all but not in legacy's own panel, including first_agg_trade_id and last_agg_trade_id; 1 of 10 candidates flips PASS→WATCH (independent_but_trending 0.7764 → 0.9340, worst partner first_agg_trade_id) | "in_legal_all_not_legacy_own": [ findings/evolution/audits/2026-07-29-probe-hardening-loop/harness/legacy_trade_id_panel_evidence.json |
| Both probes now carry a shuffled-null negative control with a positive control, and both bands sit well above their measured noise floors. MEASURED rotation: 60/60 reps pass, null p95 0.0342 against the 0.50 line, positive control BAN at ξ_worst 0.9815 over 62 cells; legacy: 400/400 reps pass, null p95 0.1106 against the 0.85 WATCH line (7.7× headroom, 7 panel columns), positive control BAN at |ρ| 0.9925 | "worst_rho": 0.9925298128881603 findings/evolution/audits/2026-07-29-probe-hardening-loop/harness/structural_gates_evidence.json |
| The probe's volatility-regime labels are computed with lookahead — a bar's regime is assigned against terciles of volatility that had not happened yet — so every re-run silently relabels a few percent of historical rows. MEASURED n_bars=20,000, 30 reps, WIN=200/STRIDE=20: 2.25% of past value rows relabelled on a +5% append, 2.47% at +10%, 5.24% at +25% and +50%; stress arm 65.74% mean, 81.43% max; session and trend labels 0.00% changed, ξ sub-slice reading bit-identical | "vol_lab: 65.74% of PAST value rows relabelled when only FUTURE bars changed (max 81.43%)" findings/evolution/audits/2026-07-29-probe-hardening-loop/harness/future_perturbation_invariance_evidence.json |
| ξ's attainable maximum is length-dependent, and that only bites the BAN line: a quarter-slice cannot reach 0.95 until far above the declared minimum bar count. MEASURED formula (n−2)/(n+1) reproduced exactly at 6 n values (max abs diff 1.1e-16); PASS 0.50 reachable by a quarter from 20 rows = 400 bar-equivalents; BAN 0.95 needs 236 rows = 4,720 bar-equivalents versus the declared 1,000-bar floor; full-minus-quarter ceiling gap 0.0023 at n=3,990 but 0.172 at the 50-row floor | "min_bar_equiv": 4720, findings/evolution/audits/2026-07-29-probe-hardening-loop/harness/attainable_ceiling_evidence.json |
| The contested sub-slice question resolves in the middle: standardising a half/third/quarter by its own length does recover comparability, but the standardising constant is not universal — it fails badly for coarse discrete data. MEASURED 2,000 reps at cell n=3,990: √k scaling holds in all 4 arms; the asymptotic τ²=2/5 collapse holds for continuous (z_sd 0.974–0.998) and 30.9%-tied Y (0.991–1.009) but fails for 5-level Y (1.069–1.092) and binary Y (1.580–1.605) | "asymptotic_collapse": false, findings/evolution/audits/2026-07-29-probe-hardening-loop/harness/subslice_null_pivotality_evidence.json |
| The plug-in variance estimator that the proposed repair route depends on is calibrated in practice, with one soft miss. MEASURED 1,000 reps × 4 regimes × n ∈ {997, 3990}: reported/empirical sd ratios 0.976–1.057, p-values uniform in every cell (p<0.05 rate 0.040–0.057); single recorded failure is tied31 at n=3,990 with ratio 1.057 | "tied31_3990: reported/empirical sd = 1.057" findings/evolution/audits/2026-07-29-probe-hardening-loop/harness/tau_hat_calibration_evidence.json |
| A permutation reference for the worst-cell maximum has usable size and detects dependence the fixed 0.50 line never sees, but the observed size ran above nominal. MEASURED 62 cells, cell n=500, B=99, α=0.05: observed size 0.0917 vs nominal 0.05 (se 0.0199) over 120 null reps; at dependence c=0.45 the permutation test rejects 100% of 40 reps while the fixed line rejects 0% (mean xi_worst only 0.1637); first_fixed_90 is null — the fixed line never reached 90% power at any level tested | "observed": 0.09166666666666666, findings/evolution/audits/2026-07-29-probe-hardening-loop/harness/permutation_reference_for_max_evidence.json |
| Verdicts are invariant to row order, and the PR #674 fetch reordering moved no number. CONFIRMED n=3,990: content-addressed tie-break invariant at every tie rate tested (0, 0.1, 0.309, 0.95, diff 0.0) while the positional tie-break moved by up to 0.0053; PR #674 reorder before 0.3538974852462653, after 0.3538974852462653, abs_diff 0.0; cell verdict identical | "pr674_reorder": { findings/evolution/audits/2026-07-29-probe-hardening-loop/harness/row_order_invariance_evidence.json |
| The research report underpinning the loop records its own limits: 63 explicit gaps, several load-bearing and unverified, including the citation that would justify deleting the sub-slice rule. OPEN 63 gaps listed in section 9; 5 flagged load-bearing and unverified; the fixed-b refutation rests on one second-hand sentence about the sample mean (Lahiri 2001, not retrieved) and ξ is not a sample mean | **Before implementing:** items 950, 951, 954, 962 and 992 in section 9 are flagged as load-bearing findings/evolution/audits/2026-07-29-probe-hardening-loop/SOTA-RESEARCH-CASCADE-DECIDABILITY.md |
| The report is explicitly not a declaration and changed no number. ASSERTED produced by a 15-agent retrieval+synthesis workflow, 11 primary-source topics, ~1.67M subagent tokens across two runs | **Status:** research artifact. NOT a declaration, NOT pre-registration, no numbers changed. findings/evolution/audits/2026-07-29-probe-hardening-loop/SOTA-RESEARCH-CASCADE-DECIDABILITY.md |
Source of record: findings/evolution/audits/2026-07-29-probe-hardening-loop/ — not published, so these are listed rather than linked.
| File | Role |
|---|---|
PRE-REGISTRATION-DRAFT-regime-tercile-lookahead.md | Non-binding draft remedy (4 options) for the lookahead in volatility-regime terciles, from iteration 8; explicitly not in force. |
PRE-REGISTRATION-DRAFT-xi-stability-guard.md | Non-binding draft remedy (4 options) for the unanimity guard's measured false-veto defect, from iteration 4; explicitly not in force. |
SOTA-RESEARCH-CASCADE-DECIDABILITY.md | The loop's research SSoT: 11-source retrieval, 14 premise refutations, 6 recorded inter-section disagreements, FOSS inventory, and 63 explicit unretrieved/unconfirmed gaps. |
findings/dashboard/campaigns/2026-07-29-probe-hardening-loop
findings/dashboard/build_audits.py from findings/evolution/audits/2026-07-29-probe-hardening-loop/AUDIT_LEDGER.json — never hand-edited. Each quote was verified to occur in the file named beside it when the ledger was written.