Two rival tools for deciding whether a proposed new market measurement merely repeats something the project already uses were run against seven measurements the project already ships and trusts: the newer tool certified none of them and the older one certified six, each is broken in a place the other is not, and the follow-up re-run that appeared to fix the problem was later found to have been reading a database table in which roughly half the rows were duplicates.
Lifecycle, not result. This says where the audit sits in its process โ never whether what it found was good.
* **Re-running R2/R3.** Compute, and gated on the operator.
2026-08-17. Derived from folder evidence, then adversarially challenged; the challenge pass was upheld.
The audit's own conclusion, reproduced in full from the source below. Not a summary โ this is the document, rendered. Links inside it that point at unpublished files are shown as plain text rather than as links that would 404 here.
Status: HALF-SETTLED. Both probes measured; both found defective; neither usable as-is. Operator decision (2026-07-27): repair BOTH probes. Do not pick a winner.
| Claim | Verdict | Evidence |
|---|---|---|
| The rotational cascade can certify a real production feature | โ NO โ 0/7 shipped PASS features pass, all eliminated at panel_r2 | telemetry/step1b_summary.json |
| The legacy probe can certify a real production feature | โ YES โ 6/7 controls PASS under leave-one-out | telemetry/step2_real_legacy.json |
The panel_r2 PASS band is mis-set | โ
YES โ PASS needs r2_worst<0.50; shipped features sit 0.65โ0.91 | Step 1b + panel_r2_evidence.json |
| Legacy ranks substrate artifacts as redundancy | โ
YES โ 4/27 have a price level or first_agg_trade_id as worst driver | step2_real_legacy.json |
| Real legacy is stricter than the emulated legacy | โ
YES, narrowly โ 2/27 numbers moved, 1/27 verdict flipped (parkinson WATCHโBAN) | step2_real_legacy.json |
| The 7 controls' BAN verdicts were self-correlation artifacts | โ
YES โ self_correlation_driven: true at 7/7 | step2_real_legacy.json |
Each probe holds exactly the half the other is missing:
panel hygiene threshold calibration
ROTATIONAL โ
correct โ PASS unreachable
LEGACY โ broken โ
can certify
This is why the arbitration was never resolvable by running the two probes against each other. It also means the repair is small and well-specified rather than a redesign.
Dial 1 โ A1 hygiene applied to legacy. Exclude the 8 substrate columns from the legacy panel: open, high, low, close, vwap, lookback_vwap_raw, first_agg_trade_id, last_agg_trade_id. A monotone row counter must never be a redundancy signal.
Dial 2 โ re-declare the panel_r2 PASS band. The current r2_worst < 0.50 sits below where the project's own shipped features live (0.65โ0.91). The band must be re-derived from the empirical distribution of shipped production columns, not asserted. Applying the current band to the 66 shipped columns yields 5 PASS / 47 WATCH / 14 BAN โ 92.4% non-PASS โ which is the same defect measured a second way.
Both dials must be pre-registered before any re-run. The band in particular is a free parameter; choosing it after seeing candidate verdicts would invalidate every downstream label.
Repairing the probes fixes redundancy measurement. It does not establish that orthogonality predicts tradeability at all. Legacy is |rho| + h_norm โ structurally blind to multivariate redundancy, which is precisely what panel_r2 was introduced to catch. "Legacy passes everything the project ships" is equally consistent with legacy is correct and legacy is too weak to be a gate.
That question needs an outside referee, and the realness/usefulness battery (19 instruments GROUNDED 2026-07-24, on main) is the only measurement asset in the stack independent of redundancy entirely.
Next: realness arbitration on cand-0021 dhvg_indeg_outdeg_kld and cand-0008 arcsine_occupation โ the only 2 of 20 that nothing found redundant under either probe โ plus the 7 shipped controls as calibration. The decisive output is the cross-tab of probe verdict x realness verdict.
If realness verdicts turn out uncorrelated with both probes, orthogonality should be demoted from a promotion gate to a hygiene tag, with realness x usefulness becoming the actual bar. Instrument #14 (spanning intercept) tests incrementality directly on returns and is arguably a better orthogonality measure than either probe.
axis.orthogonal stays deferred for all 20 evaluated candidates. Writing labels now would bake a measured-defective calibration into a hash-chained append-only SSoT. three_axis_gate.py remains wired to legacy โ now known to rank trade IDs as redundancy โ and must not be treated as authoritative until Dial 1 lands.
Current distribution unchanged: 88 null ยท 12 PASS ยท 5 BAN ยท 4 WATCH ยท 1 PENDING; 20 rows carry orthogonal_eval evidence with the label withheld.
Every row pairs a claim with the file it came from and the verbatim text in that file. The sources sit above the deploy root, so the quote is embedded and the path is printed as text rather than linked โ a link would resolve on a laptop and 404 here.
| Claim | Evidence |
|---|---|
| The audit's own status is half-settled, not concluded: both probes were measured, both found defective, and neither is usable as it stands. ASSERTED | **Status:** HALF-SETTLED. Both probes measured; both found defective; neither usable as-is. findings/evolution/audits/2026-07-24-probe-arbitration/verdict.md |
| The operator decided not to pick a winner between the two probes but to repair both. ASSERTED | **Operator decision (2026-07-27): repair BOTH probes.** Do not pick a winner. findings/evolution/audits/2026-07-24-probe-arbitration/verdict.md |
| The question was posed before any result was seen: the rotational probe had returned zero passes in 22 full evaluations, and nothing run so far could distinguish 'the probe is strict' from 'the probe cannot certify anything'. MEASURED 0 PASS in 22 full-cascade evaluations (batch-01 ร10, batch-02 ร10, iters 12/17 smoke) | The rotational probe has returned **0 PASS in 22 full-cascade evaluations** (batch-01 ร10, findings/evolution/audits/2026-07-24-probe-arbitration/PRE-REGISTRATION.md |
| Run against seven features the project already ships and labels PASS, the rotational cascade certified none of them and the legacy probe certified all seven, with every rotational elimination happening at the same stage. CONFIRMED rotational 0/7 PASS, legacy 7/7 PASS, n=7 shipped controls, 62 grounded crypto cells, 80,000 bars per cell | **rotational PASS 0/7 ยท legacy PASS 7/7 ยท every elimination at `panel_r2`, unanimously.** findings/evolution/audits/2026-07-24-probe-arbitration/RESULTS.md |
| The mechanism is a threshold set below where real production features actually sit: the pass line demanded a score under 0.50 while the shipped features measure 0.65โ0.91. CONFIRMED pass line 0.50 vs shipped-feature r2_worst 0.65โ0.91 (n=7 controls); applying the same line to all 66 shipped columns gives 5 PASS / 47 WATCH / 14 BAN = 92.4% non-PASS | `panel_r2` PASS requires `r2_worst < 0.50`. The project's own shipped features sit at **0.65โ0.91**. findings/evolution/audits/2026-07-24-probe-arbitration/RESULTS.md |
| A first attempt was invalidated by a confound: every shipped control was correlated against its own copy in the panel, so all seven read as redundant purely with themselves. CONFIRMED rho with own twin 0.9868โ0.9996 across 7 controls; after dropping the twin, 6 of 7 PASS under legacy | `self_correlation_driven: true` for **7/7**. A feature cannot be "redundant with the shipped set" findings/evolution/audits/2026-07-24-probe-arbitration/RESULTS.md |
| The legacy probe, run for the first time rather than emulated, was measured to rank price levels and a monotone trade-ID counter as top redundancy signals. CONFIRMED 4 of 27 candidates driven by a price level or first_agg_trade_id; legacy panel 120 columns vs A1-hygiened 112; 8 columns dropped by A1 | Under leave-one-out, **4 of 27** candidates have their worst-correlation driver be a price level or a findings/evolution/audits/2026-07-24-probe-arbitration/RESULTS.md |
| The pre-registered expectation that the real legacy probe would be stricter than the emulated one was confirmed but the effect was small โ two numbers moved and one verdict flipped out of 27. CONFIRMED 2 of 27 rho_delta != 0 (parkinson 0.1289, edge_spread_bps 0.0450); 1 of 27 verdicts flipped (parkinson WATCH โ BAN) | - **2/27** have `rho_delta != 0` โ the price/ID columns actually moved the worst-correlation number findings/evolution/audits/2026-07-24-probe-arbitration/RESULTS.md |
| The structural finding is that each probe holds exactly the half the other lacks โ rotational has correct panel hygiene but an unreachable pass line, legacy can certify but ranks substrate artifacts as redundancy. CONFIRMED | Each probe holds exactly the half the other is missing: findings/evolution/audits/2026-07-24-probe-arbitration/verdict.md |
| A separate durable catch: five of twenty-seven kernels silently returned all-NaN because optional Python dependencies were missing, which is indistinguishable from a legitimate fail-safe. CONFIRMED 5 of 27 kernels all-NaN; a 27/27 finite-value preflight is now mandatory and passed 27/27 at R0 | **2. Silent kernel death from missing dependencies.** 5 of 27 kernels returned all-NaN in the first run findings/evolution/audits/2026-07-24-probe-arbitration/RESULTS.md |
| The repair declaration diagnosed the real defect as an asymmetry: the reject rule used two dials (score and breadth) while the pass rule used only one, so a feature proven not broadly redundant could still be denied on a single worst cell. CONFIRMED max breadth 0.0323 across 7 controls vs a BAN breadth requirement of 0.80; new PASS breadth dial B_pass frozen at 0.20 | Maximum breadth across all seven: **0.0323**. The BAN dial requires **0.80**. findings/evolution/audits/2026-07-24-probe-arbitration/PROBE-REPAIR-DECLARATION.md |
| The new pass line was declared as an output of a pre-committed rule โ the median worst-cell score across the 66 shipped production columns โ with the number deliberately not computed before the rule was frozen. CONFIRMED k=50 frozen; P_50 later computed at 0.925944 over n=66 shipped columns (min 0.077566, p25 0.788568, median 0.925944, p75 0.99874, max 1.0) | > **`k = 50` (the median). FROZEN HERE.** findings/evolution/audits/2026-07-24-probe-arbitration/PROBE-REPAIR-DECLARATION.md |
| The nominal five-core five-gigabyte compute cap bounded only the Python client; the database server executing the queries was unbounded, and one probe cell drove it to 32.97 GiB on a shared production host. MEASURED over 741 completed probe fetches: p50 1.61 GiB, p95 12.47 GiB, max 32.97 GiB; 17 of 63 cells exceeded 5 GiB; max_memory_usage=0 (unlimited), max_threads=auto(32) | | max query memory | **32.97 GiB** (`DOGEUSDT@100`) | findings/evolution/audits/2026-07-24-probe-arbitration/DECLARATION-AMENDMENT-2026-07-27a-RESOURCE-ENVELOPE.md |
| A feasibility sweep showed the binding constraint is thread count, not sorting: two threads survive the worst cell at a hard 5 GiB in 25 seconds while five threads fail outright. MEASURED threads=5 FAIL (Code 241) in 3s; threads=2 PASS at 25s; threads=1 PASS at 46s; worst cell DOGEUSDT@100, 122-column query, 5 GiB ceiling | | **`threads=2`, default block** | **PASS** | **25s** | findings/evolution/audits/2026-07-24-probe-arbitration/DECLARATION-AMENDMENT-2026-07-27a-RESOURCE-ENVELOPE.md |
| The cap-enforcing shim had been installed but was never on the execution path, so it had never capped a single query โ and the values it carried were a guess the feasibility sweep had already disproved. CONFIRMED observed readonly=0, max_memory_usage=0, max_threads=auto(32) vs declared 2 / 5368709120 / 2; shim carried max_threads=5, the exact config the sweep proved fails with Code 241 | | **What the probe actually got** (no shim on PATH) | `0` | `0` (unlimited) | `auto(32)` | findings/evolution/audits/2026-07-24-probe-arbitration/DECLARATION-AMENDMENT-2026-07-28a-ENVELOPE-ENFORCEMENT.md |
| After enforcement, the re-run's compute envelope was verified server-side: peak query memory fell 25-fold and no query escaped the thread cap. CONFIRMED peak query memory 32.97 GiB โ 1.31 GiB; p95 12.47 โ 1.28 GiB; queries over 5 GiB 17 of 63 โ 0; failures โ 0; cells silently skipped โ 0; run 41 min, exit 0 | `groupUniqArray(Settings['max_threads'])` returns exactly `['2']`: **no probe query escaped the cap.** findings/evolution/audits/2026-07-24-probe-arbitration/R2-RESULTS.md |
| Under the repaired dials the re-run's pre-registered rule fired for 'the cascade can certify': two of seven controls passed, with all three gates satisfied including a regression self-test proving comparability with the earlier run. CONFIRMED n_pass 2 of 7; n_grounded 62 for every candidate, cells_short=[]; old band 0/7 vs new band 7/7 at the panel_r2 stage alone | **2 of 7 PASS โ H1.** All three gates fired together; none can be read without the other two. findings/evolution/audits/2026-07-24-probe-arbitration/R2-RESULTS.md |
| The repair moved the binding constraint rather than removing it: the xi stage now binds for five of seven controls, and four of those are pending because the estimator could not return a usable value. CONFIRMED xi binding for 5 of 7; 4 PENDING with unstable-cell counts 32 / 6 / 2 / 1; sibling stage None for all 7 (singleton slates) | `xi` is now binding for 5 of 7. Four of those are **PENDING from unstable cells, not from failing a findings/evolution/audits/2026-07-24-probe-arbitration/R2-RESULTS.md |
| The database fetch was shown not to be deterministic: the sort column is heavily tied, so which 80,000 rows come back is not uniquely determined, and four of seven scores moved between runs on static data. CONFIRMED BTCUSDT@250: 80,000 rows, 39,946 distinct close_time_us, 40,054 ties; run-to-run r2_worst deltas โ0.0516 / +0.0451 / +0.0379 / โ0.0251 with 3 of 7 bit-identical; newest bar predates both runs | **That is false as stated.** It holds only if the query's `ORDER BY` is a total order. It is not. findings/evolution/audits/2026-07-24-probe-arbitration/DECLARATION-AMENDMENT-2026-07-28b-SUBSTRATE-INTEGRITY.md |
| One control's apparent pass was measured to sit inside the noise: it cleared the new line by 0.0117 while its own run-to-run drift is 0.0516. CONFIRMED l_kurtosis_tau4 margin 0.0117 vs drift 0.0516 = 4.4ร; the count of 2 passes is unaffected because it is PENDING at xi anyway | **The measured run-to-run drift for that kernel is 0.0516 โ 4.4ร the margin.** Its `panel_r2` PASS is findings/evolution/audits/2026-07-24-probe-arbitration/DECLARATION-AMENDMENT-2026-07-28b-SUBSTRATE-INTEGRITY.md |
| Every cell audited was about half duplicate rows, so the probe has never read 80,000 distinct bars. CONFIRMED 40 of 40 cells; 0 clean cells; mean duplication 50.4%; range 49.8% (ADAUSDT@500) to 66.7% (DOGEUSDT@100, 3ร copies) | **40 of 40 cells audited: 80 000 rows, ~40 000 distinct bars.** findings/evolution/audits/2026-07-24-probe-arbitration/DECLARATION-AMENDMENT-2026-07-28b-SUBSTRATE-INTEGRITY.md |
| The follow-up amendment withdrew the earlier reasoning that duplication barely mattered: duplicates land adjacent in the sort, so the candidate values themselves are wrong, not just their multiplicity. REFUTED fraction of consecutive close-price increments exactly zero, raw vs deduplicated: BTCUSDT@250 50.97% โ 1.07%; LINKUSDT@100 65.97% โ 30.93%; DOGEUSDT@100 66.67% โ 0.05% | | BTCUSDT@250 | **50.97%** | 1.07% | findings/evolution/audits/2026-07-24-probe-arbitration/DECLARATION-AMENDMENT-2026-07-29-SUBSTRATE-DEDUPLICATION.md |
| Contamination varies wildly by cell, so cross-cell comparability was never established even though the comparability gate reported green. CONFIRMED SOLUSDT@100 49.1% duplicated table-wide but 0% inside the probe window; effective lookback 80,000 real bars for SOLUSDT, ~40,000 for BTCUSDT, 26,667 for DOGEUSDT โ all labelled 80,000 | `SOLUSDT@100` is **49.1% duplicated table-wide** (14,467,206 rows / 7,363,604 bars) yet findings/evolution/audits/2026-07-24-probe-arbitration/DECLARATION-AMENDMENT-2026-07-29-SUBSTRATE-DEDUPLICATION.md |
| The tie-breaking used by the rank-based xi statistic was silently degraded by the duplication, reinstating a defect previously measured at +0.244 systematic null bias, on the stage that now binds. CONFIRMED previously measured +0.244 systematic null bias; post-dedup close_time_us unique again at 80,000/80,000 on BTCUSDT@250, LINKUSDT@100 and DOGEUSDT@100 | they hashed **identically**, `np.lexsort` fell back to positional order, and that reinstated findings/evolution/audits/2026-07-24-probe-arbitration/DECLARATION-AMENDMENT-2026-07-29-SUBSTRATE-DEDUPLICATION.md |
| The forex path's identity column was also found mildly non-unique, for a different reason that deduplication cannot fix. MEASURED EURUSD@10: 79,696 distinct close_time_us across 80,000 rows; XAUUSD@25: 79,703 โ ~0.4% collision rate | `fxview_cache.forex_bars` returns **79,696 distinct `close_time_us` across 80,000 rows** on findings/evolution/audits/2026-07-24-probe-arbitration/DECLARATION-AMENDMENT-2026-07-29-SUBSTRATE-DEDUPLICATION.md |
| Because of the duplication, every measurement this campaign took โ including the two-of-seven pass result โ is marked superseded rather than edited. CONFIRMED 5 measurement sets superseded (step1b, step2b, batch-01, batch-02, R2) | step1b, step2b, batch-01, batch-02 and **R2 (including `2/7 PASS โ H1`)** were all measured on findings/evolution/audits/2026-07-24-probe-arbitration/DECLARATION-AMENDMENT-2026-07-29-SUBSTRATE-DEDUPLICATION.md |
| The deduplication fix was measured to cost almost nothing and to halve the legacy probe's row count exactly, confirming it had been equally affected. MEASURED worst cell DOGEUSDT@100 (38.75M rows): 26.65 s with FINAL vs 25.86 s without, +3%; legacy BTCUSDT@250 slice01 15,028 rows โ 7,514, exactly 2ร | legacy slice01 `BTCUSDT@250` 7,514 / 7,514 / 7,514 โ **against 15,028 / 7,514 before the findings/evolution/audits/2026-07-24-probe-arbitration/DECLARATION-AMENDMENT-2026-07-29-SUBSTRATE-DEDUPLICATION.md |
| No registry labels were written at any point: the orthogonal axis stays deferred for all 20 evaluated candidates and the distribution is unchanged. CONFIRMED 88 null ยท 12 PASS ยท 5 BAN ยท 4 WATCH ยท 1 PENDING โ unchanged throughout | **Untouched, as declared.** `axis.orthogonal` stays deferred for all 20 evaluated candidates. findings/evolution/audits/2026-07-24-probe-arbitration/R2-RESULTS.md |
| The gate script the project wires to the legacy probe cannot express the repaired rules and is declared not authoritative. CONFIRMED gate input schema reads only 3 fields (worst_abs_rho, max_r2, min_h_norm); the xi and sibling stages have no representation at all | Until then, **`three_axis_gate.py` output is not authoritative for `axis.orthogonal`.** findings/evolution/audits/2026-07-24-probe-arbitration/PROBE-REPAIR-DECLARATION.md |
| A third stage โ sibling leave-one-out R-squared โ was found being graded on the wrong scale and was suspended from verdict authority rather than re-banded on a guess. CONFIRMED measured exposure: 0 of 10 batch-02 verdicts change; 1 batch-01 verdict flips (roll_spread BAN โ WATCH, pairwise rho 0.6042) | loo_worst := computed, recorded, REPORTED โ but NOT graded findings/evolution/audits/2026-07-24-probe-arbitration/PROBE-REPAIR-DECLARATION.md |
| The audit is explicit that repairing redundancy measurement says nothing about whether orthogonality predicts tradeability, and that settling it needs an outside referee. OPEN realness/usefulness battery = 19 instruments grounded 2026-07-24; cohort for the cross-tab = 20 evaluated candidates + 7 controls | Repairing the probes fixes *redundancy measurement*. It does not establish that **orthogonality predicts findings/evolution/audits/2026-07-24-probe-arbitration/verdict.md |
| A separate SOTA research document was produced as a non-binding artifact that changes no number and lists 63 explicit unconfirmed gaps. ASSERTED 15-agent retrieval workflow, 11 primary-source topics, ~1.67M subagent tokens across two runs; 63 gaps listed in section 9 | **Status:** research artifact. NOT a declaration, NOT pre-registration, no numbers changed. findings/evolution/audits/2026-07-24-probe-arbitration/SOTA-RESEARCH-2026-07-29-CASCADE-DECIDABILITY.md |
| That research refutes one of the campaign's own premises: a candidate said to sit 0.0008 from the threshold has no stable side of the line, because the statistic is randomised when the input has ties. REFUTED 200 calls, n=2000, 30.9% point mass in x: spread 0.043200 = 54ร the 0.0008 margin, sd 0.007965 = 10 margins; with continuous x the spread was exactly 0 | > "real xicor() over 200 calls: min=0.210049 max=0.253249 sd=0.007965 SPREAD=0.043200 -> spread is 54x the reported 0.0008 margin; 1 sd = 0.0080 = 10.0 margins" findings/evolution/audits/2026-07-24-probe-arbitration/SOTA-RESEARCH-2026-07-29-CASCADE-DECIDABILITY.md |
Source of record: findings/evolution/audits/2026-07-24-probe-arbitration/ โ not published, so these are listed rather than linked.
| File | Role |
|---|---|
DECLARATION-AMENDMENT-2026-07-27a-RESOURCE-ENVELOPE.md | Amends the re-run protocol after two silent-failure defects: the cap bounded only the client, and a failed cell was silently skipped. Records R0 pass, R1 P_50, R2 abort. |
DECLARATION-AMENDMENT-2026-07-28a-ENVELOPE-ENFORCEMENT.md | Enforcement-only amendment โ the shim was installed but inert, its verifier bypassed it, and a duplicated cap would have silently skipped cells; declares a single envelope source plus a fail-closed preflight. |
DECLARATION-AMENDMENT-2026-07-28b-SUBSTRATE-INTEGRITY.md | Records D6 (the fetch is not deterministic; withdraws a frozen claim) and D7 (every cell ~50% duplicated); suspends the xi leg from label authority and blocks R3. |
DECLARATION-AMENDMENT-2026-07-29-SUBSTRATE-DEDUPLICATION.md | Withdraws 28b's out-of-scope reasoning and fixes the read (FINAL, total ORDER BY, per-substrate identity, fail-closed dedup gate); marks step1b/step2b/batch-01/batch-02/R2 superseded. |
PRE-REGISTRATION.md | Written before any result: the H1-vs-H2 question, the two negative-control arms, the decision rule, the real-legacy head-to-head, and the honest limits. |
PROBE-REPAIR-DECLARATION.md | Frozen repair declaration โ three dials (legacy panel hygiene, a two-dial derived PASS band, suspending the sibling LOO-Rยฒ leg), the gate schema gap, and the R0โR5 re-run protocol. |
R2-RESULTS.md | The repaired negative-control re-run โ 2 of 7 PASS firing for H1, the new binding stage, the verified compute envelope, and the two substrate findings that qualify the result. |
RESULTS.md | The step-1b negative control and first-ever real-legacy execution โ 0/7 rotational vs 7/7 legacy, the self-correlation confound, and two durable methodological catches. |
SOTA-RESEARCH-2026-07-29-CASCADE-DECIDABILITY.md | Non-binding literature synthesis on making the cascade decidable โ a ranked recommendation table, premise refutations, per-decision sections, a FOSS inventory, and 63 unconfirmed gaps. |
verdict.md | Root verdict โ HALF-SETTLED; the six grounded claims table, the structural finding, the two declared repair dials, what is still open, and the registry consequences. |
findings/dashboard/build_audits.py from findings/evolution/audits/2026-07-24-probe-arbitration/AUDIT_LEDGER.json โ never hand-edited. Each quote was verified to occur in the file named beside it when the ledger was written.