This gave a redundancy test its own cut-off for the first time: by measuring how well each of 66 existing bar columns can be rebuilt from the others across 62 market slices, the team replaced a threshold borrowed from a different statistic with one anchored to reference features whose answer was known in advance โ and wrote the whole plan down before computing a single number.
Lifecycle, not result. This says where the audit sits in its process โ never whether what it found was good.
**Status**: **DECLARED (G5) โ operator-directed 2026-07-16, landed in PR #639 (`c1d5ca4b`).** Bands
2026-08-17. Derived from folder evidence, then adversarially challenged; the challenge pass was upheld.
The audit's own conclusion, reproduced in full from the source below. Not a summary โ this is the document, rendered. Links inside it that point at unpublished files are shown as plain text rather than as links that would 404 here.
Redundancy bands for the panel-Rยฒ gate of the rotational orthogonality probe, declared via Terry's banlist method (compute metric vs the shipped panel -> evidence -> tiers/anchors -> operator declares) on the worst-cell statistic with the A2 breadth dial โ structurally the ฮพ ยงB two-dial gate. Replaces the borrowed Spearman line (BAN_HI=0.95 / WATCH_LO=0.85) the panel-Rยฒ stage currently reuses. Evidence: panel_r2_evidence.json (G1, 62 cells, 66 feature columns, post-A1 hygiene).
> Band + statistic travel together. These bands bind ONLY to r2_worst (max over grounded cells > of leave-one-column-out raw OLS panel-Rยฒ, on the hygiene-cleaned panel) + the breadth dial. Do NOT > port them onto any other statistic.
| Band | Rule (worst-cell + breadth) | Meaning |
|---|---|---|
| PASS / orthogonal | r2_worst < 0.50 | the panel cannot rebuild the candidate even in its worst cell โ genuinely new information |
| BAN / redundant | r2_worst > 0.95 AND breadth >= 0.80 | near-perfectly reconstructable from a panel combination in >=80% of cells โ a broad dup/derivative |
| WATCH | everything else | partially reconstructable; document + re-probe, do NOT auto-ban |
where breadth = share of a candidate's grounded cells with per-cell panel-Rยฒ >= 0.90 (VIF >= 10).
| Number | Status | Anchor (from G1 evidence) |
|---|---|---|
PASS < 0.50 | anchored (cluster boundary) | clean-orthogonal ceiling on worst-cell = 0.38 (bar_cecp_velocity), permutation-null 0.07; empty moat 0.38 -> 0.64 to the next feature (intra_bull_cv); 0.50 sits in the moat (5 cols PASS). Terry cluster-anchoring. |
BAN peak > 0.95 | anchored convention | the redundant cluster sits at ~1.0; 0.95 = near-perfect linear reconstruction (VIF > 20). The 0.95 value mirrors the ฯ/ฮพ gate โ a labeled convention, not a derived cliff. |
BAN breadth >= 0.80 at the 0.90 level | anchored | 0.90 = VIF >= 10 (standard high-multicollinearity flag = the ฯ<->Rยฒ bridge). 0.80 breadth (ฮพ valley-edge convention) separates true dups (breadth ~1.0) from partial/intra-twin redundancy (breadth 0.62 -> WATCH). |
| WATCH residual | backed | the distribution is a smooth descent (largest gap 0.101); any internal cliff would be invented. |
| Anchor | r2_worst | breadth | band | correct? |
|---|---|---|---|---|
| volume / ofi / individual_trade_count (redundant) | 1.0 | ~1.0 | BAN | yes โ genuine dup / linear-combo |
| intra_ twins (intra_kyle_lambda, intra_ofi, ...) | 1.0 | 1.0 | BAN | yes โ de-dups the twin families |
| kyle_lambda_proxy / trade_intensity | 1.0 | 0.62 | WATCH | yes โ intra-twin redundancy -> WATCH, not false-BAN |
| bar_cecp_velocity (orthogonal) | 0.38 | 0 | PASS | yes |
| permutation-null floor | 0.07 | 0 | PASS | yes โ independence baseline |
| bar_dispersion_entropy | 0.75 | 0 | WATCH | yes โ mid, document |
< 0.10 (tight cluster-anchor): would split the clean-orthogonal cluster (bar_cecp_velocity 0.38 would fail); anchor n=4 too thin.< 0.85: 0.85 is the ฯ WATCH floor โ wrong scale (|ฯ|=0.85 <-> Rยฒ~0.72).> 0.90: 0.90 is the per-cell VIF>=10 breadth level (the second dial); the peak line is deliberately stricter (near-perfect reconstruction).>= 0.50: would false-BAN kyle_lambda_proxy/trade_intensity (0.62), whose redundancy exists only where their intra-twin ships โ better as WATCH. Not >= 1.0: too strict (individual_trade_count 0.98 escapes on one cell).The breadth level 0.90 = VIF >= 10 = the ฯ<->Rยฒ bridge (ฯ=0.95 <-> Rยฒ~0.90). Doubly anchored.
A median-axis framing (PASS r2_median<0.80 / WATCH 0.80-0.90 / BAN >=0.90, on the breadth-robust median) was evaluated at G4 โ more permissive, VIF-anchored. Because raw r2_worst is inflated (median worst-cell 0.926), median is arguably the more informative axis. The operator chose the worst-cell + breadth axis (this declaration) for fidelity to the pre-registered statistic and the ฮพ ยงB precedent; the breadth dial recovers the discrimination worst-cell alone lacks. The median view is retained in panel_r2_evidence.json as a cross-reference.
band_from_worst_rho(r2_worst) call with this two-dial rule is a separate follow-up, operator-gated.Every row pairs a claim with the file it came from and the verbatim text in that file. The sources sit above the deploy root, so the quote is embedded and the path is printed as text rather than linked โ a link would resolve on a laptop and 404 here.
| Claim | Evidence |
|---|---|
| The redundancy gate had no threshold of its own โ it was reusing constants calibrated for a different statistic. ASSERTED borrowed constants BAN_HI = 0.95 and WATCH_LO = 0.85, frozen 2026-04-27 for absolute Spearman correlation | **Terry never declared a panel-Rยฒ threshold** โ `CH_FEATURE_BANLIST.md` authorizes only ฯ (0.95/0.85) and h_norm (0.05). findings/evolution/audits/2026-07-16-panel-r2-threshold-calibration/CLAUDE.md |
| The method was frozen before any result existed: the statistic, the anchor set and the cut rule were committed with zero measured numbers in the document. CONFIRMED 6 ordered gates G0โG5; the pre-registration landed as PR #638 and the declaration as PR #639 | **This document locks the plan; it contains ZERO panel-Rยฒ numbers. The threshold value is declared last (G5), by the operator, from the evidence.** findings/evolution/audits/2026-07-16-panel-r2-threshold-calibration/00-PRE-REGISTRATION.md |
| The measurement covered 66 feature columns across the 62-cell crypto grid, after the hygiene amendment removed price levels and identifier columns. MEASURED 66 columns ร 62-cell grid; 80,000 bars fetched per cell, window 200, stride 20, n = 3,990 value rows per cell | Evidence: `panel_r2_evidence.json` (G1, 62 cells, 66 feature columns, post-A1 hygiene). findings/evolution/audits/2026-07-16-panel-r2-threshold-calibration/PANEL-R2-THRESHOLD-DECLARATION.md |
| The run actually completed 61 of the 62 cells โ one cell failed on a server-side memory limit and is absent from every per-column count. MEASURED 61 of 62 cells measured (every column in panel_r2_evidence.json reports n_cells = 61); the skip was ClickHouse Code 241, 35.11 GiB requested against a 40.00 GiB ceiling | skip DOGEUSDT@100: clickhouse-client failed: Received exception from server (version 25.12.1): findings/evolution/audits/2026-07-16-panel-r2-threshold-calibration/panel_r2_run.log |
| The operator declared a two-dial band: a candidate passes below 0.50, and is banned only if it exceeds 0.95 in its worst slice and is broadly reconstructable across slices. ASSERTED PASS below 0.50; BAN above 0.95 with breadth โฅ 0.80; breadth = share of grounded cells at per-cell Rยฒ โฅ 0.90 | | **BAN / redundant** | `r2_worst > 0.95` AND breadth >= 0.80 | near-perfectly reconstructable from a panel combination in >=80% of cells โ a broad dup/derivative | findings/evolution/audits/2026-07-16-panel-r2-threshold-calibration/PANEL-R2-THRESHOLD-DECLARATION.md |
| The three deliberately-redundant reference features behaved exactly as pre-registered, landing at the top of the scale in essentially every slice. CONFIRMED volume, ofi and individual_trade_count: worst-cell 1.0 with breadth โ 1.0 (60โ61 of 61 cells at Rยฒ โฅ 0.90) | | volume / ofi / individual_trade_count (redundant) | 1.0 | ~1.0 | **BAN** | yes โ genuine dup / linear-combo | findings/evolution/audits/2026-07-16-panel-r2-threshold-calibration/PANEL-R2-THRESHOLD-DECLARATION.md |
| The pre-registered expectation that the independent reference features would read low was contradicted for two of the four โ they read at the very top of the scale, indistinguishable from the redundant ones on the worst-slice axis alone. REFUTED kyle_lambda_proxy and trade_intensity: worst-cell 1.0 with breadth 0.62 (38 of 61 cells at Rยฒ โฅ 0.90); only the breadth dial separates them from the redundant anchors | | kyle_lambda_proxy / trade_intensity | 1.0 | 0.62 | **WATCH** | yes โ intra-twin redundancy -> WATCH, not false-BAN | findings/evolution/audits/2026-07-16-panel-r2-threshold-calibration/PANEL-R2-THRESHOLD-DECLARATION.md |
| The pass line at 0.50 sits inside an empty gap in the measured distribution rather than cutting through a cluster. MEASURED clean-orthogonal ceiling 0.38; permutation-null floor 0.07; empty interval 0.38 โ 0.64; 5 of 66 columns fall below 0.50 | clean-orthogonal ceiling on worst-cell = 0.38 (`bar_cecp_velocity`), permutation-null 0.07; empty moat 0.38 -> 0.64 to the next feature (`intra_bull_cv`); 0.50 sits in the moat (5 cols PASS). findings/evolution/audits/2026-07-16-panel-r2-threshold-calibration/PANEL-R2-THRESHOLD-DECLARATION.md |
| A three-cell smoke run exposed that the candidate universe was admitting price levels and identifier columns, which read as near-perfectly redundant through trivial arithmetic identities; they were excluded by pre-registered amendment before the full measurement. CONFIRMED 6 price-level columns plus all identifier and timestamp columns excluded; counts and durations retained as genuine features | **Amendment:** the panel AND candidate universe EXCLUDE {open,high,low,close,vwap,lookback_vwap_raw} + any column ending `_id`/`_time_us` or containing `trade_id`. findings/evolution/audits/2026-07-16-panel-r2-threshold-calibration/00-PRE-REGISTRATION.md |
| The same smoke run showed the worst-slice statistic can be driven to its maximum by a single slice, which is why a second breadth dial was added to the cut rule before any threshold was declared. CONFIRMED kyle_lambda_proxy 1.0 worst versus 0.0072 second-worst; trade_intensity 1.0 versus 0.1059; single offending cell AAVEUSDT@500 | labeled-orthogonal anchors `kyle_lambda_proxy` (`r2_worst=1.0`, `r2_2nd=0.0072`) and `trade_intensity` (`r2_worst=1.0`, `r2_2nd=0.1059`) spike to 1.0 in one cell (`AAVEUSDT@500`), near 0 elsewhere. findings/evolution/audits/2026-07-16-panel-r2-threshold-calibration/00-PRE-REGISTRATION.md |
| The breadth level was cross-checked against a standard multicollinearity rule of thumb rather than being chosen freely. ASSERTED per-cell Rยฒ โฅ 0.90 corresponds to VIF โฅ 10; correlation 0.95 corresponds to Rยฒ โ 0.90 | The breadth level 0.90 = **VIF >= 10** = the ฯ<->Rยฒ bridge (ฯ=0.95 <-> Rยฒ~0.90). Doubly anchored. findings/evolution/audits/2026-07-16-panel-r2-threshold-calibration/PANEL-R2-THRESHOLD-DECLARATION.md |
| No internal middle-band cut was invented, because the measured distribution declines smoothly with no cliff. MEASURED largest gap in the 66-column distribution = 0.101 | the distribution is a smooth descent (largest gap 0.101); any internal cliff would be invented. findings/evolution/audits/2026-07-16-panel-r2-threshold-calibration/PANEL-R2-THRESHOLD-DECLARATION.md |
| A more permissive alternative framing built on the median rather than the worst slice was evaluated and recorded, and the operator chose the worst-slice axis knowing that statistic is inflated. MEASURED median worst-cell across the 66 columns = 0.926 | Because raw `r2_worst` is inflated (median worst-cell 0.926), median is arguably the more informative axis. findings/evolution/audits/2026-07-16-panel-r2-threshold-calibration/PANEL-R2-THRESHOLD-DECLARATION.md |
| Declaring the band deliberately did not change the running gate; wiring it in was left as a separate operator-gated step. ASSERTED | - Does NOT wire the band into the live gate โ replacing the borrowed `band_from_worst_rho(r2_worst)` call with this two-dial rule is a separate follow-up, operator-gated. findings/evolution/audits/2026-07-16-panel-r2-threshold-calibration/PANEL-R2-THRESHOLD-DECLARATION.md |
Source of record: findings/evolution/audits/2026-07-16-panel-r2-threshold-calibration/ โ not published, so these are listed rather than linked.
| File | Role |
|---|---|
00-PRE-REGISTRATION.md | The firewall spoke: locked statistic, frozen anchor set, cut rule, sanity rail, measurement plan, the G0โG5 gate sequence, a falsifiable expected-shape statement, and Amendment A (panel hygiene plus the lone-spike breadth dial). |
CLAUDE.md | Hub: what the calibration is, why the borrowed line was wrong, the method in one line, the DECLARED (G5) status with its wiring follow-up, spoke reading order, guardrails and provenance. |
PANEL-R2-THRESHOLD-DECLARATION.md | The terminal artifact standing in for verdict.md: declared bands, the anchor behind each number, anchor-validation table, why other lines were rejected, the recorded median-axis alternative, reconsider-when triggers and scope limits. |
findings/dashboard/build_audits.py from findings/evolution/audits/2026-07-16-panel-r2-threshold-calibration/AUDIT_LEDGER.json โ never hand-edited. Each quote was verified to occur in the file named beside it when the ledger was written.