โ€บNavigation

โ† Audits

Chatterjee ฮพ threshold calibration

SETTLED2026-06-30

Asked why the line that decides whether two market measurements are effectively the same thing sat at 0.50, measured 5,229 column pairs across 766 real data slices, and replaced the guessed number with two evidence-anchored cut lines โ€” plus a written ban on the averaging rule that had been hiding redundancy.

Why it carries this status

Lifecycle, not result. This says where the audit sits in its process โ€” never whether what it found was good.

**Status: DECLARED โ€” ยงA median statistic (2026-06-30) + ยงB worst-cell statistic (2026-07-02).**
findings/evolution/audits/2026-06-30-chatterjee-xi-threshold-calibration/verdict.md

2026-08-17. Derived from folder evidence, then adversarially challenged; the challenge pass was upheld.

The verdict

The audit's own conclusion, reproduced in full from the source below. Not a summary โ€” this is the document, rendered. Links inside it that point at unpublished files are shown as plain text rather than as links that would 404 here.

VERDICT โ€” Chatterjee ฮพ thresholds DECLARED on both statistics (2026-06-30 / 2026-07-02)

Status: DECLARED โ€” ยงA median statistic (2026-06-30) + ยงB worst-cell statistic (2026-07-02). Full declaration + evidence: CHATTERJEE-THRESHOLD-DECLARATION.md. Process + reusable recipe: 07-DECLARATION-PROCESS-AND-PRECEDENT.md.

ยงB โ€” the promotion gate's bands (worst-cell statistic; USE THESE for candidate promotion)

  • PASS: xi_worst < 0.50 โ€” cluster-anchored just above the certified-orthogonal ceiling (0.494), 0.43 moat to the nearest known-related pair (0.923). 3,355 pairs (65.1%).
  • WATCH: everything else โ€” the moat + smooth middle; document + re-probe. 1,568 (30.4%).
  • BAN: xi_worst > 0.95 AND breadth โ‰ฅ 0.80 โ€” two-dial, data-forced (raw worst-cell separation 0.0001); breadth cut at the empty-valley edge. Captures all 50 median-BANs + 180 broad-redundancy pairs the median diluted. 230 (4.5%).

Band and statistic travel together: ยงA binds only to median readings, ยงB only to worst-cell readings. The uniform re-run (G1) + consistency gate (G2: 0 band changes) verified ยงA stands.

ยงA โ€” median-statistic bands (HISTORICAL RECORD โ€” median rule BANNED for promotion, spoke 08)

> The median rule is BANNED as a candidate-promotion axis (2026-07-02, > 08-MEDIAN-RULE-PROMOTION-BAN.md). Evidence: its BAN line misses 180 broad-redundancy pairs, > contradicts two of Terry's standing ฯ Tier-3 bans, and fails a frozen known-related labeled > pair (price_impact|volume). ยงA below is kept append-only as the calibration record.

The declared bands

  • PASS / orthogonal: ฮพ < 0.30 โ€” the independent mass (82.5% of pairs); strong support.
  • WATCH: 0.30 โ‰ค ฮพ โ‰ค 0.95 โ€” watchlist (document + re-probe; don't auto-ban). Wide on purpose: the ฮพ distribution is a smooth slope with no cliff, so no sharp sub-cut is fabricated here.
  • BAN: ฮพ > 0.95 โ€” functional duplicate/derivative; the trusted hard gate (confirmed dups
  • ฯ cross-validated).

Supersedes the legacy single cut ฮพ > 0.50 (self-declared; this calibration shows it was far too loose โ€” it would absorb clearly non-duplicate pairs like bar_bartels~bar_petrosian 0.70).

Results (evidence)

  • Grounded: crypto 600/620 + forex 166/190 cell-slices โ†’ 5,229 pairs (5,153 feature-only).
  • Feature-only distribution: >0.95 = 50 (1.0%) ยท 0.85โ€“0.95 = 155 ยท 0.70โ€“0.85 = 127 ยท 0.50โ€“0.70 = 117 ยท 0.30โ€“0.50 = 453 ยท <0.30 = 4,251 (82.5%).
  • ฯ cross-validation: Terry's Tier-1 ฯ-duplicates reappear at ฮพ>0.95 โ€” two independent metrics agree.

One-line defense

> "Why BAN at ฮพ>0.95 and PASS at ฮพ<0.30 (WATCH between)? Because the mechanism-confirmed duplicates > sit above 0.95 โ€” cross-validated by Terry's ฯ banlist โ€” while 82.5% of pairs sit below 0.30 as > independent; the 0.30โ€“0.95 slope has no natural cliff, so it is WATCH, not a fabricated cut."

Honest caveats

  • ฮพ โ‰  "% predictable" โ€” it is the degree of functional dependence (0 = independent, 1 = exact function).
  • The 0.30 orthogonal line is mass-anchored, not a cliff; tighten it if the goal shifts from "avoid duplicates" to "demand strong novelty" (a research-goal stance, not a data reading).
  • The first 391 crypto cell-slices pre-date the memory-safe fetch-cap fix (sampled from the full slice vs a bounded window); affects only the densest @100 cells, and ฮพ is window-stable so the tiers are unaffected. A uniform re-run is cheap if strict consistency is later required.

Next

  • Wire the ยงB worst-cell bands into the ฮพ gate / rotation-probe verdict (replacing the legacy 0.50 cut) โ€” a separate, operator-gated change.
  • Optional: uniform re-run for perfect per-cell sampling consistency. DONE 2026-07-02 (spoke 06 G1; G2 verified 0 band changes).
source: findings/evolution/audits/2026-06-30-chatterjee-xi-threshold-calibration/verdict.md

Still owed 4

What it claims, and what backs each claim 20

Every row pairs a claim with the file it came from and the verbatim text in that file. The sources sit above the deploy root, so the quote is embedded and the path is printed as text rather than linked โ€” a link would resolve on a laptop and 404 here.

ClaimEvidence
The ฮพ > 0.50 redundancy cut that was in use came from this project itself, not from the Chatterjee (2021) paper and not from Terry's authorized Spearman/h_norm ban-list.
CONFIRMED
legacy single cut ฮพ > 0.50
Provenance trace (sibling audit `2026-06-26-rotation-probe-meta-evaluation-audit`, `MEDIAN-VS-WORST-REGIME.md`) found the ฮพ 0.50 cut is **our** effect-size declaration โ€” NOT from the Chatterjee (2021) paper and NOT from Terry's `CH_FEATURE_BANLIST` (which only authorizes the Spearman 0.95 / 0.85 and h_norm 0.05 cuts).
findings/evolution/audits/2026-06-30-chatterjee-xi-threshold-calibration/CLAUDE.md
The measurement covered 766 real cell-slices across crypto and forex, producing 5,229 column pairs (5,153 after removing raw price collinearity).
MEASURED
crypto 600/620 + forex 166/190 grounded cell-slices; n = 5,229 pairs (5,153 feature-only)
Grounded: crypto 600/620 + forex 166/190 cell-slices โ†’ 5,229 pairs (5,153 feature-only).
findings/evolution/audits/2026-06-30-chatterjee-xi-threshold-calibration/verdict.md
On the median statistic (ยงA) the declared bands are PASS below 0.30, WATCH 0.30โ€“0.95, BAN above 0.95; 82.5% of pairs sit in PASS.
MEASURED
>0.95 = 50 (1.0%) ยท 0.85โ€“0.95 = 155 ยท 0.70โ€“0.85 = 127 ยท 0.50โ€“0.70 = 117 ยท 0.30โ€“0.50 = 453 ยท <0.30 = 4,251 (82.5%); n = 5,153 feature-only pairs
- **PASS / orthogonal: ฮพ < 0.30** โ€” the independent mass (82.5% of pairs); strong support. - **WATCH: 0.30 โ‰ค ฮพ โ‰ค 0.95** โ€” watchlist (document + re-probe; don't auto-ban).
findings/evolution/audits/2026-06-30-chatterjee-xi-threshold-calibration/verdict.md
On the worst-cell statistic (ยงB) โ€” the axis the promotion gate actually reads โ€” the declared bands are PASS xi_worst < 0.50, BAN xi_worst > 0.95 AND breadth โ‰ฅ 0.80, WATCH everything else.
MEASURED
PASS 3,355 (65.1%) ยท WATCH 1,568 (30.4%) ยท BAN 230 (4.5%); n = 5,153 feature-only pairs
3,355 pairs (65.1%). - **WATCH: everything else** โ€” the moat + smooth middle; document + re-probe. 1,568 (30.4%). - **BAN: `xi_worst > 0.95 AND breadth โ‰ฅ 0.80`** โ€” two-dial, data-forced (raw worst-cell separation 0.0001); breadth cut at the empty-valley edge.
findings/evolution/audits/2026-06-30-chatterjee-xi-threshold-calibration/verdict.md
The 0.50 PASS line is anchored on a cluster boundary in the frozen labelled pairs: the top certified-orthogonal pair reads 0.494, the nearest known-related pair 0.923, with nothing labelled in between.
CONFIRMED
certified-orthogonal ceiling 0.494 ยท known-related floor 0.923 ยท moat width 0.43 ยท independent-mass p75 = 0.469; anchor set n = 3 certified-orthogonal pairs
**cluster boundary** โ€” just above the certified-orthogonal ceiling (top frozen known-orthogonal pair = 0.494); 0.43-wide empty moat to the nearest known-related pair (0.923); corroborated by independent-mass p75 = 0.469
findings/evolution/audits/2026-06-30-chatterjee-xi-threshold-calibration/CHATTERJEE-THRESHOLD-DECLARATION.md
A single-number BAN line does not exist on the worst-cell axis: the duplicate floor and the independent tail are separated by 0.0001, which is why a second dial (breadth) was added.
CONFIRMED
dup floor 0.9995 vs independent tail 0.9994 โ†’ separation 0.0001
**BAN gained a second dial** (`breadth = share of cell-slices with ฮพ โ‰ฅ 0.50`). Data-forced: > raw worst-cell separation between the dup floor (0.9995) and the independent tail (0.9994) > is 0.0001 โ€” a single-number BAN line does not exist on this axis
findings/evolution/audits/2026-06-30-chatterjee-xi-threshold-calibration/06-WORST-CELL-RECALIBRATION.md
The pre-registered expected-shape statement was contradicted by the data, and the contradiction was recorded rather than suppressed.
REFUTED
(the ยง7 expected-shape > statement was contradicted, and per its own terms the contradiction was surfaced, not hidden).
findings/evolution/audits/2026-06-30-chatterjee-xi-threshold-calibration/06-WORST-CELL-RECALIBRATION.md
Re-running all 620 crypto cell-slices on one uniform code path did not move the median bands at all.
CONFIRMED
0 band changes across n = 5,229 pairs; max |ฮ”ฮพ| = 0.0092; mean |ฮ”ฮพ| = 4.9e-05 (evidence/g2_median_consistency.json)
**Consistency gate** (G2): v2 median tiers vs committed tiers โ€” **0 band changes / 5,229 pairs**, max \|ฮ”ฮพ\| 0.0092. The ยงA declaration's "does not move the tiers" claim converted from assertion to verified fact
findings/evolution/audits/2026-06-30-chatterjee-xi-threshold-calibration/07-DECLARATION-PROCESS-AND-PRECEDENT.md
The instrument was sanity-anchored on the new axis before any banding: all five mechanism-confirmed duplicate pairs read at or above 0.9995 and the labelled distinct pairs read low.
CONFIRMED
5/5 dups xi_worst โ‰ฅ 0.9995 (ofi|turnover_imbalance 0.9995, agg_record_count|individual_trade_count 1.0); distinct anchors 0.1381 / 0.4236 / 0.4936, breadth 0/600 each
**Worst-axis sanity** (G3): all 5 mechanism dups read `xi_worst โ‰ฅ 0.9995`, distinct anchors low โ€” instrument trusted on the new axis
findings/evolution/audits/2026-06-30-chatterjee-xi-threshold-calibration/07-DECLARATION-PROCESS-AND-PRECEDENT.md
Porting the median bands onto worst-cell readings would have been destructive โ€” it would have evicted about half the median-PASS pairs and wrongly banned 63.
MEASURED
2,130 of 4,250 median-PASS pairs evicted; 63 false BANs
(spoke 06: reading ยงA on worst-cell ฮพ > would evict 2,130 of 4,250 median-PASS pairs and false-BAN 63).
findings/evolution/audits/2026-06-30-chatterjee-xi-threshold-calibration/CHATTERJEE-THRESHOLD-DECLARATION.md
The worst-cell BAN set is a strict superset of the median BAN set โ€” it captures all 50 median BANs and adds 180 pairs that are redundant in at least 80% of windows.
CONFIRMED
50 captured / 0 missed / +180 added; total BAN 230 of 5,153
**Safety check (superset property):** the ยงB BAN captures **all 50** ยงA BAN pairs (0 missed) and adds **180** genuinely broad-redundancy pairs (bad in โ‰ฅ 80% of windows) that the median statistic diluted below its own BAN line
findings/evolution/audits/2026-06-30-chatterjee-xi-threshold-calibration/CHATTERJEE-THRESHOLD-DECLARATION.md
The median aggregation rule is banned as a candidate-promotion axis because it un-bans two pairs Terry's Spearman ban-list already banned.
CONFIRMED
rel_ask_lifetime_bar_meanโ†”bid_only_rate_bar: ฯ=0.9837 BAN vs ฮพ_med 0.946 WATCH vs worst 0.987 breadth 166/166; rel_bid_lifetime_bar_meanโ†”ask_only_rate_bar: ฯ=0.9836, ฮพ_med 0.945, worst 0.987
The median rule contradicts the standing banlist; the worst-cell rule agrees with it. A promotion gate that un-bans declared bans is disqualified on its face.
findings/evolution/audits/2026-06-30-chatterjee-xi-threshold-calibration/08-MEDIAN-RULE-PROMOTION-BAN.md
The median rule also fails a frozen known-related labelled pair, reading price_impact vs volume as WATCH when it is elevated in every single window.
REFUTED
ฮพ_med 0.880 (WATCH) vs xi_worst 0.998, breadth 600/600 cell-slices
| `price_impact` โ†” `volume` (crypto) โ€” **a frozen KNOWN-RELATED labeled pair (P2, ฯ = 0.9872)** | 0.880 โ†’ WATCH | 0.998 | 600/600 |
findings/evolution/audits/2026-06-30-chatterjee-xi-threshold-calibration/08-MEDIAN-RULE-PROMOTION-BAN.md
An independent metric corroborates the ban line: Terry's Spearman Tier-1 duplicates reappear in the ฮพ > 0.95 cluster.
CONFIRMED
5 Tier-1 / watch pairs cross-checked, ฮพ 0.99โ€“1.00
Terry's ฯ-declared Tier-1 duplicates reappear in the ฮพ > 0.95 cluster โ€” two independent metrics agree on what is a duplicate: - `garman_klass_vol โ‰ก intra_garman_klass_vol` (ฯ banlist Tier-1) โ†’ ฮพ = 1.00
findings/evolution/audits/2026-06-30-chatterjee-xi-threshold-calibration/CHATTERJEE-THRESHOLD-DECLARATION.md
The declared cuts are not knife-edge โ€” sweeping either dial changes the ban count smoothly.
MEASURED
breadth sweep โ†’ 230/217/186/172/159 BANs; ฮพ sweep โ†’ 261/213/186/105/58 BANs
**Sensitivity (reported, cuts not tuned on it):** BAN count under breadth cut 0.80/0.85/0.90/0.95/1.00 = 230/217/186/172/159; under ฮพ cut 0.90/0.93/0.95/0.97/0.99 (breadth โ‰ฅ 0.90 held) = 261/213/186/105/58. Smooth in both directions โ€” no knife-edge.
findings/evolution/audits/2026-06-30-chatterjee-xi-threshold-calibration/CHATTERJEE-THRESHOLD-DECLARATION.md
The 0.95 value inside the BAN rule is openly labelled as a chosen convention rather than a derived number.
ASSERTED
dup cluster โ‰ฅ 0.9995; any cut in 0.95โ€“0.99 clears it
the two-dial BAN structure, the 0.80 valley edge, the 0.50 cluster boundary, and the WATCH residual are **empirically anchored**. The 0.95 value is an **anchored convention** (any cut in 0.95โ€“0.99 clears the dup cluster; 0.95 chosen for ฯ-gate consistency). Nothing here is presented as derived that is chosen.
findings/evolution/audits/2026-06-30-chatterjee-xi-threshold-calibration/CHATTERJEE-THRESHOLD-DECLARATION.md
Pre-flight and worked-example sanity ran before the grid and both passed on real capped-resource runs.
CONFIRMED
sanity on BTCUSDT@250/S08, n=5,036 bars: dup ฮพ=0.9983, distinct ฮพ=0.4415, all_ok true; preflight 9/9 checks ok under 5-core/5 GiB cgroup cap
- `ofi โ‰ก turnover_imbalance` โ†’ expect **ฮพ โ‰ˆ 1.0** (functional dup). - `ofi ~ kyle_lambda_proxy` โ†’ expect **ฮพ low**. If the dup isn't โ‰ˆ1 or the distinct pair isn't low โ‡’ pipeline wrong โ‡’ STOP.
findings/evolution/audits/2026-06-30-chatterjee-xi-threshold-calibration/05-FIRST-RUN-DESIGN.md
The originally planned ROC/AUC/permutation-null calibration was dropped mid-audit on an operator decision in favour of replicating Terry's ban-list method exactly.
CONFIRMED
We are **not** using a labeled-set / ROC / permutation-null calibration โ€” those were dropped 2026-06-30 per operator decision (they were the AI's additions, not Terry's method).
findings/evolution/audits/2026-06-30-chatterjee-xi-threshold-calibration/00-PRE-REGISTRATION.md
The PASS anchor rests on only three certified-orthogonal pairs and is flagged as thin, with a pre-registered rule for re-anchoring it.
OPEN
n = 3 certified-orthogonal anchor pairs
**Reconsider-when (ยงB):** the PASS anchor rests on **three** certified-orthogonal pairs โ€” real but thin. If a future **pre-registered, frozen-before-measurement** label expansion moves the certified-orthogonal ceiling, the 0.50 line re-anchors to the new boundary.
findings/evolution/audits/2026-06-30-chatterjee-xi-threshold-calibration/CHATTERJEE-THRESHOLD-DECLARATION.md
Wiring the newly declared worst-cell bands into the live gate is explicitly not authorized by this audit.
OPEN
## ยง6 โ€” What this spoke does NOT authorize - Wiring ยงB into the live ฮพ gate (`multislice` 0.50 cut) โ€” separate follow-up, operator-gated.
findings/evolution/audits/2026-06-30-chatterjee-xi-threshold-calibration/07-DECLARATION-PROCESS-AND-PRECEDENT.md

The audit folder 12 markdown files

Source of record: findings/evolution/audits/2026-06-30-chatterjee-xi-threshold-calibration/ โ€” not published, so these are listed rather than linked.

FileRole
00-PRE-REGISTRATION.mdFrozen plan before results: Terry's four-step ban-list method transposed onto ฮพ; what was dropped and why.
01-LABELED-PAIR-SET.mdThe frozen labelled pairs (mechanism duplicates, known-related, permutation-null, distinct-family) with a stated reason each.
02-METHOD-SEPARATION-CALIBRATION.mdFive-step method spec for the full-grid ฮพ probe, tiering and operator declaration.
03-CONSISTENCY-VS-SPEARMAN-GATE.mdCoherence check against the blessed Spearman gate; two pre-registered exits โ€” post-run blanks left unfilled.
04-SENSITIVITY-ANALYSIS.mdPlanned flip-curve analysis over cut 0.40โ†’0.60 to test whether the cut is load-bearing โ€” post-run blanks left unfilled.
05-FIRST-RUN-DESIGN.mdGated first-run design: 10-slice grid, cell scope, pre-flight checks, worked-example sanity, full-grid probe.
06-WORST-CELL-RECALIBRATION.mdPre-registered recalibration on the worst-cell statistic: locked statistic, cut rule, gates G0โ€“G5, amendment record; marked COMPLETE.
07-DECLARATION-PROCESS-AND-PRECEDENT.mdThe full process story including the wrong turns, comparison to Terry's ฯ declaration, per-number backing, and a 10-step reusable recipe.
08-MEDIAN-RULE-PROMOTION-BAN.mdBan-list-style entry banning the median aggregation rule as a promotion axis, with E1โ€“E4 evidence and reconsider-when.
CHATTERJEE-THRESHOLD-DECLARATION.mdThe declaration itself: both band tables, anchors per number, worked examples, sensitivity, cross-validation, evidence artifact index.
CLAUDE.mdHub: what the audit is, why, both declared band sets, guardrails, spoke reading order.
verdict.mdPlain-English verdict: bands declared on both statistics, results, caveats, next steps.
Generated by findings/dashboard/build_audits.py from findings/evolution/audits/2026-06-30-chatterjee-xi-threshold-calibration/AUDIT_LEDGER.json โ€” never hand-edited. Each quote was verified to occur in the file named beside it when the ledger was written.