โ€บNavigation

โ† Audits

Orthogonal-probe registry sweep

HALTED2026-07-22

A batch-by-batch effort to check whether each of 110 proposed new measurements actually adds anything the shipped set does not already contain: 20 have been checked, none were certified as genuinely new, the newer test earned its place mainly by catching candidates that duplicate each other rather than the existing columns, and the second batch's results are recorded but deliberately not written into the registry until it is settled which of two competing tests counts.

Why it carries this status

Lifecycle, not result. This says where the audit sits in its process โ€” never whether what it found was good.

| **Probe arbitration (legacy vs rotational)** | โŒ **UNRESOLVED** โ€” blocks labelling; `three_axis_gate.py` still wired to **legacy** |
findings/evolution/audits/2026-07-22-orthogonal-probe-registry-sweep/CLAUDE.md

2026-08-17. Derived from folder evidence, then adversarially challenged; the challenge pass was upheld.

Blocked on An unresolved operator/probe arbitration: which test โ€” the legacy correlation banlist or the newer rotational cascade โ€” governs the orthogonal axis. Until that lands, axis.orthogonal stays null for batch-02's 10 candidates, the live promotion gate stays wired to the legacy probe, and batch-03 is not started (the hub also states each new batch needs a separate operator go before running on bigblack).

The verdict

The audit's own conclusion, reproduced in full from the source below. Not a summary โ€” this is the document, rendered. Links inside it that point at unpublished files are shown as plain text rather than as links that would 404 here.

batches/batch-01-10candidate-head-to-head/verdict.md

verdict โ€” 10-candidate head-to-head

Does the rotational orthogonality probe earn its keep over the legacy banlist? YES โ€” but for a specific reason, and with none of the 10 promotable.

Bottom line

Across all 10 candidates, scored fairly (legacy = |ฯ| and h_norm, its true banlist):

  • 0 PASS โ€” nothing is certified orthogonal.
  • 4 WATCH โ€” l_skewness_tau3, l_cv_tau2, arcsine_argmax, longest_excursion (document + re-probe, do not auto-ban).
  • 5 BAN โ€” edge_spread_bps, roll_spread, corwin_schultz, parkinson (mutual duplicates) + abdi_ranaldo (degenerate).
  • 1 PENDING โ€” arcsine_occupation (the probe could not certify; fail-safe).

Why it earns its keep

The four effective-spread / range-vol estimators are duplicates of one another โ€” promote all four and you ship three redundant columns. The legacy probe cannot see this: it only compares each candidate to the already-shipped panel, never candidate-to-candidate. The rotational probe's sibling stage catches it. That is genuine, valuable orthogonality work legacy structurally lacks.

Where it does not add value (honest)

Redundancy against the shipped panel (the panel-Rยฒ two-dial gate) is WATCH for every candidate โ€” breadth < 0.80 everywhere (lone-cell spikes), so nothing is broadly reconstructable from existing columns. On that axis the rotational probe matches legacy. Its edge here is mutual-duplicate detection (sibling), not shipped-panel redundancy.

The fairness correction that mattered

Scoring legacy as |ฯ|-only (the harness's original bug) made abdi_ranaldo look like a rotational-only BAN. With the fair legacy (|ฯ| OR h_norm โ€” Terry's actual banlist), legacy BANs it too (h_normโ‰ˆ0). The correction removed the one spurious "rotational wins" case; the surviving 7 divergences are real.

Promotion decision

None of the 10 is promotable now (0 PASS on the orthogonal axis; parameterless + agnostic axes not yet evaluated). The spread/vol family should be de-duplicated (keep at most one, if any); the two arcsine WATCHes + the two L-moment WATCHes warrant re-probe; arcsine_occupation needs a cleaner substrate to resolve its PENDING.

Next steps (open)

  • Optional: run the same slate on forex (--asset forex, single-source probe post mql5#149) โ€” panel-Rยฒ would need a forex calibration first.
  • Optional: bidask-R byte-parity for the 3 spread reimpls (blocked here by the sandbox R install; the kernels are validated by ground-truth spread recovery + canonical formula โ€” see FOSS-CROSSCHECK.md).
  • The rotational probe is trustworthy as a crypto orthogonality grader; this run is its first full 10-candidate exercise and confirms the sibling stage's practical value.
source: findings/evolution/audits/2026-07-22-orthogonal-probe-registry-sweep/batches/batch-01-10candidate-head-to-head/verdict.md

batches/batch-02-10candidate-head-to-head/verdict.md

VERDICT โ€” batch-02, 10-candidate rotational-vs-legacy head-to-head

Date run: 2026-07-23 (bigblack) ยท Date landed: 2026-07-24 ยท Operator: nasimubd Substrate: 62/62 grounded crypto cells, 80k bars/cell, read-only ClickHouse Candidates: cand-0012, 0014, 0017, 0018, 0021, 0022, 0023, 0026, 0027, 0028 (10 of the 13 curated=RUNNABLE rows from the coverage reconciliation)


Result

probePASSWATCHBANPENDING
**Legacy (fair: \ฯ\and h_norm)**505โ€”
Rotational (5-stage cascade)0271

Divergences: 5/10.

Bottom line

The sibling stage earned its keep on this batch, and did so cleanly. Batch-02 surfaced two genuine mutual-duplicate pairs that the legacy banlist structurally cannot see:

  • psd_wiener_spectral_flatness โ†” psd_spectral_centroid_normfreq โ€” ฯ = 0.9783
  • guzik_increment_magnitude_asymmetry โ†” porta_increment_sign_asymmetry โ€” ฯ = 0.9954

Both pairs are theoretically near-identical by construction. Legacy rated both PSD candidates PASS, because neither is redundant against the shipped panel โ€” only against each other. That is exactly the blind spot the candidate-vs-candidate check exists to close.

Critically, and unlike batch-01, these BANs do not depend on the contested LOO-Rยฒ leg: both pairs clear 0.95 on the pairwise |ฯ| leg alone. A counterfactual re-fusion dropping the LOO-Rยฒ leg entirely changes 0 of 10 verdicts (batch-01, by contrast, would flip roll_spread BAN โ†’ WATCH). Batch-02's verdicts therefore rest on the operator-frozen |ฯ| line only, and are evidentially stronger.

What is NOT settled by this batch

axis.orthogonal is deliberately left null for all 10. The evaluation is recorded in full; the label is deferred. Reasons, all pre-existing and none introduced by this batch:

  1. The rotational cascade has never certified anything. 0 PASS in 22 full-cascade evaluations (batch-01 10, batch-02 10, plus iters 12/17 smoke runs). Its PASS path has never been exercised on real data โ€” noted in the campaign's own iter-17 falsifier verdict as F3 honestly CONDITIONAL-UNTESTABLE.
  2. The panel-Rยฒ PASS band is contradicted by its own calibration anchors. In 2026-07-16-panel-r2-threshold-calibration/panel_r2_evidence.json, two of the four features designated orthogonal anchors โ€” kyle_lambda_proxy and trade_intensity โ€” score r2_worst = 1.0000, numerically indistinguishable from the redundant anchors (volume, ofi). Applying the declared band to the 66 already-shipped production columns yields 5 PASS ยท 47 WATCH ยท 14 BAN โ€” 92.4% non-PASS.
  3. The ฮพ stage quarantines heavily. PENDING on 10 of 20 candidates across both batches; it is the sole reason dhvg_indeg_outdeg_kld (this batch) and arcsine_occupation (batch-01) are not PASS, despite both clearing every substantive gate including panel-Rยฒ.
  4. No pre-registered decision rule exists for choosing between the two probes. The three-axis promotion gate (three_axis_gate.py) is still wired to the legacy probe; the rotational probe is not referenced in it. The campaign's own loop prompt forbids self-promotion (Do NOT declare operational).

Writing 7 BANs into the registry as authoritative would bake a contested calibration into the SSoT โ€” a hash-chained, append-only structure. The evidence is preserved in full and the label is a one-line set once the probe arbitration lands. Operator decision, 2026-07-24.

The two candidates worth carrying forward

Across both batches, exactly two of 20 clear every substantive stage (h_norm, sibling, ฯ, panel-Rยฒ) and fail only on the ฮพ fail-safe:

candidatebatchpanel-Rยฒ worstฮพ unstable cellslegacyrotational
cand-0021 dhvg_indeg_outdeg_kld020.4752 (PASS)3/62PASSPENDING
cand-0008 arcsine_occupation010.1671 (PASS)โ€”PASSPENDING

Neither was found redundant by anything. They are the natural first inputs to the realness/usefulness battery (19 instruments, GROUNDED 2026-07-24) โ€” the only measurement asset in the stack that can arbitrate the probe question from outside either probe.

Next

  1. Negative control โ€” run the rotational cascade on promoted/shipped features (l_kurtosis_tau4, petrosian_fd, bartels_rank_vn_ratio, โ€ฆ) to establish whether PASS is reachable at all.
  2. Re-band or drop the sibling LOO-Rยฒ leg โ€” a two-dial declaration mirroring the 2026-07-16 panel-Rยฒ fix. No effect on batch-02; flips roll_spread in batch-01.
  3. Realness-battery arbitration on cand-0021 + cand-0008.
  4. Then set axis.orthogonal for all 20 in one pass, under whichever probe wins.

Provenance

Ran on bigblack 2026-07-23 00:05โ†’00:14 PDT; artifacts left untracked in a detached-HEAD TEMP worktree and rescued verbatim on 2026-07-24. Nothing recomputed in transit. Full per-stage detail: RESULTS.md ยท reproduction: REPRODUCE.md ยท raw: telemetry/.

source: findings/evolution/audits/2026-07-22-orthogonal-probe-registry-sweep/batches/batch-02-10candidate-head-to-head/verdict.md

Still owed 10

What it claims, and what backs each claim 21

Every row pairs a claim with the file it came from and the verbatim text in that file. The sources sit above the deploy root, so the quote is embedded and the path is printed as text rather than linked โ€” a link would resolve on a laptop and 404 here.

ClaimEvidence
This is an append-only campaign to evaluate the whole candidate registry one batch at a time, and the work is predominantly writing the measurement code rather than running it.
ASSERTED
110 candidates against ~84 existing kernels; batch-01 hand-authored 8 of its 10 kernels
**Why a campaign, not one audit**: 110 candidates, ~84 kernels โ€” evaluating them all is predominantly a **kernel-authoring** effort (batch-01 had to hand-author 8 of its 10 kernels).
findings/evolution/audits/2026-07-22-orthogonal-probe-registry-sweep/CLAUDE.md
Batch-01 evaluated 10 candidates and certified none; the newer probe and the legacy banlist disagreed on 7 of 10.
MEASURED
rotational 0 PASS ยท 4 WATCH ยท 5 BAN ยท 1 PENDING versus fair legacy 4 PASS ยท 5 WATCH ยท 1 BAN; 7 of 10 divergent; 62 of 62 crypto cells grounded at 80,000 bars per cell
**Divergences: 8/10 (ฯ-only) โ†’ 7/10 (fair).** The h_norm fairness fix removed `abdi_ranaldo`.
findings/evolution/audits/2026-07-22-orthogonal-probe-registry-sweep/batches/batch-01-10candidate-head-to-head/RESULTS.md
Batch-01's bans are mostly mutual duplicates among four spread and range-volatility estimators, which the legacy check structurally cannot see because it only compares candidates to the already-shipped panel.
CONFIRMED
4 of 5 bans driven by the candidate-versus-candidate stage (edge_spread_bps, roll_spread, corwin_schultz, parkinson); the fifth, abdi_ranaldo, is degenerate with h_norm_min โ‰ˆ 0
The **legacy probe cannot see this**: it only compares each candidate to the *already-shipped* panel, never candidate-to-candidate.
findings/evolution/audits/2026-07-22-orthogonal-probe-registry-sweep/batches/batch-01-10candidate-head-to-head/verdict.md
Against the shipped panel, none of batch-01's candidates was broadly redundant โ€” that axis produced only middle-band verdicts, so the newer probe's advantage is duplicate detection, not shipped-panel redundancy.
MEASURED
breadth โ‰ค 0.242 across all 10 batch-01 candidates, most at 0.016โ€“0.032
Redundancy **against the shipped panel** (the panel-Rยฒ two-dial gate) is **WATCH for every candidate** โ€” breadth < 0.80 everywhere (lone-cell spikes), so nothing is broadly reconstructable from existing columns.
findings/evolution/audits/2026-07-22-orthogonal-probe-registry-sweep/batches/batch-01-10candidate-head-to-head/verdict.md
Batch-02 evaluated a further 10 candidates and also certified none; the two probes disagreed on 5 of 10.
MEASURED
rotational 0 PASS ยท 2 WATCH ยท 7 BAN ยท 1 PENDING versus fair legacy 5 PASS ยท 0 WATCH ยท 5 BAN; 5 of 10 divergent; 62 of 62 cells at 80,000 bars per cell in about 9 minutes on bigblack
**Tally: 0 PASS ยท 2 WATCH ยท 7 BAN ยท 1 PENDING.**
findings/evolution/audits/2026-07-22-orthogonal-probe-registry-sweep/batches/batch-02-10candidate-head-to-head/RESULTS.md
Batch-02 found two pairs of candidates that are near-copies of each other, both of which the legacy check rated as fine because neither duplicates the shipped columns.
MEASURED
pairwise correlations 0.9783 and 0.9954, both above the frozen 0.95 ban line; both spectral candidates were rated PASS by the legacy check
- `psd_wiener_spectral_flatness` โ†” `psd_spectral_centroid_normfreq` โ€” **ฯ = 0.9783** - `guzik_increment_magnitude_asymmetry` โ†” `porta_increment_sign_asymmetry` โ€” **ฯ = 0.9954**
findings/evolution/audits/2026-07-22-orthogonal-probe-registry-sweep/batches/batch-02-10candidate-head-to-head/verdict.md
Batch-02's verdicts do not depend on the contested part of the duplicate check โ€” removing that leg entirely changes none of the ten, whereas batch-01 would lose one ban.
CONFIRMED
0 of 10 batch-02 verdicts change under the counterfactual; batch-01's roll_spread would flip from BAN to WATCH (its pairwise correlation 0.6042 sits inside the pass band)
**0 of 10 verdicts change.** Batch-02 is fully robust to the LOO-Rยฒ defect
findings/evolution/audits/2026-07-22-orthogonal-probe-registry-sweep/batches/batch-02-10candidate-head-to-head/RESULTS.md
The rotational cascade has never issued a pass on real data across every run to date, so it is unknown whether its pass path is reachable at all.
MEASURED
0 passes in 22 full-cascade evaluations (10 + 10 + 2 smoke iterations)
- **0 PASS in 22 full-cascade evaluations** (batch-01 ร—10, batch-02 ร—10, iters 12/17 smoke). The cascade's PASS path has never been exercised on real data.
findings/evolution/audits/2026-07-22-orthogonal-probe-registry-sweep/CLAUDE.md
The declared redundancy band the cascade relies on is contradicted by its own calibration references, and applying it to the already-shipped columns would fail 92.4% of them.
CONFIRMED
5 PASS ยท 47 WATCH ยท 14 BAN of 66 shipped columns = 92.4% non-pass; recomputed independently from panel_r2_evidence.json during this extraction with identical results
Applying the declared band to the **66 already-shipped production columns** gives 5 PASS ยท 47 WATCH ยท 14 BAN = **92.4% non-PASS**.
findings/evolution/audits/2026-07-22-orthogonal-probe-registry-sweep/CLAUDE.md
The final stage of the cascade declines to certify very often, which is the sole reason the two best candidates are not passes.
MEASURED
PENDING on 10 of 20 candidates across both batches; per-candidate unstable cells range from 0 to 27 of 62; moors_octile_kurtosis 24 of 59
`moors_octile_kurtosis` is the extreme: **24 of 59** grounded cells quarantined as unstable.
findings/evolution/audits/2026-07-22-orthogonal-probe-registry-sweep/batches/batch-02-10candidate-head-to-head/RESULTS.md
Exactly two candidates out of 20 cleared every substantive stage and were held back only by the fail-safe stage โ€” nothing found them redundant.
MEASURED
cand-0021 panel-Rยฒ worst 0.4752 with 3 of 62 unstable cells; cand-0008 arcsine_occupation panel-Rยฒ worst 0.1671 โ€” 2 of 20 candidates
| **cand-0021** | 02 | `dhvg_indeg_outdeg_kld` | **0.4752 (PASS)** | PASS | PENDING |
findings/evolution/audits/2026-07-22-orthogonal-probe-registry-sweep/CLAUDE.md
The registry labels for batch-02 were deliberately withheld by operator decision rather than written as authoritative, because doing so would bake a contested calibration into an append-only hash-chained record.
ASSERTED
10 of 10 batch-02 candidates carry a full evaluation record with a null axis label; batch-01's 10 were labelled
**`axis.orthogonal` is deliberately left `null` for all 10.**
findings/evolution/audits/2026-07-22-orthogonal-probe-registry-sweep/batches/batch-02-10candidate-head-to-head/verdict.md
No rule exists for deciding between the two probes, and the live promotion gate still uses the older one.
OPEN
**No pre-registered decision rule exists** for choosing between the two probes. The three-axis promotion gate (`three_axis_gate.py`) is still wired to the **legacy** probe; the rotational probe is not referenced in it.
findings/evolution/audits/2026-07-22-orthogonal-probe-registry-sweep/batches/batch-02-10candidate-head-to-head/verdict.md
A read-only reconciliation established how much of the remaining backlog is even runnable: most of it needs measurement code written first.
MEASURED
of 88 remaining candidates: 13 runnable now, 9 reuse-confirm, 66 need a kernel across 25 families; 84 registry kernels parsed with 50 unclaimed; the 13 is stated as a lower bound
| **NEEDS a kernel** | **66** | no registry kernel โ€” author one (+ FOSS cross-check), the batch-01 pattern |
findings/evolution/audits/2026-07-22-orthogonal-probe-registry-sweep/COVERAGE.md
Overall progress is under a third of the registry, and the previously-promoted candidates rest on weaker evidence than the newly-evaluated ones.
ASSERTED
32 of 110 candidates touched (12 promoted plus 20 evaluated); 78 unevaluated
Note the 12 promoted and the 20 evaluated are **not on the same evidential footing** โ€” the 12 carry no `orthogonal_eval` record, method string, or probe attribution.
findings/evolution/audits/2026-07-22-orthogonal-probe-registry-sweep/CLAUDE.md
Batch-01's ten measurement kernels were checked against state-of-the-art open-source implementations, with exact or near-exact agreement wherever a Python reference existed.
CONFIRMED
10 kernels checked; l_cv_tau2 agrees with lmoments3 to about 1e-16; three tsfresh comparisons exact; parkinson difference about 0
| `l_cv_tau2` | 0003 | **lmoments3** 1.0.8 | L2/L1 on \|r\| vs mine | **MATCH, \|ฮ”\| โ‰ˆ 1e-16** (bit-exact) |
findings/evolution/audits/2026-07-22-orthogonal-probe-registry-sweep/batches/batch-01-10candidate-head-to-head/FOSS-CROSSCHECK.md
The three spread estimators with no Python reference were validated by recovering a known injected spread rather than by trusting the formula, and reproduce the published behaviour including one estimator's known downward bias.
CONFIRMED
injected true spread 0.10%; recovered Roll 0.094%, Abdi-Ranaldo 0.071%, Corwin-Schultz 0.031%
| Roll | **0.094%** | ~exact |
findings/evolution/audits/2026-07-22-orthogonal-probe-registry-sweep/batches/batch-01-10candidate-head-to-head/FOSS-CROSSCHECK.md
Duplicate findings are a property of which candidates were submitted together โ€” re-running any subset will not reproduce them.
ASSERTED
each candidate is regressed on the other 9 in the same slate of 10
> **Slate-dependence warning.** The `sibling` stage regresses each candidate on the *other nine > candidates in the same slate*.
findings/evolution/audits/2026-07-22-orthogonal-probe-registry-sweep/batches/batch-02-10candidate-head-to-head/REPRODUCE.md
Batch-02's results were nearly lost โ€” they were computed in a temporary untracked working directory and rescued unchanged a day later, then re-verified purely from the committed telemetry.
CONFIRMED
run 2026-07-23 00:05โ†’00:14 PDT and landed 2026-07-24; re-deriving the verdict fusion from the recorded stage values reproduces 10 of 10
artifacts left **untracked** in a detached-HEAD TEMP worktree and rescued verbatim on 2026-07-24. Nothing recomputed in transit.
findings/evolution/audits/2026-07-22-orthogonal-probe-registry-sweep/batches/batch-02-10candidate-head-to-head/verdict.md
Two batch-02 kernels produced no usable values on some slices, so one candidate is graded on fewer slices than the rest.
MEASURED
2 kernels ร— 2 cells produced no finite values; moors is graded on 59 rather than 62 cells
- **Kernel failures:** `bowley_quantile_skew` and `moors_octile_kurtosis` each returned all-NaN on 2 cells (`AVAXUSDT@100`, `LINKUSDT@100`)
findings/evolution/audits/2026-07-22-orthogonal-probe-registry-sweep/batches/batch-02-10candidate-head-to-head/RESULTS.md
All measurement ran under explicit resource caps against a read-only database with no writes.
CONFIRMED
5 CPU cores, 5 GB memory cap (ulimit -v 5242880), readonly=2, 62 cells ร— 80,000 bars
Run: 62/62 crypto cells grounded, 80k bars/cell, read-only ClickHouse, bigblack, `nasimubd`,
findings/evolution/audits/2026-07-22-orthogonal-probe-registry-sweep/batches/batch-01-10candidate-head-to-head/RESULTS.md

The audit folder 9 markdown files

Source of record: findings/evolution/audits/2026-07-22-orthogonal-probe-registry-sweep/ โ€” not published, so these are listed rather than linked.

FileRole
CLAUDE.mdCampaign hub: running status table, why batch-02's labels are deferred, folder layout, batch log with combined tallies, the two candidates worth carrying forward, coverage plan, provenance chain and how to add a batch.
COVERAGE.mdKernel-coverage reconciliation for the 88 unevaluated candidates: runnable / reuse-confirm / needs-a-kernel buckets, the 25-family breakdown, the honest lower-bound caveat and the proposed batch ordering.
batches/batch-01-10candidate-head-to-head/FOSS-CROSSCHECK.mdCorrectness check of batch-01's 10 kernels against state-of-the-art open-source implementations, including ground-truth spread recovery for the three with no Python reference, and what was not done.
batches/batch-01-10candidate-head-to-head/REPRODUCE.mdBatch-01 reproduction recipe: prerequisites, throwaway venv, worktrees, the capped read-only run command, fair-legacy post-processing, cleanup and a determinism note.
batches/batch-01-10candidate-head-to-head/RESULTS.mdBatch-01 per-stage detail: the five-stage cascade verdict for each of 10 candidates, the legacy-versus-rotational comparison table and a reading of what drove each divergence.
batches/batch-01-10candidate-head-to-head/verdict.mdBatch-01 verdict: whether the rotational probe earns its keep, where it does and does not add value, the legacy fairness correction, the promotion decision and open next steps.
batches/batch-02-10candidate-head-to-head/REPRODUCE.mdBatch-02 reproduction record: exact host, time, working directory, venv, resource caps, re-run commands, the slate-dependence warning, the two verifications performed on the landed artifacts, and the registry updater's deliberate no-label behaviour.
batches/batch-02-10candidate-head-to-head/RESULTS.mdBatch-02 per-stage detail: candidate-to-kernel mapping, the 10-row cascade table, the legacy comparison, the leave-one-out counterfactual, driver reading and carried caveats.
batches/batch-02-10candidate-head-to-head/verdict.mdBatch-02 verdict: results table, the two genuine duplicate pairs, the four reasons the registry labels were withheld, the two candidates to carry forward, next actions and provenance.
Generated by findings/dashboard/build_audits.py from findings/evolution/audits/2026-07-22-orthogonal-probe-registry-sweep/AUDIT_LEDGER.json โ€” never hand-edited. Each quote was verified to occur in the file named beside it when the ledger was written.