Every research audit in this repository: where it sits in its lifecycle, the evidence that state was derived from, and โ on each audit's own page โ every claim it makes beside the file that backs it. Statuses are lifecycle, the same vocabulary the campaigns hub uses. They say nothing about whether an audit's findings were good; a result is reported as numbers with its uncertainty.
Not one unfinished audit is running. Every audit that has not reached a terminal state is stopped on a named human decision, so each of those rows says who must decide.
A stop rule fired, an operator stopped it, or it was abandoned mid-flight. Each row names what would unblock it.
| Audit | What it turned up |
|---|---|
| Orthogonal features for BTCUSD โ candidate set with two-sided provenance 2026-05-08 ยท 21 claims ยท 51 files ยท verdict: verdict.md findings/evolution/audits/2026-05-08-orthogonal-features-btcusd/ Blocked on Terry โ the folder issues an explicit stop-order pending his accuracy review of FRAMEWORK.md: "Do not initiate per-candidate Level-0-through-Level-8 analysis until Terry confirms the framework." Phase 1 evidence closed; the per-candidate analysis it exists to enable has never started. | A catalogue of 64 published "complexity" measurements that could become new columns of Bitcoin bar data โ each tied to a specific paper and a specific open-source implementation, and each checked against what the project already ships โ followed by four measurement runs on real exchange data that kept three of them (Petrosian, Katz, dispersion entropy), shelved several on a house rule against tunable knobs, and left 41 of 65 candidates untested. |
| Orthogonality measurement instruments โ SOTA catalog 2026-05-16 ยท 18 claims ยท 38 files ยท verdict: verdict.md findings/evolution/audits/2026-05-16-orthogonality-measurement-instruments/ Blocked on Terry โ "Deployment is BLOCKED awaiting Terry's review." The Phase-1 catalogue of ~95 instruments is complete, but six Phase-2 spike folders have since appeared under candidates/ that the folder's own verdict says are out of scope, so the verdict no longer describes the folder. | A survey of about 95 published ways to measure whether one market feature is really telling you the same thing as another, each checked for whether usable open-source code exists โ followed by test runs on real Bitcoin data showing that a cheap rank-based method detects genuine curved relationships in 118 of 171 feature pairs that the project's current correlation check calls independent. |
| Forward-looking orthogonality prediction โ methodology catalogue with stable-ID provenance 2026-05-26 ยท 22 claims ยท 23 files ยท verdict: verdict.md findings/evolution/audits/2026-05-26-forward-orthogonality-prediction/ Blocked on Terry โ the ADR-2026-07-15 live declaration engine is PROPOSED and needs his refine/approve before any build; separately the B-05 change-detector family is GATED on a Frontier-#4 review that only the operator can lift, which blocks the automatic regime-turn trigger. | An attempt to answer whether a feature that looks independent today will still look independent tomorrow: across 384 slices of six years of crypto and currency history it found that about 94โ95% of today's independent features stay independent into the next market regime, that today's redundancy margin ranks the ones about to flip with 0.88โ0.95 accuracy, and that not one of 92 features tested is independent in every regime โ so every such statement has to name the market conditions it applies to. |
| Rotation probe meta-evaluation audit 2026-06-26 ยท 20 claims ยท 13 files ยท verdict: verdict.md findings/evolution/audits/2026-06-26-rotation-probe-meta-evaluation-audit/ Blocked on Operator โ the no-synthetic-data conversion has one unfinished row and it is gated on a decision: "The self-test is removed; the rest need operator-directed conversion to frozen-real." No firing in the folder since late June. | This investigation asked whether the tool that judges candidate measurements is anywhere secretly grading itself, found that it is not and that its two main cut-off numbers were written down months earlier and never moved, but also found that the safety machinery meant to keep those judgements trustworthy had been written and never switched on โ and that the forex and crypto copies of the tool were quietly using different arithmetic. |
| Matrix admission paradox loop โ standing admission gate 2026-07-02 ยท 23 claims ยท 17 files ยท verdict: verdict.md findings/evolution/audits/2026-07-02-matrix-admission-paradox-loop/ Blocked on Operator (nasimubd / Nasim). The decision queue in HANDOVER-PROMPT.md must be steered before the loop resumes: the Frontier-#4 review that would unGATE batch B-05 and the ADR's L3 sentinel, ratification of ADR-2026-07-14 (4 open questions) and its hardened successor ADR-2026-07-15 (6 open questions), the tri-tier BAN/WATCH/ORTHOGONAL ratification (LEDGER row 105), the 73 cycle-1 draft statuses, and the PR #599 merge. | Built and sealed a standing lab that decides, one candidate at a time, whether a new statistical instrument earns a place in the trusted measurement stack; across about a hundred sittings it admitted 16 instruments, rejected 45, and then stopped and handed a list of decisions back to the operator. |
| Matrix harvest and status-grounding loop 2026-07-04 ยท 16 claims ยท 9 files ยท verdict: verdict.md findings/evolution/audits/2026-07-04-matrix-harvest-status-grounding-loop/ Blocked on The operator (nasimubd). The 20-minute /loop cron was cancelled mid-campaign by an explicit operator instruction; LEDGER row 38 records the resume path ('re-launch per LOOP-PROMPT.md'). H0 ยง7 additionally leaves the operator the choice to resume, convert to slow standing-refill, or close the campaign. A secondary blocker is the firecrawl host on littleblack being off the tailnet, which parks 5 NO-IMPL revisits. | An automated literature-sweeping loop ran 32 times to collect and paper-screen published statistical tests that could tell whether a data feature stays independent from one market period to the next; it reached 108 vetted candidates against a planned 300 and handed 68 of them to the sister evaluation lab before the operator told it to stop. |
| Feature realness and usefulness โ census, gap map and instrument grounding 2026-07-07 ยท 24 claims ยท 38 files ยท verdict: verdict.md findings/evolution/audits/2026-07-07-feature-realness-usefulness/ Blocked on Operator โ the instrument-building phase finished; the feature campaign the audit exists to run has not started. Its own plan defers it: "Only then the feature campaign: HYPOTHESIS-REGISTRY.md (F0) for the 30 promoted columns ... Separate, later PR." | A sweep of three codebases found the team could prove a feature was computed correctly and was not a duplicate of another feature, but had nothing checking whether it predicted anything or fired often enough to be worth a slot โ so the audit then built and stress-tested twenty statistical instruments on real Bitcoin bar data, keeping nineteen and discarding one. |
| Crypto cost-realism labels investigation and ADR pre-lock 2026-07-08 ยท 22 claims ยท 23 files ยท verdict: verdict.md findings/evolution/audits/2026-07-08-crypto-cost-realism-labels/ Blocked on Terry Li โ the folder's declared blocker is his sign-off on the four open questions before the ADR locks (OPEN-QUESTIONS-FOR-TERRY.md Q1โQ4), and PILOT-BTCUSDT.md Step 4 adds a second gate: no extension beyond BTCUSDT without his explicit go. Note the folder was never updated after Terry's 2026-07-14 directive; the dashboard twin records the block as subsequently cleared, but no file inside this folder says so. | Eleven read-only investigation passes worked out how to record, for every crypto bar, the worst price a real order would actually have filled at in the seconds after the bar closed โ landing on a separate database table filled in afterwards from the stored tick history โ and left four design questions for the supervisor before any code could be locked in. |
| Orthogonal-probe registry sweep 2026-07-22 ยท 21 claims ยท 9 files ยท verdict: none findings/evolution/audits/2026-07-22-orthogonal-probe-registry-sweep/ Blocked on An unresolved operator/probe arbitration: which test โ the legacy correlation banlist or the newer rotational cascade โ governs the orthogonal axis. Until that lands, axis.orthogonal stays null for batch-02's 10 candidates, the live promotion gate stays wired to the legacy probe, and batch-03 is not started (the hub also states each new batch needs a separate operator go before running on bigblack). | A batch-by-batch effort to check whether each of 110 proposed new measurements actually adds anything the shipped set does not already contain: 20 have been checked, none were certified as genuinely new, the newer test earned its place mainly by catching candidates that duplicate each other rather than the existing columns, and the second batch's results are recorded but deliberately not written into the registry until it is settled which of two competing tests counts. |
| Regime-invariant orthogonality โ campaign hub 2026-07-22 ยท 29 claims ยท 27 files ยท verdict: verdict.md findings/evolution/audits/2026-07-22-regime-invariant-orthogonality-campaign/ Blocked on Operator (nasimubd). Three things are named as awaiting a decision: ยง8 of CAMPAIGN-STATUS-NOT-OPERATIONAL.md (repair the harness vs. stop and run a single throwaway one-candidate pilot), A17 (replace the positive-control anchor that measures 0.887 rank / 0.960 Pearson with its own target), and A19 (whether the block length should be pre-certified at all). Resuming additionally requires lifting the stop banner in LOOP-PROMPT.md. D45 (appended 2026-08-14) separately needs an operator ruling plus an amendment. The hub's own verdict.md also asks the operator to confirm the alignment statement before the campaign is made live. | This umbrella folder renames the project's whole feature-selection direction โ a feature counts as genuinely new information only if it keeps its meaning as market conditions change, not merely because it looks unrelated today โ and its one live piece of new work, a 120-candidate hunt for liquidity and fragility measures, was stopped by the operator with none of the 120 judged, because eighteen attempts to run the judging machinery kept turning up problems in the machinery itself. |
| Probe arbitration 2026-07-24 ยท 33 claims ยท 10 files ยท verdict: verdict.md findings/evolution/audits/2026-07-24-probe-arbitration/ Blocked on Operator (nasimubd). Re-running R2 and R3 on the deduplicated substrate is compute work explicitly gated on the operator; R3 is additionally blocked because it makes the suspended xi and sibling stages live. R4 needs the separate realness/usefulness arbitration to act as the outside referee, and R5 (writing axis.orthogonal into the registry) is gated on R2 passing and R4 being available. The table-side remedy (min_age_to_force_merge_seconds) is a production ALTER and is called the operator's call. | Two rival tools for deciding whether a proposed new market measurement merely repeats something the project already uses were run against seven measurements the project already ships and trusts: the newer tool certified none of them and the older one certified six, each is broken in a place the other is not, and the follow-up re-run that appeared to fix the problem was later found to have been reading a database table in which roughly half the rows were duplicates. |
| Probe hardening loop โ making the rotation cascade decidable 2026-07-29 ยท 20 claims ยท 3 files ยท verdict: none findings/evolution/audits/2026-07-29-probe-hardening-loop/ Blocked on Operator โ sign-off on the two parameter-bearing pre-registration drafts (ฮพ stability guard, volatility-regime terciles) before any dial moves, plus a supervised full-width 5 GiB bigblack read for the queued items (R2 re-run on the deduplicated substrate, P_50 re-derivation, ฮพ-jitter confirmation at real per-cell n). Item L-3/L-4/L-5 also need a further firing of the loop, whose cron was cancelled. | A fourteen-round loop stress-tested the two feature-screening probes and found that most of what had been blamed on the data is actually a property of the measuring tool and its pass rule โ including a rule that refused 72% of candidates it should have accepted โ but the loop hit its own round cap with three items untouched and every proposed dial change left unsigned for the operator. |
| Worst-of-N basis, grid, and probe read-path 2026-07-30 ยท 19 claims ยท 6 files ยท verdict: verdict.md findings/evolution/audits/2026-07-30-worst-of-n-basis-and-grid/ Blocked on Operator โ two named decisions: what to do about the 2 of 62 cells that exhaust the 5 GiB envelope (chunked-column fallback versus raising the envelope), and whether the comparison panel should keep the intra_* near-siblings of the anchor features. Phases 2โ4 are planned but not started, and any change to the 0.50 line needs its own pre-registration and sign-off. | Someone checked whether the 0.50 pass mark used to judge candidate features still means what it meant when it was set, and found it was calibrated against a different measurement than the one it now governs โ so far apart that the very feature chosen to represent "definitely orthogonal" scores as a near-duplicate under the live rule. |
| Substrate computability โ can a candidate be computed from the pinned slice at all? 2026-08-14 ยท 19 claims ยท 2 files ยท verdict: verdict.md findings/evolution/audits/2026-08-14-substrate-computability-audit/ Blocked on Operator โ section 9 is titled "Open items (not self-fixable)" and item 1 is D45: "R7 amendment + operator ruling on whether to widen the axis-2 slice now that 21 symbols carry labels." The offline gate shipped; the measurement is only partial. | Before spending compute on roughly 230 research ideas, someone checked which ones the data on hand can actually support โ then measured the check itself and found that 41 of its 69 confident "impossible" verdicts were wrong, leaving the missing order book as the only genuine data gap. |
Reached a terminal state on its own rules. Says nothing about whether the findings were good.
| Audit | What it turned up |
|---|---|
| Three-axis probe formalization โ orthogonal, parameterless, agnostic 2026-06-09 ยท 20 claims ยท 11 files ยท verdict: verdict.md findings/evolution/audits/2026-06-09-three-axis-probe-formalization/ | This audit turned the three house rules for accepting a new data column โ it must add information nothing else captures, it must contain no hand-tuned dials, and any two correct programs computing it must agree โ into three programs that check those rules automatically plus one combined yes/no, and it corrected a mistranslation of the third rule that had turned 'any correct implementation agrees' into 'works on any asset'. |
| Crypto and forex parameterless candidate discovery (target 100) 2026-06-17 ยท 10 claims ยท 1 files ยท verdict: none findings/evolution/audits/2026-06-17-crypto-forex-parameterless-discovery/ | This folder sets the rules for an automated hunt for one hundred brand-new candidate measurements that carry no hand-tuned settings and come with real open-source code, and the shared candidate registry now holds exactly one hundred entries filed under this hunt โ but the folder itself never recorded a conclusion and none of the per-candidate paperwork it promised to keep here was ever written. |
| Chatterjee ฮพ threshold calibration 2026-06-30 ยท 20 claims ยท 12 files ยท verdict: verdict.md findings/evolution/audits/2026-06-30-chatterjee-xi-threshold-calibration/ | Asked why the line that decides whether two market measurements are effectively the same thing sat at 0.50, measured 5,229 column pairs across 766 real data slices, and replaced the guessed number with two evidence-anchored cut lines โ plus a written ban on the averaging rule that had been hiding redundancy. |
| Batch-6 Chatterjee ฮพ filter 2026-07-02 ยท 16 claims ยท 3 files ยท verdict: verdict.md findings/evolution/audits/2026-07-02-batch6-chatterjee-xi-filter/ | Ran the newly declared redundancy test on the eight candidate features that had already cleared an earlier screen โ four came out clearly independent and four landed in the grey zone โ and along the way showed that the obvious way of computing this test on rolling-window features produces false answers. |
| Crypto cost-realism labels: BTC backfill execution, incidents and fixes 2026-07-15 ยท 21 claims ยท 10 files ยท verdict: none findings/evolution/audits/2026-07-15-crypto-cost-realism-labels-backfill/ | A record of what actually happened when the cost-of-trading labels for Bitcoin bars were computed and written on the production machine โ the finished table holds 5,958,856 unique labels covering essentially all 5,958,874 distinct bars with no duplicates, but reaching it took two aborted runs, a bug caught in review before it shipped, and five code fixes, several of whose follow-ups are still being worked through. |
| Panel-Rยฒ threshold calibration 2026-07-16 ยท 14 claims ยท 3 files ยท verdict: none findings/evolution/audits/2026-07-16-panel-r2-threshold-calibration/ | This gave a redundancy test its own cut-off for the first time: by measuring how well each of 66 existing bar columns can be rebuilt from the others across 62 market slices, the team replaced a threshold borrowed from a different statistic with one anchored to reference features whose answer was known in advance โ and wrote the whole plan down before computing a single number. |
This hub covers nasimubd's audits. Folders excluded for authorship are named rather than dropped, so the count on this page can be reconciled against the folders on disk.
2026-07-14-tessara-fade-crypto-provenance โ authored entirely by terrylica (3 commits, terry@eonlabs.com); no nasimubd commitsfindings/dashboard/build_audits.py from each audit's AUDIT_LEDGER.json โ never hand-edited, so it cannot drift from the folders the way the hand-written version did the day a 21st audit landed. Coverage is enforced by tests/test_dashboard_audits.py: an audit with no derivable status fails the build rather than rendering a reassuring default.