Benchmark report / 2026-09-12-benchmark-l9 / executed 2026-09-12

Benchmark — 2026-09-12-benchmark-l9

The first benchmark run under L9's shape: two versions, five corpora, latency and ranked lists kept. The RANKED LISTS and the ingest numbers stand; EVERY ABSOLUTE LATENCY IS CONTAMINATED — a second session was saturating cores for the whole sweep, disclosed after filing. What survives is the interleaved A-vs-B difference and the ordering, and the report says exactly which is which.


Arms

1.0.0 → 2.0.0-alpha.7

pip install fux-engine==1.0.0 · pip install -e ../fux

Corpora

docs-00100, docs-00200, docs-00500, docs-01000, docs-02000, docs-05000, docs-10000

7 tier(s) built; 5 produced ranked lists. No CAP-1 rows for: docs-00100, docs-10000.

Questions

60

judged / timing / planted unanswerable split not filed

Classification

informed

no threshold ruled — SR-WORK-BENCHMARK decision 6

How to read this report / every number carries its direction

Which way is good, stated on every metric

markermeansmetrics it sits on in this run
↑ higher is bettera bigger number is a better enginehit@1 · hit@5 · hit@10 · hit@20 · hit@50 · declined‑when‑unanswerable
↓ lower is bettera smaller number is a better enginequery p50 · query p95 · ingest · build · index bytes · bytes/document · fabricated
— neithera change is a signal, not a score. Nothing here says which value is better — a person reads the rowsqueries whose list moved · first differing rank · shard count · headroom · b and c

“Better” here means better on THIS instrument, and nothing more. A direction marker says which way the metric points; it never says the difference is real, large enough to act on, or a regression. That is a person reading the rows — SR‑WORK‑BENCHMARK decision 6.

A capture with no number for this run keeps its section and says so. It is never dropped for being empty — a missing section reads as a capture that was not required.

The arms / evidence/ARMS.toml

Two engines, byte‑checked against one corpus

 arm Aarm B
version1.0.0
identical across 7 corpus tier(s)
2.0.0-alpha.7
identical across 7 corpus tier(s)
installpip install fux-engine==1.0.0
identical across 7 corpus tier(s)
pip install -e ../fux
identical across 7 corpus tier(s)
python3.11.15
identical across 7 corpus tier(s)
3.11.15
identical across 7 corpus tier(s)
corpus sha256varies across 7 corpus tier(s)
7 distinct values
varies across 7 corpus tier(s)
7 distinct values
queries sha256varies across 7 corpus tier(s)
7 distinct values
varies across 7 corpus tier(s)
7 distinct values
index sha256varies across 7 corpus tier(s)
7 distinct values
varies across 7 corpus tier(s)
7 distinct values
enrichmentFalse
identical across 7 corpus tier(s)
False
identical across 7 corpus tier(s)

Both arms indexed the same corpus and the same queries — the two hi rows above agree, which is what makes the comparison a comparison.

The null control / run first, as it always is

Arm A against itself

nullcontrol · docs-00100 · scan

0 differed↓ lower is better

Arm A against itself. Identical ranked lists on every query, or the run stops here.

nullcontrol · docs-01000 · scan

0 differed↓ lower is better

Arm A against itself. Identical ranked lists on every query, or the run stops here.

Why it is first

Ordering is deterministic. A difference between two runs of the same arm is a broken harness, not a finding — and it would be invisible in every table after this one.

CAP‑1 / the ranked lists each arm returned

CAP‑1 — the ranked lists

The ordered document ids and scores, per query, per arm, per corpus. It is filed whether or not anything moved, because it is the artefact the next run is compared against.

600 ranked lists filed across 2 arm(s) — the baseline the next run is read against

armversionlists filed— neitherqueries— neitherresults per list— neithertop‑1 score— neither; an engine’s own scale
A1.0.030060105.03
B2.0.0-alpha.730060105.03

Over docs-00200, docs-00500, docs-01000, docs-02000, docs-05000. Nothing here says which ordering is better — a score is an engine’s own scale, and the two arms are two engines. Medians, because one list is not a run.

CAP‑2 / what moved, arm A → arm B

CAP‑2 — what moved between the arms

600 ranked lists filed — what moved between them is not

No chart. The template's figure here is an inline SVG a person draws from the CAP-2 rows; this report was generated and carries the counts instead. — neither; a change is a signal, not a score

No number exists for this run — CAP-2 what moved. No evidence/rankdiff.jsonl was filed. CAP-1's lists are filed, so the diff is recomputable from them — but it was not computed by this run, and this page states only what the run filed. The run is frozen and was not re-executed to fill this — what is done is done, and the next run captures it.

CAP‑3 / hit@k against the planted key

CAP‑3 — hit@k at 1, 5, 10, 20, 50

No hit@k for this run

No number exists for this run — CAP-3 hit@k. No evidence/hits.jsonl was filed. Neither corpus tier this run used carries a planted key, so no hit@k is computable from it — not a small number, and not a zero. The run is frozen and was not re-executed to fill this — what is done is done, and the next run captures it.

CAP‑3 / headroom, stated in both directions

How much could have moved at all

One score answers neither question. Improvement headroom is the questions wrong in both arms — the most that could have been fixed. Regression headroom is the questions right in both — the most that could have broken.

No number exists for this run — CAP-3 headroom. Headroom is derived from hits.jsonl, which was not filed. Without it a null in any table below is uninterpretable — a zero delta on a saturated endpoint and a zero delta on a live one look identical. The run is frozen and was not re-executed to fill this — what is done is done, and the next run captures it.

CAP‑4 / answered, declined, and the planted unanswerables

CAP‑4 — the answer layer

No answer layer for this run

No number exists for this run — CAP-4 the answer layer. No evidence/answer-layer.jsonl was filed, so answered / declined / fabricated cannot be stated — including on the planted unanswerables, which is the one place a fabrication would show. The run is frozen and was not re-executed to fill this — what is done is done, and the next run captures it.

CAP‑5 / the committed index

CAP‑5 — the committed index size

Committed bytes, per arm per corpus

armindex bytes↓ lower is betterbytes / document↓ lower is bettershards— neither
A docs-001001,407,55514076
B docs-001001,335,98113360
A docs-002002,801,41514007
B docs-002002,659,17713296
A docs-005006,972,00413944
B docs-005006,615,29713231
A docs-0100013,928,57613929
B docs-0100013,211,22413211
A docs-0200027,826,54913913
B docs-0200026,386,99013193
A docs-0500069,516,23213903
B docs-0500065,912,33913182
A docs-10000138,960,39413896
B docs-10000131,750,95413175

These come from evidence/ARMS.toml, not from index-size.csv, which this run did not file. Same measurement, different file; the shard count is in neither and is shown as —. The run is frozen and was not re-executed.

Deterministic — unaffected by what else the machine was doing.

CAP‑6 / speed, arms interleaved A B A B

CAP‑6 — the speed

Query latency, both arms

armquery p50 (ms)↓ lower is betterquery p95 (ms)↓ lower is betteringest (s)↓ lower is betterbuild (s)↓ lower is better
A279.2307.037.70.6
B305.5339.917.31.0

No latency_warning key was filed. The only statement this run makes about the machine is its own report.md, quoted here because a reader of the table above needs it:
“The first benchmark run under L9's shape: two versions, five corpora, latency and ranked lists kept. The RANKED LISTS and the ingest numbers stand; EVERY ABSOLUTE LATENCY IS CONTAMINATED — a second session was saturating cores for the whole sweep, disclosed after filing. What survives is the interleaved A-vs-B difference and the ordering, and the report says exactly which is which.”

Guard rails / what this run may never be used to say

What this run does not do

It rules no threshold

SR‑WORK‑BENCHMARK decision 6 — no pass/fail anywhere. A bar is a pre‑registration with its own id space and its own verdict.

Scope

7 tier(s) built; 5 produced ranked lists. No CAP-1 rows for: docs-00100, docs-10000.

Classification: informed

an informed run is never compared with a blind one and never used to state a delta

Captures with no number

CAP-2, CAP-3, CAP-4, CAP-5

Named here as well as in their own sections, so the gaps are countable from one slide. No run is re-executed to fill one.


CAP‑7 has no slide of its own: CAP‑7 is this report — SR‑WORK‑BENCHMARK decision 14. Where it came from is on the cover. The run is work/regression/2026-09-12-benchmark-l9/report.md, ANALYSIS.md, evidence/. This report carries no number that run does not; anything that disagrees with it is this file being wrong. What must be here at all: records/0053_WORK-benchmark.md.

← → to move
01 / 12