Benchmark report / 2026-09-12-benchmark-l9 / executed 2026-09-12
Benchmark — 2026-09-12-benchmark-l9
The first benchmark run under L9's shape: two versions, five corpora, latency and ranked lists kept. The RANKED LISTS and the ingest numbers stand; EVERY ABSOLUTE LATENCY IS CONTAMINATED — a second session was saturating cores for the whole sweep, disclosed after filing. What survives is the interleaved A-vs-B difference and the ordering, and the report says exactly which is which.
Arms
1.0.0 → 2.0.0-alpha.7
pip install fux-engine==1.0.0 · pip install -e ../fux
Corpora
docs-00100, docs-00200, docs-00500, docs-01000, docs-02000, docs-05000, docs-10000
7 tier(s) built; 5 produced ranked lists. No CAP-1 rows for: docs-00100, docs-10000.
Questions
60
judged / timing / planted unanswerable split not filed
Classification
informed
no threshold ruled — SR-WORK-BENCHMARK decision 6
How to read this report / every number carries its direction
Which way is good, stated on every metric
| marker | means | metrics it sits on in this run |
|---|---|---|
| ↑ higher is better | a bigger number is a better engine | hit@1 · hit@5 · hit@10 · hit@20 · hit@50 · declined‑when‑unanswerable |
| ↓ lower is better | a smaller number is a better engine | query p50 · query p95 · ingest · build · index bytes · bytes/document · fabricated |
| — neither | a change is a signal, not a score. Nothing here says which value is better — a person reads the rows | queries whose list moved · first differing rank · shard count · headroom · b and c |
“Better” here means better on THIS instrument, and nothing more. A direction marker says which way the metric points; it never says the difference is real, large enough to act on, or a regression. That is a person reading the rows — SR‑WORK‑BENCHMARK decision 6.
A capture with no number for this run keeps its section and says so. It is never dropped for being empty — a missing section reads as a capture that was not required.
The arms / evidence/ARMS.toml
Two engines, byte‑checked against one corpus
| arm A | arm B | |
|---|---|---|
| version | 1.0.0 identical across 7 corpus tier(s) | 2.0.0-alpha.7 identical across 7 corpus tier(s) |
| install | pip install fux-engine==1.0.0 identical across 7 corpus tier(s) | pip install -e ../fux identical across 7 corpus tier(s) |
| python | 3.11.15 identical across 7 corpus tier(s) | 3.11.15 identical across 7 corpus tier(s) |
| corpus sha256 | varies across 7 corpus tier(s) 7 distinct values | varies across 7 corpus tier(s) 7 distinct values |
| queries sha256 | varies across 7 corpus tier(s) 7 distinct values | varies across 7 corpus tier(s) 7 distinct values |
| index sha256 | varies across 7 corpus tier(s) 7 distinct values | varies across 7 corpus tier(s) 7 distinct values |
| enrichment | False identical across 7 corpus tier(s) | False identical across 7 corpus tier(s) |
Both arms indexed the same corpus and the same queries — the two hi rows above agree, which is what makes the comparison a comparison.
The null control / run first, as it always is
Arm A against itself
nullcontrol · docs-00100 · scan
0 differed↓ lower is better
Arm A against itself. Identical ranked lists on every query, or the run stops here.
nullcontrol · docs-01000 · scan
0 differed↓ lower is better
Arm A against itself. Identical ranked lists on every query, or the run stops here.
Why it is first
Ordering is deterministic. A difference between two runs of the same arm is a broken harness, not a finding — and it would be invisible in every table after this one.
CAP‑1 / the ranked lists each arm returned
CAP‑1 — the ranked lists
The ordered document ids and scores, per query, per arm, per corpus. It is filed whether or not anything moved, because it is the artefact the next run is compared against.
600 ranked lists filed across 2 arm(s) — the baseline the next run is read against
| arm | version | lists filed— neither | queries— neither | results per list— neither | top‑1 score— neither; an engine’s own scale |
|---|---|---|---|---|---|
| A | 1.0.0 | 300 | 60 | 10 | 5.03 |
| B | 2.0.0-alpha.7 | 300 | 60 | 10 | 5.03 |
Over docs-00200, docs-00500, docs-01000, docs-02000, docs-05000. Nothing here says which ordering is better — a score is an engine’s own scale, and the two arms are two engines. Medians, because one list is not a run.
CAP‑2 / what moved, arm A → arm B
CAP‑2 — what moved between the arms
600 ranked lists filed — what moved between them is not
No number exists for this run — CAP-2 what moved. No evidence/rankdiff.jsonl was filed. CAP-1's lists are filed, so the diff is recomputable from them — but it was not computed by this run, and this page states only what the run filed. The run is frozen and was not re-executed to fill this — what is done is done, and the next run captures it.
CAP‑3 / hit@k against the planted key
CAP‑3 — hit@k at 1, 5, 10, 20, 50
No hit@k for this run
No number exists for this run — CAP-3 hit@k. No evidence/hits.jsonl was filed. Neither corpus tier this run used carries a planted key, so no hit@k is computable from it — not a small number, and not a zero. The run is frozen and was not re-executed to fill this — what is done is done, and the next run captures it.
CAP‑3 / headroom, stated in both directions
How much could have moved at all
One score answers neither question. Improvement headroom is the questions wrong in both arms — the most that could have been fixed. Regression headroom is the questions right in both — the most that could have broken.
No number exists for this run — CAP-3 headroom. Headroom is derived from hits.jsonl, which was not filed. Without it a null in any table below is uninterpretable — a zero delta on a saturated endpoint and a zero delta on a live one look identical. The run is frozen and was not re-executed to fill this — what is done is done, and the next run captures it.
CAP‑4 / answered, declined, and the planted unanswerables
CAP‑4 — the answer layer
No answer layer for this run
No number exists for this run — CAP-4 the answer layer. No evidence/answer-layer.jsonl was filed, so answered / declined / fabricated cannot be stated — including on the planted unanswerables, which is the one place a fabrication would show. The run is frozen and was not re-executed to fill this — what is done is done, and the next run captures it.
CAP‑5 / the committed index
CAP‑5 — the committed index size
Committed bytes, per arm per corpus
| arm | index bytes↓ lower is better | bytes / document↓ lower is better | shards— neither |
|---|---|---|---|
| A docs-00100 | 1,407,555 | 14076 | — |
| B docs-00100 | 1,335,981 | 13360 | — |
| A docs-00200 | 2,801,415 | 14007 | — |
| B docs-00200 | 2,659,177 | 13296 | — |
| A docs-00500 | 6,972,004 | 13944 | — |
| B docs-00500 | 6,615,297 | 13231 | — |
| A docs-01000 | 13,928,576 | 13929 | — |
| B docs-01000 | 13,211,224 | 13211 | — |
| A docs-02000 | 27,826,549 | 13913 | — |
| B docs-02000 | 26,386,990 | 13193 | — |
| A docs-05000 | 69,516,232 | 13903 | — |
| B docs-05000 | 65,912,339 | 13182 | — |
| A docs-10000 | 138,960,394 | 13896 | — |
| B docs-10000 | 131,750,954 | 13175 | — |
These come from evidence/ARMS.toml, not from index-size.csv, which this run did not file. Same measurement, different file; the shard count is in neither and is shown as —. The run is frozen and was not re-executed.
Deterministic — unaffected by what else the machine was doing.
CAP‑6 / speed, arms interleaved A B A B
CAP‑6 — the speed
Query latency, both arms
| arm | query p50 (ms)↓ lower is better | query p95 (ms)↓ lower is better | ingest (s)↓ lower is better | build (s)↓ lower is better |
|---|---|---|---|---|
| A | 279.2 | 307.0 | 37.7 | 0.6 |
| B | 305.5 | 339.9 | 17.3 | 1.0 |
No latency_warning key was filed. The only statement this run makes about the machine is its own report.md, quoted here because a reader of the table above needs it:
“The first benchmark run under L9's shape: two versions, five corpora, latency and ranked lists kept. The RANKED LISTS and the ingest numbers stand; EVERY ABSOLUTE LATENCY IS CONTAMINATED — a second session was saturating cores for the whole sweep, disclosed after filing. What survives is the interleaved A-vs-B difference and the ordering, and the report says exactly which is which.”
Guard rails / what this run may never be used to say
What this run does not do
It rules no threshold
SR‑WORK‑BENCHMARK decision 6 — no pass/fail anywhere. A bar is a pre‑registration with its own id space and its own verdict.
Scope
7 tier(s) built; 5 produced ranked lists. No CAP-1 rows for: docs-00100, docs-10000.
Classification: informed
an informed run is never compared with a blind one and never used to state a delta
Captures with no number
CAP-2, CAP-3, CAP-4, CAP-5
Named here as well as in their own sections, so the gaps are countable from one slide. No run is re-executed to fill one.
CAP‑7 has no slide of its own: CAP‑7 is this report — SR‑WORK‑BENCHMARK decision 14. Where it came from is on the cover. The run is work/regression/2026-09-12-benchmark-l9/ — report.md, ANALYSIS.md, evidence/. This report carries no number that run does not; anything that disagrees with it is this file being wrong. What must be here at all: records/0053_WORK-benchmark.md.