Benchmark report / {{RUN}} / executed {{DATE}}
{{TITLE}}
{{ONE_PARAGRAPH: what this run was for, and what it is not. A benchmark rules no threshold.}}
Arms
{{ARM_A}} → {{ARM_B}}
{{how each was installed}}
Corpora
{{CORPORA}}
{{tiers run, and tiers not run}}
Questions
{{N}}
{{judged / timing / planted unanswerable}}
Classification
{{CLASSIFICATION}}
no threshold ruled — SR-WORK-BENCHMARK decision 6
How to read this report / every number carries its direction
Which way is good, stated on every metric
| marker | means | metrics it sits on in this run |
|---|---|---|
| ↑ higher is better | a bigger number is a better engine | hit@1 · hit@5 · hit@10 · hit@20 · hit@50 · declined‑when‑unanswerable |
| ↓ lower is better | a smaller number is a better engine | query p50 · query p95 · ingest · build · index bytes · bytes/document · fabricated |
| — neither | a change is a signal, not a score. Nothing here says which value is better — a person reads the rows | queries whose list moved · first differing rank · shard count · headroom · b and c |
“Better” here means better on THIS instrument, and nothing more. A direction marker says which way the metric points; it never says the difference is real, large enough to act on, or a regression. That is a person reading the rows — SR‑WORK‑BENCHMARK decision 6.
A capture with no number for this run keeps its section and says so. It is never dropped for being empty — a missing section reads as a capture that was not required.
The arms / evidence/ARMS.toml
Two engines, byte‑checked against one corpus
| arm A | arm B | |
|---|---|---|
| version | {{ARM_A}} | {{ARM_B}} |
| install | {{...}} | {{...}} |
| python | {{...}} | {{...}} |
| corpus sha256 | {{...}} | {{... identical, or the run is void}} |
| queries sha256 | {{...}} | {{...}} |
| enrichment | {{present|absent}} | {{present|absent}} |
{{Anything that makes this run non-reproducible -- an editable install of a dirty tree, an uncommitted change, a machine that was not quiet -- is stated HERE, not discovered later.}}
The null control / run first, as it always is
Arm A against itself
nullcontrol · {{corpus}} · {{path}}
{{n}} differed↓ lower is better
{{Identical ranked lists on every query, or the run stops here.}}
Why it is first
Ordering is deterministic. A difference between two runs of the same arm is a broken harness, not a finding — and it would be invisible in every table after this one.
CAP‑1 / the ranked lists each arm returned
CAP‑1 — the ranked lists
The ordered document ids and scores, per query, per arm, per corpus. It is filed whether or not anything moved, because it is the artefact the next run is compared against.
{{the one-line finding, or: no number exists for this run}}
| arm | version | lists filed— neither | queries— neither | results per list— neither | top‑1 score— neither |
|---|---|---|---|---|---|
| A | {{ARM_A}} | {{}} | {{}} | {{}} | {{}} |
| B | {{ARM_B}} | {{}} | {{}} | {{}} | {{}} |
Nothing on this slide says which ordering is better — scores from two engines are not comparable to each other. What moved is CAP‑2's; this is what each arm returned.
CAP‑2 / what moved, arm A → arm B
CAP‑2 — what moved between the arms
{{n}} of {{N}} lists differ — {{and where}}
{{entered / left counts, and the one sentence a reader should take away}}
CAP‑3 / hit@k against the planted key
CAP‑3 — hit@k at 1, 5, 10, 20, 50
{{the one-line finding, or: no number exists for this run}}
| arm | version | hit@1↑ higher is better | hit@5↑ higher is better | hit@10↑ higher is better | hit@20↑ higher is better | hit@50↑ higher is better |
|---|---|---|---|---|---|---|
| A | {{ARM_A}} | {{}} | {{}} | {{}} | {{}} | {{}} |
| B | {{ARM_B}} | {{}} | {{}} | {{}} | {{}} | {{}} |
hit@5 is the headline; the other four are context and are never dropped for being undramatic.
CAP‑3 / headroom, stated in both directions
How much could have moved at all
One score answers neither question. Improvement headroom is the questions wrong in both arms — the most that could have been fixed. Regression headroom is the questions right in both — the most that could have broken.
| k | n | improvement headroom— neither | regression headroom— neither | b A wrong → B right↑ higher is better | c A right → B wrong↓ lower is better |
|---|---|---|---|---|---|
| 1 | {{}} | {{}} | {{}} | {{}} | {{}} |
| 5 | {{}} | {{}} | {{}} | {{}} | {{}} |
| 10 | {{}} | {{}} | {{}} | {{}} | {{}} |
Zero headroom is not agreement. Where the improvement column is 0 no delta is representable, so equal scores there report a ceiling, never a match. Say which rows are ceilings.
CAP‑4 / answered, declined, and the planted unanswerables
CAP‑4 — the answer layer
{{the one-line finding}}
| arm | answered— neither | declined↑ higher is better | fabricated↓ lower is better |
|---|---|---|---|
| A {{ARM_A}} | {{}} | {{}} | {{}} |
| B {{ARM_B}} | {{}} | {{}} | {{}} |
fabricated = answered a question planted to have no answer. Higher declined is better only on the planted unanswerables — on the judged set, declining is a miss. Say which set each column is over.
CAP‑5 / the committed index
CAP‑5 — the committed index size
{{one line}}
| arm | index bytes↓ lower is better | bytes / document↓ lower is better | shards— neither |
|---|---|---|---|
| A | {{}} | {{}} | {{}} |
| B | {{}} | {{}} | {{}} |
Deterministic — unaffected by what else the machine was doing.
CAP‑6 / speed, arms interleaved A B A B
CAP‑6 — the speed
{{one line}} {{do not quote when the machine was not quiet}}
| arm | query p50↓ lower is better | query p95↓ lower is better | ingest↓ lower is better | build↓ lower is better |
|---|---|---|---|---|
| A | {{}} | {{}} | {{}} | {{}} |
| B | {{}} | {{}} | {{}} | {{}} |
State whether the machine was quiet. A loaded machine does not produce noise — it produces a clean, localised anomaly that reads like a finding. Interleaving A B A B protects the difference and never the absolute number.
Reader parity / Node against Python, over ONE index
Do the two readers agree?
🔴 Read this the opposite way from every other table on this report. Everywhere else a difference between two arms is the finding. Here the two arms are two readers of one committed index claiming identical output — same ids, same order, same locators, same band — so 0 discordant is the expected result and anything else is a defect in one of them. SR‑WORK‑BENCHMARK decision 16.
{{the one-line finding, or: no number exists for this run}}
| pair | queries— neither | identical lists↑ higher is better | discordant↓ lower is better; 0 is the only clean value | earliest differing rank— neither | max |Δscore|↓ lower is better |
|---|---|---|---|---|---|
| B | B‑node | {{}} | {{}} | {{}} | {{}} | {{}} |
max |Δscore| is the number that had never been taken — two readers can agree on ORDER while their scores differ in the last places, which is what a log()/libm divergence looks like before it changes an ordering.
⚠ CAP‑5 and the ingest/build half of CAP‑6 have no Node column, by construction — the Node reader writes no index. The arms are A · B · B‑node: a benchmark measures what ships, and the graph tier ships on, so there is no tier‑off column.
Guard rails / what this run may never be used to say
What this run does not do
It rules no threshold
SR‑WORK‑BENCHMARK decision 6 — no pass/fail anywhere. A bar is a pre‑registration with its own id space and its own verdict.
{{scope limit}}
{{tiers not run, paths not run, arms not compared}}
{{classification limit}}
{{what informed costs this run, or what blind buys it}}
{{captures with no number}}
{{named here as well as in their own section, so the gaps are countable from one slide}}
CAP‑7 has no slide of its own: CAP‑7 is this report — SR‑WORK‑BENCHMARK decision 14. Where it came from is on the cover. The run is work/regression/{{RUN}}/ — report.md, ANALYSIS.md, evidence/. This report carries no number that run does not; anything that disagrees with it is this file being wrong. What must be here at all: records/0053_WORK-benchmark.md.