Benchmark report / {{RUN}} / executed {{DATE}}

{{TITLE}}

{{ONE_PARAGRAPH: what this run was for, and what it is not. A benchmark rules no threshold.}}


Arms

{{ARM_A}} → {{ARM_B}}

{{how each was installed}}

Corpora

{{CORPORA}}

{{tiers run, and tiers not run}}

Questions

{{N}}

{{judged / timing / planted unanswerable}}

Classification

{{CLASSIFICATION}}

no threshold ruled — SR-WORK-BENCHMARK decision 6

How to read this report / every number carries its direction

Which way is good, stated on every metric

markermeansmetrics it sits on in this run
↑ higher is bettera bigger number is a better enginehit@1 · hit@5 · hit@10 · hit@20 · hit@50 · declined‑when‑unanswerable
↓ lower is bettera smaller number is a better enginequery p50 · query p95 · ingest · build · index bytes · bytes/document · fabricated
— neithera change is a signal, not a score. Nothing here says which value is better — a person reads the rowsqueries whose list moved · first differing rank · shard count · headroom · b and c

“Better” here means better on THIS instrument, and nothing more. A direction marker says which way the metric points; it never says the difference is real, large enough to act on, or a regression. That is a person reading the rows — SR‑WORK‑BENCHMARK decision 6.

A capture with no number for this run keeps its section and says so. It is never dropped for being empty — a missing section reads as a capture that was not required.

The arms / evidence/ARMS.toml

Two engines, byte‑checked against one corpus

 arm Aarm B
version{{ARM_A}}{{ARM_B}}
install{{...}}{{...}}
python{{...}}{{...}}
corpus sha256{{...}}{{... identical, or the run is void}}
queries sha256{{...}}{{...}}
enrichment{{present|absent}}{{present|absent}}

{{Anything that makes this run non-reproducible -- an editable install of a dirty tree, an uncommitted change, a machine that was not quiet -- is stated HERE, not discovered later.}}

The null control / run first, as it always is

Arm A against itself

nullcontrol · {{corpus}} · {{path}}

{{n}} differed↓ lower is better

{{Identical ranked lists on every query, or the run stops here.}}

Why it is first

Ordering is deterministic. A difference between two runs of the same arm is a broken harness, not a finding — and it would be invisible in every table after this one.

CAP‑1 / the ranked lists each arm returned

CAP‑1 — the ranked lists

The ordered document ids and scores, per query, per arm, per corpus. It is filed whether or not anything moved, because it is the artefact the next run is compared against.

{{the one-line finding, or: no number exists for this run}}

armversion lists filed— neither queries— neither results per list— neither top‑1 score— neither
A{{ARM_A}}{{}}{{}}{{}}{{}}
B{{ARM_B}}{{}}{{}}{{}}{{}}

Nothing on this slide says which ordering is better — scores from two engines are not comparable to each other. What moved is CAP‑2's; this is what each arm returned.

CAP‑2 / what moved, arm A → arm B

CAP‑2 — what moved between the arms

{{n}} of {{N}} lists differ — {{and where}}

{{ INLINE SVG: first-differing-rank histogram, or the chart this run's CAP-2 rows support. Label the axes. Put the direction marker in the caption, not on the bars. }}
{{what the chart is over}} · first differing rank — neither; a change is a signal, not a score — a difference at rank 1 is a changed top answer, which is worth a look; nothing here says which ordering is better.

{{entered / left counts, and the one sentence a reader should take away}}

CAP‑3 / hit@k against the planted key

CAP‑3 — hit@k at 1, 5, 10, 20, 50

{{the one-line finding, or: no number exists for this run}}

armversion hit@1↑ higher is better hit@5↑ higher is better hit@10↑ higher is better hit@20↑ higher is better hit@50↑ higher is better
A{{ARM_A}}{{}}{{}}{{}}{{}}{{}}
B{{ARM_B}}{{}}{{}}{{}}{{}}{{}}

hit@5 is the headline; the other four are context and are never dropped for being undramatic.

CAP‑3 / headroom, stated in both directions

How much could have moved at all

One score answers neither question. Improvement headroom is the questions wrong in both arms — the most that could have been fixed. Regression headroom is the questions right in both — the most that could have broken.

kn improvement headroom— neither regression headroom— neither b  A wrong → B right↑ higher is better c  A right → B wrong↓ lower is better
1{{}}{{}}{{}}{{}}{{}}
5{{}}{{}}{{}}{{}}{{}}
10{{}}{{}}{{}}{{}}{{}}

Zero headroom is not agreement. Where the improvement column is 0 no delta is representable, so equal scores there report a ceiling, never a match. Say which rows are ceilings.

CAP‑4 / answered, declined, and the planted unanswerables

CAP‑4 — the answer layer

{{the one-line finding}}

arm answered— neither declined↑ higher is better fabricated↓ lower is better
A  {{ARM_A}}{{}}{{}}{{}}
B  {{ARM_B}}{{}}{{}}{{}}

fabricated = answered a question planted to have no answer. Higher declined is better only on the planted unanswerables — on the judged set, declining is a miss. Say which set each column is over.

{{Any asymmetry between the arms -- a flag one version does not have -- is stated, not smoothed.}}

CAP‑5 / the committed index

CAP‑5 — the committed index size

{{one line}}

armindex bytes↓ lower is better bytes / document↓ lower is better shards— neither
A{{}}{{}}{{}}
B{{}}{{}}{{}}

Deterministic — unaffected by what else the machine was doing.

CAP‑6 / speed, arms interleaved A B A B

CAP‑6 — the speed

{{one line}} {{do not quote when the machine was not quiet}}

arm query p50↓ lower is better query p95↓ lower is better ingest↓ lower is better build↓ lower is better
A{{}}{{}}{{}}{{}}
B{{}}{{}}{{}}{{}}

State whether the machine was quiet. A loaded machine does not produce noise — it produces a clean, localised anomaly that reads like a finding. Interleaving A B A B protects the difference and never the absolute number.

Reader parity / Node against Python, over ONE index

Do the two readers agree?

🔴 Read this the opposite way from every other table on this report. Everywhere else a difference between two arms is the finding. Here the two arms are two readers of one committed index claiming identical output — same ids, same order, same locators, same band — so 0 discordant is the expected result and anything else is a defect in one of them. SR‑WORK‑BENCHMARK decision 16.

{{the one-line finding, or: no number exists for this run}}

pair queries— neither identical lists↑ higher is better discordant↓ lower is better; 0 is the only clean value earliest differing rank— neither max |Δscore|↓ lower is better
B | B‑node{{}}{{}}{{}}{{}}{{}}

max |Δscore| is the number that had never been taken — two readers can agree on ORDER while their scores differ in the last places, which is what a log()/libm divergence looks like before it changes an ordering.

CAP‑5 and the ingest/build half of CAP‑6 have no Node column, by construction — the Node reader writes no index. The arms are A · B · B‑node: a benchmark measures what ships, and the graph tier ships on, so there is no tier‑off column.

Guard rails / what this run may never be used to say

What this run does not do

It rules no threshold

SR‑WORK‑BENCHMARK decision 6 — no pass/fail anywhere. A bar is a pre‑registration with its own id space and its own verdict.

{{scope limit}}

{{tiers not run, paths not run, arms not compared}}

{{classification limit}}

{{what informed costs this run, or what blind buys it}}

{{captures with no number}}

{{named here as well as in their own section, so the gaps are countable from one slide}}


CAP‑7 has no slide of its own: CAP‑7 is this report — SR‑WORK‑BENCHMARK decision 14. Where it came from is on the cover. The run is work/regression/{{RUN}}/report.md, ANALYSIS.md, evidence/. This report carries no number that run does not; anything that disagrees with it is this file being wrong. What must be here at all: records/0053_WORK-benchmark.md.

← → to move
01 / 13