Benchmark report / 2026-09-16-node-column / executed 2026-09-16
Benchmark — 2026-09-16-node-column
[not filled by the emitter — this slot needs a human sentence]
Arms
1.0.0 → 2.0.1
pip install fux-engine==1.0.0 · pip install -e ../fux
Corpora
docs-00100
1 tier(s) built; 1 produced ranked lists. Every built tier filed lists.
Questions
60
judged / timing / planted unanswerable split not filed
Classification
not stated
no threshold ruled — SR-WORK-BENCHMARK decision 6
How to read this report / every number carries its direction
Which way is good, stated on every metric
| marker | means | metrics it sits on in this run |
|---|---|---|
| ↑ higher is better | a bigger number is a better engine | hit@1 · hit@5 · hit@10 · hit@20 · hit@50 · declined‑when‑unanswerable |
| ↓ lower is better | a smaller number is a better engine | query p50 · query p95 · ingest · build · index bytes · bytes/document · fabricated |
| — neither | a change is a signal, not a score. Nothing here says which value is better — a person reads the rows | queries whose list moved · first differing rank · shard count · headroom · b and c |
“Better” here means better on THIS instrument, and nothing more. A direction marker says which way the metric points; it never says the difference is real, large enough to act on, or a regression. That is a person reading the rows — SR‑WORK‑BENCHMARK decision 6.
A capture with no number for this run keeps its section and says so. It is never dropped for being empty — a missing section reads as a capture that was not required.
The arms / evidence/ARMS.toml
Two engines, byte‑checked against one corpus
| arm A | arm B | |
|---|---|---|
| version | 1.0.0 | 2.0.1 |
| install | pip install fux-engine==1.0.0 | pip install -e ../fux |
| python | 3.11.15 | 3.11.15 |
| corpus sha256 | 670fb9513e8ae4828461a2533b441ed691476c836930902ddce682415ed5a54e | 670fb9513e8ae4828461a2533b441ed691476c836930902ddce682415ed5a54e |
| queries sha256 | 544e8f2337a90faea7be809ca4cf26160c32cb5171192d7ed626065cfbe92017 | 544e8f2337a90faea7be809ca4cf26160c32cb5171192d7ed626065cfbe92017 |
| index sha256 | 9b861bb8aa245b80b19901f3c4cf198241f9fc3113fa118eaccb9bbf5abf099f | 2dbccd163a74cd4f7cdc0362afb8ac247fe36818e9c1cbdbe227cd327c0d0d99 |
| enrichment | False | False |
Both arms indexed the same corpus and the same queries — the two hi rows above agree, which is what makes the comparison a comparison.
The null control / run first, as it always is
Arm A against itself
No number exists for this run — the null control. No nullcontrol-*.json was filed. Every number in this run rests on a control that cannot be read here, which is a stronger caveat than any table below. The run is frozen and was not re-executed to fill this — what is done is done, and the next run captures it.
Why it is first
Ordering is deterministic. A difference between two runs of the same arm is a broken harness, not a finding — and it would be invisible in every table after this one.
CAP‑1 / the ranked lists each arm returned
CAP‑1 — the ranked lists
The ordered document ids and scores, per query, per arm, per corpus. It is filed whether or not anything moved, because it is the artefact the next run is compared against.
180 ranked lists filed across 3 arm(s) — the baseline the next run is read against
| arm | version | lists filed— neither | queries— neither | results per list— neither | top‑1 score— neither; an engine’s own scale |
|---|---|---|---|---|---|
| A | 1.0.0 | 60 | 60 | 10 | 4.96 |
| B | 2.0.1 | 60 | 60 | 10 | 4.96 |
| B-node | — | 60 | 60 | 10 | 4.96 |
Over docs-00100. Nothing here says which ordering is better — a score is an engine’s own scale, and the two arms are two engines. Medians, because one list is not a run.
CAP‑2 / what moved, arm A → arm B
CAP‑2 — what moved between the arms
29 of 60 lists differ
| queries— neither | queries that moved— neither; a change is a signal, not a score | documents entered (B only)— neither | left (A only)— neither |
|---|---|---|---|
| 60 | 29 | 7 | 7 |
A difference at rank 1 is a changed top answer, which is worth a look. Nothing here says which ordering is better.
CAP‑3 / hit@k against the planted key
CAP‑3 — hit@k at 1, 5, 10, 20, 50
126 graded rows across 3 arm(s)
| arm | version | hit@1↑ higher is better | hit@5↑ higher is better | hit@10↑ higher is better | hit@20↑ higher is better | hit@50↑ higher is better |
|---|---|---|---|---|---|---|
| A | 1.0.0 | 0.833 | 0.952 | 1.000 | 1.000 | 1.000 |
| B | 2.0.1 | 0.833 | 0.976 | 1.000 | 1.000 | 1.000 |
| B-node | fux 2.0.1 (node 24.13.0) — the read plane | 0.833 | 0.976 | 1.000 | 1.000 | 1.000 |
CAP‑3 / headroom, stated in both directions
How much could have moved at all
One score answers neither question. Improvement headroom is the questions wrong in both arms — the most that could have been fixed. Regression headroom is the questions right in both — the most that could have broken.
| k | n | improvement headroom— neither | regression headroom— neither | b A wrong → B right↑ higher is better | c A right → B wrong↓ lower is better |
|---|---|---|---|---|---|
| 1 | 42 | 7 | 35 | 0 | 0 |
| 5 | 42 | 1 | 40 | 1 | 0 |
| 10 | 42 | 0 | 42 | 0 | 0 |
Zero headroom is not agreement. Where the improvement column is 0 no delta is representable, so equal scores there report a ceiling, never a match.
CAP‑4 / answered, declined, and the planted unanswerables
CAP‑4 — the answer layer
156 answer verdicts filed
| arm | answered— neither | declined↑ higher is better | fabricated↓ lower is better |
|---|---|---|---|
| A | 42 | 0 | 10 |
| B | 24 | 18 | 10 |
| B-node | 24 | 18 | 10 |
30 fabrication(s) filed — an answer to a question planted with none. A/j043; B/j043; B-node/j043; A/j044; B/j044; B-node/j044
CAP‑5 / the committed index
CAP‑5 — the committed index size
Committed bytes, per arm per corpus
| arm | index bytes↓ lower is better | bytes / document↓ lower is better | shards— neither |
|---|---|---|---|
| A docs-00100 | 1,407,555 | 14075.5 | 84 |
| B docs-00100 | 1,335,981 | 13359.8 | 84 |
Deterministic — unaffected by what else the machine was doing.
CAP‑6 / speed, arms interleaved A B A B
CAP‑6 — the speed
Query latency, both arms
| arm | query p50 (ms)↓ lower is better | query p95 (ms)↓ lower is better | ingest (s)↓ lower is better | build (s)↓ lower is better |
|---|---|---|---|---|
| A | 52.4 | 53.7 | 3.8 | 0.1 |
| B | 78.4 | 80.0 | 1.8 | 0.2 |
| B-node | 63.7 | 64.7 | — | — |
No quietness statement was filed with this run. Absence is not a claim that the machine was quiet: these numbers are comparable within this run and not across machines.
Reader parity / Node against Python, over ONE index
Do the two readers agree?
🔴 Read this the opposite way from every other table on this report. Everywhere else a difference between two arms is the finding. Here the two arms are two readers of one committed index claiming identical output — same ids, same order, same locators, same band — so 0 discordant is the expected result and anything else is a defect in one of them. SR‑WORK‑BENCHMARK decision 16.
60 query comparisons across 1 reader pair(s)
| pair | queries— neither | identical lists↑ higher is better | discordant↓ lower is better; 0 is the only clean value | earliest differing rank— neither | max |Δscore|↓ lower is better |
|---|---|---|---|---|---|
| B|B-node | 60 | 60 | 0 | — | 0.00e+00 |
Every reader pair agreed on every query. That is the expected result, not a good one: it is the claim holding, and it is the only reading under which the Node column elsewhere on this report can be compared with Python’s at all.
max |Δscore| is the number that had never been taken. Two readers can agree on ORDER while their scores differ in the last places — which is what a log()/libm divergence looks like before it changes an ordering.
⚠ CAP‑5 and the ingest/build half of CAP‑6 have no Node column, by construction — the Node reader writes no index. The arms are A · B · B‑node: a benchmark measures what ships, and the graph tier ships on, so there is no tier‑off column.
Guard rails / what this run may never be used to say
What this run does not do
It rules no threshold
SR‑WORK‑BENCHMARK decision 6 — no pass/fail anywhere. A bar is a pre‑registration with its own id space and its own verdict.
Scope
1 tier(s) built; 1 produced ranked lists. Every built tier filed lists.
Classification: not stated
SR-RS decisions 11-15 govern what this label permits
Captures with no number
none — every capture filed a number
Named here as well as in their own sections, so the gaps are countable from one slide. No run is re-executed to fill one.
CAP‑7 has no slide of its own: CAP‑7 is this report — SR‑WORK‑BENCHMARK decision 14. Where it came from is on the cover. The run is work/regression/2026-09-16-node-column/ — report.md, ANALYSIS.md, evidence/. This report carries no number that run does not; anything that disagrees with it is this file being wrong. What must be here at all: records/0053_WORK-benchmark.md.