Benchmark report / 2026-09-16-node-column / executed 2026-09-16

Benchmark — 2026-09-16-node-column

[not filled by the emitter — this slot needs a human sentence]


Arms

1.0.0 → 2.0.1

pip install fux-engine==1.0.0 · pip install -e ../fux

Corpora

docs-00100

1 tier(s) built; 1 produced ranked lists. Every built tier filed lists.

Questions

60

judged / timing / planted unanswerable split not filed

Classification

not stated

no threshold ruled — SR-WORK-BENCHMARK decision 6

How to read this report / every number carries its direction

Which way is good, stated on every metric

markermeansmetrics it sits on in this run
↑ higher is bettera bigger number is a better enginehit@1 · hit@5 · hit@10 · hit@20 · hit@50 · declined‑when‑unanswerable
↓ lower is bettera smaller number is a better enginequery p50 · query p95 · ingest · build · index bytes · bytes/document · fabricated
— neithera change is a signal, not a score. Nothing here says which value is better — a person reads the rowsqueries whose list moved · first differing rank · shard count · headroom · b and c

“Better” here means better on THIS instrument, and nothing more. A direction marker says which way the metric points; it never says the difference is real, large enough to act on, or a regression. That is a person reading the rows — SR‑WORK‑BENCHMARK decision 6.

A capture with no number for this run keeps its section and says so. It is never dropped for being empty — a missing section reads as a capture that was not required.

The arms / evidence/ARMS.toml

Two engines, byte‑checked against one corpus

 arm Aarm B
version1.0.02.0.1
installpip install fux-engine==1.0.0pip install -e ../fux
python3.11.153.11.15
corpus sha256670fb9513e8ae4828461a2533b441ed691476c836930902ddce682415ed5a54e670fb9513e8ae4828461a2533b441ed691476c836930902ddce682415ed5a54e
queries sha256544e8f2337a90faea7be809ca4cf26160c32cb5171192d7ed626065cfbe92017544e8f2337a90faea7be809ca4cf26160c32cb5171192d7ed626065cfbe92017
index sha2569b861bb8aa245b80b19901f3c4cf198241f9fc3113fa118eaccb9bbf5abf099f2dbccd163a74cd4f7cdc0362afb8ac247fe36818e9c1cbdbe227cd327c0d0d99
enrichmentFalseFalse

Both arms indexed the same corpus and the same queries — the two hi rows above agree, which is what makes the comparison a comparison.

The null control / run first, as it always is

Arm A against itself

No number exists for this run — the null control. No nullcontrol-*.json was filed. Every number in this run rests on a control that cannot be read here, which is a stronger caveat than any table below. The run is frozen and was not re-executed to fill this — what is done is done, and the next run captures it.

Why it is first

Ordering is deterministic. A difference between two runs of the same arm is a broken harness, not a finding — and it would be invisible in every table after this one.

CAP‑1 / the ranked lists each arm returned

CAP‑1 — the ranked lists

The ordered document ids and scores, per query, per arm, per corpus. It is filed whether or not anything moved, because it is the artefact the next run is compared against.

180 ranked lists filed across 3 arm(s) — the baseline the next run is read against

armversionlists filed— neitherqueries— neitherresults per list— neithertop‑1 score— neither; an engine’s own scale
A1.0.06060104.96
B2.0.16060104.96
B-node6060104.96

Over docs-00100. Nothing here says which ordering is better — a score is an engine’s own scale, and the two arms are two engines. Medians, because one list is not a run.

CAP‑2 / what moved, arm A → arm B

CAP‑2 — what moved between the arms

29 of 60 lists differ

queries— neitherqueries that moved— neither; a change is a signal, not a scoredocuments entered (B only)— neitherleft (A only)— neither
602977
No chart. The template's figure here is an inline SVG a person draws from the CAP-2 rows; this report was generated and carries the counts instead. — neither; a change is a signal, not a score

A difference at rank 1 is a changed top answer, which is worth a look. Nothing here says which ordering is better.

CAP‑3 / hit@k against the planted key

CAP‑3 — hit@k at 1, 5, 10, 20, 50

126 graded rows across 3 arm(s)

armversionhit@1↑ higher is betterhit@5↑ higher is betterhit@10↑ higher is betterhit@20↑ higher is betterhit@50↑ higher is better
A1.0.00.8330.9521.0001.0001.000
B2.0.10.8330.9761.0001.0001.000
B-nodefux 2.0.1 (node 24.13.0) — the read plane0.8330.9761.0001.0001.000

CAP‑3 / headroom, stated in both directions

How much could have moved at all

One score answers neither question. Improvement headroom is the questions wrong in both arms — the most that could have been fixed. Regression headroom is the questions right in both — the most that could have broken.

knimprovement headroom— neitherregression headroom— neitherb  A wrong → B right↑ higher is betterc  A right → B wrong↓ lower is better
14273500
54214010
104204200

Zero headroom is not agreement. Where the improvement column is 0 no delta is representable, so equal scores there report a ceiling, never a match.

CAP‑4 / answered, declined, and the planted unanswerables

CAP‑4 — the answer layer

156 answer verdicts filed

armanswered— neitherdeclined↑ higher is betterfabricated↓ lower is better
A42010
B241810
B-node241810

30 fabrication(s) filed — an answer to a question planted with none. A/j043; B/j043; B-node/j043; A/j044; B/j044; B-node/j044

CAP‑5 / the committed index

CAP‑5 — the committed index size

Committed bytes, per arm per corpus

armindex bytes↓ lower is betterbytes / document↓ lower is bettershards— neither
A docs-001001,407,55514075.584
B docs-001001,335,98113359.884

Deterministic — unaffected by what else the machine was doing.

CAP‑6 / speed, arms interleaved A B A B

CAP‑6 — the speed

Query latency, both arms

armquery p50 (ms)↓ lower is betterquery p95 (ms)↓ lower is betteringest (s)↓ lower is betterbuild (s)↓ lower is better
A52.453.73.80.1
B78.480.01.80.2
B-node63.764.7

No quietness statement was filed with this run. Absence is not a claim that the machine was quiet: these numbers are comparable within this run and not across machines.

Reader parity / Node against Python, over ONE index

Do the two readers agree?

🔴 Read this the opposite way from every other table on this report. Everywhere else a difference between two arms is the finding. Here the two arms are two readers of one committed index claiming identical output — same ids, same order, same locators, same band — so 0 discordant is the expected result and anything else is a defect in one of them. SR‑WORK‑BENCHMARK decision 16.

60 query comparisons across 1 reader pair(s)

pairqueries— neitheridentical lists↑ higher is betterdiscordant↓ lower is better; 0 is the only clean valueearliest differing rank— neithermax |Δscore|↓ lower is better
B|B-node606000.00e+00

Every reader pair agreed on every query. That is the expected result, not a good one: it is the claim holding, and it is the only reading under which the Node column elsewhere on this report can be compared with Python’s at all.

max |Δscore| is the number that had never been taken. Two readers can agree on ORDER while their scores differ in the last places — which is what a log()/libm divergence looks like before it changes an ordering.

CAP‑5 and the ingest/build half of CAP‑6 have no Node column, by construction — the Node reader writes no index. The arms are A · B · B‑node: a benchmark measures what ships, and the graph tier ships on, so there is no tier‑off column.

Guard rails / what this run may never be used to say

What this run does not do

It rules no threshold

SR‑WORK‑BENCHMARK decision 6 — no pass/fail anywhere. A bar is a pre‑registration with its own id space and its own verdict.

Scope

1 tier(s) built; 1 produced ranked lists. Every built tier filed lists.

Classification: not stated

SR-RS decisions 11-15 govern what this label permits

Captures with no number

none — every capture filed a number

Named here as well as in their own sections, so the gaps are countable from one slide. No run is re-executed to fill one.


CAP‑7 has no slide of its own: CAP‑7 is this report — SR‑WORK‑BENCHMARK decision 14. Where it came from is on the cover. The run is work/regression/2026-09-16-node-column/report.md, ANALYSIS.md, evidence/. This report carries no number that run does not; anything that disagrees with it is this file being wrong. What must be here at all: records/0053_WORK-benchmark.md.

← → to move
01 / 13