Benchmark report / 2026-09-15-node-column / executed 2026-09-15

Benchmark — 2026-09-15-node-column

Node's query latency, measured for the first time, as a COLUMN beside Python's. The Node reader is FLAT in corpus size — 26.3 ms at 100 documents and 26.4 ms at 1 000 — and every bit of its growth is W-161's in-memory graph rebuild, which takes a Node ask to 269.5 ms at 1 000 and about 2.4 s at 10 000.


Arms

arm A → arm B

install not filed · install not filed

Corpora

docs-00100, docs-01000

tiers run are those named above; any not run are not stated in the evidence

Questions

0

judged / timing / planted unanswerable split not filed

Classification

surface capture

no threshold ruled — SR-WORK-BENCHMARK decision 6

How to read this report / every number carries its direction

Which way is good, stated on every metric

markermeansmetrics it sits on in this run
↑ higher is bettera bigger number is a better enginehit@1 · hit@5 · hit@10 · hit@20 · hit@50 · declined‑when‑unanswerable
↓ lower is bettera smaller number is a better enginequery p50 · query p95 · ingest · build · index bytes · bytes/document · fabricated
— neithera change is a signal, not a score. Nothing here says which value is better — a person reads the rowsqueries whose list moved · first differing rank · shard count · headroom · b and c

“Better” here means better on THIS instrument, and nothing more. A direction marker says which way the metric points; it never says the difference is real, large enough to act on, or a regression. That is a person reading the rows — SR‑WORK‑BENCHMARK decision 6.

A capture with no number for this run keeps its section and says so. It is never dropped for being empty — a missing section reads as a capture that was not required.

The arms / evidence/ARMS.toml

Two engines, byte‑checked against one corpus

No number exists for this run — the arms. No evidence/ARMS.toml was filed, so what was compared cannot be stated from the evidence. The run is frozen and was not re-executed to fill this — what is done is done, and the next run captures it.

The null control / run first, as it always is

Arm A against itself

No number exists for this run — the null control. No nullcontrol-*.json was filed. Every number in this run rests on a control that cannot be read here, which is a stronger caveat than any table below. The run is frozen and was not re-executed to fill this — what is done is done, and the next run captures it.

Why it is first

Ordering is deterministic. A difference between two runs of the same arm is a broken harness, not a finding — and it would be invisible in every table after this one.

CAP‑1 / the ranked lists each arm returned

CAP‑1 — the ranked lists

The ordered document ids and scores, per query, per arm, per corpus. It is filed whether or not anything moved, because it is the artefact the next run is compared against.

No number exists for this run — CAP-1 the ranked lists. No evidence/ranked-lists.jsonl was filed, so what each arm returned cannot be shown — and the next run has nothing to compare against, which is the one thing this capture exists for. The run is frozen and was not re-executed to fill this — what is done is done, and the next run captures it.

CAP‑2 / what moved, arm A → arm B

CAP‑2 — what moved between the arms

No ranked lists and no rank diff

No chart. The template's figure here is an inline SVG a person draws from the CAP-2 rows; this report was generated and carries the counts instead. — neither; a change is a signal, not a score

No number exists for this run — CAP-2 what moved. No evidence/rankdiff.jsonl was filed. Neither capture is present. The run is frozen and was not re-executed to fill this — what is done is done, and the next run captures it.

CAP‑3 / hit@k against the planted key

CAP‑3 — hit@k at 1, 5, 10, 20, 50

No hit@k for this run

No number exists for this run — CAP-3 hit@k. No evidence/hits.jsonl was filed. Neither corpus tier this run used carries a planted key, so no hit@k is computable from it — not a small number, and not a zero. The run is frozen and was not re-executed to fill this — what is done is done, and the next run captures it.

CAP‑3 / headroom, stated in both directions

How much could have moved at all

One score answers neither question. Improvement headroom is the questions wrong in both arms — the most that could have been fixed. Regression headroom is the questions right in both — the most that could have broken.

No number exists for this run — CAP-3 headroom. Headroom is derived from hits.jsonl, which was not filed. Without it a null in any table below is uninterpretable — a zero delta on a saturated endpoint and a zero delta on a live one look identical. The run is frozen and was not re-executed to fill this — what is done is done, and the next run captures it.

CAP‑4 / answered, declined, and the planted unanswerables

CAP‑4 — the answer layer

No answer layer for this run

No number exists for this run — CAP-4 the answer layer. No evidence/answer-layer.jsonl was filed, so answered / declined / fabricated cannot be stated — including on the planted unanswerables, which is the one place a fabrication would show. The run is frozen and was not re-executed to fill this — what is done is done, and the next run captures it.

CAP‑5 / the committed index

CAP‑5 — the committed index size

No committed index size for this run

No number exists for this run — CAP-5 committed index size. No evidence/index-size.csv was filed and no ARMS.toml carries the bytes, so the committed size is unknown for both arms. The run is frozen and was not re-executed to fill this — what is done is done, and the next run captures it.

CAP‑6 / speed, arms interleaved A B A B

CAP‑6 — the speed

Query latency, both arms

armquery p50 (ms)↓ lower is betterquery p95 (ms)↓ lower is betteringest (s)↓ lower is betterbuild (s)↓ lower is better
A129.1132.3
B154.1157.8
B-node166.3169.2
B-node-nograph26.426.6

No latency_warning key was filed. The only statement this run makes about the machine is its own report.md, quoted here because a reader of the table above needs it:
“Node's query latency, measured for the first time, as a COLUMN beside Python's. The Node reader is FLAT in corpus size — 26.3 ms at 100 documents and 26.4 ms at 1 000 — and every bit of its growth is W-161's in-memory graph rebuild, which takes a Node ask to 269.5 ms at 1 000 and about 2.4 s at 10 000.”

These come from latency-docs-00100-scan.csv, latency-docs-01000-scan.csv, not from latency.csv, which this run did not file. Same columns, same rows, split one file per tier. Nothing was re-executed.

Guard rails / what this run may never be used to say

What this run does not do

It rules no threshold

SR‑WORK‑BENCHMARK decision 6 — no pass/fail anywhere. A bar is a pre‑registration with its own id space and its own verdict.

Scope

tiers run are those named above; any not run are not stated in the evidence

Classification: surface capture

SR-RS decisions 11-15 govern what this label permits

Captures with no number

CAP-2, CAP-3, CAP-4, CAP-5

Named here as well as in their own sections, so the gaps are countable from one slide. No run is re-executed to fill one.


CAP‑7 has no slide of its own: CAP‑7 is this report — SR‑WORK‑BENCHMARK decision 14. Where it came from is on the cover. The run is work/regression/2026-09-15-node-column/report.md, ANALYSIS.md, evidence/. This report carries no number that run does not; anything that disagrees with it is this file being wrong. What must be here at all: records/0053_WORK-benchmark.md.

← → to move
01 / 12