rag—jev EXPLORATORY / CORPUS RETRIEVALAnswer study ↗

BEIR / SciFact

Retrieve first.
Rerank the same candidates.

PipelineNDCG@10Recall@10MRR@10
bm2566.4778.4963.28
jev75.1382.4273.64
ettin72.1180.2670.64

Exploratory retrieval-only result: all 300 BEIR/SciFact test queries and all 5,183 corpus documents. Custom BM25 retrieves 20 candidates; BM25, Jev and Ettin each return 10 from that identical pool. No gold passages are inserted. All table values are percentages.

First-stage Recall@20: 82.59%. Queries with missing relevant papers remain in every denominator.

jev vs bm25: NDCG@10 difference +8.66 points; descriptive 95% interval [+5.85, +11.59]. Bootstrap units: 247 connected groups of queries sharing relevant papers; the estimate remains query-weighted.

jev vs ettin: NDCG@10 difference +3.01 points; descriptive 95% interval [+0.61, +5.41]. Bootstrap units: 247 connected groups of queries sharing relevant papers; the estimate remains query-weighted.

Jev scoring API usage estimate: $0.13188 for 300 queries ($0.4396 per 1,000). Ettin inference: 780.1 process CPU seconds; its hosting cost is unknown. Common retrieval cost is excluded; these figures do not establish a total-cost advantage over Ettin.

All 2,700 per-query metric values match pytrec-eval-terrier 0.5.10 within 1e-12. The exported run contains only ten results per query, so reciprocal rank is truncated at ten.

This is one retrieval dataset with a custom BM25 first stage, not the full BEIR suite, an answer-accuracy result, or a state-of-the-art claim. The separate HotpotQA/MuSiQue answer study keeps its original primary endpoints.

Ettin's published training recipe selects checkpoints on NanoBEIR, including SciFact and HotpotQA; overlap with those selection examples has not been ruled out. Jev pretraining exposure is unknown. These are not demonstrably benchmark-naive models. No dense retriever or query expansion is evaluated.

Results · Frozen protocol · Metric verification · Baseline training recipe