rag—jev RESEARCH / 400 NEW QUESTIONSRound 1 results ↗

Evidence retention, tested again

Keep the links.
Measure the tradeoff.

A lower cutoff and relevance ordering, chosen on development data and evaluated once on fresh questions.

Answer F1, total API cost and complete evidence retention
Dataset (200 questions)PipelineAnswer EMAnswer F1API cost / 100Support recallAll support retained
HotpotQAAll context58.0075.08$0.033994100.00%100.00%
HotpotQAOld Jev profile62.0077.57$0.02502297.75%95.50%
HotpotQAResearch v259.5076.87$0.02774699.25%98.50%
MuSiQue-AnsAll context55.0066.52$0.064869100.00%100.00%
MuSiQue-AnsOld Jev profile50.0062.24$0.04779891.79%80.00%
MuSiQue-AnsResearch v254.0063.92$0.05382396.50%91.00%

The predeclared accuracy-and-cost gate versus all context is not met.

HotpotQA, research v2 versus all context: F1 +1.80 points; 95% paired interval [-1.34, +4.78]. Estimated API cost 18.4% lower.

HotpotQA, research v2 versus old jev profile: F1 -0.70 points; 95% paired interval [-3.67, +2.29]. Estimated API cost 10.9% higher.

MuSiQue-Ans, research v2 versus all context: F1 -2.60 points; 95% paired interval [-7.36, +2.42]. Estimated API cost 17.0% lower.

MuSiQue-Ans, research v2 versus old jev profile: F1 +1.68 points; 95% paired interval [-2.54, +6.08]. Estimated API cost 12.6% higher.

400 fresh validation questions, 200 per dataset, disjoint from all 240 development questions. Supplied candidate passages; Jev 1.13.0 and GPT-5.6 Luna, low reasoning. At most one generation per branch; 1,200 evaluation branches. Costs include Jev and generation with reported cache discounts; common retrieval, hosting and networking excluded. These are subsets, not full benchmark or leaderboard results.

0 errors; 0 unknown costs; 0 unknown cache-usage records; 0 truncations

Answer F1 measures normalized token overlap, so wording changes can affect scores. Supporting-passage retention is a separate diagnostic. Model training overlap is unknown. No query expansion, full-corpus retrieval, or modern neural reranker comparison was tested. MuSiQue uses the pinned community mirror described in the protocol.

Download measured results · Download frozen protocol · Open workbench

Profile: passages together, filter & rerank, cutoff 0.20, no top-N cap. Test on your own retrieval results before adopting it.