rag—jev CONTROLLED STUDY / ROUND 3 Earlier results ↗

A frozen test with stronger controls

Separate the effects.
Show the uncertainty.

400 isolated questions, repeated answers, a neural baseline, and every pipeline measured against the same evidence.

Primary paired F1 confidence intervals and secondary factorial contrasts
DatasetPipelineAnswer F1EMAll support retainedUncached API $ / 1,000CPU seconds / questionAnswer disagreement
HotpotQAall_context76.7062.00100.0%0.34720.00016.5%
HotpotQAjev_0.2_original78.5364.8399.0%0.27970.00013.5%
HotpotQAjev_0.2_rank (frozen candidate)77.2063.0099.0%0.27880.00014.0%
HotpotQAjev_0.35_original78.8864.8396.5%0.25260.00013.5%
HotpotQAjev_0.35_rank76.3361.3396.5%0.25160.00012.5%
HotpotQAettin_top876.2462.8395.0%0.28100.40912.5%
MuSiQue-Ansall_context70.3160.83100.0%0.61970.00020.0%
MuSiQue-Ansjev_0.2_original68.7959.5095.5%0.47980.00022.5%
MuSiQue-Ansjev_0.2_rank (frozen candidate)72.8063.3395.5%0.47830.00021.0%
MuSiQue-Ansjev_0.35_original67.2959.1785.5%0.42270.00021.5%
MuSiQue-Ansjev_0.35_rank67.5359.8385.5%0.42440.00021.0%
MuSiQue-Ansettin_top854.3246.3363.5%0.32620.55820.5%

The predeclared accuracy-and-cost superiority gate is not met.

400 isolated evaluation questions, three samples per pipeline/question, 7,200 evaluation branches. The independent sample size is 200 questions per dataset, not the number of generations. All-context, a full cutoff/order factorial, and development-tuned Ettin are evaluated on the same questions. Errors: 0; truncations: 0; unknown-usage records: 0.

HotpotQA: frozen candidate versus all context, F1 difference +0.50 points, simultaneous 97.5% interval [-2.72, +3.69]; Holm-adjusted one-sided p=0.3595. Normalized API cost 19.7% lower.

HotpotQA: candidate versus Ettin, descriptive F1 difference +0.96 points, unadjusted 95% interval [-1.31, +3.35]. CPU-cost break-even: $-0.0195 per CPU-hour, excluding other hosting costs. A negative break-even means Ettin generation alone already costs more than the candidate's API bill.

HotpotQA: lowering the cutoff, averaged over both orders, descriptive F1 difference +0.26 points; unadjusted 95% interval [-1.19, +1.65].

HotpotQA: relevance ordering, averaged over both cutoffs, descriptive F1 difference -1.94 points; unadjusted 95% interval [-3.91, -0.15].

HotpotQA evaluation composition — bridge: 166, comparison: 34.

MuSiQue-Ans: frozen candidate versus all context, F1 difference +2.50 points, simultaneous 97.5% interval [-1.31, +6.37]; Holm-adjusted one-sided p=0.1531. Normalized API cost 22.8% lower.

MuSiQue-Ans: candidate versus Ettin, descriptive F1 difference +18.48 points, unadjusted 95% interval [+13.24, +23.92]. CPU-cost break-even: $0.9808 per CPU-hour, excluding other hosting costs. A negative break-even means Ettin generation alone already costs more than the candidate's API bill.

MuSiQue-Ans: lowering the cutoff, averaged over both orders, descriptive F1 difference +3.39 points; unadjusted 95% interval [+1.22, +5.78].

MuSiQue-Ans: relevance ordering, averaged over both cutoffs, descriptive F1 difference +2.13 points; unadjusted 95% interval [+0.27, +4.17].

MuSiQue-Ans evaluation composition — 2hop: 170, 3hop1: 23, 3hop2: 2, 4hop1: 3, 4hop3: 2.

Primary intervals use question-paired bootstrap (10,000 draws) with simultaneous Bonferroni coverage for two endpoints. One-sided sign-flip tests use 20,000 draws and Holm correction. Secondary factorial, neural and subgroup intervals are descriptive. A nonsignificant difference does not establish equivalence.

Costs are returned-usage estimates at uncached token rates, including Jev. Ettin's listed dollars cover GENERATION ONLY; CPU cost is separate and its total monetary cost is unknown. Observed cache-discounted API costs are in the JSON report. Common retrieval, application hosting and network costs are excluded.

Strict source-component isolation changes the validation population. These are not full benchmark or corpus-retrieval results. Answer F1 is token overlap, not semantic or citation correctness. Human adjudication has not been completed. Proprietary-model training contamination and weights cannot be independently verified. Ettin's published checkpoint-selection procedure uses NanoBEIR, including HotpotQA and SciFact; disjointness from those selection examples is not established. Scoring is fixed per question; repeats measure conditional generation variance. No query expansion claim. Secondary findings cannot be used to replace the frozen candidate and claim a new held-out win. In particular, the small higher-hop MuSiQue subgroups cannot support precise higher-hop accuracy claims.

Download measured results · Download frozen protocol · Download integrity checks

Separate corpus-retrieval check: all 300 SciFact test queries ↗