The better a system scores, the less it will admit it does not know.
IBM MTRAG, 842 human multi-turn tasks. Of 55 unanswerable ones, this is how
often each system correctly refused, judged by the benchmark's own scorer.
CORRECT REFUSALS, OF 55 UNANSWERABLE
END TO END ANSWER QUALITY, HARMONIC MEAN
Every baseline recomputed from the MTRAG release through one harness, so
these are anchored comparisons and not the published leaderboard. On MTRAG the refusal rate
is set by the generator prompt, not by retrieval.