RE-call
MCP DIRECTORY RATING

The better a system scores,
the less it will admit it does not know.

IBM MTRAG, 842 human multi-turn tasks. Of 55 unanswerable ones, this is how often each system correctly refused, judged by the benchmark's own scorer.
0% 10% 20% 30% 0.47 0.49 0.51 0.53 0.55 0.57 gpt-4o-mini c4ai-command-r+ gpt-4o qwen-2.5-7b qwen-2.5-72b mixtral-8x22b llama-3.1-8b llama-3.1-70b llama-3.1-405b RE-call 29.1%, second of ten SAME FAMILY, MORE CAPABLE llama-3.1-8b 32.7% llama-3.1-70b 29.1% llama-3.1-405b 5.5%
CORRECT REFUSALS, OF 55 UNANSWERABLE
END TO END ANSWER QUALITY, HARMONIC MEAN
Every baseline recomputed from the MTRAG release through one harness, so these are anchored comparisons and not the published leaderboard. On MTRAG the refusal rate is set by the generator prompt, not by retrieval.
github.com/GiulioDER/RE-call