How well do open-source vision-language models resist adversarial attacks? We run every model through the full Zortium battery — 18 attack suites across the visual and text channels — and score each one by Attack Success Rate. Lower is safer.
Each cell is one attack suite run against one model — darker means more of its test cases succeeded. Hover any cell for the exact rate. Diagnostic suites are hatched: they measure robustness drift, not jailbreaks, so they are never scored.
Weighted ASR is a severity-weighted score across all scored suites — a breach on a HIGH suite counts 3× a LOW one (HIGH×3, MID×2, LOW×1). Diagnostic suites (robustness probes, not jailbreaks) are shown in each breakdown but never affect the score. Click any row to see the per-suite breakdown.
Each model is pointed at through its OpenRouter endpoint and run through the identical set of 18 suites in fast mode — no per-model tuning, no cherry-picked prompts.
Per suite, ASR is the fraction of harmful cases the model complied with. Critical suites are graded by an LLM judge; the rest use refusal-aware matching. 0% means it resisted every case.
Every suite carries a severity tier — High Mid Low Diag — so a breach on a high-impact attack reads differently from a diagnostic drift signal.
Point Zortium at any OpenAI-compatible endpoint and get the same 18-suite report for your own model in minutes.