Solutions · AI Teams → Solutions · Banking → Benchmarks → Integrations → Docs → Pricing → Sign in Get Early Access

OSS VLM Security Leaderboard

How well do open-source vision-language models resist adversarial attacks? We run every model through the full Zortium battery — 18 attack suites across the visual and text channels — and score each one by Attack Success Rate. Lower is safer.

18 attack suites 137–143 test cases per model Fast scan mode Tested via OpenRouter Updated Jul 2026
Resistant Partial Vulnerable

Weighted ASR is a severity-weighted score across all scored suites — a breach on a HIGH suite counts 3× a LOW one (HIGH×3, MID×2, LOW×1). Diagnostic suites (robustness probes, not jailbreaks) are shown in each breakdown but never affect the score. Click any row to see the per-suite breakdown.

How these numbers are produced

Same battery, every model

Each model is pointed at through its OpenRouter endpoint and run through the identical set of 18 suites in fast mode — no per-model tuning, no cherry-picked prompts.

Attack Success Rate

Per suite, ASR is the fraction of harmful cases the model complied with. Critical suites are graded by an LLM judge; the rest use refusal-aware matching. 0% means it resisted every case.

Severity, not just score

Every suite carries a severity tier — High Mid Low Diag — so a breach on a high-impact attack reads differently from a diagnostic drift signal.

Test your own model

Your model isn't on the list?

Point Zortium at any OpenAI-compatible endpoint and get the same 18-suite report for your own model in minutes.

Get Early AccessComing Soon See how it works