Atlas

Benchmarks

← All benchmarks

LLM Deception Vulnerability Score

Safety · 2025-03-20

Lech Mazur's LLM Deceptiveness benchmark also ranks models by how vulnerable they are to disinformation attempts. Lower vulnerability scores indicate stronger resistance.

Top models (lower is better)

ModelScore
Opus 30.3
Claude 3.5 Sonnet (Oct 2024)0.3
o1 Preview0.3
Mistral Large 2 (Instruct 2407)0.4
o1-mini0.5
Llama 3.1 405B0.5
Qwen2.5 72B Instruct0.6
GPT-4o0.6
Gemini 1.5 Pro 0020.6
Gemini 1.5 Flash0.6
Claude 3 Haiku0.8
Grok 20.8
Llama 3.1 70B0.8
GPT-4o Mini1.1
GPT-4 Turbo1.2
DeepSeek-V2.51.4
Gemma 2 27B1.4
Loading Atlas data…