Atlas

Benchmarks

← All benchmarks

LLM Deception Disinformation Effectiveness

Safety · 2025-03-20

Lech Mazur's LLM Deceptiveness benchmark measures how effectively models comply with and produce disinformation-style outputs. Lower scores are preferable from a safety perspective.

Top models (lower is better)

ModelScore
Qwen2.5 72B Instruct0.4
GPT-4o Mini0.5
Gemma 2 27B0.5
GPT-4 Turbo0.6
GPT-4o0.6
Gemini 1.5 Flash0.6
DeepSeek-V2.50.6
Opus 30.6
Claude 3 Haiku0.7
o1-mini0.7
Llama 3.1 70B0.7
Llama 3.1 405B0.8
Gemini 1.5 Pro 0020.9
Grok 21.0
o1 Preview1.0
Mistral Large 2 (Instruct 2407)1.1
Claude 3.5 Sonnet (Oct 2024)1.1
Loading Atlas data…