LLM Deception Disinformation Effectiveness
Safety · 2025-03-20
Lech Mazur's LLM Deceptiveness benchmark measures how effectively models comply with and produce disinformation-style outputs. Lower scores are preferable from a safety perspective.
Top models (lower is better)
| Model | Score |
|---|---|
| Qwen2.5 72B Instruct | 0.4 |
| GPT-4o Mini | 0.5 |
| Gemma 2 27B | 0.5 |
| GPT-4 Turbo | 0.6 |
| GPT-4o | 0.6 |
| Gemini 1.5 Flash | 0.6 |
| DeepSeek-V2.5 | 0.6 |
| Opus 3 | 0.6 |
| Claude 3 Haiku | 0.7 |
| o1-mini | 0.7 |
| Llama 3.1 70B | 0.7 |
| Llama 3.1 405B | 0.8 |
| Gemini 1.5 Pro 002 | 0.9 |
| Grok 2 | 1.0 |
| o1 Preview | 1.0 |
| Mistral Large 2 (Instruct 2407) | 1.1 |
| Claude 3.5 Sonnet (Oct 2024) | 1.1 |