LLM Deception Vulnerability Score
Safety · 2025-03-20
Lech Mazur's LLM Deceptiveness benchmark also ranks models by how vulnerable they are to disinformation attempts. Lower vulnerability scores indicate stronger resistance.
Top models (lower is better)
| Model | Score |
|---|---|
| Opus 3 | 0.3 |
| Claude 3.5 Sonnet (Oct 2024) | 0.3 |
| o1 Preview | 0.3 |
| Mistral Large 2 (Instruct 2407) | 0.4 |
| o1-mini | 0.5 |
| Llama 3.1 405B | 0.5 |
| Qwen2.5 72B Instruct | 0.6 |
| GPT-4o | 0.6 |
| Gemini 1.5 Pro 002 | 0.6 |
| Gemini 1.5 Flash | 0.6 |
| Claude 3 Haiku | 0.8 |
| Grok 2 | 0.8 |
| Llama 3.1 70B | 0.8 |
| GPT-4o Mini | 1.1 |
| GPT-4 Turbo | 1.2 |
| DeepSeek-V2.5 | 1.4 |
| Gemma 2 27B | 1.4 |