Confabulations Leaderboard Non-Response Rate
Safety · 2024-10-10
This metric is the percentage of leaderboard prompts that receive no substantive response. Lower is better, and it contributes 50% of the published weighted score.
Top models (lower is better)
| Model | Score |
|---|---|
| MiniMax-Text-01 | 3.3 |
| ChatGPT-4o Latest observed 2025-03-27 | 3.5 |
| o3 | 4.0 |
| GPT-5 Mini | 4.8 |
| O4 Mini | 4.8 |
| o3-pro | 5.1 |
| Nova Pro | 5.6 |
| Mistral Medium 3 | 5.7 |
| QwQ-32B | 5.9 |
| Qwen2.5 72B Instruct | 6.0 |
| o3-mini | 6.2 |
| Phi-4 | 6.4 |
| ChatGPT-4o Latest observed 2025-01-29 | 6.5 |
| ERNIE 4.5 300B A47B | 6.7 |
| Gemma 2 27B | 7.2 |
| Haiku 3.5 | 7.6 |
| o1 Preview | 7.8 |
| DeepSeek-R1 | 8.0 |
| Qwen3-235B-A22B | 8.0 |
| gpt-oss-120b | 8.0 |
| GPT-4o (2024-11-20) | 8.2 |
| GPT-4o (2024-08-06) | 8.4 |
| GPT-5 | 9.8 |
| Gemini 2.0 Flash Thinking Experimental 01-21 | 10.0 |
| Gemini 1.5 Pro 002 | 10.2 |
| Grok 3 | 10.6 |
| Kimi K2 Instruct | 10.6 |
| Mistral Large 2 (Instruct 2407) | 10.6 |
| o1-mini | 10.9 |
| Claude 3 Haiku | 11.5 |
| Qwen3-30B-A3B | 11.7 |
| Mistral Small 3 24B Instruct 2501 | 11.8 |
| Qwen2.5-Max | 12.4 |
| O1 | 12.6 |
| DeepSeek-V3-0324 | 13.2 |
| GPT-4o Mini | 13.5 |
| Gemma 3 27B IT | 14.2 |
| Claude 3.7 Sonnet | 14.3 |
| Grok 2 | 14.5 |
| Grok 3 Mini | 14.7 |