Atlas

Benchmarks

← All benchmarks

Confabulations Leaderboard Non-Response Rate

Safety · 2024-10-10

This metric is the percentage of leaderboard prompts that receive no substantive response. Lower is better, and it contributes 50% of the published weighted score.

Top models (lower is better)

ModelScore
MiniMax-Text-013.3
ChatGPT-4o Latest observed 2025-03-273.5
o34.0
GPT-5 Mini4.8
O4 Mini4.8
o3-pro5.1
Nova Pro5.6
Mistral Medium 35.7
QwQ-32B5.9
Qwen2.5 72B Instruct6.0
o3-mini6.2
Phi-46.4
ChatGPT-4o Latest observed 2025-01-296.5
ERNIE 4.5 300B A47B6.7
Gemma 2 27B7.2
Haiku 3.57.6
o1 Preview7.8
DeepSeek-R18.0
Qwen3-235B-A22B8.0
gpt-oss-120b8.0
GPT-4o (2024-11-20)8.2
GPT-4o (2024-08-06)8.4
GPT-59.8
Gemini 2.0 Flash Thinking Experimental 01-2110.0
Gemini 1.5 Pro 00210.2
Grok 310.6
Kimi K2 Instruct10.6
Mistral Large 2 (Instruct 2407)10.6
o1-mini10.9
Claude 3 Haiku11.5
Qwen3-30B-A3B11.7
Mistral Small 3 24B Instruct 250111.8
Qwen2.5-Max12.4
O112.6
DeepSeek-V3-032413.2
GPT-4o Mini13.5
Gemma 3 27B IT14.2
Claude 3.7 Sonnet14.3
Grok 214.5
Grok 3 Mini14.7
Loading Atlas data…