Atlas

Benchmarks

← All benchmarks

Confabulations Leaderboard Confabulation Rate

Safety · 2024-10-10

This metric is the percentage of leaderboard answers judged to be confabulated. Lower is better, and it contributes 50% of the published weighted score.

Top models (lower is better)

ModelScore
Sonnet 42.5
Opus 42.5
Opus 4.12.5
Gemini 2.5 Pro Preview 03-254.0
Gemini 2.5 Pro4.0
Grok 44.0
Gemini 2.5 Flash Preview 04-174.5
Gemini 2.5 Pro Preview 05-065.9
Grok 3 Mini6.9
GLM-4.57.9
Claude 3.7 Sonnet7.9
GPT-510.9
O110.9
GPT-4.511.9
Qwen3-30B-A3B12.9
DeepSeek-R1-052812.9
Claude 3.5 Sonnet (Oct 2024)12.9
Qwen3 235B A22B Thinking 250713.9
Llama 3.1 405B14.4
Gemini 2.0 Flash Thinking Experimental 01-2114.9
Gemini 2.0 Pro Experimental 02-0515.8
Gemini 1.5 Pro 00216.8
DeepSeek-R117.3
Grok 317.8
Llama 3.3 70B Instruct17.8
o1 Preview18.3
GPT-5 Mini21.8
GPT-4o (2024-08-06)22.3
Qwen3-235B-A22B23.3
gpt-oss-120b23.3
o3-pro23.4
Gemini 2.0 Flash24.3
o324.8
QwQ-32B25.2
ERNIE 4.5 300B A47B25.2
Grok 225.7
GPT-4o (2024-11-20)26.2
o1-mini26.2
O4 Mini26.7
ChatGPT-4o Latest observed 2025-01-2926.7
Loading Atlas data…