Atlas

Benchmarks

← All benchmarks

Confabulations Leaderboard Weighted Score

Safety · 2024-10-10

Weighted score combining confabulation and non-response.

Top models (lower is better)

ModelScore
GPT-510.3
Gemini 2.5 Pro Preview 05-0610.6
Grok 3 Mini10.8
Gemini 2.5 Pro Preview 03-2510.8
GLM-4.511.3
O111.7
Qwen3-30B-A3B12.3
Gemini 2.5 Pro12.4
Grok 412.4
Gemini 2.0 Flash Thinking Experimental 01-2112.4
DeepSeek-R112.7
o1 Preview13.0
Sonnet 413.2
GPT-5 Mini13.3
Gemini 1.5 Pro 00213.5
GPT-4.513.6
Grok 314.2
o3-pro14.2
o314.4
DeepSeek-R1-052814.6
Claude 3.7 Sonnet14.7
GPT-4o (2024-08-06)15.3
QwQ-32B15.6
Qwen3-235B-A22B15.6
gpt-oss-120b15.7
O4 Mini15.8
Opus 415.9
ERNIE 4.5 300B A47B16.0
ChatGPT-4o Latest observed 2025-01-2916.6
Qwen3 235B A22B Thinking 250716.8
Gemini 2.5 Flash Preview 04-1716.8
Opus 4.117.1
GPT-4o (2024-11-20)17.2
Llama 3.1 405B17.6
o3-mini17.9
Gemini 2.0 Pro Experimental 02-0518.4
o1-mini18.6
Qwen2.5 72B Instruct19.1
Claude 3.5 Sonnet (Oct 2024)19.9
Grok 220.1
Loading Atlas data…