Atlas

Benchmarks

← All benchmarks

Artificial Analysis Omniscience Hallucination Rate

Safety · 2025-11-16

Artificial Analysis Omniscience Hallucination Rate is reported over the evaluation's 6,000 questions and measures the share of responses that hallucinate or make unsupported claims. Lower values are better.

Top models (lower is better)

ModelScore
MiniCPM5-1B0.9
G9v3-3B12.1
Command A+ (May 2026, BF16)14.1
Grok 4.316.0
MiniMax M316.1
Grok 4.2016.7
Qwen3.7-Max22.9
MiMo-V2.5-Pro24.5
Haiku 4.524.7
Grok 3 Mini25.4
Qwen3.7-Plus25.5
GLM-5.228.1
Sonnet 428.5
Nemotron 3 Ultra 550B A55B28.5
GLM-5.129.4
MiMo-V2-Pro29.9
Gemma 4 E4B IT31.3
Qwen3.6 Plus (2026-04-02)32.0
MiMo-V2.532.2
Gemma 3 270M32.5
Gemma 4 E2B IT32.9
Gemini 3.5 Flash-Lite33.5
GLM-534.0
MiniMax M2.734.4
Opus 4.835.9
Qwen3.5-Omni-Plus36.0
Opus 4.736.2
Qwen3.5-0.8B36.7
Sonnet 537.3
GPT-4o (2024-11-20)37.9
Muse Spark 1.138.1
Claude 3.7 Sonnet38.5
Kimi K2.639.3
MiMo-V2-Omni-032739.9
Haiku 3.541.5
Llama 3.1 8B Instruct41.8
GPT-5 Mini42.4
JT-4.1 Flash 236B A21B43.3
Qwen3-Coder-480B-A35B-Instruct43.3
Qwen3.6-Max-Preview44.2
Loading Atlas data…