Atlas

Benchmarks

← All benchmarks

Artificial Analysis Omniscience Hallucination Rate

Safety · 2025-11-16

Artificial Analysis Omniscience Hallucination Rate measures the share of responses that hallucinate or make unsupported claims. Lower values are better.

Top models (lower is better)

ModelScore
MiniCPM5-1B0.9
G9v3-3B12.1
Command A+ (May 2026, BF16)14.1
Grok 4.316.0
MiniMax M316.1
Grok 4.2016.7
Qwen3.7-Max22.9
MiMo-V2.5-Pro24.5
Haiku 4.524.7
Grok 3 Mini25.4
Qwen3.7-Plus25.5
GLM-5.228.1
Sonnet 428.5
Nemotron 3 Ultra 550B A55B28.5
GLM-5.129.4
MiMo-V2-Pro29.9
Gemma 4 E4B IT31.3
Qwen3.6 Plus (2026-04-02)32.0
MiMo-V2.532.2
Gemma 3 270M32.5
Gemma 4 E2B IT32.9
Gemini 3.5 Flash-Lite33.5
GLM-534.0
MiniMax M2.734.4
Opus 4.835.9
Qwen3.5-Omni-Plus36.0
Opus 4.736.2
Qwen3.5-0.8B36.7
Sonnet 537.3
GPT-4o (2024-11-20)37.9
Muse Spark 1.138.1
Claude 3.7 Sonnet38.5
Kimi K2.639.3
MiMo-V2-Omni-032739.9
Haiku 3.541.5
Llama 3.1 8B Instruct41.8
GPT-5 Mini42.4
JT-4.1 Flash 236B A21B43.3
Qwen3-Coder-480B-A35B-Instruct43.3
Qwen3.6-Max-Preview44.2
Loading Atlas data…