Artificial Analysis Omniscience Hallucination Rate
Safety · 2025-11-16
Artificial Analysis Omniscience Hallucination Rate is reported over the evaluation's 6,000 questions and measures the share of responses that hallucinate or make unsupported claims. Lower values are better.
Top models (lower is better)
| Model | Score |
|---|---|
| MiniCPM5-1B | 0.9 |
| G9v3-3B | 12.1 |
| Command A+ (May 2026, BF16) | 14.1 |
| Grok 4.3 | 16.0 |
| MiniMax M3 | 16.1 |
| Grok 4.20 | 16.7 |
| Qwen3.7-Max | 22.9 |
| MiMo-V2.5-Pro | 24.5 |
| Haiku 4.5 | 24.7 |
| Grok 3 Mini | 25.4 |
| Qwen3.7-Plus | 25.5 |
| GLM-5.2 | 28.1 |
| Sonnet 4 | 28.5 |
| Nemotron 3 Ultra 550B A55B | 28.5 |
| GLM-5.1 | 29.4 |
| MiMo-V2-Pro | 29.9 |
| Gemma 4 E4B IT | 31.3 |
| Qwen3.6 Plus (2026-04-02) | 32.0 |
| MiMo-V2.5 | 32.2 |
| Gemma 3 270M | 32.5 |
| Gemma 4 E2B IT | 32.9 |
| Gemini 3.5 Flash-Lite | 33.5 |
| GLM-5 | 34.0 |
| MiniMax M2.7 | 34.4 |
| Opus 4.8 | 35.9 |
| Qwen3.5-Omni-Plus | 36.0 |
| Opus 4.7 | 36.2 |
| Qwen3.5-0.8B | 36.7 |
| Sonnet 5 | 37.3 |
| GPT-4o (2024-11-20) | 37.9 |
| Muse Spark 1.1 | 38.1 |
| Claude 3.7 Sonnet | 38.5 |
| Kimi K2.6 | 39.3 |
| MiMo-V2-Omni-0327 | 39.9 |
| Haiku 3.5 | 41.5 |
| Llama 3.1 8B Instruct | 41.8 |
| GPT-5 Mini | 42.4 |
| JT-4.1 Flash 236B A21B | 43.3 |
| Qwen3-Coder-480B-A35B-Instruct | 43.3 |
| Qwen3.6-Max-Preview | 44.2 |