Atlas

Benchmarks

← All benchmarks

AA-Omniscience Index

Indexes · 2025-11-16

Factual recall and hallucination benchmark across economically relevant domains; higher scores mean more correct answers and fewer hallucinations.

Top models (higher is better)

ModelScore
Claude Fable 540.1
Gemini 3.1 Pro Preview32.9
Claude Opus 531.3
Opus 4.827.4
Grok 4.526.4
Opus 4.726.2
Gemini 3.6 Flash23.5
Gemini 3.5 Flash22.7
GPT-5.6 Sol21.7
GPT-5.520.1
Kimi K318.4
Grok 4.318.3
Muse Spark 1.118.0
Gemini 3 Pro Preview15.8
Grok 4.2015.3
Sonnet 515.3
Qwen3.7-Max14.1
Opus 4.613.5
Opus 4.513.3
Sonnet 4.612.4
Gemini 3 Flash Preview11.6
GPT-5.5 Instant10.3
Qwen3.6-Max-Preview10.2
GPT-5.3-Codex9.9
Grok Build 0.17.4
Gemini 3.5 Flash-Lite6.9
Kimi K2.66.4
GPT-5.45.7
GPT-5.15.5
MiMo-V2-Pro4.9
Muse Spark4.1
GLM-5.24.0
Grok 43.8
MiMo-V2.5-Pro3.6
GPT-5.5 Instant (2026-06-25 hosted snapshot)3.5
Qwen3.6 Plus (2026-04-02)2.6
Qwen3.7-Plus2.4
Inkling2.0
GLM-52.0
GLM-5.11.9
Loading Atlas data…