Atlas

Benchmarks

← All benchmarks

Artificial Analysis Omniscience Accuracy

General QA · 2025-11-16

Artificial Analysis Omniscience Accuracy measures how often a model answers fact-checking or knowledge questions correctly in the Omniscience evaluation. It is paired with a hallucination-rate metric.

Top models (higher is better)

ModelScore
Claude Fable 561.4
GPT-5.6 Sol58.5
GPT-5.556.9
Gemini 3 Pro Preview55.9
Gemini 3.1 Pro Preview55.3
Claude Opus 554.2
Gemini 3 Flash Preview54.0
Grok 4.552.0
Gemini 3.5 Flash51.9
GPT-5.3-Codex51.8
Grok Build 0.151.4
Gemini 3.6 Flash50.2
GPT-5.450.0
Opus 4.846.6
Opus 4.646.4
Kimi K346.0
GPT-5.6 Terra45.9
Opus 4.745.8
Opus 4.545.7
GPT-5.5 Instant45.5
Muse Spark44.6
GPT-5.5 Instant (2026-06-25 hosted snapshot)44.1
GPT-5.243.8
DeepSeek-V4-Pro43.3
GPT-5.6 Luna41.5
Grok 441.4
GPT-5.2-Codex40.7
GPT-540.6
Muse Spark 1.140.6
Sonnet 4.640.0
Inkling40.0
GPT-5.1-Codex39.2
Gemini 2.5 Pro39.0
GPT-5-Codex38.7
Kimi K2.7 Code38.6
o338.4
DeepSeek-V3.2-Speciale38.4
Sonnet 538.3
Qwen3.6-Max-Preview37.7
GPT-5.137.6
Loading Atlas data…