Atlas

Benchmarks

← All benchmarks

FACTS Benchmark Suite

General QA · 2025-12-11

The FACTS Benchmark Suite is Google DeepMind's aggregate factuality score, the unweighted average of four component leaderboards: FACTS Grounding, FACTS Parametric, FACTS Search, and FACTS Multimodal. It is a later and broader construct than FACTS Grounding alone, which it contains as one component, and the two are not interchangeable.

Top models (higher is better)

ModelScore
Gemini 3 Pro Preview70.5
Gemini 3.1 Pro Preview67.3
GPT-5.566.6
Gemini 3.6 Flash65.3
Gemini 3 Flash Preview63.4
Gemini 3.5 Flash62.3
GPT-5.261.4
GPT-5 (2025-08-07)56.5
Opus 4.853.5
Gemini 3.5 Flash-Lite53.5
GPT-5.6 Sol52.4
GPT-5.6 Terra52.0
Gemini 3.1 Flash-Lite Preview51.4
GPT-5.150.8
Grok 4.550.2
GPT-5.2 (2025-12-11)48.5
Grok 446.3
GPT-4.1 (2025-04-14)45.8
Opus 4.745.4
Gemma 4 31B IT45.3
O3 (2025-04-16)44.4
Opus 4.543.9
Gemma 4 26B A4B IT42.7
Grok 4.1 Fast42.1
Opus 4.641.4
Sonnet 4.540.6
Opus 4.140.4
GPT-5 Mini (2025-08-07)38.6
Sonnet 4.637.0
Sonnet 435.0
GPT-5.6 Luna34.8
GPT-5 Mini33.7
Gemma 3 27B IT29.1
O4 Mini (2025-04-16)28.1
Grok 4 Fast25.5
Haiku 4.518.6
Gemini 2.5 Flash-Lite17.9
Loading Atlas data…