Atlas

Benchmarks

← All benchmarks

Phare Factuality (English)

Safety · 2025-02-19

Phare task testing factual answer reliability and hallucination resistance. This row is the English language split.

Top models (higher is better)

ModelScore
GPT-5.590.8
Gemini 3.1 Pro Preview89.7
Opus 4.688.6
Gemini 3 Pro Preview88.6
Kimi K2.587.5
Gemini 3.5 Flash87.2
Opus 4.586.5
Kimi K2.686.5
DeepSeek-V4-Pro86.1
Sonnet 4.685.8
GPT-585.8
GPT-5.185.8
GPT-5.285.4
Claude 3.7 Sonnet85.0
Gemini 3.1 Flash-Lite85.0
Grok 484.7
Opus 4.184.0
Gemini 2.5 Pro84.0
GLM-5.284.0
Sonnet 4.583.6
GPT-4.183.6
Grok 4.383.3
Claude 3.5 Sonnet (Oct 2024)83.2
GPT-4o82.6
Qwen3.7 Max Preview82.6
Grok 381.5
Sonnet 580.8
DeepSeek-V4-Flash80.1
DeepSeek-R1-052880.0
Gemini 1.5 Pro79.4
Gemini 2.5 Flash79.4
Mistral Large 2 (Instruct 2407)79.4
GPT-5 Mini78.3
Grok 278.3
DeepSeek-V377.9
DeepSeek-V3-032477.9
Gemini 2.0 Flash77.9
Qwen3 Max (rolling alias)77.9
Mistral Large 3 675B Instruct 251277.6
Qwen2.5-Max77.6
Loading Atlas data…