Atlas

Benchmarks

← All benchmarks

Phare Factuality (French)

Safety · 2025-02-19

Phare task testing factual answer reliability and hallucination resistance. This row is the French language split.

Top models (higher is better)

ModelScore
Gemini 3.1 Pro Preview82.3
GPT-5.578.9
Gemini 3.5 Flash78.2
Gemini 3 Pro Preview76.5
Claude 3.5 Sonnet (Oct 2024)73.8
GPT-573.5
Grok 373.5
Grok 473.5
GPT-5.272.1
Opus 4.571.9
Opus 4.671.8
Gemini 2.5 Pro71.8
GPT-4.171.1
Claude 3.7 Sonnet70.8
GPT-4o70.8
GPT-5.170.8
Kimi K2.570.8
Gemini 3.1 Flash-Lite69.7
Kimi K2.669.7
Sonnet 568.4
DeepSeek-V4-Pro68.4
Grok 4.368.4
Sonnet 4.668.3
DeepSeek-V3-032468.0
DeepSeek-V4-Flash68.0
Opus 4.167.3
Gemini 1.5 Pro67.2
Gemini 2.0 Flash67.0
Qwen3.7 Max Preview67.0
GLM-5.266.3
DeepSeek-V366.0
Sonnet 4.565.5
Mistral Large 2 (Instruct 2407)64.3
Qwen3.7-Plus62.9
Mistral Large 3 675B Instruct 251262.6
DeepSeek-R1-052862.2
GPT-5 Mini62.2
Mistral Medium 3.162.1
Grok 3 Mini61.9
Gemini 2.5 Flash61.6
Loading Atlas data…