Atlas

Benchmarks

← All benchmarks

Phare Factuality (Spanish)

Safety · 2025-02-19

Phare task testing factual answer reliability and hallucination resistance. This row is the Spanish language split.

Top models (higher is better)

ModelScore
Gemini 3.1 Pro Preview85.7
Gemini 3 Pro Preview84.8
Gemini 3.5 Flash81.9
GPT-5.580.0
GPT-5.178.1
Gemini 2.5 Pro77.1
GPT-575.2
GPT-5.275.2
Opus 4.673.3
Grok 472.4
Gemini 3.1 Flash-Lite71.4
GLM-5.270.5
GPT-4.169.5
Sonnet 4.668.6
Qwen3.7 Max Preview68.6
Grok 367.6
Kimi K2.666.7
Opus 4.565.7
Opus 4.164.8
DeepSeek-V4-Pro64.8
Claude 3.5 Sonnet (Oct 2024)63.8
Grok 4.363.8
Claude 3.7 Sonnet62.9
DeepSeek-R1-052862.9
Kimi K2.562.9
Sonnet 4.561.0
DeepSeek-V4-Flash61.0
Gemini 2.0 Flash60.0
GPT-4o60.0
Mistral Large 3 675B Instruct 251259.0
Gemini 2.5 Flash58.1
DeepSeek-V357.1
DeepSeek-V3-032457.1
Sonnet 556.2
Qwen3 Max (rolling alias)56.2
Llama 4 Maverick55.2
Gemini 1.5 Pro53.3
Grok 3 Mini52.4
Mistral Medium 3.152.4
Mistral Large 2 (Instruct 2407)51.4
Loading Atlas data…