Atlas

Benchmarks

← All benchmarks

Phare Misinformation (Spanish)

Safety · 2025-02-19

Phare task testing resistance to misinformation generation or endorsement. This row is the Spanish language split.

Top models (higher is better)

ModelScore
Haiku 4.591.2
Claude 3.7 Sonnet87.8
Sonnet 584.3
DeepSeek-V4-Flash83.7
GPT-5 Nano83.0
Opus 4.581.0
Claude 3.5 Sonnet (Oct 2024)79.6
Sonnet 4.579.6
Magistral Medium 1.279.6
GPT-4.1 Nano78.2
Qwen-Plus (2025-01-25)75.5
Gemini 1.5 Pro74.8
GPT-5 Mini74.8
Opus 4.174.2
Opus 4.674.2
Haiku 3.570.8
Llama 3.1 8B Instruct69.4
Mistral Large 2 (Instruct 2407)69.4
Gemini 2.5 Flash68.0
Gemini 2.5 Flash-Lite66.7
Gemini 3.1 Pro Preview66.7
Llama 3.1 405B Instruct66.7
Sonnet 4.666.0
Kimi K2.666.0
GPT-4o65.3
Grok 4.362.6
Qwen3.7 Max Preview61.9
Gemini 3.1 Flash-Lite60.5
Command A59.9
Llama 4 Maverick59.9
Gemma 4 31B IT59.2
Gemini 3 Pro Preview57.8
GPT-5.157.1
Grok 3 Mini57.1
Gemini 2.5 Pro56.5
Mistral Small 3.1 24B Instruct 250356.5
Qwen2.5-Max56.5
Grok 453.7
Gemini 2.0 Flash53.1
Kimi K2.553.1
Loading Atlas data…