Atlas

Benchmarks

← All benchmarks

Phare Misinformation (English)

Safety · 2025-02-19

Phare task testing resistance to misinformation generation or endorsement. This row is the English language split.

Top models (higher is better)

ModelScore
Haiku 4.599.1
Opus 4.195.3
Claude 3.5 Sonnet (Oct 2024)93.2
Sonnet 4.591.9
Claude 3.7 Sonnet91.6
Opus 4.591.0
Sonnet 591.0
Llama 3.1 405B Instruct91.0
Opus 4.688.5
Haiku 3.587.0
Kimi K2.685.7
Command A84.8
Qwen-Plus (2025-01-25)84.5
Gemini 3.1 Pro Preview83.8
Gemma 4 31B IT83.2
Magistral Medium 1.283.2
Grok 482.0
Llama 3.3 70B Instruct81.7
GPT-5 Mini79.2
Sonnet 4.678.9
GPT-4.1 Nano78.9
Gemini 1.5 Pro78.6
Gemini 2.5 Pro77.6
GPT-5.177.0
GPT-5 Nano77.0
Gemini 2.0 Flash-Lite76.7
Qwen3 Max (rolling alias)76.7
Gemini 2.5 Flash-Lite76.4
GPT-4o75.2
Mistral Large 2 (Instruct 2407)75.2
Gemini 2.0 Flash74.5
Gemini 2.5 Flash74.2
Llama 4 Maverick73.9
Gemini 3.1 Flash-Lite73.3
Kimi K2.572.4
Grok 3 Mini71.4
Llama 3.1 8B Instruct70.5
Gemini 3.5 Flash69.6
GPT-4.168.9
DeepSeek-V3-032468.6
Loading Atlas data…