Atlas

Benchmarks

← All benchmarks

Phare Misinformation (French)

Safety · 2025-02-19

Phare task testing resistance to misinformation generation or endorsement. This row is the French language split.

Top models (higher is better)

ModelScore
Haiku 4.596.3
Claude 3.7 Sonnet89.2
Sonnet 4.587.0
Opus 4.185.5
Opus 4.585.5
Claude 3.5 Sonnet (Oct 2024)83.5
Llama 3.1 8B Instruct82.6
GPT-5 Nano80.1
GPT-5 Mini78.6
Sonnet 577.9
Gemini 3.1 Pro Preview77.6
Opus 4.676.2
Sonnet 4.674.7
Qwen-Plus (2025-01-25)73.7
Gemini 2.5 Flash-Lite73.5
Kimi K2.673.0
GPT-4.1 Nano72.7
Haiku 3.571.5
DeepSeek-V4-Pro71.0
Gemini 1.5 Pro69.8
DeepSeek-V4-Flash69.3
GPT-5.168.8
Gemini 3.1 Flash-Lite67.3
Magistral Medium 1.267.3
Qwen3 Max (rolling alias)66.1
Qwen3.7 Max Preview65.6
Grok 4.365.4
Gemma 4 31B IT64.1
GPT-4o63.9
Gemini 3 Pro Preview63.4
DeepSeek-V3.163.1
Llama 3.1 405B Instruct63.1
Gemini 2.5 Flash62.9
Gemini 3.5 Flash62.6
Mistral Large 2 (Instruct 2407)61.7
Gemini 2.5 Pro61.4
Llama 4 Maverick61.2
Grok 3 Mini60.9
Gemini 2.0 Flash59.7
Kimi K2.559.7
Loading Atlas data…