Atlas

Benchmarks

← All benchmarks

Phare Self-Assessed Stereotypes (French)

Safety · 2025-03-27

Phare bias task testing resistance to self-assessed stereotypes and biased outputs. This row is the French language split.

Top models (higher is better)

ModelScore
GPT-4.1 Mini87.0
Grok 4 Fast79.6
Llama 4 Maverick77.5
Llama 3.1 405B Instruct77.1
Haiku 4.575.0
Mistral Small 3.2 24B Instruct 250673.3
Opus 4.569.5
Llama 4 Scout69.1
Mistral Large 3 675B Instruct 251265.2
DeepSeek-V364.0
Qwen3 8B61.8
Kimi K2.661.7
Qwen-Plus (2025-01-25)60.3
Mistral Medium 3.560.1
DeepSeek-V3.160.0
Gemini 2.5 Flash57.3
Magistral Medium 1.256.8
GLM-5.256.7
GPT-4.155.5
Qwen3.7-Plus54.7
Qwen3.7 Max Preview53.4
Llama 3.1 8B Instruct51.7
Grok 3 Mini51.0
Qwen2.5-Max50.4
Magistral Small 1.250.1
Qwen3 VL 30B A3B Instruct49.6
Gemini 2.5 Flash-Lite49.5
Gemini 3 Pro Preview49.3
Gemini 3.1 Pro Preview47.7
Command A47.7
Gemini 2.0 Flash47.2
GPT-4o46.0
Grok 4.345.5
Mistral Large 2 (Instruct 2407)45.5
Llama 3.3 70B Instruct45.2
Sonnet 4.544.1
Opus 4.144.0
Gemini 3.1 Flash-Lite43.2
DeepSeek-V3-032441.4
Mistral Medium 3.141.0
Loading Atlas data…