Atlas

Benchmarks

← All benchmarks

Phare Self-Assessed Stereotypes (English)

Safety · 2025-03-27

Phare bias task testing resistance to self-assessed stereotypes and biased outputs. This row is the English language split.

Top models (higher is better)

ModelScore
GPT-4.1 Mini89.0
Grok 4 Fast81.6
Mistral Small 3.2 24B Instruct 250673.3
Llama 4 Scout69.0
Llama 4 Maverick68.8
Llama 3.1 405B Instruct66.7
Mistral Large 3 675B Instruct 251263.4
Mistral Medium 3.563.2
Haiku 4.562.8
Gemini 3 Pro Preview61.5
Opus 4.561.4
Qwen3 8B60.9
DeepSeek-V360.8
DeepSeek-V3.160.4
Qwen3.7 Max Preview60.0
Sonnet 4.557.6
Gemini 2.0 Flash56.9
DeepSeek-V3-032455.7
Command A55.2
GPT-4o55.2
GLM-5.253.9
Gemini 2.0 Flash-Lite52.3
GPT-4.152.0
Gemini 3.1 Pro Preview50.5
Qwen3 VL 30B A3B Instruct50.4
GPT-5.150.1
GPT-5.549.3
Qwen3.7-Plus48.8
Qwen-Plus (2025-01-25)48.7
Gemini 3.1 Flash-Lite48.5
Gemini 2.5 Flash47.8
Grok 4.347.4
Sonnet 4.646.5
Grok 3 Mini46.1
Qwen3 Max (rolling alias)45.9
Sonnet 545.8
Opus 4.145.4
GPT-5 Mini44.0
Llama 3.3 70B Instruct42.9
Kimi K2.642.4
Loading Atlas data…