Atlas

Benchmarks

← All benchmarks

Phare Bias Resistance

Safety · 2025-05-16

Phare module score for resistance to biased or stereotyped responses. Higher is better.

Top models (higher is better)

ModelScore
GPT-4.1 Mini88.1
Grok 4 Fast80.3
Llama 3.1 405B Instruct75.2
Mistral Small 3.2 24B Instruct 250673.9
Llama 4 Maverick73.7
Haiku 4.570.7
Llama 4 Scout67.1
DeepSeek-V3.165.2
Opus 4.563.2
Mistral Large 3 675B Instruct 251262.7
Mistral Medium 3.561.0
DeepSeek-V359.8
Qwen3 8B58.6
Qwen-Plus (2025-01-25)55.7
DeepSeek-V3-032455.3
Qwen3.7-Plus55.0
Gemini 3 Pro Preview53.6
Qwen3 VL 30B A3B Instruct53.6
Gemini 2.0 Flash53.5
Kimi K2.653.1
GPT-4.152.4
Qwen3.7 Max Preview52.4
Gemini 2.5 Flash51.7
GLM-5.251.3
GPT-4o50.9
Magistral Medium 1.250.8
Sonnet 4.549.1
Gemini 3.1 Pro Preview48.1
Magistral Small 1.248.0
Grok 4.347.0
GPT-5.146.8
GPT-5 Mini46.4
Grok 3 Mini46.4
Command A45.6
Gemini 2.5 Flash-Lite45.5
Llama 3.3 70B Instruct45.4
Qwen3 Max (rolling alias)44.8
Gemini 3.1 Flash-Lite44.6
Llama 3.1 8B Instruct44.1
Opus 4.143.6
Loading Atlas data…