Atlas

Benchmarks

← All benchmarks

Phare Self-Assessed Stereotypes (Spanish)

Safety · 2025-03-27

Phare bias task testing resistance to self-assessed stereotypes and biased outputs. This row is the Spanish language split.

Top models (higher is better)

ModelScore
GPT-4.1 Mini88.3
Llama 3.1 405B Instruct81.8
Grok 4 Fast79.6
Mistral Small 3.2 24B Instruct 250675.2
DeepSeek-V3.175.1
Llama 4 Maverick74.6
Haiku 4.574.2
DeepSeek-V3-032468.7
Llama 4 Scout63.2
Qwen3.7-Plus61.4
Magistral Medium 1.261.0
Qwen3 VL 30B A3B Instruct60.7
Mistral Large 3 675B Instruct 251259.6
Mistral Medium 3.559.6
Opus 4.558.7
Qwen-Plus (2025-01-25)58.1
Gemini 2.0 Flash56.4
Kimi K2.655.2
DeepSeek-V354.4
GPT-5 Mini54.2
Qwen3 8B53.2
Magistral Small 1.251.8
GPT-4o51.6
Gemini 2.5 Flash-Lite51.1
GPT-5.150.7
Llama 3.1 8B Instruct50.4
Gemini 3 Pro Preview50.2
Gemini 2.5 Flash50.1
GPT-4.149.7
Qwen3 Max (rolling alias)49.1
Grok 4.348.1
Llama 3.3 70B Instruct48.1
Qwen2.5-Max48.0
Gemini 3.1 Pro Preview46.1
Sonnet 4.545.7
Qwen3.7 Max Preview43.8
GPT-5.243.6
GLM-5.243.2
Gemma 3 27B IT43.1
gpt-oss-120b42.8
Loading Atlas data…