Atlas

Benchmarks

← All benchmarks

SpeciEval Speciesism (Refusal-Aware Retries)

Safety · 2026-09-30

Published mean 1–7 speciesism rating under SpeciEval's refusal-aware retry protocol. Four nominal prompts, ten epochs, up to fifteen additional generations for refused/unparseable responses. Valid counts and per-model language coverage are undisclosed; lower ratings indicate more animal-friendly attitudes.

Top models (lower is better)

ModelScore
Hy3 preview1.0
Gemini 2.5 Pro1.1
GPT-5.6 Sol Pro1.1
GPT-5.6 Sol1.2
DeepSeek V4.1 Flash1.2
Qwen3 30B A3B Thinking 25071.2
DeepSeek-V4-Flash-07311.2
Muse Spark 1.31.2
GPT-5.6 Terra1.2
GPT-6 Sol1.3
GPT-5.6 Terra Pro1.3
GPT-4.11.3
GPT-5.51.3
GPT-5.11.4
GPT-5 Chat (2025-08-07)1.4
o4-mini-deep-research (2025-06-26)1.4
Kimi K2 Instruct1.4
GLM 4.61.4
GPT-5 Pro1.4
GPT-6.1 Sol1.5
Inkling1.5
Qwen3.8 Flash1.5
GPT-5.2 Pro1.5
Qwen3 30B A3B Instruct 25071.5
Qwen3 Max (rolling alias)1.5
Llama 3.3 70B Instruct1.5
GPT-5.21.6
Hy4 Preview1.6
GLM-4.51.6
GPT-51.6
Nova Lite1.6
Kimi K2 Instruct 09051.6
MiniMax M2.71.6
GPT-6 Astra1.6
Grok 4.71.6
Grok 41.6
Qwen3.7 Flash1.7
Grok 3 Mini1.7
Grok 4.201.7
GLM 4.5 Air1.7
Loading Atlas data…