Atlas

Benchmarks

← All benchmarks

SpeciEval Speciesism (Pooled English Runs)

Safety · 2026-10-02

Mean speciesism rating across four question means, pooling valid responses from successful English SpeciEval runs with original wording; negated-wording runs are excluded. Refusals remain missing after retries. Lower values indicate more animal-friendly attitudes. Pooled epoch and valid-response counts are undisclosed.

Top models (lower is better)

ModelScore
Hy3 preview1.0
Gemini 2.5 Pro1.1
GPT-5.6 Sol Pro1.1
GPT-5.6 Sol1.2
DeepSeek V4.1 Flash1.2
Qwen3 30B A3B Thinking 25071.2
DeepSeek-V4-Flash-07311.2
Muse Spark 1.31.2
GPT-5.6 Terra1.2
GPT-6 Sol1.3
Qwen3 Max (rolling alias)1.3
GPT-5.6 Terra Pro1.3
GPT-4.11.3
GPT-5.51.3
GPT-5.11.4
GPT-5 Chat (2025-08-07)1.4
o4-mini-deep-research (2025-06-26)1.4
Kimi K2 Instruct1.4
GLM 4.61.4
GPT-5 Pro1.5
GPT-6.1 Sol1.5
Inkling1.5
Qwen3.8 Flash1.5
GPT-5.2 Pro1.5
Qwen3 30B A3B Instruct 25071.5
Llama 3.3 70B Instruct1.5
GPT-5.21.6
Hy4 Preview1.6
GLM-4.51.6
GPT-51.6
Nova Lite1.6
Kimi K2 Instruct 09051.6
MiniMax M2.71.6
GPT-6 Astra1.6
Grok 4.71.6
Grok 41.7
Qwen3.7 Flash1.7
Grok 3 Mini1.7
GLM 4.5 Air1.7
Grok 4.201.7
Loading Atlas data…