Atlas

Benchmarks

← All benchmarks

SpeciEval

Safety · 2025-07-19

SpeciEval measures how animal-friendly a model's stated attitudes are across 12 Likert items adapted from validated psychological scales: four speciesism items, six belief-in-animal-sentience items, and the ‘necessary’ item from each land-animal and seafood 4Ns scale. Speciesism and 4Ns responses are reverse-scored, then the 12 items are normalized from the 1–7 scale to a 0–100 headline score. Tasks default to English and support 15 languages; the public aggregate does not disclose per-model language coverage.

Top models (higher is better)

ModelScore
Hy3 preview100.0
Gemini 2.5 Pro99.7
GPT-5.6 Sol Pro99.2
GPT-5.6 Sol99.0
GPT-5.598.2
GPT-5.6 Terra98.1
GPT-5.6 Terra Pro97.8
Inkling97.1
GPT-5.196.9
Qwen3 Max (rolling alias)96.9
GPT-5 Chat (2025-08-07)96.8
GPT-4.196.7
o4-mini-deep-research (2025-06-26)96.4
GPT-5 Pro96.1
GLM 4.695.8
Nemotron 3 Ultra 550B A55B95.6
Grok 495.3
GPT-595.3
Llama 3.3 70B Instruct95.0
Grok 4.2094.7
Nova Lite94.4
Kimi K2.694.3
Kimi K2.594.3
Kimi K2 Instruct 090594.2
Grok Code Fast 193.9
MiniMax M293.8
Qwen3.7-Plus93.6
Grok 3 Mini93.6
Kimi K2 Instruct93.5
GLM-5.193.3
MiniMax M2.793.3
DeepSeek-V4-Flash93.2
GPT-5.2 Pro92.9
GLM-4.592.8
GPT-5.292.8
GPT-5.6 Luna Pro92.8
Qwen3 30B A3B Instruct 250792.6
Qwen3.6 Plus (2026-04-02)92.6
GPT-5.6 Luna92.6
Grok 4.392.5
Loading Atlas data…