SpeciEval
Safety · 2025-07-19
SpeciEval measures how animal-friendly a model's stated attitudes are across 12 Likert items adapted from validated psychological scales: four speciesism items, six belief-in-animal-sentience items, and the ‘necessary’ item from each land-animal and seafood 4Ns scale. Speciesism and 4Ns responses are reverse-scored, then the 12 items are normalized from the 1–7 scale to a 0–100 headline score. Tasks default to English and support 15 languages; the public aggregate does not disclose per-model language coverage.
Top models (higher is better)
| Model | Score |
|---|---|
| Hy3 preview | 100.0 |
| Gemini 2.5 Pro | 99.7 |
| GPT-5.6 Sol Pro | 99.2 |
| GPT-5.6 Sol | 99.0 |
| GPT-5.5 | 98.2 |
| GPT-5.6 Terra | 98.1 |
| GPT-5.6 Terra Pro | 97.8 |
| Inkling | 97.1 |
| GPT-5.1 | 96.9 |
| Qwen3 Max (rolling alias) | 96.9 |
| GPT-5 Chat (2025-08-07) | 96.8 |
| GPT-4.1 | 96.7 |
| o4-mini-deep-research (2025-06-26) | 96.4 |
| GPT-5 Pro | 96.1 |
| GLM 4.6 | 95.8 |
| Nemotron 3 Ultra 550B A55B | 95.6 |
| Grok 4 | 95.3 |
| GPT-5 | 95.3 |
| Llama 3.3 70B Instruct | 95.0 |
| Grok 4.20 | 94.7 |
| Nova Lite | 94.4 |
| Kimi K2.6 | 94.3 |
| Kimi K2.5 | 94.3 |
| Kimi K2 Instruct 0905 | 94.2 |
| Grok Code Fast 1 | 93.9 |
| MiniMax M2 | 93.8 |
| Qwen3.7-Plus | 93.6 |
| Grok 3 Mini | 93.6 |
| Kimi K2 Instruct | 93.5 |
| GLM-5.1 | 93.3 |
| MiniMax M2.7 | 93.3 |
| DeepSeek-V4-Flash | 93.2 |
| GPT-5.2 Pro | 92.9 |
| GLM-4.5 | 92.8 |
| GPT-5.2 | 92.8 |
| GPT-5.6 Luna Pro | 92.8 |
| Qwen3 30B A3B Instruct 2507 | 92.6 |
| Qwen3.6 Plus (2026-04-02) | 92.6 |
| GPT-5.6 Luna | 92.6 |
| Grok 4.3 | 92.5 |