SpeciEval (Refusal-Aware Retries)
Safety · 2026-09-30
SpeciEval animal-attitude composite under the refusal-aware retry protocol: ten epochs, up to fifteen additional generations for unparseable/refused answers, refusal-aware epoch means and sample metrics. Twelve designated Likert items contribute to the published 0–100 headline. Language and valid-response coverage are not disclosed per model. Distinct from the earlier treatment of missing answers.
Top models (higher is better)
| Model | Score |
|---|---|
| Hy3 preview | 100.0 |
| Gemini 2.5 Pro | 99.7 |
| GPT-5.6 Sol Pro | 99.2 |
| GPT-5.6 Sol | 99.0 |
| DeepSeek-V4-Flash-0731 | 98.8 |
| Muse Spark 1.3 | 98.8 |
| GPT-6 Sol | 98.5 |
| GPT-5.5 | 98.2 |
| GPT-5.6 Terra | 98.1 |
| GPT-5.6 Terra Pro | 97.8 |
| DeepSeek V4.1 Flash | 97.4 |
| GPT-6.1 Sol | 97.4 |
| Inkling | 97.1 |
| GPT-5.1 | 96.9 |
| Qwen3 Max (rolling alias) | 96.9 |
| GPT-5 Chat (2025-08-07) | 96.8 |
| GPT-6 Astra | 96.5 |
| GPT-4.1 | 96.5 |
| Grok 4.7 | 96.4 |
| o4-mini-deep-research (2025-06-26) | 96.4 |
| GPT-5 Pro | 96.1 |
| GLM 4.6 | 95.8 |
| Llama 3.3 70B Instruct | 95.6 |
| Nemotron 3 Ultra 550B A55B | 95.6 |
| Grok 4.20 | 95.4 |
| Qwen3.7 Flash | 95.4 |
| GPT-5 | 95.3 |
| Grok 4 | 95.3 |
| Hy4 Preview | 94.7 |
| GPT-6 Luna | 94.6 |
| Nova Lite | 94.4 |
| Qwen3.8 Flash | 94.4 |
| Kimi K2.5 | 94.3 |
| Kimi K2.6 | 94.3 |
| Kimi K2 Instruct 0905 | 94.2 |
| Muse Spark 1.2 | 94.0 |
| Gemini 2.5 Flash-Lite | 93.9 |
| Muse Spark 1.1 | 93.9 |
| Grok Code Fast 1 | 93.7 |
| Grok 3 Mini | 93.6 |