Atlas

Benchmarks

← All benchmarks

SpeciEval (Refusal-Aware Retries)

Safety · 2026-09-30

SpeciEval animal-attitude composite under the refusal-aware retry protocol: ten epochs, up to fifteen additional generations for unparseable/refused answers, refusal-aware epoch means and sample metrics. Twelve designated Likert items contribute to the published 0–100 headline. Language and valid-response coverage are not disclosed per model. Distinct from the earlier treatment of missing answers.

Top models (higher is better)

ModelScore
Hy3 preview100.0
Gemini 2.5 Pro99.7
GPT-5.6 Sol Pro99.2
GPT-5.6 Sol99.0
DeepSeek-V4-Flash-073198.8
Muse Spark 1.398.8
GPT-6 Sol98.5
GPT-5.598.2
GPT-5.6 Terra98.1
GPT-5.6 Terra Pro97.8
DeepSeek V4.1 Flash97.4
GPT-6.1 Sol97.4
Inkling97.1
GPT-5.196.9
Qwen3 Max (rolling alias)96.9
GPT-5 Chat (2025-08-07)96.8
GPT-6 Astra96.5
GPT-4.196.5
Grok 4.796.4
o4-mini-deep-research (2025-06-26)96.4
GPT-5 Pro96.1
GLM 4.695.8
Llama 3.3 70B Instruct95.6
Nemotron 3 Ultra 550B A55B95.6
Grok 4.2095.4
Qwen3.7 Flash95.4
GPT-595.3
Grok 495.3
Hy4 Preview94.7
GPT-6 Luna94.6
Nova Lite94.4
Qwen3.8 Flash94.4
Kimi K2.594.3
Kimi K2.694.3
Kimi K2 Instruct 090594.2
Muse Spark 1.294.0
Gemini 2.5 Flash-Lite93.9
Muse Spark 1.193.9
Grok Code Fast 193.7
Grok 3 Mini93.6
Loading Atlas data…