Atlas

Benchmarks

← All benchmarks

SpeciEval Sea Animal 4Ns (Pooled English Runs)

Safety · 2026-10-02

Mean agreement that seafood is natural, necessary, normal or nice across four question means, pooling successful English SpeciEval runs with original wording; negated-wording runs are excluded. Only the necessary question enters the headline. Refusals remain missing after retries; pooled epoch and valid-response counts are undisclosed.

Top models (lower is better)

ModelScore
GPT-5.43.6
GPT-5.3 Instant3.7
GPT-6 Luna3.9
Nova Lite4.0
Nemotron 3.5 Lightning 30B A3B (precision unspecified)4.1
GPT-5.6 Luna Pro4.1
GPT-5.6 Terra Pro4.2
Inkling4.2
Qwen3.8 Flash4.2
GPT-5.6 Terra4.3
GPT-5 Pro4.3
GPT-5.6 Luna4.3
GPT-5.2 Pro4.3
GPT-5.14.3
GPT-5.24.3
GPT-5.2 Instant4.3
DeepSeek V4.1 Flash4.3
GPT-6 Astra4.3
Kimi K2.54.3
GPT-6.1 Sol4.4
Claude Opus 54.4
MiniMax M34.4
Claude Fable 5.14.4
GPT-5.6 Sol Pro4.4
GPT-54.4
Hy4 Preview4.4
GPT-5.6 Sol4.5
GPT-6 Sol4.5
Kimi K2.64.5
Opus 4.14.5
GLM-5.3 Flash4.5
Claude 3.7 Sonnet4.5
Gemini 2.5 Flash-Lite4.5
Qwen3 30B A3B Thinking 25074.5
Sonnet 44.5
Opus 44.5
Haiku 4.54.5
GPT-5.54.5
Nemotron 3 Ultra 550B A55B4.6
Muse Spark 1.24.6
Loading Atlas data…