Atlas

Benchmarks

← All benchmarks

SpeciEval (Pooled English Runs)

Safety · 2026-10-02

SpeciEval animal-attitude composite pooling successful English runs with original wording by question. The October 4 source correction excludes negated-wording runs. Refused or unparseable answers remain missing after up to fifteen retries. All twelve designated Likert question means must be present; speciesism and necessary-meat items are reversed and the sum is rescaled to 0–100. Published bootstrap intervals resample epochs within fixed questions. Pooled epoch and valid-response counts are undisclosed.

Top models (higher is better)

ModelScore
Hy3 preview100.0
Gemini 2.5 Pro99.7
GPT-5.6 Sol Pro99.2
GPT-5.6 Sol99.0
DeepSeek-V4-Flash-073198.7
Muse Spark 1.398.7
GPT-6 Sol98.5
Qwen3 Max (rolling alias)98.3
GPT-5.598.2
GPT-5.6 Terra98.1
GPT-5.6 Terra Pro97.8
DeepSeek V4.1 Flash97.4
GPT-6.1 Sol97.4
Inkling97.1
GPT-5.196.9
GPT-5 Chat (2025-08-07)96.8
GPT-6 Astra96.5
GPT-4.196.5
Grok 4.796.4
o4-mini-deep-research (2025-06-26)96.4
GPT-5 Pro96.1
GLM 4.695.8
Llama 3.3 70B Instruct95.6
Nemotron 3 Ultra 550B A55B95.6
Grok 4.2095.4
Qwen3.7 Flash95.4
GPT-595.3
Grok 495.3
Hy4 Preview94.7
GPT-6 Luna94.6
Nova Lite94.4
Qwen3.8 Flash94.4
Kimi K2.594.3
Kimi K2.694.3
Kimi K2 Instruct 090594.2
Muse Spark 1.294.0
Gemini 2.5 Flash-Lite93.9
Muse Spark 1.193.9
Grok Code Fast 193.7
Grok 3 Mini93.6
Loading Atlas data…