Atlas

Benchmarks

← All benchmarks

EQ-Bench 3 (Rubric Score)

Chat & Writing · 2025-04-28

EQ-Bench 3 aggregate rubric score, a separate absolute 0-100 score from judge-scored scenario criteria.

Top models (higher is better)

ModelScore
Opus 4.883.3
Claude Fable 582.6
Opus 4.782.6
GPT-5.482.4
GPT-5.582.4
GPT-5.280.4
Sonnet 4.680.0
Hivemind-32B-Preview79.6
Opus 4.679.2
GLM-5.278.7
GPT-5.178.1
DeepSeek-V4-Pro77.6
Opus 4.576.8
Horizon Alpha76.8
Kimi K2.676.6
Gemini 3 Pro Preview75.8
GLM-5.175.2
GLM-575.0
Kimi K2 Instruct75.0
o374.7
Gemini 3.1 Pro Preview74.3
GPT-5 Chat (2025-08-07)73.7
DeepSeek-V4-Flash72.5
Kimi K2.572.5
Sonnet 4.572.2
Gemini 2.5 Pro Preview 06-0572.2
Qwen3.5 397B A17B71.3
Gemma 4 31B IT70.8
GPT-5.3 Instant70.7
Gemma 4 26B A4B IT70.0
GLM-4.768.9
Opus 468.2
ChatGPT-4o Latest observed 2025-03-2767.8
GLM-4.567.4
Gemini 2.5 Pro Preview 03-2567.3
ChatGPT-4o (2025-04-25 update)66.7
Sonnet 466.6
DeepSeek-R165.8
Hermes 4 405B65.2
Grok 4.1 Fast65.1
Loading Atlas data…