Atlas

Benchmarks

← All benchmarks

EQ-Bench 4

Chat & Writing · 2026-07-23

EQ-Bench 4 assesses applied emotional and social intelligence across 120 eight-turn role-play chats with simulated personas. Blind pairwise judgments on six ability dimensions are combined with a soft Bradley–Terry model and reported as anchored Elo ratings with bootstrapped confidence intervals.

Top models (higher is better)

ModelScore
Claude Opus 51385
Claude Fable 51340
Kimi K31339
GPT-5.51315
Opus 4.71311
Opus 4.81281
GPT-5.41272
Muse Spark 1.11260
GPT-5.6 Sol1250
Sonnet 51236
GPT-5.6 Terra1234
Inkling1226
Opus 4.61223
GLM-5.21222
MiMo-V2.5-Pro1208
Sonnet 4.61207
Kimi K2.61202
DeepSeek-V4-Pro1166
GPT-5.6 Luna1156
MiniMax M31150
Gemini 3.1 Pro Preview1142
Gemma 4 31B IT1120
Qwen3.7-Max1110
Gemini 3.5 Flash1087
Grok 4.31075
Haiku 4.51064
Qwen3.6 27B1026
Mistral Medium 3.5993
Loading Atlas data…