Atlas

Benchmarks

← All benchmarks

EQ-Bench 4: Emotion Sensemaking

Chat & Writing · 2026-07-23

EQ-Bench 4's independently fitted emotion-sensemaking ability rating measures how well a model helps the user understand their emotional experience. The source normalizes its per-dimension soft Bradley–Terry rating to a 1–10 scale.

Top models (higher is better)

ModelScore
Claude Opus 510.0
Kimi K39.4
Claude Fable 59.2
Opus 4.88.0
Opus 4.77.9
GPT-5.57.6
Opus 4.66.8
Sonnet 56.6
GPT-5.46.6
GLM-5.26.5
Muse Spark 1.16.5
Sonnet 4.66.5
Inkling6.4
Kimi K2.66.4
GPT-5.6 Sol6.3
MiMo-V2.5-Pro6.1
GPT-5.6 Terra5.6
DeepSeek-V4-Pro4.8
MiniMax M34.7
GPT-5.6 Luna4.4
Haiku 4.54.1
Gemini 3.1 Pro Preview3.7
Gemma 4 31B IT3.7
Qwen3.7-Max3.5
Gemini 3.5 Flash3.1
Qwen3.6 27B2.6
Grok 4.32.2
Mistral Medium 3.51.0
Loading Atlas data…