Atlas

Benchmarks

← All benchmarks

EQ-Bench 4: Meeting Preferences and Needs

Chat & Writing · 2026-07-23

EQ-Bench 4's independently fitted preferences-and-needs ability rating measures how well a model balances the user's stated preferences with broader needs. The source normalizes its per-dimension soft Bradley–Terry rating to a 1–10 scale.

Top models (higher is better)

ModelScore
Claude Opus 510.0
Claude Fable 59.2
GPT-5.59.1
Kimi K38.9
Opus 4.78.4
GPT-5.48.0
GPT-5.6 Sol7.9
Opus 4.87.6
GPT-5.6 Terra7.3
Muse Spark 1.17.2
Sonnet 56.6
Inkling6.5
Opus 4.66.3
GLM-5.26.1
MiMo-V2.5-Pro5.8
Sonnet 4.65.7
Kimi K2.65.7
GPT-5.6 Luna5.5
DeepSeek-V4-Pro4.8
Gemini 3.1 Pro Preview4.5
MiniMax M34.5
Gemma 4 31B IT3.8
Qwen3.7-Max3.8
Gemini 3.5 Flash3.4
Grok 4.33.2
Haiku 4.52.4
Qwen3.6 27B1.8
Mistral Medium 3.51.0
Loading Atlas data…