Atlas

Benchmarks

← All benchmarks

EQ-Bench 3: Pragmatic

Chat & Writing · 2025-04-28

Within EQ-Bench 3's emotional-intelligence scenarios, this subscore rates the pragmatic dimension of model responses on a 0–10 scale. Higher values are better.

Top models (higher is better)

ModelScore
Opus 4.78.8
GPT-5.48.7
Horizon Alpha8.6
Opus 4.88.5
GPT-5.58.5
GPT-5.18.2
GPT-5.28.2
Sonnet 4.68.1
Gemini 3 Pro Preview8.1
Claude Fable 58.0
Opus 4.68.0
Gemma 4 31B IT8.0
Hivemind-32B-Preview8.0
GLM-5.27.8
Gemma 4 26B A4B IT7.7
GPT-5 Chat (2025-08-07)7.7
Opus 4.57.6
o37.6
DeepSeek-V4-Pro7.5
Gemini 2.5 Pro Preview 06-057.5
GLM-57.5
Kimi K2 Instruct7.5
Gemini 2.5 Pro Preview 05-067.2
Gemini 3.1 Pro Preview7.2
Gemini 2.5 Pro Preview 03-257.1
Gemma 4 12B IT7.1
GLM-4.77.1
Qwen3.5 397B A17B7.1
DeepSeek-V4-Flash7.0
Hermes 4 405B7.0
ChatGPT-4o Latest observed 2025-03-276.8
GLM-5.16.8
GLM-4.56.7
ChatGPT-4o (2025-04-25 update)6.6
GPT-5.3 Instant6.6
Kimi K2.56.6
Sonnet 4.56.5
DeepSeek-R16.5
GPT-4.16.5
Claude 3.7 Sonnet6.4
Loading Atlas data…