Atlas

Benchmarks

← All benchmarks

EQ-Bench 3: Empathy

Chat & Writing · 2025-04-28

Within EQ-Bench 3's emotional-intelligence scenarios, this subscore rates the empathy dimension of model responses on a 0–10 scale. Higher values are better.

Top models (higher is better)

ModelScore
Opus 4.88.7
Opus 4.78.5
Sonnet 4.68.4
Hivemind-32B-Preview8.4
Claude Fable 58.2
DeepSeek-V4-Pro8.1
Gemini 2.5 Pro Preview 06-058.1
GLM-5.28.1
GPT-5.58.0
Opus 4.67.9
Gemini 3.1 Pro Preview7.9
GPT-5.47.9
Opus 4.57.8
Gemini 3 Pro Preview7.8
Gemma 4 26B A4B IT7.8
GLM-57.8
Kimi K2 Instruct7.8
Gemma 4 31B IT7.7
GPT-5.17.7
Kimi K2.67.6
ChatGPT-4o Latest observed 2025-03-277.5
DeepSeek-V4-Flash7.5
Gemini 2.5 Pro Preview 03-257.5
Gemma 4 12B IT7.5
GPT-5.27.5
Gemini 2.5 Pro Preview 05-067.4
Kimi K2.57.4
ChatGPT-4o (2025-04-25 update)7.3
Horizon Alpha7.3
Qwen3-235B-A22B7.3
Qwen3.5 397B A17B7.3
Opus 47.2
GPT-4.17.2
GLM-5.17.1
GPT-5 Chat (2025-08-07)7.1
Hermes 4 405B7.1
o37.1
Gemini 2.5 Flash Preview 05-207.0
GLM-4.77.0
Sonnet 4.56.9
Loading Atlas data…