Atlas

Benchmarks

← All benchmarks

EQ-Bench 3: Insight

Chat & Writing · 2025-04-28

Within EQ-Bench 3's emotional-intelligence scenarios, this subscore rates the insight dimension of model responses on a 0–10 scale. Higher values are better.

Top models (higher is better)

ModelScore
Claude Fable 59.4
Kimi K2.69.3
GPT-5.49.2
Opus 4.79.1
Opus 4.89.1
GPT-5.59.0
DeepSeek-V4-Pro8.9
Opus 4.68.8
Sonnet 4.58.8
GLM-5.18.8
GLM-5.28.8
GPT-5.28.8
Sonnet 4.68.7
Kimi K2.58.7
Kimi K2 Instruct8.7
Hivemind-32B-Preview8.6
GLM-58.5
Opus 4.58.4
Gemini 3.1 Pro Preview8.4
Gemini 3 Pro Preview8.4
GPT-5.18.4
Horizon Alpha8.4
o38.3
DeepSeek-V4-Flash8.1
GPT-5 Chat (2025-08-07)8.1
Qwen3.5 397B A17B8.1
Gemini 2.5 Pro Preview 06-058.0
GPT-5.3 Instant8.0
Opus 47.9
Sonnet 47.9
GLM-4.77.9
Gemma 4 26B A4B IT7.8
Gemma 4 31B IT7.7
DeepSeek-R17.6
GLM-4.57.6
Grok 4.1 Fast7.5
Gemini 2.5 Pro Preview 03-257.3
GPT-4.17.2
DeepSeek-V3-03247.1
ChatGPT-4o (2025-04-25 update)7.1
Loading Atlas data…