EQ-Bench 3
Chat & Writing · 2025-04-28
EQ-Bench 3 is a multi-turn emotional-intelligence benchmark whose primary leaderboard score is normalized Elo from pairwise LLM-judge comparisons over role-play and analysis scenarios.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Fable 5 | 2050 |
| Opus 4.8 | 2030 |
| Opus 4.7 | 1884 |
| Opus 4.6 | 1717 |
| Sonnet 4.6 | 1714 |
| Hivemind-32B-Preview | 1618 |
| GPT-5.5 | 1577 |
| GLM-5.2 | 1575 |
| DeepSeek-V4-Pro | 1570 |
| GLM-5.1 | 1567 |
| GPT-5.4 | 1563 |
| Kimi K2 Instruct | 1562 |
| Kimi K2.6 | 1561 |
| Gemini 3 Pro Preview | 1559 |
| GPT-5.2 | 1559 |
| Horizon Alpha | 1554 |
| GPT-5.1 | 1551 |
| Opus 4.5 | 1546 |
| Gemini 2.5 Pro Preview 06-05 | 1545 |
| Kimi K2.5 | 1545 |
| Gemini 3.1 Pro Preview | 1538 |
| GLM-5 | 1526 |
| Sonnet 4.5 | 1511 |
| o3 | 1500 |
| DeepSeek-V4-Flash | 1491 |
| GPT-5 Chat (2025-08-07) | 1488 |
| Gemma 4 31B IT | 1434 |
| Gemma 4 26B A4B IT | 1431 |
| GLM-4.7 | 1428 |
| Qwen3.5 397B A17B | 1413 |
| Opus 4 | 1400 |
| Gemini 2.5 Pro Preview 03-25 | 1393 |
| GPT-5.3 Instant | 1393 |
| ChatGPT-4o (2025-04-25 update) | 1390 |
| ChatGPT-4o Latest observed 2025-03-27 | 1387 |
| Hermes 4 405B | 1361 |
| Gemma 4 12B IT | 1361 |
| GLM-4.5 | 1276 |
| Sonnet 4 | 1243 |
| O4 Mini | 1233 |