Atlas

Benchmarks

← All benchmarks

EQ-Bench 4: Authenticity

Chat & Writing · 2026-07-23

EQ-Bench 4's independently fitted authenticity ability rating measures how genuine rather than scripted or performative a model seems. The source normalizes its per-dimension soft Bradley–Terry rating to a 1–10 scale.

Top models (higher is better)

ModelScore
Claude Opus 510.0
Kimi K38.8
Claude Fable 58.8
Opus 4.78.4
GPT-5.57.4
Opus 4.87.4
GPT-5.46.8
Muse Spark 1.16.7
Sonnet 56.6
Opus 4.66.4
GLM-5.26.3
Sonnet 4.66.2
Kimi K2.66.2
MiMo-V2.5-Pro6.1
Inkling6.1
GPT-5.6 Sol5.9
GPT-5.6 Terra5.7
MiniMax M35.1
DeepSeek-V4-Pro4.7
GPT-5.6 Luna3.8
Gemma 4 31B IT3.6
Gemini 3.1 Pro Preview3.6
Haiku 4.53.0
Grok 4.33.0
Qwen3.7-Max2.9
Gemini 3.5 Flash2.0
Qwen3.6 27B1.2
Mistral Medium 3.51.0
Loading Atlas data…