Atlas

Benchmarks

← All benchmarks

Arena Text — Exclude Ties (Style Controlled)

Chat & Writing

Style-controlled Text Arena rating computed with tied votes excluded. It is separate from the overall series because removing ties changes the vote sample and rating estimates.

Top models (higher is better)

ModelScore
Claude Fable 51530
Opus 4.61525
Opus 4.71519
Claude Opus 51506
Muse Spark 1.11504
Gemini 3 Pro Preview1502
Muse Spark1500
Gemini 3.1 Pro Preview1500
GPT-5.6 Sol1497
Opus 4.81496
Kimi K31496
Gemini 3.6 Flash1494
GPT-5.51493
GPT-5.41486
Gemini 3 Flash Preview1485
Gemini 3.5 Flash1484
Qwen3.7 Max Preview1482
Sonnet 4.61482
Opus 4.51477
GLM-5.21477
GLM-5.11476
Grok 4.51476
GPT-5.6 Terra1475
Grok 4.11474
ERNIE 5.11473
MiMo-V2.5-Pro1470
Sonnet 51469
Kimi K2.61464
Qwen3.7-Plus1463
DeepSeek-V4-Pro1461
Gemini 3.5 Flash-Lite1460
GLM-51460
Hy31459
Sonnet 4.51458
GPT-5.11458
GPT-5.6 Luna1454
Gemma 4 31B1454
GPT-5.3 Instant1452
GPT-5.4 Mini1451
Kimi K2.51450
Loading Atlas data…