Atlas

Benchmarks

← All benchmarks

Arena Text — Exclude Ties (No Style Control)

Chat & Writing

Raw Text Arena rating computed with tied votes excluded and without style-control adjustment. It is separate because both filters change the vote sample and rating estimates.

Top models (higher is better)

ModelScore
Claude Opus 51524
Opus 4.61519
Claude Fable 51510
Opus 4.71500
Gemini 3 Pro Preview1491
Gemini 3.1 Pro Preview1488
Gemini 3.5 Flash1488
Gemini 3.6 Flash1485
Muse Spark 1.11484
Qwen3.7 Max Preview1480
Muse Spark1478
Kimi K31477
GPT-5.51475
GPT-5.41474
Gemini 3 Flash Preview1474
ERNIE 5.11471
GLM-5.11468
GLM-5.21467
Opus 4.81466
MiMo-V2.5-Pro1463
Gemini 2.5 Pro1460
Qwen3.7-Plus1460
Sonnet 4.61458
Kimi K2.61455
GPT-5.6 Sol1454
Grok 4.51451
DeepSeek-V4-Pro1449
Opus 4.51447
GPT-5.6 Terra1446
GLM-51442
Kimi K2.51442
Sonnet 51440
Gemma 4 31B1440
GPT-5.11438
GLM 4.61438
Hy31437
Qwen3 Max Preview1436
Qwen3.5 397B A17B1434
Inkling1433
Sonnet 4.51433
Loading Atlas data…