Atlas

Benchmarks

← All benchmarks

LLM Debate Benchmark

Chat & Writing · 2026-03-22

LLM debate benchmark in which models debate both sides of contentious topics and are rated through pairwise outcomes. The leaderboard reports Bradley-Terry ratings.

Top models (higher is better)

ModelScore
Claude Fable 51757
Kimi K31741
Opus 4.71685
Muse Spark 1.11684
GPT-5.6 Sol1681
Opus 4.81668
Sonnet 51619
Sonnet 4.61597
GLM-5.21595
GPT-5.41584
GPT-5.51564
GLM-5.11555
Kimi K2.61547
MiniMax M31537
Gemini 3.1 Pro Preview1529
MiMo-V2.5-Pro1526
Grok 4.51519
Qwen3.6-Max-Preview1511
Kimi K2.51490
Doubao Seed 2.0 Pro1489
DeepSeek-V4-Pro1478
Qwen3.7-Max1470
MiniMax M2.71469
Grok 4.201448
Gemini 3.5 Flash1444
MiMo-V2-Pro1428
Qwen3.5 397B A17B1426
Step 3.7 Flash1421
ERNIE 5.11419
Hy3 preview1417
Grok 4.31406
DeepSeek-V3.21399
Mistral Medium 3.51382
Gemini 3.1 Flash-Lite Preview1366
gpt-oss-120b1307
ERNIE 5.01283
Mistral Large 3 675B Instruct 25121255
Llama 4 Maverick1071
Loading Atlas data…