Atlas

Benchmarks

← All benchmarks

Elimination Game TrueSkill Mu

Games · 2025-07-14

Lech Mazur's Elimination Game benchmark evaluates multi-agent social reasoning, coalition formation, deception, voting, and jury management. The public leaderboard reports TrueSkill mu from tournament outcomes.

Top models (higher is better)

ModelScore
GPT-5.27.5
GPT-56.0
GPT-5 Mini5.7
Opus 4.55.7
Gemini 3 Flash Preview5.7
Grok 3 Mini5.5
ChatGPT-4o Latest observed 2025-03-275.5
DeepSeek-R1-05285.3
Claude 3.7 Sonnet5.3
Opus 4.15.2
Sonnet 4.55.2
Grok 45.2
GPT-4.55.1
Claude 3.5 Sonnet (Oct 2024)5.1
Grok 35.1
Gemini 3 Pro Preview4.9
Gemini 2.5 Flash4.7
Sonnet 44.6
MiniMax M24.6
Qwen3-Max-Preview (Thinking Mode)4.5
o34.5
Gemini 2.5 Pro Preview 03-254.5
Opus 44.4
Qwen3 235B A22B Instruct 25074.4
o3-mini4.4
Kimi K2 Thinking4.3
GLM-4.54.2
Mistral Large 2 (Instruct 2407)4.1
DeepSeek-V34.1
DeepSeek-R14.1
O13.9
gpt-oss-120b3.8
Gemini 2.5 Pro Preview 06-053.7
Mistral Large 3 675B Instruct 25123.6
Llama 4 Maverick3.6
Grok 4.1 Fast3.6
Llama 3.3 70B Instruct3.6
Nova Pro3.5
Qwen3-235B-A22B3.5
ChatGPT-4o Latest observed 2025-01-293.5
Loading Atlas data…