Atlas

Benchmarks

← All benchmarks

PGG-Bench TrueSkill Mu

Games · 2025-04-10

PGG-Bench: Contribute & Punish is a multi-agent public-goods game where models decide how much to contribute and how to punish other players. The headline leaderboard ranks models by multi-pass TrueSkill mu.

Top models (higher is better)

ModelScore
Gemini 2.5 Pro Preview 03-2514.3
O113.4
Mistral Large 2 (Instruct 2407)11.9
Gemini 2.0 Pro Experimental 02-0511.1
o3-mini11.1
DeepSeek-V310.9
GPT-4.510.8
Llama 3.3 70B Instruct10.2
Grok 210.1
Claude 3.7 Sonnet10.0
ChatGPT-4o Latest observed 2025-01-299.9
DeepSeek-R19.7
Qwen2.5-Max9.5
QwQ-32B8.6
Llama 3.1 405B8.5
Gemini 2.0 Flash Thinking Experimental 01-218.3
Claude 3.5 Sonnet (Oct 2024)8.1
Gemini 2.0 Flash5.1
Loading Atlas data…