Atlas

Benchmarks

← All benchmarks

Public Goods Game Average Contribution

Games · 2025-02-22

Lech Mazur's Public Goods Game benchmark measures how much models contribute in a repeated public-goods scenario. The public README reports average percentage contribution with standard errors.

Top models (higher is better)

ModelScore
Gemini 2.0 Flash Experimental45.2
Haiku 3.541.0
Gemma 2 27B31.2
Gemini 1.5 Flash30.2
GPT-4o Mini24.8
Gemini 1.5 Pro 00224.6
Qwen2.5 72B Instruct23.6
Claude 3.5 Sonnet (Oct 2024)20.9
Gemini 2.0 Flash Thinking Experimental 01-2120.0
Grok 215.8
DeepSeek-V315.5
GPT-4o14.7
Llama 3.1 405B3.2
o1-mini2.6
O12.5
Mistral Large 2 (Instruct 2407)1.6
Llama 3.3 70B Instruct1.2
Qwen2.5-Max0.5
DeepSeek-R10.3
o3-mini0.0
Loading Atlas data…