Public Goods Game Average Contribution
Games · 2025-02-22
Lech Mazur's Public Goods Game benchmark measures how much models contribute in a repeated public-goods scenario. The public README reports average percentage contribution with standard errors.
Top models (higher is better)
| Model | Score |
|---|---|
| Gemini 2.0 Flash Experimental | 45.2 |
| Haiku 3.5 | 41.0 |
| Gemma 2 27B | 31.2 |
| Gemini 1.5 Flash | 30.2 |
| GPT-4o Mini | 24.8 |
| Gemini 1.5 Pro 002 | 24.6 |
| Qwen2.5 72B Instruct | 23.6 |
| Claude 3.5 Sonnet (Oct 2024) | 20.9 |
| Gemini 2.0 Flash Thinking Experimental 01-21 | 20.0 |
| Grok 2 | 15.8 |
| DeepSeek-V3 | 15.5 |
| GPT-4o | 14.7 |
| Llama 3.1 405B | 3.2 |
| o1-mini | 2.6 |
| O1 | 2.5 |
| Mistral Large 2 (Instruct 2407) | 1.6 |
| Llama 3.3 70B Instruct | 1.2 |
| Qwen2.5-Max | 0.5 |
| DeepSeek-R1 | 0.3 |
| o3-mini | 0.0 |