DiploBench
Games · 2025-03-04
Experimental full-press Diplomacy testbed reporting single-game LLM strategic reasoning and negotiation results.
Top models (higher is better)
| Model | Score |
|---|---|
| O1 | 91.0 |
| Claude 3.7 Sonnet | 86.0 |
| DeepSeek-R1 | 86.0 |
| o3-mini | 66.0 |
| GPT-4o (2024-11-20) | 50.0 |
| QwQ-32B | 25.8 |
| Gemini 2.0 Flash 001 | 12.0 |
| GPT-4o Mini | 5.0 |