Atlas

Benchmarks

← All benchmarks

DiploBench

Games · 2025-03-04

Experimental full-press Diplomacy testbed reporting single-game LLM strategic reasoning and negotiation results.

Top models (higher is better)

ModelScore
O191.0
Claude 3.7 Sonnet86.0
DeepSeek-R186.0
o3-mini66.0
GPT-4o (2024-11-20)50.0
QwQ-32B25.8
Gemini 2.0 Flash 00112.0
GPT-4o Mini5.0
Loading Atlas data…