GameDevBench
Games · 2026-06-30
GameDevBench evaluates LLM agents on 333 real game-development tasks in the Godot engine, derived from 88 web and video tutorials. Tasks span 2D and 3D graphics and animation, user interfaces, and gameplay logic, requiring agents to modify code and multimodal assets; the headline score is pass@1 under each model's best published multimodal-feedback configuration and harness.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Fable 5 | 67.3 |
| GPT-5.6 Sol | 63.7 |
| Opus 4.8 | 55.9 |
| GPT-5.5 | 54.7 |
| Gemini 3 Pro Preview | 53.8 |
| GPT-5.4 | 52.0 |
| Gemini 3 Flash Preview | 46.8 |
| GPT-5.4 Mini | 43.2 |
| GLM-5.2 | 38.4 |
| Sonnet 4.5 | 34.8 |
| Kimi K2.5 | 20.7 |
| Haiku 4.5 | 18.6 |
| Qwen3.5 397B A17B | 5.4 |