Atlas

Benchmarks

← All benchmarks

GameDevBench

Games · 2026-06-30

GameDevBench evaluates LLM agents on 333 real game-development tasks in the Godot engine, derived from 88 web and video tutorials. Tasks span 2D and 3D graphics and animation, user interfaces, and gameplay logic, requiring agents to modify code and multimodal assets; the headline score is pass@1 under each model's best published multimodal-feedback configuration and harness.

Top models (higher is better)

ModelScore
Claude Fable 567.3
GPT-5.6 Sol63.7
Opus 4.855.9
GPT-5.554.7
Gemini 3 Pro Preview53.8
GPT-5.452.0
Gemini 3 Flash Preview46.8
GPT-5.4 Mini43.2
GLM-5.238.4
Sonnet 4.534.8
Kimi K2.520.7
Haiku 4.518.6
Qwen3.5 397B A17B5.4
Loading Atlas data…