Atlas

Benchmarks

← All benchmarks

RuneBench

Games · 2026-02-26

RuneBench evaluates coding agents that play an emulated RuneScape world by writing and executing TypeScript against rs-sdk while consulting extracted game-wiki files. Its headline 30-minute skill score is the mean of ln(1 + peak XP/min) across 16 skills; published sub-tasks also cover peak gold accumulation from four starting conditions.

Top models (higher is better)

ModelScore
Claude Fable 55.9
GPT-5.6 Sol5.9
Grok 4.55.7
GPT-5.55.7
Gemini 3.5 Flash5.4
GPT-5.6 Luna5.3
Opus 4.85.1
Sonnet 54.8
GLM-5.24.8
Opus 4.74.7
GPT-5.44.7
Gemini 3 Flash Preview4.7
Gemini 3.1 Pro Preview4.5
Opus 4.64.4
Kimi K2.7 Code4.3
DeepSeek-V4-Pro4.3
GPT-5.3-Codex4.3
GPT-5.4 Mini4.1
Opus 4.54.1
Qwen3.7-Max4.1
Gemini 3 Pro Preview3.8
Grok 4.33.7
Kimi K2.63.6
Sonnet 4.53.2
Sonnet 4.63.2
GPT-5.4 Nano2.3
Kimi K2.52.1
GLM-51.9
Haiku 4.51.6
Qwen3 Max (rolling alias)1.4
Qwen3.5 35B A3B0.7
Loading Atlas data…