Atlas

Benchmarks

← All benchmarks

EBR-bench

Games · 2026-07-01

EBR-bench tests whether AI agents improve at the unfamiliar, complex card-based campaign game Earthborne Rangers through repeated play and persistent note-taking. Its topline score is the mean fraction of the 21 available objectives completed during the final 20% of a run's playthroughs, reported in percentage points; a default run has 10 playthroughs.

Top models (higher is better)

ModelScore
Claude Opus 545.7
Claude Fable 539.5
GPT-5.6 Sol38.6
GPT-5.425.4
Opus 4.825.2
GPT-5.223.0
GPT-5.521.0
Opus 4.719.0
Opus 4.514.3
Gemini 3.1 Pro Preview14.3
Opus 4.612.7
GPT-512.7
GLM-5.29.5
Qwen3.7-Max9.5
Opus 4.17.9
Gemini 3.5 Flash4.8
Sonnet 4.52.4
Kimi K2.62.4
Loading Atlas data…