EBR-bench
Games · 2026-07-01
EBR-bench tests whether AI agents improve at the unfamiliar, complex card-based campaign game Earthborne Rangers through repeated play and persistent note-taking. Its topline score is the mean fraction of the 21 available objectives completed during the final 20% of a run's playthroughs, reported in percentage points; a default run has 10 playthroughs.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Opus 5 | 45.7 |
| Claude Fable 5 | 39.5 |
| GPT-5.6 Sol | 38.6 |
| GPT-5.4 | 25.4 |
| Opus 4.8 | 25.2 |
| GPT-5.2 | 23.0 |
| GPT-5.5 | 21.0 |
| Opus 4.7 | 19.0 |
| Opus 4.5 | 14.3 |
| Gemini 3.1 Pro Preview | 14.3 |
| Opus 4.6 | 12.7 |
| GPT-5 | 12.7 |
| GLM-5.2 | 9.5 |
| Qwen3.7-Max | 9.5 |
| Opus 4.1 | 7.9 |
| Gemini 3.5 Flash | 4.8 |
| Sonnet 4.5 | 2.4 |
| Kimi K2.6 | 2.4 |