Atlas

Benchmarks

← All benchmarks

EBR-bench v4 (card ban, single agent)

Games · 2026-09-22

Earthborne Rangers leaderboard protocol documented September 22: an exploitable card is banned, and the default uses a single agent. Scores measure objective completion using the best final two playthroughs; compaction is 250k tokens except for Claude Fable models at 90% of context. Separate from earlier card-allowed runs and optional multi-agent experiments.

Top models (higher is better)

ModelScore
GPT-6 Astra76.2
Claude Opus 5.571.4
Claude Fable 5.157.1
GPT-6 Sol53.3
Claude Opus 545.7
GPT-5.6 Sol44.8
Loading Atlas data…