Atlas

Benchmarks

← All benchmarks

EBR-bench v2

Games · 2026-07-27

EBR-bench v2 topline: the best objective-completion fraction from the final two playthroughs, using at least ten samples per starting state. This protocol is incompatible with v1's final-20%-mean aggregation.

Top models (higher is better)

ModelScore
Claude Opus 545.7
Claude Fable 539.5
GPT-5.6 Sol38.6
Opus 4.825.2
GPT-5.521.0
Loading Atlas data…