Atlas

Benchmarks

← All benchmarks

EBR-bench v1

Games · 2026-07-01

EBR-bench v1 topline: the mean fraction of the 21 available objectives completed during the final 20% of a run's playthroughs. It is incompatible with v2's best-of-final-two aggregation and larger sampling regime.

Top models (higher is better)

ModelScore
GPT-5.425.4
GPT-5.223.0
Opus 4.719.0
Opus 4.514.3
Gemini 3.1 Pro Preview14.3
Opus 4.612.7
GPT-512.7
GLM-5.29.5
Qwen3.7-Max9.5
Opus 4.17.9
Gemini 3.5 Flash4.8
Sonnet 4.52.4
Kimi K2.62.4
Loading Atlas data…