Atlas

Benchmarks

← All benchmarks

MirrorCode (ML, +Private, 2L)

Code

Epoch and METR's leaderboard protocol: reimplement 15 medium/large programs, including private targets, in two languages for 30 tasks. Three attempts per task, each with 10 billion tokens and seven days. A solve requires all public and held-out end-to-end tests to pass. Distinct from the paper's 25-program, six-language protocol.

Top models (higher is better)

ModelScore
Claude Fable 5.173.3
Claude Fable 563.9
GPT-6 Astra46.7
Opus 4.731.1
GPT-5.6 Sol20.0
GPT-5.415.6
GPT-5.510.0
Gemini 3.1 Pro Preview8.9
Loading Atlas data…