MirrorCode (ML, +Private, 2L)
Code
Epoch and METR's leaderboard protocol: reimplement 15 medium/large programs, including private targets, in two languages for 30 tasks. Three attempts per task, each with 10 billion tokens and seven days. A solve requires all public and held-out end-to-end tests to pass. Distinct from the paper's 25-program, six-language protocol.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Fable 5.1 | 73.3 |
| Claude Fable 5 | 63.9 |
| GPT-6 Astra | 46.7 |
| Opus 4.7 | 31.1 |
| GPT-5.6 Sol | 20.0 |
| GPT-5.4 | 15.6 |
| GPT-5.5 | 10.0 |
| Gemini 3.1 Pro Preview | 8.9 |