METR Time Horizons
Code · 2025-03-18
Durations of the longest task that models can complete correctly more often than not, across a set of software engineering and related tasks.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Mythos Preview | 1045 |
| Opus 4.6 | 719 |
| Gemini 3.1 Pro Preview | 384 |
| GPT-5.2 | 352 |
| GPT-5.3-Codex | 350 |
| GPT-5.4 | 342 |
| Opus 4.5 | 293 |
| Gemini 3 Pro Preview | 224 |
| GPT-5 | 203 |
| GPT-5.1-Codex-Max | 162 |
| Sonnet 4.5 | 122 |
| Opus 4.1 | 114 |
| Grok 4 | 110 |
| Opus 4 | 100 |
| o3 | 91.3 |
| O4 Mini | 76.5 |
| Sonnet 4 | 74.9 |
| Claude 3.7 Sonnet | 60.4 |
| Kimi K2 Thinking | 54.2 |
| gpt-oss-120b | 42.0 |
| O1 | 39.2 |
| Gemini 2.5 Pro Preview 06-05 | 38.7 |
| DeepSeek-R1-0528 | 31.2 |
| Claude 3.5 Sonnet (Oct 2024) | 29.6 |
| DeepSeek-R1 | 26.9 |
| DeepSeek-V3-0324 | 23.1 |
| o1 Preview | 22.2 |
| Claude 3.5 Sonnet (June 2024) | 18.7 |
| DeepSeek-V3 | 18.5 |
| GPT-4o (2024-11-20) | 9.2 |
| GPT-4 Turbo (1106 Preview) | 8.6 |
| GPT-4o (2024-08-06) | 7.0 |
| Opus 3 | 6.4 |
| GPT-4 0125 Preview | 5.4 |
| GPT-4 | 5.4 |
| Qwen2.5 72B Instruct | 5.2 |
| GPT-4 0613 | 4.0 |
| Qwen2 72B Instruct | 2.2 |
| GPT-3.5 Turbo Instruct | 0.6 |
| GPT-2 (1.5B) | 0.0 |