METR Time Horizons (1.1, 50% Success)
Code · 2026-01-29
METR Time Horizon 1.1: human-expert task duration in minutes at which the evaluated agent is predicted to succeed 50% of the time on software engineering, machine learning, and cybersecurity tasks. Uses METR's official 1.1 results only, excluding the 1.0 baselines stitched into its download. Kept outside the capability fit for external validation; this is task difficulty, not the agent's runtime. METR cautions that measurements above 16 hours are unreliable with this suite.
Top models (higher is better)
| Model | Score |
|---|---|
| Opus 4.6 | 719 |
| Gemini 3.1 Pro Preview | 384 |
| GPT-5.2 | 352 |
| GPT-5.3-Codex | 350 |
| GPT-5.4 | 342 |
| Opus 4.5 | 293 |
| Gemini 3 Pro Preview | 224 |
| GPT-5.1-Codex-Max | 224 |
| GPT-5 (2025-08-07) | 203 |
| o3 | 120 |
| Opus 4.1 | 100 |
| Opus 4 | 100 |
| Claude 3.7 Sonnet | 60.4 |
| O1 | 38.8 |
| Claude 3.5 Sonnet (Oct 2024) | 20.5 |
| o1 Preview | 20.3 |
| Claude 3.5 Sonnet (June 2024) | 11.4 |
| GPT-4o | 7.0 |
| GPT-4 Turbo (1106 Preview) | 4.0 |
| GPT-4 (0314) | 4.0 |
| Opus 3 | 4.0 |
| GPT-4 Turbo | 3.7 |