Atlas

Benchmarks

← All benchmarks

METR Time Horizons

Code · 2025-03-18

Durations of the longest task that models can complete correctly more often than not, across a set of software engineering and related tasks.

Top models (higher is better)

ModelScore
Claude Mythos Preview1045
Opus 4.6719
Gemini 3.1 Pro Preview384
GPT-5.2352
GPT-5.3-Codex350
GPT-5.4342
Opus 4.5293
Gemini 3 Pro Preview224
GPT-5203
GPT-5.1-Codex-Max162
Sonnet 4.5122
Opus 4.1114
Grok 4110
Opus 4100
o391.3
O4 Mini76.5
Sonnet 474.9
Claude 3.7 Sonnet60.4
Kimi K2 Thinking54.2
gpt-oss-120b42.0
O139.2
Gemini 2.5 Pro Preview 06-0538.7
DeepSeek-R1-052831.2
Claude 3.5 Sonnet (Oct 2024)29.6
DeepSeek-R126.9
DeepSeek-V3-032423.1
o1 Preview22.2
Claude 3.5 Sonnet (June 2024)18.7
DeepSeek-V318.5
GPT-4o (2024-11-20)9.2
GPT-4 Turbo (1106 Preview)8.6
GPT-4o (2024-08-06)7.0
Opus 36.4
GPT-4 0125 Preview5.4
GPT-45.4
Qwen2.5 72B Instruct5.2
GPT-4 06134.0
Qwen2 72B Instruct2.2
GPT-3.5 Turbo Instruct0.6
GPT-2 (1.5B)0.0
Loading Atlas data…