Atlas

Benchmarks

← All benchmarks

METR Time Horizons (1.1, 50% Success)

Code · 2026-01-29

METR Time Horizon 1.1: human-expert task duration in minutes at which the evaluated agent is predicted to succeed 50% of the time on software engineering, machine learning, and cybersecurity tasks. Uses METR's official 1.1 results only, excluding the 1.0 baselines stitched into its download. Kept outside the capability fit for external validation; this is task difficulty, not the agent's runtime. METR cautions that measurements above 16 hours are unreliable with this suite.

Top models (higher is better)

ModelScore
Opus 4.6719
Gemini 3.1 Pro Preview384
GPT-5.2352
GPT-5.3-Codex350
GPT-5.4342
Opus 4.5293
Gemini 3 Pro Preview224
GPT-5.1-Codex-Max224
GPT-5 (2025-08-07)203
o3120
Opus 4.1100
Opus 4100
Claude 3.7 Sonnet60.4
O138.8
Claude 3.5 Sonnet (Oct 2024)20.5
o1 Preview20.3
Claude 3.5 Sonnet (June 2024)11.4
GPT-4o7.0
GPT-4 Turbo (1106 Preview)4.0
GPT-4 (0314)4.0
Opus 34.0
GPT-4 Turbo3.7
Loading Atlas data…