Atlas

Benchmarks

← All benchmarks

GSO Opt@1

Code · 2025-05-29

GSO Opt@1 estimates the percentage of 102 repository-level optimization tasks where one agent attempt both passes correctness tests and achieves at least 95% of the expert developer's speedup. The tasks span 10 real codebases and five programming languages.

Top models (higher is better)

ModelScore
Opus 4.847.1
Opus 4.744.1
GPT-5.540.2
Sonnet 537.3
Opus 4.633.3
GPT-5.431.4
Opus 4.526.5
Gemini 3.1 Pro Preview22.6
Sonnet 4.514.7
GPT-5.113.7
Gemini 3 Flash Preview9.8
Opus 46.9
Sonnet 44.9
Kimi K2 Instruct4.9
Qwen3-Coder-480B-A35B-Instruct4.9
Claude 3.5 Sonnet (Oct 2024)4.6
Gemini 2.5 Pro Preview 06-053.9
Claude 3.7 Sonnet3.8
O4 Mini3.6
GLM 4.5 Air2.9
o3-mini1.3
GPT-4o0.0
GPT-4o (2024-11-20)0.0
Loading Atlas data…