GSO Opt@1
Code · 2025-05-29
GSO Opt@1 estimates the percentage of 102 repository-level optimization tasks where one agent attempt both passes correctness tests and achieves at least 95% of the expert developer's speedup. The tasks span 10 real codebases and five programming languages.
Top models (higher is better)
| Model | Score |
|---|---|
| Opus 4.8 | 47.1 |
| Opus 4.7 | 44.1 |
| GPT-5.5 | 40.2 |
| Sonnet 5 | 37.3 |
| Opus 4.6 | 33.3 |
| GPT-5.4 | 31.4 |
| Opus 4.5 | 26.5 |
| Gemini 3.1 Pro Preview | 22.6 |
| Sonnet 4.5 | 14.7 |
| GPT-5.1 | 13.7 |
| Gemini 3 Flash Preview | 9.8 |
| Opus 4 | 6.9 |
| Sonnet 4 | 4.9 |
| Kimi K2 Instruct | 4.9 |
| Qwen3-Coder-480B-A35B-Instruct | 4.9 |
| Claude 3.5 Sonnet (Oct 2024) | 4.6 |
| Gemini 2.5 Pro Preview 06-05 | 3.9 |
| Claude 3.7 Sonnet | 3.8 |
| O4 Mini | 3.6 |
| GLM 4.5 Air | 2.9 |
| o3-mini | 1.3 |
| GPT-4o | 0.0 |
| GPT-4o (2024-11-20) | 0.0 |