GSO Hack-Adjusted Opt@1
Code · 2025-11-03
Hack-Adjusted GSO Opt@1 applies GSO's Hack Detector penalty to the one-attempt score, penalizing deceptive optimizations such as memoization or evaluation-harness hijacking after comparing patches with the oracle solution and tests.
Top models (higher is better)
| Model | Score |
|---|---|
| Opus 4.8 | 47.1 |
| Opus 4.7 | 42.2 |
| Opus 4.6 | 37.3 |
| GPT-5.5 | 37.3 |
| Sonnet 5 | 36.3 |
| GPT-5.4 | 30.4 |
| GPT-5.2 | 26.5 |
| Opus 4.5 | 24.5 |
| Gemini 3.1 Pro Preview | 21.6 |
| Gemini 3 Pro Preview | 17.6 |
| GPT-5.1 | 12.8 |
| Sonnet 4.5 | 12.7 |
| Gemini 3 Flash Preview | 7.8 |
| GPT-5 | 5.9 |
| Opus 4 | 4.9 |
| Sonnet 4 | 4.9 |
| Claude 3.5 Sonnet (Oct 2024) | 4.6 |
| o3 | 3.9 |
| Qwen3-Coder-480B-A35B-Instruct | 3.9 |
| Claude 3.7 Sonnet | 3.8 |
| O4 Mini | 3.6 |
| Kimi K2 Instruct | 2.0 |
| o3-mini | 1.3 |
| GLM 4.5 Air | 1.0 |
| Gemini 2.5 Pro | 0.0 |
| GPT-4o | 0.0 |