Atlas

Benchmarks

← All benchmarks

GSO Hack-Adjusted Opt@1

Code · 2025-11-03

Hack-Adjusted GSO Opt@1 applies GSO's Hack Detector penalty to the one-attempt score, penalizing deceptive optimizations such as memoization or evaluation-harness hijacking after comparing patches with the oracle solution and tests.

Top models (higher is better)

ModelScore
Opus 4.847.1
Opus 4.742.2
Opus 4.637.3
GPT-5.537.3
Sonnet 536.3
GPT-5.430.4
GPT-5.226.5
Opus 4.524.5
Gemini 3.1 Pro Preview21.6
Gemini 3 Pro Preview17.6
GPT-5.112.8
Sonnet 4.512.7
Gemini 3 Flash Preview7.8
GPT-55.9
Opus 44.9
Sonnet 44.9
Claude 3.5 Sonnet (Oct 2024)4.6
o33.9
Qwen3-Coder-480B-A35B-Instruct3.9
Claude 3.7 Sonnet3.8
O4 Mini3.6
Kimi K2 Instruct2.0
o3-mini1.3
GLM 4.5 Air1.0
Gemini 2.5 Pro0.0
GPT-4o0.0
Loading Atlas data…