Atlas

Benchmarks

← All benchmarks

EDIT-Bench Complete Hard

Code · 2025-11-06

The EDIT-Bench Complete hard slice reports pass@1 on the 245 tasks classified as hard by the benchmark authors.

Top models (higher is better)

ModelScore
Sonnet 430.2
Claude 3.5 Sonnet (Oct 2024)25.7
GPT-523.3
Kimi K2 Instruct 090522.9
DeepSeek-V3.122.0
Gemini 2.5 Pro22.0
GLM 4.622.0
Claude 3.7 Sonnet21.2
Qwen3-Coder-480B-A35B-Instruct20.4
Qwen3 Coder Flash (rolling alias)20.0
GPT-4o (2024-08-06)19.6
GPT-5 Mini19.6
o3-mini19.6
Sonnet 4.519.2
Llama 4 Maverick Instruct15.9
O4 Mini15.9
Grok Code Fast 115.1
Llama 3.1 405B Instruct15.1
gpt-oss-20b14.7
Grok 4 Fast14.7
Mistral Small 3.2 24B Instruct 250614.7
DeepSeek-R1-052814.3
GPT-5 Nano14.3
Llama 3.3 70B Instruct13.9
Codestral 25.0813.5
gpt-oss-120b13.5
GPT-4o Mini13.1
Llama 4 Scout Instruct12.2
Qwen2.5 72B Instruct11.4
Qwen2.5 Coder 32B Instruct11.4
Gemma 3 27B IT10.6
Devstral Medium 1.09.4
GLM-4.59.0
Llama 3.1 8B Instruct9.0
Devstral Small 1.18.2
Qwen3-30B-A3B7.8
Gemma 3n E4B Instruct6.9
Gemma 3 12B IT6.5
Qwen3 14B6.1
Kimi-Dev-72B5.7
Loading Atlas data…