Atlas

Benchmarks

← All benchmarks

EDIT-Bench Core Hard

Code · 2025-11-06

The EDIT-Bench Core hard slice reports pass@1 on the 46 distinct tasks classified as hard by the benchmark authors.

Top models (higher is better)

ModelScore
DeepSeek-V3.130.4
Claude 3.5 Sonnet (Oct 2024)28.3
Sonnet 426.1
Qwen3 Coder Flash (rolling alias)26.1
GLM 4.623.9
GPT-4o (2024-08-06)23.9
Gemini 2.5 Flash21.7
Gemini 2.5 Pro21.7
GPT-521.7
o3-mini21.7
Qwen3-Coder-480B-A35B-Instruct21.7
Claude 3.7 Sonnet21.7
Llama 4 Maverick Instruct19.6
Qwen2.5 Coder 32B Instruct19.6
Sonnet 4.517.4
GPT-5 Mini17.4
Grok Code Fast 117.4
Kimi K2 Instruct 090517.4
O4 Mini17.4
Qwen2.5 72B Instruct17.4
Codestral 25.0815.2
Devstral Small 1.115.2
DeepSeek-R1-052813.0
Devstral Medium 1.013.0
GPT-5 Nano13.0
gpt-oss-20b13.0
Grok 4 Fast13.0
Llama 3.1 8B Instruct13.0
Llama 3.3 70B Instruct13.0
Llama 4 Scout Instruct13.0
Mistral Small 3.2 24B Instruct 250613.0
GPT-4o Mini10.9
Llama 3.1 405B Instruct10.9
Gemma 3 27B IT8.7
GLM-4.58.7
gpt-oss-120b8.7
Gemma 3 12B IT6.5
Gemma 3n E4B Instruct4.3
Kimi-Dev-72B4.3
Qwen3 14B4.3
Loading Atlas data…