Atlas

Benchmarks

← All benchmarks

EDIT-Bench Complete

Code · 2025-11-06

EDIT-Bench Complete evaluates whole-file code editing from a natural-language instruction and highlighted code on 540 real-world developer tasks. The Complete set expands 108 distinct editing problems across five natural-language instruction variants; the headline score is pass@1 over all tasks.

Top models (higher is better)

ModelScore
Sonnet 464.8
Sonnet 4.559.8
Claude 3.7 Sonnet59.3
Claude 3.5 Sonnet (Oct 2024)59.1
GLM 4.656.5
Kimi K2 Instruct 090556.5
o3-mini56.3
DeepSeek-V3.154.3
GPT-5 Mini54.1
Qwen3-Coder-480B-A35B-Instruct53.9
O4 Mini53.7
GPT-4o (2024-08-06)53.3
GPT-552.8
Grok 4 Fast52.0
Gemini 2.5 Pro51.3
Grok Code Fast 150.9
Qwen3 Coder Flash (rolling alias)50.7
Llama 3.3 70B Instruct49.6
Llama 4 Maverick Instruct49.4
Llama 3.1 405B Instruct48.7
gpt-oss-20b48.1
GPT-4o Mini47.8
Mistral Small 3.2 24B Instruct 250646.3
Qwen3 14B45.9
GPT-5 Nano45.7
Qwen2.5 72B Instruct45.2
Codestral 25.0844.8
DeepSeek-R1-052844.4
Llama 4 Scout Instruct43.3
Qwen3-30B-A3B43.1
gpt-oss-120b41.3
Devstral Medium 1.041.1
Qwen2.5 Coder 32B Instruct40.0
Gemma 3 27B IT37.0
Devstral Small 1.136.7
Llama 3.1 8B Instruct34.1
Kimi-Dev-72B31.7
Gemma 3 12B IT30.0
Gemma 3n E4B Instruct29.3
GLM-4.529.1
Loading Atlas data…