Atlas

Benchmarks

← All benchmarks

EDIT-Bench Core

Code · 2025-11-06

EDIT-Bench Core evaluates whole-file code editing on the 108 distinct real-world developer tasks, using one instruction-language variant per task. Scores report pass@1 over all Core tasks.

Top models (higher is better)

ModelScore
Sonnet 466.7
Claude 3.5 Sonnet (Oct 2024)63.9
o3-mini63.0
Claude 3.7 Sonnet62.0
Sonnet 4.560.2
DeepSeek-V3.159.3
O4 Mini59.3
Kimi K2 Instruct 090558.3
GPT-556.5
GLM 4.655.6
Qwen3-Coder-480B-A35B-Instruct55.6
Gemini 2.5 Pro54.6
GPT-4o (2024-08-06)53.7
Grok Code Fast 153.7
Qwen2.5 72B Instruct53.7
Qwen2.5 Coder 32B Instruct53.7
GPT-5 Mini52.8
Grok 4 Fast52.8
Gemini 2.5 Flash51.9
Llama 3.3 70B Instruct51.9
Qwen3 Coder Flash (rolling alias)51.9
Llama 4 Maverick Instruct50.9
Devstral Medium 1.050.0
GPT-4o Mini50.0
gpt-oss-20b50.0
Devstral Small 1.148.1
Llama 3.1 405B Instruct48.1
GPT-5 Nano47.2
Qwen3 14B47.2
Llama 4 Scout Instruct45.4
gpt-oss-120b44.4
Codestral 25.0843.5
Mistral Small 3.2 24B Instruct 250643.5
Qwen3-30B-A3B43.5
DeepSeek-R1-052841.7
Llama 3.1 8B Instruct38.0
Kimi-Dev-72B33.3
Gemma 3n E4B Instruct31.5
Gemma 3 27B IT29.6
GLM-4.529.6
Loading Atlas data…