Atlas

Benchmarks

← All benchmarks

EDIT-Bench Complete Easy

Code · 2025-11-06

The EDIT-Bench Complete easy slice reports pass@1 on the 295 tasks classified as easy by the benchmark authors.

Top models (higher is better)

ModelScore
Sonnet 493.6
Sonnet 4.593.6
Claude 3.7 Sonnet90.8
Claude 3.5 Sonnet (Oct 2024)86.8
o3-mini86.8
GLM 4.685.1
O4 Mini85.1
Kimi K2 Instruct 090584.4
Grok 4 Fast83.0
GPT-5 Mini82.7
Qwen3-Coder-480B-A35B-Instruct81.7
GPT-4o (2024-08-06)81.4
DeepSeek-V3.181.0
Grok Code Fast 180.7
Llama 3.3 70B Instruct79.3
Qwen3 14B79.0
GPT-577.3
Llama 4 Maverick Instruct77.3
GPT-4o Mini76.6
Llama 3.1 405B Instruct76.6
Qwen3 Coder Flash (rolling alias)76.3
gpt-oss-20b75.9
Gemini 2.5 Pro75.6
Qwen2.5 72B Instruct73.2
Mistral Small 3.2 24B Instruct 250672.5
Qwen3-30B-A3B72.5
GPT-5 Nano71.9
Codestral 25.0870.8
DeepSeek-R1-052869.5
Llama 4 Scout Instruct69.2
Devstral Medium 1.067.5
gpt-oss-120b64.4
Qwen2.5 Coder 32B Instruct63.7
Devstral Small 1.160.3
Gemma 3 27B IT59.0
Llama 3.1 8B Instruct54.9
Kimi-Dev-72B53.2
Gemma 3 12B IT49.5
Gemma 3n E4B Instruct47.8
GLM-4.545.8
Loading Atlas data…