Atlas

Benchmarks

← All benchmarks

EDIT-Bench Core Easy

Code · 2025-11-06

The EDIT-Bench Core easy slice reports pass@1 on the 62 distinct tasks classified as easy by the benchmark authors.

Top models (higher is better)

ModelScore
Sonnet 496.8
o3-mini93.5
Sonnet 4.591.9
Claude 3.7 Sonnet91.9
Claude 3.5 Sonnet (Oct 2024)90.3
O4 Mini90.3
Kimi K2 Instruct 090588.7
GPT-582.3
Grok 4 Fast82.3
DeepSeek-V3.180.7
Grok Code Fast 180.7
Llama 3.3 70B Instruct80.7
Qwen2.5 72B Instruct80.7
Qwen3-Coder-480B-A35B-Instruct80.7
Gemini 2.5 Pro79.0
GLM 4.679.0
GPT-4o Mini79.0
GPT-5 Mini79.0
Qwen2.5 Coder 32B Instruct79.0
Qwen3 14B79.0
Devstral Medium 1.077.4
gpt-oss-20b77.4
GPT-4o (2024-08-06)75.8
Llama 3.1 405B Instruct75.8
Gemini 2.5 Flash74.2
Llama 4 Maverick Instruct74.2
Devstral Small 1.172.6
GPT-5 Nano72.6
Qwen3-30B-A3B72.6
gpt-oss-120b71.0
Qwen3 Coder Flash (rolling alias)71.0
Llama 4 Scout Instruct69.3
Mistral Small 3.2 24B Instruct 250666.1
Codestral 25.0864.5
DeepSeek-R1-052862.9
Llama 3.1 8B Instruct56.5
Kimi-Dev-72B54.8
Gemma 3n E4B Instruct51.6
Gemma 3 27B IT45.2
GLM-4.545.2
Loading Atlas data…