EDIT-Bench Core
Code · 2025-11-06
EDIT-Bench Core evaluates whole-file code editing on the 108 distinct real-world developer tasks, using one instruction-language variant per task. Scores report pass@1 over all Core tasks.
Top models (higher is better)
| Model | Score |
|---|---|
| Sonnet 4 | 66.7 |
| Claude 3.5 Sonnet (Oct 2024) | 63.9 |
| o3-mini | 63.0 |
| Claude 3.7 Sonnet | 62.0 |
| Sonnet 4.5 | 60.2 |
| DeepSeek-V3.1 | 59.3 |
| O4 Mini | 59.3 |
| Kimi K2 Instruct 0905 | 58.3 |
| GPT-5 | 56.5 |
| GLM 4.6 | 55.6 |
| Qwen3-Coder-480B-A35B-Instruct | 55.6 |
| Gemini 2.5 Pro | 54.6 |
| GPT-4o (2024-08-06) | 53.7 |
| Grok Code Fast 1 | 53.7 |
| Qwen2.5 72B Instruct | 53.7 |
| Qwen2.5 Coder 32B Instruct | 53.7 |
| GPT-5 Mini | 52.8 |
| Grok 4 Fast | 52.8 |
| Gemini 2.5 Flash | 51.9 |
| Llama 3.3 70B Instruct | 51.9 |
| Qwen3 Coder Flash (rolling alias) | 51.9 |
| Llama 4 Maverick Instruct | 50.9 |
| Devstral Medium 1.0 | 50.0 |
| GPT-4o Mini | 50.0 |
| gpt-oss-20b | 50.0 |
| Devstral Small 1.1 | 48.1 |
| Llama 3.1 405B Instruct | 48.1 |
| GPT-5 Nano | 47.2 |
| Qwen3 14B | 47.2 |
| Llama 4 Scout Instruct | 45.4 |
| gpt-oss-120b | 44.4 |
| Codestral 25.08 | 43.5 |
| Mistral Small 3.2 24B Instruct 2506 | 43.5 |
| Qwen3-30B-A3B | 43.5 |
| DeepSeek-R1-0528 | 41.7 |
| Llama 3.1 8B Instruct | 38.0 |
| Kimi-Dev-72B | 33.3 |
| Gemma 3n E4B Instruct | 31.5 |
| Gemma 3 27B IT | 29.6 |
| GLM-4.5 | 29.6 |