EDIT-Bench Complete
Code · 2025-11-06
EDIT-Bench Complete evaluates whole-file code editing from a natural-language instruction and highlighted code on 540 real-world developer tasks. The Complete set expands 108 distinct editing problems across five natural-language instruction variants; the headline score is pass@1 over all tasks.
Top models (higher is better)
| Model | Score |
|---|---|
| Sonnet 4 | 64.8 |
| Sonnet 4.5 | 59.8 |
| Claude 3.7 Sonnet | 59.3 |
| Claude 3.5 Sonnet (Oct 2024) | 59.1 |
| GLM 4.6 | 56.5 |
| Kimi K2 Instruct 0905 | 56.5 |
| o3-mini | 56.3 |
| DeepSeek-V3.1 | 54.3 |
| GPT-5 Mini | 54.1 |
| Qwen3-Coder-480B-A35B-Instruct | 53.9 |
| O4 Mini | 53.7 |
| GPT-4o (2024-08-06) | 53.3 |
| GPT-5 | 52.8 |
| Grok 4 Fast | 52.0 |
| Gemini 2.5 Pro | 51.3 |
| Grok Code Fast 1 | 50.9 |
| Qwen3 Coder Flash (rolling alias) | 50.7 |
| Llama 3.3 70B Instruct | 49.6 |
| Llama 4 Maverick Instruct | 49.4 |
| Llama 3.1 405B Instruct | 48.7 |
| gpt-oss-20b | 48.1 |
| GPT-4o Mini | 47.8 |
| Mistral Small 3.2 24B Instruct 2506 | 46.3 |
| Qwen3 14B | 45.9 |
| GPT-5 Nano | 45.7 |
| Qwen2.5 72B Instruct | 45.2 |
| Codestral 25.08 | 44.8 |
| DeepSeek-R1-0528 | 44.4 |
| Llama 4 Scout Instruct | 43.3 |
| Qwen3-30B-A3B | 43.1 |
| gpt-oss-120b | 41.3 |
| Devstral Medium 1.0 | 41.1 |
| Qwen2.5 Coder 32B Instruct | 40.0 |
| Gemma 3 27B IT | 37.0 |
| Devstral Small 1.1 | 36.7 |
| Llama 3.1 8B Instruct | 34.1 |
| Kimi-Dev-72B | 31.7 |
| Gemma 3 12B IT | 30.0 |
| Gemma 3n E4B Instruct | 29.3 |
| GLM-4.5 | 29.1 |