CursorBench 4.0
Code · 2026-09-10
CursorBench 4.0 evaluates agents on ambiguous, multi-file tasks from real Cursor sessions. Released September 10, 2026, this version introduces long-horizon edit, refactor, investigation, intent-understanding, job-management, and design-adherence problems. It is a distinct task set from CursorBench 3.2; the public table reports task scores in percentage points without disclosing an exact evaluation item count.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Opus 5.5 | 57.8 |
| Claude Fable 5.1 | 51.8 |
| Claude Opus 5 | 46.6 |
| Grok 4.7 | 46.3 |
| GPT-5.6 Sol | 41.7 |
| Muse Spark 1.3 | 41.6 |
| Grok 4.6 | 41.4 |
| GPT-5.6 Terra | 41.3 |
| Gemini 3.8 Flash | 39.6 |
| GPT-5.6 Luna | 35.9 |
| Sonnet 5 | 34.1 |
| Composer 2.5 | 27.7 |