Atlas

Benchmarks

← All benchmarks

CursorBench 4.0

Code · 2026-09-10

CursorBench 4.0 evaluates agents on ambiguous, multi-file tasks from real Cursor sessions. Released September 10, 2026, this version introduces long-horizon edit, refactor, investigation, intent-understanding, job-management, and design-adherence problems. It is a distinct task set from CursorBench 3.2; the public table reports task scores in percentage points without disclosing an exact evaluation item count.

Top models (higher is better)

ModelScore
Claude Opus 5.557.8
Claude Fable 5.151.8
Claude Opus 546.6
Grok 4.746.3
GPT-5.6 Sol41.7
Muse Spark 1.341.6
Grok 4.641.4
GPT-5.6 Terra41.3
Gemini 3.8 Flash39.6
GPT-5.6 Luna35.9
Sonnet 534.1
Composer 2.527.7
Loading Atlas data…