Atlas

Benchmarks

← All benchmarks

CursorBench 3.2

Code · 2026-07-08

CursorBench 3.2 is the July 2026 successor to CursorBench 3.1, adding instruction-following and advanced tool-use problems. Scores are task resolve rates in percentage points.

Top models (higher is better)

ModelScore
Claude Fable 570.5
Claude Opus 570.0
GPT-5.6 Sol67.2
Grok 4.566.7
GPT-5.6 Terra64.9
Opus 4.862.3
Sonnet 561.5
GPT-5.6 Luna61.1
Kimi K360.8
GPT-5.558.4
Composer 2.556.1
GLM-5.255.0
Gemini 3.6 Flash53.5
Kimi K2.7 Code49.7
Gemini 3.5 Flash48.8
Loading Atlas data…