Atlas

Benchmarks

← All benchmarks

SWE-bench Verified (Vals Index 102-task subset)

Code · 2026-07-01

SWE-bench Verified evaluated on the 102-task Vals Index subset: 33 sampled tasks from each of the first three human-duration bands plus all three tasks in the hardest band. It uses the Vals mini-swe-agent bash harness and reports issue-resolution rate.

Top models (higher is better)

ModelScore
GPT-5.6 Sol99.0
Claude Fable 597.1
Claude Opus 595.1
Kimi K395.1
GPT-5.6 Luna93.1
Opus 4.889.2
Grok 4.586.3
GLM-5.283.3
GPT-5.582.6
Opus 4.782.4
Muse Spark 1.178.4
Sonnet 4.677.5
Gemini 3.6 Flash77.5
Sonnet 575.5
DeepSeek-V4-Pro75.5
Gemini 3.5 Flash75.5
Inkling75.5
Inkling-Small74.5
Kimi K2.674.5
GPT-5.6 Terra73.5
Gemini 3.1 Pro Preview72.5
GLM-5.172.5
Grok 4.370.6
Qwen3.6 Plus (2026-04-02)70.6
MiniMax M2.769.6
MiniMax M369.6
Gemini 3.5 Flash-Lite68.6
Gemini 3 Flash Preview68.6
GPT-5.4 Mini67.6
GPT-5.4 Nano67.6
Nemotron 3 Ultra 550B A55B67.6
Grok 4.2066.7
MiMo-V2.5-Pro66.7
Qwen3.7-Max66.7
Qwen3.7-Plus66.7
Haiku 4.564.7
MiMo-V2.564.7
Mistral Medium 3.564.7
Gemini 3.1 Flash-Lite Preview57.8
Laguna M.151.0
Loading Atlas data…