Atlas

Benchmarks

← All benchmarks

DeepSWE v1.1

Code · 2026-06-14

DeepSWE v1.1 evaluates frontier coding agents on the same 113 original, long-horizon software-engineering tasks as v1, with committed patches graded in clean verifier containers and structured per-test reports. Pass@1 is the pooled pass rate over scored rollout attempts; context-window failures and agent timeouts count as failures, while provider, verifier, and network errors are excluded.

Top models (higher is better)

ModelScore
Claude Opus 573.6
GPT-5.6 Sol72.7
Claude Fable 570.0
GPT-5.6 Terra69.6
Kimi K369.0
GPT-5.6 Luna67.2
GPT-5.567.0
Opus 4.859.0
Sonnet 554.0
Grok 4.554.0
Muse Spark 1.153.3
GPT-5.451.8
GLM-5.246.2
Laguna S 2.140.4
Kimi K2.7 Code30.5
Sonnet 4.629.9
Gemini 3.1 Pro Preview12.0
DeepSeek-V4-Pro9.0
Loading Atlas data…