Atlas

Benchmarks

← All benchmarks

DeepSWE v1

Code · 2026-05-26

DeepSWE v1 evaluates frontier coding agents on 113 original, long-horizon software-engineering tasks across 91 repositories. Pass@1 is the pooled pass rate over scored rollout attempts; context-window failures and agent timeouts count as failures, while provider, verifier, and network errors are excluded.

Top models (higher is better)

ModelScore
GPT-5.570.0
Opus 4.858.2
GPT-5.455.5
Opus 4.754.2
GLM-5.241.5
Sonnet 4.631.8
Gemini 3.5 Flash28.3
Opus 4.627.6
GPT-5.4 Mini24.3
Kimi K2.623.9
MiniMax M320.4
MiMo-V2.5-Pro19.5
Qwen3.7-Max17.7
GLM-5.117.5
Grok Build 0.113.1
Gemini 3.1 Pro Preview9.7
DeepSeek-V4-Pro7.5
Gemini 3 Flash Preview5.1
Qwen3.6 Plus (2026-04-02)2.7
Haiku 4.50.2
MiniMax M2.70.2
Loading Atlas data…