Atlas

Benchmarks

← All benchmarks

DeepSWE v1 pass@4

Code · 2026-05-26

DeepSWE v1 pass@4 is the percentage of attempted tasks with at least one passing scored rollout among four nominal runs. Tasks with only excluded provider, verifier, or network errors do not enter the denominator.

Top models (higher is better)

ModelScore
GPT-5.588.3
Opus 4.785.8
Opus 4.880.5
GPT-5.477.0
GLM-5.269.9
Sonnet 4.661.9
Gemini 3.5 Flash56.6
Opus 4.650.4
Kimi K2.648.7
MiniMax M348.7
GPT-5.4 Mini46.0
MiMo-V2.5-Pro45.1
Qwen3.7-Max41.6
GLM-5.138.9
Grok Build 0.129.2
Gemini 3.1 Pro Preview24.8
DeepSeek-V4-Pro18.6
Gemini 3 Flash Preview15.0
Qwen3.6 Plus (2026-04-02)9.7
Haiku 4.50.9
MiniMax M2.70.9
Loading Atlas data…