Atlas

Benchmarks

← All benchmarks

DeepSWE v1 (Best of up to 3 attempts)

Code · 2026-05-26

This DeepSWE v1 aggregation reports the best whole-benchmark score from up to three attempts on its 113 long-horizon software-engineering tasks. The aggregation qualifier is kept separate from the benchmark's standard pooled pass@1 and pass@4 metrics.

Top models (higher is better)

ModelScore
GPT-5.570.0
Macaron-V1-Venti58.4
Opus 4.858.0
GLM-5.254.9
MiniMax M320.0
Qwen3.7-Max18.0
Gemini 3.1 Pro Preview10.0
Loading Atlas data…