Atlas

Benchmarks

← All benchmarks

SWE-bench Verified

Code · 2024-08-13

SWE-bench Verified is a human-validated subset of 500 SWE-bench tasks drawn from real GitHub issues in Python repositories. It measures the percentage of issues resolved by a model-generated patch that passes the repository tests.

Top models (higher is better)

ModelScore
Claude Opus 597.0
GPT-5.6 Sol96.2
Claude Mythos 595.5
Claude Fable 595.0
Claude Mythos Preview93.9
Kimi K393.4
GPT-5.6 Luna93.0
Opus 4.888.6
Grok 4.586.6
Macaron-V1-Venti85.6
Sonnet 585.2
GPT-5.582.9
GLM-5.282.8
Opus 4.782.0
Muse Spark 1.182.0
Opus 4.680.8
Opus 4.580.6
Gemini 3.1 Pro Preview80.6
MiniMax M380.5
Inkling-Small80.2
GPT-5.280.0
Sonnet 4.679.6
Composer 2.579.6
Gemini 3.6 Flash79.6
Gemini 3.5 Flash79.3
DeepSeek-V4-Flash79.0
GPT-5.478.2
Kimi K2.7 Code78.2
GPT-5.3-Codex78.0
DeepSeek-V4-Pro77.6
Inkling77.6
Inkling-Small Preview77.4
Sonnet 4.577.2
Kimi K2.676.7
Qwen3.6-Max-Preview76.7
GLM-5.176.4
Gemini 3 Pro Preview76.4
Qwen3.5 397B A17B76.4
GPT-5.6 Terra75.2
Gemini 3.5 Flash-Lite75.0
Loading Atlas data…