Atlas

Benchmarks

← All benchmarks

Snorkel SWE-bench CLI pass@5

Code

Published pass@5 on the 200-task SWE-bench CLI release; a correlated diagnostic of the pass@1 leaderboard.

Top models (higher is better)

ModelScore
Claude Opus 5.529.7
Claude Fable 5.126.2
Claude Opus 524.7
GPT-6 Astra18.8
Grok 4.717.1
Gemini 3.8 Flash14.1
Grok 4.613.0
Kimi K311.9
GLM-5.310.8
Muse Spark 1.35.7
Loading Atlas data…