Atlas

Benchmarks

← All benchmarks

Snorkel SWE-bench CLI

Code

Pass@1 on 200 repository engineering tasks spanning eleven languages, ten technical domains and eight task types. Harbor runs fail-to-pass and pass-to-pass tests in reproducible network-disabled Docker environments; credit requires all required tests and substantive changes spanning at least two files. Native effort and agent scaffold are not disclosed.

Top models (higher is better)

ModelScore
Claude Opus 5.516.6
Claude Fable 5.114.5
Claude Opus 514.0
GPT-6 Astra12.4
Gemini 3.8 Flash6.3
Grok 4.76.1
Grok 4.65.7
Kimi K34.1
GLM-5.33.5
Muse Spark 1.32.2
Loading Atlas data…