Atlas

Benchmarks

← All benchmarks

Snorkel Agentic Coding 2.0

Code

Pass@1 on Snorkel's 200 frontier terminal tasks across nine task types and nine languages. Harbor evaluates deterministic tests, required outputs and trajectory rubrics in network-disabled Docker environments, with 30-minute agent and verifier limits and five attempts per task. Native model effort is not disclosed.

Top models (higher is better)

ModelScore
GPT-6 Astra47.6
Claude Opus 5.540.4
Claude Fable 5.139.6
Claude Opus 538.9
Grok 4.633.6
Gemini 3.8 Flash32.6
Grok 4.731.9
Kimi K326.1
Muse Spark 1.325.0
GLM-5.324.9
Loading Atlas data…