Atlas

Benchmarks

← All benchmarks

Snorkel Agentic Coding 2.0 pass@5

Code

Published pass@5 on the 200-task Agentic Coding 2.0 release; a correlated diagnostic of the pass@1 leaderboard.

Top models (higher is better)

ModelScore
Claude Opus 562.0
GPT-6 Astra58.7
Claude Opus 5.556.9
Grok 4.756.0
Claude Fable 5.154.8
Gemini 3.8 Flash53.1
Grok 4.653.1
Kimi K352.4
GLM-5.346.3
Muse Spark 1.340.3
Loading Atlas data…