Atlas

Benchmarks

← All benchmarks

Snorkel Agentic Coding

Code · 2026-06-05

Snorkel Agentic Coding evaluates coding agents on 100 multi-step software-engineering tasks in sandboxed execution environments. Tasks span four difficulty tiers and use human-validated reference solutions, unit tests, and trajectory-aware rubrics; the leaderboard reports the aggregate rubric score.

Top models (higher is better)

ModelScore
Opus 4.665.2
Opus 4.558.0
Sonnet 4.557.6
Gemini 3 Pro Preview51.6
GPT-5.249.4
GPT-545.2
Kimi K2 Thinking36.8
Devstral 233.2
Grok 4.1 Fast25.2
Qwen3-Coder-480B-A35B-Instruct18.8
Mistral Large 3 675B Instruct 251213.8
Loading Atlas data…