Snorkel Agentic Coding
Code · 2026-06-05
Snorkel Agentic Coding evaluates coding agents on 100 multi-step software-engineering tasks in sandboxed execution environments. Tasks span four difficulty tiers and use human-validated reference solutions, unit tests, and trajectory-aware rubrics; the leaderboard reports the aggregate rubric score.
Top models (higher is better)
| Model | Score |
|---|---|
| Opus 4.6 | 65.2 |
| Opus 4.5 | 58.0 |
| Sonnet 4.5 | 57.6 |
| Gemini 3 Pro Preview | 51.6 |
| GPT-5.2 | 49.4 |
| GPT-5 | 45.2 |
| Kimi K2 Thinking | 36.8 |
| Devstral 2 | 33.2 |
| Grok 4.1 Fast | 25.2 |
| Qwen3-Coder-480B-A35B-Instruct | 18.8 |
| Mistral Large 3 675B Instruct 2512 | 13.8 |