Snorkel Agentic Coding 2.0
Code
Pass@1 on Snorkel's 200 frontier terminal tasks across nine task types and nine languages. Harbor evaluates deterministic tests, required outputs and trajectory rubrics in network-disabled Docker environments, with 30-minute agent and verifier limits and five attempts per task. Native model effort is not disclosed.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-6 Astra | 47.6 |
| Claude Opus 5.5 | 40.4 |
| Claude Fable 5.1 | 39.6 |
| Claude Opus 5 | 38.9 |
| Grok 4.6 | 33.6 |
| Gemini 3.8 Flash | 32.6 |
| Grok 4.7 | 31.9 |
| Kimi K3 | 26.1 |
| Muse Spark 1.3 | 25.0 |
| GLM-5.3 | 24.9 |