Atlas

Benchmarks

← All benchmarks

Senior SWE-Bench

Code · 2026-06-30

Senior SWE-Bench is a Harbor-compatible software-engineering agent benchmark from Snorkel AI that evaluates agents on senior-engineer-style tasks. The initial v2026.06 release has 100 real-PR-derived tasks, with 50 public and 50 private, spanning under-specified feature and migration work plus bug and performance investigations. The leaderboard reports tasteful solve rate (pass@1), requiring runtime verifiers, validation, rubric, patch-bloat, practice, and relative-taste criteria.

Top models (higher is better)

ModelScore
Claude Fable 534.7
Claude Opus 534.7
GPT-5.6 Sol34.7
Opus 4.830.5
GPT-5.529.5
Sonnet 528.0
GPT-5.6 Terra27.4
Kimi K320.2
Opus 4.718.1
GLM-5.217.9
GPT-5.416.8
GPT-5.6 Luna8.1
Gemini 3.5 Flash6.3
Gemini 3.1 Pro Preview2.1
Loading Atlas data…