Atlas

Benchmarks

← All benchmarks

FrontierSWE: Git to Zig (Mean@5)

Code · 2026-04-16

Reimplement git v2.47.0 in Zig. FrontierSWE Mean@5 is the average percentage of verifier tests passed across up to five independent long-horizon agent runs.

Top models (higher is better)

ModelScore
Grok 4.523.0
Claude Fable 521.0
Opus 4.621.0
Opus 4.721.0
Opus 4.821.0
GPT-5.519.0
Composer 2.515.0
GLM-5.114.0
Kimi K2.613.0
GPT-5.412.0
DeepSeek-V4-Pro11.0
GLM-5.211.0
Kimi K2.59.2
Gemini 3.1 Pro Preview2.7
Qwen3.6 Plus (2026-04-02)0.0
Loading Atlas data…