Atlas

Benchmarks

← All benchmarks

FrontierSWE: Git to Zig (Best@5)

Code · 2026-04-16

Reimplement git v2.47.0 in Zig. FrontierSWE Best@5 is the highest percentage of verifier tests passed among up to five independent long-horizon agent runs.

Top models (higher is better)

ModelScore
GLM-5.228.0
Grok 4.525.0
Opus 4.623.0
Claude Fable 522.0
Opus 4.722.0
Opus 4.822.0
DeepSeek-V4-Pro22.0
GPT-5.522.0
Composer 2.521.0
GLM-5.119.0
Kimi K2.517.0
Kimi K2.617.0
GPT-5.416.0
Gemini 3.1 Pro Preview13.0
Qwen3.6 Plus (2026-04-02)0.1
Loading Atlas data…