Atlas

Benchmarks

← All benchmarks

FrontierSWE: Dart → Haskell (Mean@5)

Code · 2026-04-16

Reimplement the dart_style formatter as a standalone Haskell executable. FrontierSWE Mean@5 is the average percentage of verifier tests passed across up to five independent long-horizon agent runs.

Top models (higher is better)

ModelScore
Claude Fable 599.0
Grok 4.523.0
Opus 4.822.0
GPT-5.519.0
Opus 4.716.0
GLM-5.215.0
GPT-5.414.0
Composer 2.53.1
Opus 4.63.0
Kimi K2.62.9
GLM-5.12.3
Gemini 3.1 Pro Preview2.2
Kimi K2.51.1
DeepSeek-V4-Pro0.4
Qwen3.6 Plus (2026-04-02)0.2
Loading Atlas data…