Atlas

Benchmarks

← All benchmarks

FrontierSWE: Dependent Type Checker (Mean@5)

Code · 2026-04-16

Implement a fast dependent type checker for a Martin-Löf-style core language. FrontierSWE Mean@5 is the average official task score across up to five independent long-horizon agent runs.

Top models (higher is better)

ModelScore
Claude Fable 545.1
GLM-5.24.3
GPT-5.52.3
Grok 4.52.1
Opus 4.81.8
GPT-5.40.7
Opus 4.70.5
Opus 4.60.5
Gemini 3.1 Pro Preview0.4
DeepSeek-V4-Pro0.4
Composer 2.50.3
Kimi K2.60.3
Qwen3.6 Plus (2026-04-02)0.3
Kimi K2.50.3
GLM-5.10.2
Loading Atlas data…