Atlas

Benchmarks

← All benchmarks

FrontierSWE: Dependent Type Checker (Best@5)

Code · 2026-04-16

Implement a fast dependent type checker for a Martin-Löf-style core language. FrontierSWE Best@5 is the highest official task score among up to five independent long-horizon agent runs.

Top models (higher is better)

ModelScore
Claude Fable 550.5
GLM-5.29.9
GPT-5.58.6
Opus 4.84.2
Grok 4.53.5
GPT-5.41.5
Opus 4.70.5
Opus 4.60.5
DeepSeek-V4-Pro0.5
Gemini 3.1 Pro Preview0.5
GLM-5.10.5
Composer 2.50.4
Kimi K2.60.4
Qwen3.6 Plus (2026-04-02)0.4
Kimi K2.50.3
Loading Atlas data…