Atlas

Benchmarks

← All benchmarks

FrontierSWE: Pyright Type Checking Optimization (Mean@5)

Code · 2026-04-16

Make pyright faster without changing diagnostics. FrontierSWE Mean@5 is the average official task score across up to five independent long-horizon agent runs.

Top models (higher is better)

ModelScore
Claude Fable 51.2
Grok 4.51.1
GLM-5.21.1
Gemini 3.1 Pro Preview1.1
GPT-5.51.1
GPT-5.41.0
Qwen3.6 Plus (2026-04-02)1.0
Kimi K2.51.0
GLM-5.10.9
Opus 4.80.9
Opus 4.60.8
DeepSeek-V4-Pro0.8
Opus 4.70.5
Composer 2.50.4
Kimi K2.60.4
Loading Atlas data…