Atlas

Benchmarks

← All benchmarks

FrontierSWE: Pyright Type Checking Optimization (Best@5)

Code · 2026-04-16

Make pyright faster without changing diagnostics. FrontierSWE Best@5 is the highest official task score among up to five independent long-horizon agent runs.

Top models (higher is better)

ModelScore
Claude Fable 51.3
Grok 4.51.2
Opus 4.61.1
GLM-5.21.1
GLM-5.11.1
Opus 4.71.1
Qwen3.6 Plus (2026-04-02)1.1
Gemini 3.1 Pro Preview1.1
GPT-5.51.1
Opus 4.81.1
Kimi K2.61.1
GPT-5.41.1
DeepSeek-V4-Pro1.1
Kimi K2.51.0
Composer 2.50.5
Loading Atlas data…