Atlas

Benchmarks

← All benchmarks

FrontierSWE: Cranelift Codegen Optimization (Best@5)

Code · 2026-04-16

Speed up Wasmtime's Cranelift backend without breaking correctness. FrontierSWE Best@5 is the highest official task score among up to five independent long-horizon agent runs.

Top models (higher is better)

ModelScore
Gemini 3.1 Pro Preview1.0
GPT-5.51.0
Opus 4.71.0
Grok 4.51.0
Opus 4.81.0
Composer 2.51.0
Claude Fable 51.0
Opus 4.61.0
DeepSeek-V4-Pro1.0
GLM-5.21.0
Kimi K2.51.0
GPT-5.41.0
GLM-5.10.0
Kimi K2.60.0
Qwen3.6 Plus (2026-04-02)0.0
Loading Atlas data…