Atlas

Benchmarks

← All benchmarks

FrontierSWE: Granite Mamba2 Inference Optimization (Mean@5)

Code · 2026-04-16

Make a pinned Granite hybrid Mamba2 layer faster on B200 without changing semantics. FrontierSWE Mean@5 is the average official task score across up to five independent long-horizon agent runs.

Top models (higher is better)

ModelScore
GPT-5.51.3
Claude Fable 51.2
Opus 4.80.8
Grok 4.50.7
GPT-5.40.7
Gemini 3.1 Pro Preview0.7
DeepSeek-V4-Pro0.6
Composer 2.50.6
GLM-5.10.6
Qwen3.6 Plus (2026-04-02)0.6
Opus 4.60.6
Kimi K2.50.6
GLM-5.20.6
Kimi K2.60.6
Opus 4.70.3
Loading Atlas data…