Atlas

Benchmarks

← All benchmarks

FrontierSWE: Granite Mamba2 Inference Optimization (Best@5)

Code · 2026-04-16

Make a pinned Granite hybrid Mamba2 layer faster on B200 without changing semantics. FrontierSWE Best@5 is the highest official task score among up to five independent long-horizon agent runs.

Top models (higher is better)

ModelScore
GPT-5.51.7
Claude Fable 51.4
Opus 4.81.2
GLM-5.11.2
GLM-5.21.1
DeepSeek-V4-Pro1.1
GPT-5.41.0
Composer 2.51.0
Kimi K2.61.0
Gemini 3.1 Pro Preview1.0
Grok 4.51.0
Opus 4.70.7
Opus 4.60.6
Qwen3.6 Plus (2026-04-02)0.6
Kimi K2.50.6
Loading Atlas data…