Atlas

Benchmarks

← All benchmarks

FrontierSWE: SGLang Inference System Optimization (Best@5)

Code · 2026-04-16

Make SGLang serving for Qwen3.5-4B faster on a B200 GPU. FrontierSWE Best@5 is the highest official task score among up to five independent long-horizon agent runs.

Top models (higher is better)

ModelScore
Opus 4.81.2
Claude Fable 50.4
GLM-5.20.4
Opus 4.70.4
Kimi K2.50.4
Gemini 3.1 Pro Preview0.4
Composer 2.50.4
GLM-5.10.4
GPT-5.50.4
Qwen3.6 Plus (2026-04-02)0.4
Opus 4.60.4
DeepSeek-V4-Pro0.3
Grok 4.50.3
Kimi K2.60.3
GPT-5.40.0
Loading Atlas data…