Atlas

Benchmarks

← All benchmarks

FrontierSWE: FrogsGame Post-Training (Best@5)

Code · 2026-04-16

Post-train Qwen3-8B to solve FrogsGame boards through tool use. FrontierSWE Best@5 is the highest solve score among up to five independent long-horizon agent runs.

Top models (higher is better)

ModelScore
Claude Fable 568.0
Grok 4.54.4
GLM-5.24.0
GPT-5.54.0
Opus 4.83.8
GPT-5.43.4
Opus 4.63.0
Opus 4.72.4
Composer 2.52.4
DeepSeek-V4-Pro2.0
GLM-5.11.8
Kimi K2.61.8
Gemini 3.1 Pro Preview0.0
Kimi K2.50.0
Qwen3.6 Plus (2026-04-02)0.0
Loading Atlas data…