Atlas

Benchmarks

← All benchmarks

FrontierSWE: Wan 2.1 on MAX/Mojo (Best@5)

Code · 2026-04-16

Implement Wan 2.1 text-to-video inference on Modular's MAX/Mojo stack. FrontierSWE Best@5 is the highest verifier score among up to five independent long-horizon agent runs.

Top models (higher is better)

ModelScore
Claude Fable 575.0
Opus 4.875.0
GLM-5.275.0
Grok 4.575.0
Opus 4.650.0
GPT-5.450.0
Opus 4.70.0
Composer 2.50.0
DeepSeek-V4-Pro0.0
Gemini 3.1 Pro Preview0.0
GLM-5.10.0
GPT-5.50.0
Kimi K2.50.0
Kimi K2.60.0
Qwen3.6 Plus (2026-04-02)0.0
Loading Atlas data…