Atlas

Benchmarks

← All benchmarks

FrontierSWE: Lua Native Compiler (Mean@5)

Code · 2026-04-16

Build a real AOT compiler from Lua 5.4 source to native x86-64 ELF. FrontierSWE Mean@5 is the average percentage of verifier tests passed across up to five independent long-horizon agent runs.

Top models (higher is better)

ModelScore
Claude Fable 599.0
Opus 4.786.0
Opus 4.871.0
GLM-5.271.0
GPT-5.469.0
GPT-5.552.0
Opus 4.649.0
DeepSeek-V4-Pro7.8
Composer 2.57.5
Gemini 3.1 Pro Preview7.5
Kimi K2.50.5
GLM-5.10.0
Grok 4.50.0
Kimi K2.60.0
Qwen3.6 Plus (2026-04-02)0.0
Loading Atlas data…