Atlas

Benchmarks

← All benchmarks

FrontierSWE: Notebook Compression (Mean@5)

Code · 2026-04-16

Build a lossless domain-specific compressor for canonicalized Jupyter notebooks. FrontierSWE Mean@5 is the average official task score across up to five independent long-horizon agent runs.

Top models (higher is better)

ModelScore
Claude Fable 50.7
GLM-5.20.7
GPT-5.50.7
Grok 4.50.7
Gemini 3.1 Pro Preview0.7
Opus 4.80.6
Opus 4.70.5
GPT-5.40.5
Qwen3.6 Plus (2026-04-02)0.5
Opus 4.60.4
Kimi K2.60.4
Kimi K2.50.4
GLM-5.10.3
Composer 2.50.0
DeepSeek-V4-Pro0.0
Loading Atlas data…