Atlas

Benchmarks

← All benchmarks

Excel Modeling Benchmark Scratch

Professional Work · 2026-07-01

Scratch-mode score on Excel Modeling Benchmark, where agents build a financial model without a template workbook and are judged against the task rubric.

Top models (higher is better)

ModelScore
Claude Opus 572.3
Claude Fable 570.0
GPT-5.6 Sol69.1
GPT-5.6 Luna68.8
Kimi K368.2
Opus 4.867.4
Grok 4.567.3
GPT-5.566.8
GLM-5.266.7
Gemini 3.6 Flash65.7
Sonnet 565.0
Sonnet 4.664.8
GPT-5.6 Terra64.7
Gemini 3.5 Flash64.0
Muse Spark 1.163.3
Qwen3.7-Max61.4
MiniMax M361.0
Kimi K2.659.7
MiMo-V2.5-Pro57.9
Gemini 3.1 Pro Preview57.7
Qwen3.7-Plus57.0
DeepSeek-V4-Pro56.7
GPT-5.4 Mini54.7
Inkling-Small45.1
Gemini 3.5 Flash-Lite44.9
GPT-5.4 Nano43.8
Inkling43.0
Grok 4.329.4
Gemini 3.1 Flash-Lite Preview9.1
Loading Atlas data…