Atlas

Benchmarks

← All benchmarks

GDP.xlsx

Professional Work

Surge AI benchmark for professional spreadsheet reasoning. Models answer real workflow questions from workbooks, including finance, engineering and clinical data dictionaries. The public leaderboard reports rubric-based percentage scores; task count and execution-tool settings are not disclosed on the leaderboard.

Top models (higher is better)

ModelScore
Gemini 4 Argon38.3
Claude Opus 5.530.3
Claude Sonnet 5.529.1
Claude Fable 5.123.1
GPT-6 Astra22.9
Muse Spark 1.322.6
Grok 4.722.3
GPT-6.1 Sol22.0
GLM-5.321.4
Gemini 3.8 Flash18.3
GPT-6 Sol17.1
Sonnet 516.9
GPT-6 Luna14.9
Kimi K313.4
Gemini 3.1 Pro Preview10.3
Mistral Large 3 675B Instruct 25120.0
Loading Atlas data…