GDP.xlsx
Professional Work
Surge AI benchmark for professional spreadsheet reasoning. Models answer real workflow questions from workbooks, including finance, engineering and clinical data dictionaries. The public leaderboard reports rubric-based percentage scores; task count and execution-tool settings are not disclosed on the leaderboard.
Top models (higher is better)
| Model | Score |
|---|---|
| Gemini 4 Argon | 38.3 |
| Claude Opus 5.5 | 30.3 |
| Claude Sonnet 5.5 | 29.1 |
| Claude Fable 5.1 | 23.1 |
| GPT-6 Astra | 22.9 |
| Muse Spark 1.3 | 22.6 |
| Grok 4.7 | 22.3 |
| GPT-6.1 Sol | 22.0 |
| GLM-5.3 | 21.4 |
| Gemini 3.8 Flash | 18.3 |
| GPT-6 Sol | 17.1 |
| Sonnet 5 | 16.9 |
| GPT-6 Luna | 14.9 |
| Kimi K3 | 13.4 |
| Gemini 3.1 Pro Preview | 10.3 |
| Mistral Large 3 675B Instruct 2512 | 0.0 |