Atlas

Benchmarks

← All benchmarks

SpreadsheetBench 2

Professional Work

SpreadsheetBench 2 evaluates agents on end-to-end business spreadsheet workflows spanning financial modeling and template completion, debugging, and visualization in complex multi-sheet workbooks. Unlike the isolated manipulations of SpreadsheetBench v1, its 321 tasks require multi-step coordinated operations, cross-sheet reasoning, and deliverable-level outcomes.

Top models (higher is better)

ModelScore
Kimi K334.8
Claude Fable 534.7
GPT-5.6 Sol32.4
Opus 4.831.6
GPT-5.529.1
GLM-5.228.1
Loading Atlas data…