Atlas

Benchmarks

← All benchmarks

OfficeQA

Professional Work · 2026-03-09

OfficeQA is a public benchmark from Databricks that evaluates end-to-end grounded reasoning over a large corpus of historical U.S. Treasury Bulletin documents. Models must locate the relevant tables across the corpus and perform precise numerical reasoning over them.

Top models (higher is better)

ModelScore
Opus 4.786.3
Claude Mythos 579.0
Claude Opus 578.1
Opus 4.877.6
Sonnet 573.3
GPT-5.468.1
GPT-5.3-Codex65.1
GPT-5.263.1
Loading Atlas data…