OfficeQA
Professional Work · 2026-03-09
OfficeQA is a public benchmark from Databricks that evaluates end-to-end grounded reasoning over a large corpus of historical U.S. Treasury Bulletin documents. Models must locate the relevant tables across the corpus and perform precise numerical reasoning over them.
Top models (higher is better)
| Model | Score |
|---|---|
| Opus 4.7 | 86.3 |
| Claude Mythos 5 | 79.0 |
| Claude Opus 5 | 78.1 |
| Opus 4.8 | 77.6 |
| Sonnet 5 | 73.3 |
| GPT-5.4 | 68.1 |
| GPT-5.3-Codex | 65.1 |
| GPT-5.2 | 63.1 |