Atlas

Benchmarks

← All benchmarks

TaxCalcBench TY24

Professional Work · 2025-06-22

TaxCalcBench TY24 evaluates whether a language model can calculate complete U.S. federal tax returns from structured taxpayer data. Its headline strict score is the percentage of 51 cases for which every evaluated Form 1040 line exactly matches the expected return.

Top models (higher is better)

ModelScore
GPT-5.4 Pro62.8
GPT-5.462.8
Opus 4.652.9
Gemini 3.1 Pro Preview49.0
GPT-541.7
GPT-5.2 Pro41.2
Sonnet 4.637.3
Opus 4.536.3
Gemini 3 Pro Preview36.3
GPT-5.233.8
Gemini 2.5 Pro32.4
Sonnet 4.531.4
Opus 4.128.4
Opus 427.4
Gemini 2.5 Flash26.0
Sonnet 423.0
Haiku 4.513.7
Loading Atlas data…