TaxCalcBench TY24
Professional Work · 2025-06-22
TaxCalcBench TY24 evaluates whether a language model can calculate complete U.S. federal tax returns from structured taxpayer data. Its headline strict score is the percentage of 51 cases for which every evaluated Form 1040 line exactly matches the expected return.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-5.4 Pro | 62.8 |
| GPT-5.4 | 62.8 |
| Opus 4.6 | 52.9 |
| Gemini 3.1 Pro Preview | 49.0 |
| GPT-5 | 41.7 |
| GPT-5.2 Pro | 41.2 |
| Sonnet 4.6 | 37.3 |
| Opus 4.5 | 36.3 |
| Gemini 3 Pro Preview | 36.3 |
| GPT-5.2 | 33.8 |
| Gemini 2.5 Pro | 32.4 |
| Sonnet 4.5 | 31.4 |
| Opus 4.1 | 28.4 |
| Opus 4 | 27.4 |
| Gemini 2.5 Flash | 26.0 |
| Sonnet 4 | 23.0 |
| Haiku 4.5 | 13.7 |