Atlas

Benchmarks

← All benchmarks

TaxBench Total pass^5

Professional Work · 2026-05-04

Overall TaxBench pass^5 reliability score, computed across tax knowledge, tax calculations, and data retrieval task groups. The metric emphasizes consistent success across repeated attempts.

Top models (higher is better)

ModelScore
GPT-5.5 Pro29.3
GPT-5.4 Pro27.0
GPT-5.524.4
Opus 4.621.4
Gemini 3.1 Pro Preview20.1
Grok 4.1 Fast19.6
GPT-5.2 Pro17.8
Gemini 3.1 Flash (benchmark label)15.9
Grok 4.2015.2
Opus 4.714.4
Sonnet 4.611.2
GPT-5.49.3
Gemini 2.5 Pro9.0
Sonnet 4.58.0
GPT-5.24.6
Loading Atlas data…