Atlas

Benchmarks

← All benchmarks

Tax Agent Bench

Professional Work · 2026-09-23

193 held-out tax-research questions evaluated with web search, authority lookup, document fetching, information retrieval and calculation tools. GPT-5.4 judges the final answer against weighted expert rubrics, with zero credit for a failed must-pass item. Rubric credit is multiplied by 0.7 + 0.3 times independently verified citation quality. The five public and 193 private-validation questions are excluded.

Top models (higher is better)

ModelScore
Claude Fable 5.177.6
Claude Opus 575.1
GLM-5.373.1
Muse Spark 1.371.9
Grok 4.670.8
Claude Opus 5.570.5
Kimi K368.7
GPT-5.6 Sol68.0
Gemini 3.8 Flash66.8
Grok 4.765.6
GPT-5.6 Terra65.2
MiMo-V2.6-Pro64.9
Hy4 Preview63.7
GPT-6 Astra63.3
DeepSeek V4.1 Flash62.5
Sonnet 562.3
GPT-5.6 Luna60.8
GPT-5.560.5
MiMo-V2.6-Flash59.9
GPT-6 Luna58.9
DeepSeek V4 Pro 081358.7
Gemini 3.7 Flash57.7
GPT-6 Sol53.0
MiniMax M349.7
Inkling43.5
Mercury 2.512.8
Loading Atlas data…