Atlas

Benchmarks

← All benchmarks

Tax Agent Bench — Numbers & Calculations

Professional Work · 2026-09-23

Weighted rubric and citation credit on 30 quantitative tax questions, using the parent evaluation protocol.

Top models (higher is better)

ModelScore
Claude Opus 5.588.9
Claude Fable 5.188.0
Claude Opus 585.2
GLM-5.384.5
Kimi K384.0
Gemini 3.8 Flash83.1
Muse Spark 1.381.8
GPT-5.6 Terra81.6
GPT-5.6 Sol81.4
Hy4 Preview81.0
Grok 4.679.8
Gemini 3.7 Flash78.8
GPT-6 Astra77.9
GPT-5.576.8
Grok 4.776.7
DeepSeek V4.1 Flash76.6
MiMo-V2.6-Pro75.5
GPT-6 Luna73.8
GPT-5.6 Luna71.8
Sonnet 570.6
DeepSeek V4 Pro 081370.2
GPT-6 Sol69.5
MiMo-V2.6-Flash62.4
Inkling61.9
MiniMax M354.9
Mercury 2.531.1
Loading Atlas data…