Atlas

Benchmarks

← All benchmarks

Tax Agent Bench — Forms & Filings

Professional Work · 2026-09-23

Weighted rubric and citation credit on 30 tax forms-and-filings questions, using the parent evaluation protocol.

Top models (higher is better)

ModelScore
Claude Fable 5.173.2
GPT-5.6 Sol72.5
Grok 4.672.1
Muse Spark 1.371.7
GLM-5.371.7
GPT-5.6 Terra68.7
Kimi K368.6
Hy4 Preview68.5
Claude Opus 567.2
Sonnet 565.2
Gemini 3.8 Flash65.1
Claude Opus 5.563.8
Grok 4.763.8
GPT-6 Luna62.9
GPT-6 Astra60.8
GPT-5.559.9
MiMo-V2.6-Pro59.1
GPT-5.6 Luna58.5
DeepSeek V4.1 Flash56.8
MiMo-V2.6-Flash54.9
GPT-6 Sol53.7
MiniMax M353.1
DeepSeek V4 Pro 081352.4
Gemini 3.7 Flash48.8
Inkling42.4
Mercury 2.56.9
Loading Atlas data…