Atlas

Benchmarks

← All benchmarks

Tax Agent Bench — All-Pass

Professional Work · 2026-09-23

Percentage of the 193 held-out questions passing every rubric item, without partial credit.

Top models (higher is better)

ModelScore
Claude Fable 5.149.2
Claude Opus 5.545.6
Claude Opus 545.6
GLM-5.339.9
Muse Spark 1.336.8
Grok 4.734.2
Kimi K333.7
Grok 4.633.2
Hy4 Preview32.6
MiMo-V2.6-Pro32.6
Gemini 3.8 Flash32.1
GPT-5.6 Sol31.1
MiMo-V2.6-Flash30.6
Sonnet 530.1
GPT-5.6 Terra30.1
DeepSeek V4 Pro 081328.5
DeepSeek V4.1 Flash27.5
GPT-5.6 Luna24.9
Gemini 3.7 Flash22.8
GPT-5.521.8
MiniMax M321.2
GPT-6 Astra20.7
GPT-6 Luna20.2
Inkling17.6
GPT-6 Sol15.0
Mercury 2.53.6
Loading Atlas data…