Atlas

Benchmarks

← All benchmarks

Tax Agent Bench — Fact-Pattern Analysis

Professional Work · 2026-09-23

Weighted rubric and citation credit on 43 fact-pattern tax questions, using the parent evaluation protocol.

Top models (higher is better)

ModelScore
Claude Fable 5.179.3
Claude Opus 577.3
Muse Spark 1.375.1
Claude Opus 5.572.9
Gemini 3.8 Flash68.0
GPT-6 Astra66.0
MiMo-V2.6-Flash66.0
Grok 4.666.0
Grok 4.764.3
GPT-5.6 Terra63.3
Kimi K362.2
DeepSeek V4.1 Flash60.6
GPT-5.6 Sol60.4
GPT-6 Luna60.2
Sonnet 559.9
DeepSeek V4 Pro 081359.7
GPT-5.6 Luna59.6
GPT-5.559.2
MiMo-V2.6-Pro58.4
GLM-5.358.2
Hy4 Preview55.8
GPT-6 Sol49.9
Gemini 3.7 Flash49.3
Inkling44.1
MiniMax M333.8
Mercury 2.511.7
Loading Atlas data…