Atlas

Benchmarks

← All benchmarks

Tax Agent Bench — Controversy & Precedence

Professional Work · 2026-09-23

Weighted rubric and citation credit on 30 tax controversy-and-precedence questions, using the parent evaluation protocol.

Top models (higher is better)

ModelScore
GLM-5.371.1
Claude Fable 5.170.6
Claude Opus 570.0
Kimi K368.8
GPT-5.6 Sol68.5
Grok 4.668.1
Claude Opus 5.566.9
MiMo-V2.6-Pro65.7
GPT-5.6 Terra64.8
Grok 4.762.7
GPT-5.6 Luna61.2
Muse Spark 1.360.9
Hy4 Preview59.1
Gemini 3.8 Flash59.0
GPT-6 Astra58.6
Sonnet 557.4
MiMo-V2.6-Flash55.7
GPT-5.554.9
GPT-6 Luna54.4
DeepSeek V4.1 Flash52.8
MiniMax M352.4
DeepSeek V4 Pro 081351.5
GPT-6 Sol50.7
Inkling46.7
Gemini 3.7 Flash46.0
Mercury 2.55.9
Loading Atlas data…