Atlas

Benchmarks

← All benchmarks

Tax Agent Bench — Fact-Pattern All-Pass

Professional Work · 2026-09-23

All-rubric-item pass rate on the 43 fact-pattern questions.

Top models (higher is better)

ModelScore
Claude Fable 5.151.2
Claude Opus 548.8
Claude Opus 5.546.5
Muse Spark 1.346.5
MiMo-V2.6-Flash41.9
GPT-5.6 Terra37.2
Gemini 3.8 Flash34.9
GLM-5.334.9
Grok 4.734.9
Kimi K332.6
DeepSeek V4 Pro 081330.2
GPT-6 Luna30.2
Hy4 Preview30.2
MiMo-V2.6-Pro30.2
Sonnet 527.9
DeepSeek V4.1 Flash27.9
GPT-5.527.9
GPT-6 Astra27.9
Grok 4.627.9
Inkling27.9
GPT-5.6 Luna25.6
GPT-5.6 Sol23.3
Gemini 3.7 Flash18.6
GPT-6 Sol16.3
MiniMax M316.3
Mercury 2.52.3
Loading Atlas data…