Atlas

Benchmarks

← All benchmarks

Tax Agent Bench — Rule & Source Lookup

Professional Work · 2026-09-23

Weighted rubric and citation credit on 30 rule-and-source tax questions, using the parent evaluation protocol.

Top models (higher is better)

ModelScore
GLM-5.381.3
Claude Opus 578.0
Claude Fable 5.172.9
Muse Spark 1.367.7
Grok 4.667.5
Sonnet 561.5
Gemini 3.8 Flash60.2
MiMo-V2.6-Pro59.8
Kimi K358.4
Gemini 3.7 Flash57.9
MiniMax M357.9
Hy4 Preview57.5
DeepSeek V4.1 Flash56.9
MiMo-V2.6-Flash56.3
Claude Opus 5.556.2
GPT-5.6 Luna54.3
GPT-5.6 Sol53.8
GPT-5.6 Terra53.3
Grok 4.752.2
DeepSeek V4 Pro 081349.8
GPT-6 Astra49.4
GPT-6 Luna45.9
GPT-5.544.2
GPT-6 Sol39.8
Inkling31.3
Mercury 2.58.3
Loading Atlas data…