Atlas

Benchmarks

← All benchmarks

Tax Agent Bench — Current & Temporal Analysis

Professional Work · 2026-09-23

Weighted rubric and citation credit on 30 current-and-temporal tax questions, using the parent evaluation protocol.

Top models (higher is better)

ModelScore
Claude Fable 5.181.2
GLM-5.378.2
Grok 4.774.5
GPT-5.6 Sol74.4
MiMo-V2.6-Pro73.9
Grok 4.673.4
Claude Opus 5.573.3
Muse Spark 1.373.1
Kimi K372.9
DeepSeek V4.1 Flash71.9
Claude Opus 571.6
Gemini 3.7 Flash68.8
GPT-5.568.2
DeepSeek V4 Pro 081367.9
GPT-6 Astra66.2
Gemini 3.8 Flash64.8
Hy4 Preview63.7
MiMo-V2.6-Flash61.4
GPT-5.6 Terra60.3
GPT-5.6 Luna60.1
Sonnet 560.0
GPT-6 Sol56.1
GPT-6 Luna55.3
MiniMax M352.9
Inkling34.4
Mercury 2.513.2
Loading Atlas data…