Atlas

Benchmarks

← All benchmarks

TaxEval v2 - Stepwise Reasoning

Professional Work · 2025-04-18

The Stepwise Reasoning split of TaxEval v2. This child benchmark separates a source-reported subtask or subtrack from the parent aggregate so scores at different grains do not share one benchmark_slug.

Top models (higher is better)

ModelScore
Muse Spark 1.192.6
Muse Spark89.8
Sonnet 4.688.4
Claude Opus 588.3
Claude Fable 588.1
Grok 387.9
Sonnet 587.2
Opus 4.687.2
Inkling87.1
Qwen3.7-Max87.1
Kimi K2.586.9
Gemini 3.5 Flash86.8
Opus 4.786.8
Opus 4.886.8
Kimi K2.686.6
GPT-5.6 Terra86.5
GPT-4.186.3
Gemini 3.6 Flash86.2
Kimi K386.1
Opus 4.585.9
GPT-5.6 Luna85.8
GPT-5 Mini85.7
Qwen3 Max Preview85.5
Gemini 3 Flash Preview85.4
MiniMax M385.4
Qwen3.6 Plus (2026-04-02)85.3
Qwen3 Max (rolling alias)85.3
Sonnet 4.585.2
GPT-5.285.2
o385.1
O4 Mini85.1
Gemini 3.1 Pro Preview85.0
GPT-5.184.9
Grok 4.2084.9
Claude 3.7 Sonnet84.8
GLM-5.284.6
Grok 4.1 Fast84.6
GPT-5.6 Sol84.5
Mistral Large 3 675B Instruct 251284.5
Nemotron 3 Ultra 550B A55B84.5
Loading Atlas data…