Atlas

Benchmarks

← All benchmarks

MortgageTax - Semantic Extraction

Professional Work · 2025-03-05

The Semantic Extraction split of MortgageTax. This child benchmark separates a source-reported subtask or subtrack from the parent aggregate so scores at different grains do not share one benchmark_slug.

Top models (higher is better)

ModelScore
Claude Opus 562.7
Gemini 3.1 Pro Preview62.4
Gemini 3 Pro Preview62.2
Opus 4.761.8
Gemini 3 Flash Preview61.7
Gemini 3.5 Flash-Lite61.5
Sonnet 561.4
GPT-5.561.4
Claude Fable 561.1
Sonnet 4.561.1
Grok 4.560.8
Qwen3.6 27B60.8
GPT-5.460.7
Sonnet 4.660.5
GPT-5.6 Terra60.3
Gemini 2.5 Pro60.3
Opus 4.660.1
Opus 4.860.1
Opus 4.559.9
Qwen3.5-Flash59.8
Gemini 2.5 Pro Experimental 03-2559.7
Claude 3.7 Sonnet59.6
GPT-5.6 Sol59.5
MiniMax M359.5
Gemini 3.1 Flash-Lite Preview59.1
GPT-5.6 Luna59.0
Grok 4.359.0
Opus 4.158.8
Gemini 3.5 Flash58.7
GPT-4.1 Mini58.7
O4 Mini58.7
Qwen3.5 Plus (2026-02-15)58.7
Sonnet 458.4
Gemini 3.6 Flash58.4
GPT-5 Mini58.3
Opus 458.2
Kimi K2.558.1
Muse Spark 1.157.8
Qwen3.6 Plus (2026-04-02)57.8
Gemini 2.5 Flash Preview 04-1757.7
Loading Atlas data…