Atlas

Benchmarks

← All benchmarks

Vals Index v2.1

Indexes · 2026-09-25

Vals Index v2.1 combines GDP-weighted sector means: finance (Finance Agent v2 and EMB, weight 8.0), coding (Terminal-Bench 4.0, Vibe Code Bench and the fixed Code Migration subset, weight 5.6), legal (Legal Research and HLAB, weight 1.2), and Tax Agent Bench (weight 0.5), divided by 15.3. It replaces Terminal-Bench 2.1 and adds tax relative to v2. Component runs use their respective harnesses. Mixed-model fallback aggregates are excluded from single-model observations.

Top models (higher is better)

ModelScore
Gemini 4 Argon68.9
GPT-6 Astra63.1
GPT-6.1 Sol61.2
Muse Spark 1.358.2
GPT-5.6 Sol58.0
GPT-6 Sol57.5
MiMo-V2.6-Pro55.2
Opus 4.855.1
Grok 4.754.9
Gemini 3.8 Flash54.8
GLM-5.353.5
MiMo-V2.6-Flash53.2
GPT-5.6 Terra53.1
Grok 4.652.1
Sonnet 551.8
GPT-5.6 Luna51.7
DeepSeek V4.1 Flash51.3
Gemini 3.7 Flash51.3
GPT-6 Luna51.2
Kimi K350.7
Hy4 Preview49.9
Muse Spark 1.249.3
Qwen3.8-Max48.3
DeepSeek-V4-Flash-073148.0
DeepSeek V4 Pro 081347.6
Gemini 3.5 Flash44.8
Grok 4.544.7
DeepSeek-V4-Pro38.6
MiniMax M336.5
MiMo-V2.5-Pro33.8
Gemini 3.1 Pro Preview33.4
GPT-5.4 Mini33.2
Inkling28.7
Inkling-Small25.5
Mercury 2.58.9
Loading Atlas data…