Atlas

Benchmarks

← All benchmarks

Vals Index v2

Indexes · 2026-08-13

Vals Index v2 combines sector means with GDP weights 8.0 for finance, 5.6 for coding and 1.2 for legal, divided by 14.8. Finance averages Finance Agent v2 and EMB; coding averages Terminal-Bench 2.1, Vibe Code Bench and the Code Migration subset; legal averages Legal Research and HLAB. The changed composition is kept separate from the earlier index.

Top models (higher is better)

ModelScore
Claude Opus 5.569.7
Claude Fable 5.168.8
Claude Opus 567.2
GPT-6 Astra66.6
Claude Fable 566.0
Muse Spark 1.364.5
GPT-5.6 Sol63.7
GPT-6 Sol62.6
Gemini 3.8 Flash62.3
Opus 4.860.9
Grok 4.760.2
GPT-5.6 Luna59.9
Sonnet 559.6
GPT-5.6 Terra59.6
MiMo-V2.6-Flash59.6
MiMo-V2.6-Pro59.5
Gemini 3.7 Flash59.3
Grok 4.659.2
GPT-6 Luna58.5
DeepSeek V4.1 Flash57.9
Kimi K357.8
GPT-5.557.4
Muse Spark 1.257.1
GLM-5.357.0
Opus 4.756.1
Hy4 Preview55.4
Gemini 3.6 Flash55.4
Muse Spark 1.154.8
DeepSeek-V4-Flash-073153.6
GLM-5.253.1
Gemini 3.5 Flash53.1
DeepSeek V4 Pro 081352.4
Qwen3.8-Max51.8
Grok 4.551.5
Sonnet 4.650.6
Qwen3.8 27B48.5
GLM-5.3 Flash47.2
Qwen3.7-Max44.8
Kimi K2.643.5
DeepSeek-V4-Pro42.9
Loading Atlas data…