Atlas

Benchmarks

← All benchmarks

Finance Agent v2 (Vals Index subset)

Professional Work · 2026-07-01

Finance Agent v2 evaluated on the Vals Index subset using the benchmark's shared six-tool agent harness and partial-credit metric. Vals reports the component as an average over three model runs; it remains a protocol-specific child rather than being relabeled as the full benchmark.

Top models (higher is better)

ModelScore
Gemini 3.5 Flash58.1
Muse Spark 1.157.1
Claude Opus 556.6
Claude Fable 556.3
Kimi K355.9
GPT-5.6 Luna55.6
Gemini 3.6 Flash55.3
GPT-5.6 Sol55.0
Opus 4.853.9
GPT-5.6 Terra52.6
Sonnet 552.0
GPT-5.551.8
Opus 4.751.5
Sonnet 4.651.0
GLM-5.250.8
Grok 4.548.8
Qwen3.7-Max48.4
MiniMax M348.0
Gemini 3.5 Flash-Lite46.6
Inkling46.0
GPT-5.4 Mini45.4
Kimi K2.644.9
GLM-5.144.8
DeepSeek-V4-Pro44.1
Gemini 3.1 Pro Preview43.0
Gemini 3 Flash Preview42.5
Inkling-Small42.1
Qwen3.7-Plus41.2
MiMo-V2.5-Pro41.2
Qwen3.6 Plus (2026-04-02)40.9
GPT-5.4 Nano38.2
Nemotron 3 Ultra 550B A55B38.2
Grok 4.337.8
MiMo-V2.537.6
Mistral Medium 3.531.5
Haiku 4.531.0
Gemini 3.1 Flash-Lite Preview30.1
Grok 4.2028.6
MiniMax M2.727.9
Laguna M.124.6
Loading Atlas data…