Atlas

Benchmarks

← All benchmarks

Finance Agent v2 - Financial Modeling

Professional Work · 2026-05-19

The Financial Modeling split of Finance Agent v2. This child benchmark separates a source-reported subtask or subtrack from the parent aggregate so scores at different grains do not share one benchmark_slug.

Top models (higher is better)

ModelScore
Muse Spark 1.134.1
Gemini 3.5 Flash29.4
Claude Opus 526.6
Gemini 3.6 Flash25.1
Sonnet 524.2
Opus 4.823.3
Kimi K323.0
GPT-5.522.9
GPT-5.6 Luna21.6
GPT-5.6 Terra21.4
GPT-5.6 Sol20.7
Opus 4.720.1
GPT-5.4 Mini19.6
Gemini 3.5 Flash-Lite19.3
Claude Fable 517.6
Sonnet 4.617.6
Qwen3.7-Max17.0
Qwen3.6 Plus (2026-04-02)16.7
Inkling16.7
Grok 4.515.6
Gemini 3.1 Pro Preview15.1
Gemini 3 Flash Preview14.4
Kimi K2.614.2
MiniMax M314.0
GPT-5.4 Nano13.5
GLM-5.113.5
GLM-5.212.9
DeepSeek-V4-Pro10.5
Qwen3.7-Plus10.1
MiMo-V2.5-Pro9.9
Nemotron 3 Ultra 550B A55B7.1
Grok 4.35.9
MiMo-V2.55.3
Haiku 4.55.1
Mistral Medium 3.55.1
Gemini 3.1 Flash-Lite Preview4.1
MiniMax M2.74.1
Laguna M.12.6
Laguna XS.22.6
Grok 4.202.4
Loading Atlas data…