Atlas

Benchmarks

← All benchmarks

Finance Agent v2 - Comparables

Professional Work · 2026-05-19

The Comparables split of Finance Agent v2. This child benchmark separates a source-reported subtask or subtrack from the parent aggregate so scores at different grains do not share one benchmark_slug.

Top models (higher is better)

ModelScore
Gemini 3.5 Flash48.8
GPT-5.6 Luna48.3
Gemini 3.6 Flash47.8
Claude Opus 547.4
Muse Spark 1.147.2
Kimi K346.7
Sonnet 545.8
Opus 4.745.6
GPT-5.6 Sol45.3
GPT-5.6 Terra45.2
Gemini 3.5 Flash-Lite45.2
Opus 4.844.9
GLM-5.244.8
Claude Fable 544.0
Qwen3.6 Plus (2026-04-02)43.7
GPT-5.543.7
Inkling42.8
Qwen3.7-Max42.3
Grok 4.541.8
MiniMax M341.7
GLM-5.141.6
Kimi K2.641.1
Sonnet 4.641.0
GPT-5.4 Mini40.3
MiMo-V2.5-Pro40.1
DeepSeek-V4-Pro39.6
Qwen3.7-Plus38.8
Gemini 3 Flash Preview38.6
Nemotron 3 Ultra 550B A55B37.2
GPT-5.4 Nano37.2
Gemini 3.1 Pro Preview36.7
Grok 4.335.9
Gemini 3.1 Flash-Lite Preview34.1
MiMo-V2.531.7
Mistral Medium 3.531.5
Haiku 4.531.5
Grok 4.2030.2
MiniMax M2.728.4
Laguna M.128.0
Laguna XS.216.4
Loading Atlas data…