Atlas

Benchmarks

← All benchmarks

Professional Reasoning Benchmark - Finance

Professional Work · 2025-11-13

Weighted score on finance reasoning tasks drawn from real professional practice.

Top models (higher is better)

ModelScore
Muse Spark 1.155.0
Claude Fable 553.9
Opus 4.653.3
Muse Spark52.4
GPT-551.3
GPT-5 Pro51.1
GPT-5.6 Sol50.5
o3-pro49.1
GPT-5.148.0
o347.7
Kimi K2.546.5
GPT-5.2 Pro46.3
Opus 4.546.2
GPT-5.445.6
gpt-oss-120b43.8
Sonnet 4.543.8
Kimi K2 Thinking43.4
Gemini 3.1 Pro Preview41.9
Mistral Medium 339.4
O4 Mini39.2
Gemini 3 Pro Preview39.2
Qwen3 235B A22B Instruct 250739.1
Gemini 2.5 Pro38.9
Gemini 2.5 Flash38.4
Kimi K2 Instruct38.3
Opus 4.135.2
DeepSeek-V3.135.1
GPT-4.134.3
DeepSeek-R1-052832.7
GPT-4.1 Mini30.4
Llama 4 Maverick Instruct22.4
Loading Atlas data…