Atlas

Benchmarks

← All benchmarks

FinanceQA - Assumption-Based

Professional Work · 2025-01-29

The assumption-based FinanceQA component contains 46 tactical financial-analysis questions where information is incomplete and the model must make logical, defensible assumptions while following professional accounting and valuation conventions. The AfterQuery FinanceArena leaderboard reports accuracy on this component. Reported scores land off the 46-item lattice (fine-grained decimals), indicating partial credit or run averaging, so the reported number is a mean of per-item scores and a binomial noise model is not licensed.

Top models (higher is better)

ModelScore
o354.1
Grok 449.3
O4 Mini48.6
Gemini 2.5 Pro45.3
Llama 4 Maverick44.6
Opus 444.6
Grok 344.6
Sonnet 443.9
Phi-4-reasoning-plus43.2
DeepSeek-R142.9
QwQ-32B42.6
GPT-4.1 Mini41.9
Llama 3.1 Nemotron Ultra 253B V141.9
Qwen3-30B-A3B37.2
Kimi K2 Instruct33.8
Gemini 2.5 Flash32.4
Command A27.7
Nova Pro20.3
Loading Atlas data…