FinanceQA - Assumption-Based
Professional Work · 2025-01-29
The assumption-based FinanceQA component contains 46 tactical financial-analysis questions where information is incomplete and the model must make logical, defensible assumptions while following professional accounting and valuation conventions. The AfterQuery FinanceArena leaderboard reports accuracy on this component. Reported scores land off the 46-item lattice (fine-grained decimals), indicating partial credit or run averaging, so the reported number is a mean of per-item scores and a binomial noise model is not licensed.
Top models (higher is better)
| Model | Score |
|---|---|
| o3 | 54.1 |
| Grok 4 | 49.3 |
| O4 Mini | 48.6 |
| Gemini 2.5 Pro | 45.3 |
| Llama 4 Maverick | 44.6 |
| Opus 4 | 44.6 |
| Grok 3 | 44.6 |
| Sonnet 4 | 43.9 |
| Phi-4-reasoning-plus | 43.2 |
| DeepSeek-R1 | 42.9 |
| QwQ-32B | 42.6 |
| GPT-4.1 Mini | 41.9 |
| Llama 3.1 Nemotron Ultra 253B V1 | 41.9 |
| Qwen3-30B-A3B | 37.2 |
| Kimi K2 Instruct | 33.8 |
| Gemini 2.5 Flash | 32.4 |
| Command A | 27.7 |
| Nova Pro | 20.3 |