Atlas

Benchmarks

← All benchmarks

Snorkel Finance Reasoning

Professional Work · 2026-06-05

Snorkel Finance Reasoning evaluates tool-using agents on multi-step financial questions grounded in 10-K filings. Claude Sonnet 3.7 grades final-answer correctness and completeness against ground truth, and the leaderboard reports accuracy across all traces, including failed or prematurely terminated runs.

Top models (higher is better)

ModelScore
Grok 453.1
GPT-5.452.0
Claude 3.7 Sonnet51.9
GPT-551.0
Sonnet 449.4
Opus 448.1
Gemini 3 Pro Preview46.8
GPT-5 Mini46.8
O4 Mini45.6
Opus 4.145.6
GPT-4.144.3
o343.0
Grok 341.8
Grok 4 Fast40.5
Llama 3.3 Nemotron Super 49B V1.535.4
Kimi K2 Thinking35.0
Gemini 2.5 Pro34.6
Nova Premier34.2
Gemini 2.5 Flash32.0
gpt-oss-120b31.6
o3-mini30.4
GPT-5 Nano26.6
Qwen3-235B-A22B17.7
Magistral Medium 1.013.9
Nova Pro12.7
Mistral Large 1.010.1
Loading Atlas data…