Atlas

Benchmarks

← All benchmarks

SnorkelFinance

Professional Work · 2026-06-08

SnorkelFinance is an expert-verified financial question-answering benchmark from Snorkel AI. It evaluates ReAct-style agents on 290 financial QA tasks derived from real reports using SQL, code, and other tools, with scores reported as overall accuracy from full execution traces.

Top models (higher is better)

ModelScore
GPT-581.0
o381.0
Gemini 3 Pro Preview80.3
Opus 4.180.3
GPT-5 Mini79.3
Opus 478.3
Claude 3.7 Sonnet77.9
Sonnet 476.6
O4 Mini76.6
Grok 474.0
Grok 4 Fast73.5
Kimi K2 Thinking71.7
gpt-oss-120b66.6
Grok 365.9
o3-mini63.8
GPT-4.162.7
Nova Premier62.1
Gemini 2.5 Pro60.6
Gemini 2.5 Flash53.1
Qwen3-235B-A22B51.4
GPT-5 Nano50.0
Llama 3.3 Nemotron Super 49B V1.544.0
Nova Pro40.3
Codestral 25.0127.6
Nova Lite16.9
Magistral Medium 1.016.2
Nova Micro14.5
Mistral Large 1.013.4
Loading Atlas data…