Atlas

Benchmarks

← All benchmarks

SnorkelUnderwrite

Professional Work · 2026-01-31

SnorkelUnderwrite is an expert-verified frontier benchmark with multi-turn conversations focused on agentic reasoning and tool use in commercial underwriting settings. It evaluates agents in a LangGraph/MCP/ReAct tool environment across six underwriting task types and reports overall accuracy from full execution traces.

Top models (higher is better)

ModelScore
GPT-5.491.0
Opus 4.186.3
GPT-583.3
Grok 483.3
Grok 4 Fast81.3
Grok 378.0
O4 Mini78.0
Opus 477.0
o377.0
Claude 3.7 Sonnet74.6
Sonnet 472.3
GPT-5 Mini71.7
Kimi K2 Thinking71.3
GPT-4.170.6
Gemini 2.5 Flash61.0
Nova Premier57.0
Gemini 2.5 Pro56.3
Nova Pro52.3
GPT-5 Nano47.0
Llama 3.3 70B Instruct46.3
Llama 4 Maverick46.3
Llama 4 Scout44.3
o3-mini44.3
Nova Lite40.0
Mistral Large 1.038.3
Codestral 25.0134.0
Nova Micro31.0
gpt-oss-120b30.0
Magistral Medium 1.029.3
C4AI Command R+25.7
Qwen3-235B-A22B21.3
Llama 3.1 405B Instruct20.0
Command R15.3
Loading Atlas data…