SnorkelUnderwrite
Professional Work · 2026-01-31
SnorkelUnderwrite is an expert-verified frontier benchmark with multi-turn conversations focused on agentic reasoning and tool use in commercial underwriting settings. It evaluates agents in a LangGraph/MCP/ReAct tool environment across six underwriting task types and reports overall accuracy from full execution traces.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-5.4 | 91.0 |
| Opus 4.1 | 86.3 |
| GPT-5 | 83.3 |
| Grok 4 | 83.3 |
| Grok 4 Fast | 81.3 |
| Grok 3 | 78.0 |
| O4 Mini | 78.0 |
| Opus 4 | 77.0 |
| o3 | 77.0 |
| Claude 3.7 Sonnet | 74.6 |
| Sonnet 4 | 72.3 |
| GPT-5 Mini | 71.7 |
| Kimi K2 Thinking | 71.3 |
| GPT-4.1 | 70.6 |
| Gemini 2.5 Flash | 61.0 |
| Nova Premier | 57.0 |
| Gemini 2.5 Pro | 56.3 |
| Nova Pro | 52.3 |
| GPT-5 Nano | 47.0 |
| Llama 3.3 70B Instruct | 46.3 |
| Llama 4 Maverick | 46.3 |
| Llama 4 Scout | 44.3 |
| o3-mini | 44.3 |
| Nova Lite | 40.0 |
| Mistral Large 1.0 | 38.3 |
| Codestral 25.01 | 34.0 |
| Nova Micro | 31.0 |
| gpt-oss-120b | 30.0 |
| Magistral Medium 1.0 | 29.3 |
| C4AI Command R+ | 25.7 |
| Qwen3-235B-A22B | 21.3 |
| Llama 3.1 405B Instruct | 20.0 |
| Command R | 15.3 |