Atlas

Benchmarks

← All benchmarks

SnorkelUnderwrite 2.0

Professional Work

Pass@1 on 200 commercial property-and-casualty underwriting tasks with partial information, cross-source evidence, role permissions and stateful workflows. Experts review scenarios, tools, rubrics and verifiers. Native effort and agent scaffold are not disclosed.

Top models (higher is better)

ModelScore
Grok 4.730.5
GLM-5.330.4
Kimi K327.9
Claude Fable 5.124.0
Claude Opus 5.522.7
GPT-6 Astra21.2
Muse Spark 1.319.6
Gemini 3.8 Flash19.1
Claude Opus 518.3
Grok 4.617.8
Loading Atlas data…