Atlas

Benchmarks

← All benchmarks

SnorkelLegal

Professional Work

Pass@1 on 200 simulated mid-market law-firm tasks requiring matter-specific grounding, procedural discipline, long-context evidence and governed updates. Native effort and agent scaffold are not disclosed.

Top models (higher is better)

ModelScore
Grok 4.640.7
GLM-5.337.8
Muse Spark 1.336.5
Kimi K336.1
Claude Fable 5.135.9
Grok 4.732.8
Claude Opus 5.531.4
Claude Opus 526.4
Gemini 3.8 Flash25.3
GPT-6 Astra20.9
Loading Atlas data…