Atlas

Benchmarks

← All benchmarks

Harvey HLAB - Criteria Pass Rate - Insurance

Professional Work · 2026-05-06

The Insurance practice-area split of Harvey's Legal Agent Benchmark held-out evaluation criteria pass rate. This reports pooled rubric criteria passed within that split and is distinct from the all-pass task-score split.

Top models (higher is better)

ModelScore
Claude Fable 593.4
Claude Opus 591.6
Kimi K390.1
Grok 4.589.4
Muse Spark 1.187.9
Opus 4.887.2
GLM-5.287.2
MiniMax M387.2
Sonnet 586.8
Gemini 3.6 Flash86.4
Sonnet 4.684.6
GPT-5.6 Sol84.2
GPT-5.582.8
Qwen3.7-Max82.1
GPT-5.6 Luna81.0
DeepSeek-V4-Pro80.6
GLM-5.180.6
Inkling80.6
Kimi K2.680.2
GPT-5.6 Terra78.0
GPT-5.473.3
Gemini 3.5 Flash-Lite67.0
Gemini 3.5 Flash63.7
Laguna M.158.2
Qwen3.7-Plus57.5
Grok 4.348.0
Laguna XS.244.8
Loading Atlas data…