Atlas

Benchmarks

← All benchmarks

Harvey HLAB - Criteria Pass Rate - White Collar Defense Investigations

Professional Work · 2026-05-06

The White Collar Defense Investigations practice-area split of Harvey's Legal Agent Benchmark held-out evaluation criteria pass rate. This reports pooled rubric criteria passed within that split and is distinct from the all-pass task-score split.

Top models (higher is better)

ModelScore
Grok 4.590.6
Muse Spark 1.189.6
MiniMax M388.6
Kimi K387.9
Claude Fable 587.2
GLM-5.284.8
Sonnet 584.5
Sonnet 4.684.2
Claude Opus 583.8
Gemini 3.6 Flash83.8
DeepSeek-V4-Pro83.5
Opus 4.882.5
Qwen3.7-Max80.8
GPT-5.6 Sol78.5
GPT-5.6 Luna78.1
Gemini 3.5 Flash73.7
GPT-5.472.4
GPT-5.6 Terra70.0
GLM-5.167.0
GPT-5.565.3
Kimi K2.665.0
Gemini 3.5 Flash-Lite64.6
Grok 4.360.9
Inkling59.6
Qwen3.7-Plus46.5
Laguna M.144.4
Laguna XS.235.0
Loading Atlas data…