Atlas

Benchmarks

← All benchmarks

Harvey HLAB — Criteria Pass Rate

Professional Work · 2026-05-06

Harvey's Legal Agent Benchmark held-out evaluation criteria pass rate (0–100). Split from the coarse task score, which is on a different scale.

Top models (higher is better)

ModelScore
Muse Spark 1.192.9
Kimi K390.8
Grok 4.590.5
Claude Fable 590.5
Opus 4.887.9
Claude Opus 587.7
Sonnet 4.686.7
Sonnet 586.3
MiniMax M386.3
GLM-5.285.6
DeepSeek-V4-Pro84.4
Gemini 3.6 Flash84.4
Qwen3.7-Max83.5
GPT-5.6 Luna83.4
GPT-5.6 Sol80.5
GPT-5.580.0
Gemini 3.5 Flash79.3
GPT-5.6 Terra78.5
GPT-5.474.5
Kimi K2.674.1
GLM-5.171.7
Gemini 3.5 Flash-Lite69.5
Inkling65.7
Grok 4.364.8
Qwen3.7-Plus62.7
Laguna M.154.9
Laguna XS.241.1
Loading Atlas data…