Atlas

Benchmarks

← All benchmarks

Harvey LAB-AA

Professional Work · 2026-07-07

Artificial Analysis' implementation of Harvey's Legal Agent Benchmark (LAB) evaluates AI agents on 120 private legal-work tasks spanning 24 practice areas. Agents use the Stirrup harness to read case documents in a sandbox and produce legal deliverables such as memos, disclosure schedules, and deposition summaries. The headline score is the all-pass rate: the percentage of tasks where every rubric criterion passes with no partial credit.

Top models (higher is better)

ModelScore
Kimi K326.7
Claude Fable 514.2
Grok 4.513.3
Muse Spark 1.18.3
Opus 4.87.5
GLM-5.27.5
MiniMax M36.7
Sonnet 55.0
GPT-5.6 Luna5.0
Sonnet 4.64.2
GPT-5.54.2
DeepSeek-V4-Pro3.3
Nemotron 3 Ultra 550B A55B3.3
GPT-5.6 Terra2.5
Qwen3.6 27B2.5
DeepSeek-V4-Flash1.7
Gemini 3.5 Flash1.7
GPT-5.6 Sol1.7
Qwen3.7-Plus1.7
Kimi K2.7 Code0.8
Mistral Medium 3.50.8
Haiku 4.50.0
Gemini 3.1 Flash-Lite Preview0.0
Gemini 3.1 Pro Preview0.0
Gemma 4 31B IT0.0
GPT-5.4 Mini0.0
GPT-5.4 Nano0.0
gpt-oss-120b0.0
Grok 4.30.0
Kimi K2.60.0
MiMo-V2.5-Pro0.0
Qwen3.5 397B A17B0.0
Qwen3.7-Max0.0
Step 3.7 Flash0.0
Loading Atlas data…