Atlas

Benchmarks

← All benchmarks

Harvey's Legal Agent Benchmark

Professional Work · 2026-05-06

Harvey's Legal Agent Benchmark held-out evaluation tests whether agents can complete realistic legal work using documents, spreadsheets, presentations, and file-system tools. Scores report the percentage of tasks completed successfully.

Top models (higher is better)

ModelScore
Claude Opus 523.6
Muse Spark 1.120.0
Claude Fable 513.3
Grok 4.512.9
Kimi K310.8
Opus 4.810.4
GLM-5.27.1
Sonnet 55.8
Sonnet 4.65.0
MiniMax M34.2
DeepSeek-V4-Pro3.8
GPT-5.53.8
Gemini 3.6 Flash3.3
Gemini 3.5 Flash2.5
GPT-5.6 Sol2.5
Inkling2.1
Kimi K2.61.7
Qwen3.7-Max1.7
GPT-5.6 Luna1.3
GPT-5.6 Terra0.4
Grok 4.30.4
Gemini 3.5 Flash-Lite0.0
GLM-5.10.0
GPT-5.40.0
Laguna M.10.0
Laguna XS.20.0
Qwen3.7-Plus0.0
Gemini 3.1 Pro Preview0.0
Loading Atlas data…