Atlas

Benchmarks

← All benchmarks

Harvey HLAB - Criteria Pass Rate - Tax

Professional Work · 2026-05-06

The Tax practice-area split of Harvey's Legal Agent Benchmark held-out evaluation criteria pass rate. This reports pooled rubric criteria passed within that split and is distinct from the all-pass task-score split.

Top models (higher is better)

ModelScore
Muse Spark 1.191.3
Sonnet 4.686.0
GLM-5.285.6
MiniMax M385.0
Claude Opus 584.6
Kimi K384.6
Grok 4.584.3
Claude Fable 583.6
Sonnet 582.3
Opus 4.880.6
DeepSeek-V4-Pro77.9
Gemini 3.6 Flash77.6
Qwen3.7-Max77.6
Kimi K2.675.6
GPT-5.6 Luna74.9
Gemini 3.5 Flash74.6
GPT-5.572.6
GPT-5.6 Sol70.9
GPT-5.6 Terra69.9
GLM-5.167.9
Grok 4.363.2
Gemini 3.5 Flash-Lite62.9
Qwen3.7-Plus60.9
GPT-5.457.5
Laguna M.156.5
Inkling48.8
Laguna XS.225.6
Loading Atlas data…