Atlas

Benchmarks

← All benchmarks

Harvey's Legal Agent Benchmark - Arbitration International Dispute Resolution

Professional Work · 2026-05-06

The Arbitration International Dispute Resolution split of Harvey's Legal Agent Benchmark held-out evaluation. This child benchmark separates a source-reported subtask or subtrack from the parent aggregate so scores at different grains do not share one benchmark_slug.

Top models (higher is better)

ModelScore
Claude Fable 50.0
Opus 4.80.0
Claude Opus 50.0
Sonnet 4.60.0
Sonnet 50.0
DeepSeek-V4-Pro0.0
Gemini 3.5 Flash0.0
Gemini 3.5 Flash-Lite0.0
Gemini 3.6 Flash0.0
GLM-5.10.0
GLM-5.20.0
GPT-5.40.0
GPT-5.50.0
GPT-5.6 Luna0.0
GPT-5.6 Sol0.0
GPT-5.6 Terra0.0
Grok 4.30.0
Grok 4.50.0
Inkling0.0
Kimi K2.60.0
Kimi K30.0
Laguna M.10.0
Laguna XS.20.0
MiniMax M30.0
Muse Spark 1.10.0
Qwen3.7-Max0.0
Qwen3.7-Plus0.0
Loading Atlas data…