Atlas

Benchmarks

← All benchmarks

CaseLaw

Professional Work · 2025-08-19

CaseLaw is a legal analysis benchmark developed with Jurisage that evaluates models' ability to analyze and reason about recent family and criminal case law across US and Canadian jurisdictions.

Top models (higher is better)

ModelScore
Grok 4.379.3
GPT-5.173.4
GPT-4.169.9
GPT-5 Mini68.5
Opus 4.768.4
GPT-566.5
GPT-5.566.2
GPT-5.266.0
Grok 465.8
Grok 4 Fast65.7
Kimi K2 Thinking65.7
Gemini 3.1 Pro Preview64.8
Command A64.5
Sonnet 4.664.0
Gemini 2.5 Pro63.9
GPT-5.463.8
Muse Spark63.1
Opus 4.562.6
Sonnet 4.562.2
Opus 4.662.1
Mistral Large 3 675B Instruct 251261.4
Kimi K2.661.2
MiniMax M2.760.9
Grok 4.1 Fast60.5
GPT-4o (2024-11-20)59.7
Qwen3.5 Plus (2026-02-15)59.7
DeepSeek-V4-Pro59.4
Kimi K2.558.7
Trinity Large Thinking57.9
Haiku 4.556.5
Qwen3.5-Flash55.9
Gemini 3 Flash Preview55.8
MiniMax M2.155.8
DeepSeek-V3.255.4
Gemini 3.1 Flash-Lite Preview55.0
Qwen3 Max Thinking (2026-01-23)55.0
GLM-4.754.9
Grok 4.2054.4
DeepSeek-V3.153.9
MiniMax M2.553.5
Loading Atlas data…