CaseLaw
Professional Work · 2025-08-19
CaseLaw is a legal analysis benchmark developed with Jurisage that evaluates models' ability to analyze and reason about recent family and criminal case law across US and Canadian jurisdictions.
Top models (higher is better)
| Model | Score |
|---|---|
| Grok 4.3 | 79.3 |
| GPT-5.1 | 73.4 |
| GPT-4.1 | 69.9 |
| GPT-5 Mini | 68.5 |
| Opus 4.7 | 68.4 |
| GPT-5 | 66.5 |
| GPT-5.5 | 66.2 |
| GPT-5.2 | 66.0 |
| Grok 4 | 65.8 |
| Grok 4 Fast | 65.7 |
| Kimi K2 Thinking | 65.7 |
| Gemini 3.1 Pro Preview | 64.8 |
| Command A | 64.5 |
| Sonnet 4.6 | 64.0 |
| Gemini 2.5 Pro | 63.9 |
| GPT-5.4 | 63.8 |
| Muse Spark | 63.1 |
| Opus 4.5 | 62.6 |
| Sonnet 4.5 | 62.2 |
| Opus 4.6 | 62.1 |
| Mistral Large 3 675B Instruct 2512 | 61.4 |
| Kimi K2.6 | 61.2 |
| MiniMax M2.7 | 60.9 |
| Grok 4.1 Fast | 60.5 |
| GPT-4o (2024-11-20) | 59.7 |
| Qwen3.5 Plus (2026-02-15) | 59.7 |
| DeepSeek-V4-Pro | 59.4 |
| Kimi K2.5 | 58.7 |
| Trinity Large Thinking | 57.9 |
| Haiku 4.5 | 56.5 |
| Qwen3.5-Flash | 55.9 |
| Gemini 3 Flash Preview | 55.8 |
| MiniMax M2.1 | 55.8 |
| DeepSeek-V3.2 | 55.4 |
| Gemini 3.1 Flash-Lite Preview | 55.0 |
| Qwen3 Max Thinking (2026-01-23) | 55.0 |
| GLM-4.7 | 54.9 |
| Grok 4.20 | 54.4 |
| DeepSeek-V3.1 | 53.9 |
| MiniMax M2.5 | 53.5 |