Atlas

Benchmarks

← All benchmarks

Professional Reasoning Benchmark - Legal

Professional Work · 2025-11-13

Weighted score on legal reasoning tasks drawn from real professional practice.

Top models (higher is better)

ModelScore
Muse Spark 1.157.0
Claude Fable 552.6
Muse Spark52.3
Opus 4.652.3
GPT-5.6 Sol50.5
GPT-5 Pro49.9
o3-pro49.7
GPT-5.149.3
GPT-549.0
o348.6
GPT-5.2 Pro45.4
GPT-5.444.4
Opus 4.544.2
Gemini 3.1 Pro Preview44.0
Kimi K2.543.8
Gemini 2.5 Pro41.4
Gemini 2.5 Flash41.0
Kimi K2 Thinking40.9
Sonnet 4.540.8
Gemini 3 Pro Preview40.6
gpt-oss-120b40.2
Mistral Medium 339.5
Qwen3 235B A22B Instruct 250738.3
O4 Mini38.1
DeepSeek-V3.137.6
DeepSeek-R1-052836.6
GPT-4.136.5
Kimi K2 Instruct36.4
Opus 4.134.0
GPT-4.1 Mini30.4
Llama 4 Maverick Instruct24.8
Loading Atlas data…