Atlas

Benchmarks

← All benchmarks

Diligence Stack Agent — Analysis

Professional Work · 2026-07-09

Diligence Stack Agent Analysis is the normalized percentage of available rubric points earned for turning evidence into useful, well-supported judgment, averaged across evaluated work.

Top models (higher is better)

ModelScore
GPT-5.6 Sol97.0
Sonnet 590.7
Claude Fable 589.7
GPT-5.6 Luna89.4
Grok 4.588.0
Kimi K388.0
GPT-5.585.4
GPT-5.6 Terra85.0
Muse Spark 1.183.3
Opus 4.882.4
DeepSeek-V4-Pro81.0
Gemini 3.5 Flash75.0
Kimi K2.656.7
Grok 4.356.7
Loading Atlas data…