Atlas

Benchmarks

← All benchmarks

Diligence Stack Agent — Completion

Professional Work · 2026-07-09

Diligence Stack Agent Completion is the normalized percentage of available rubric points earned for following the requested scope, format, and decision criteria, averaged across evaluated work.

Top models (higher is better)

ModelScore
GPT-5.6 Sol94.8
GPT-5.592.8
GPT-5.6 Luna92.0
Claude Fable 590.5
Kimi K390.0
Grok 4.589.5
Sonnet 588.8
Opus 4.888.0
GPT-5.6 Terra86.5
Muse Spark 1.185.0
DeepSeek-V4-Pro80.0
Gemini 3.5 Flash76.8
Kimi K2.675.0
Grok 4.340.0
Loading Atlas data…