Atlas

Benchmarks

← All benchmarks

Diligence Stack Agent — Accuracy

Professional Work · 2026-07-09

Diligence Stack Agent Accuracy is the normalized percentage of available rubric points earned for material factual correctness and avoidance of unsupported claims, averaged across evaluated work.

Top models (higher is better)

ModelScore
GPT-5.6 Sol87.0
GPT-5.584.0
GPT-5.6 Terra83.0
GPT-5.6 Luna82.0
Claude Fable 577.0
Sonnet 573.0
Grok 4.570.0
Kimi K368.0
Opus 4.866.0
Muse Spark 1.158.0
Gemini 3.5 Flash56.0
DeepSeek-V4-Pro53.0
Grok 4.342.0
Kimi K2.638.0
Loading Atlas data…