Atlas

Benchmarks

← All benchmarks

Diligence Stack Agent Bench

Professional Work · 2026-07-09

Diligence Stack Agent Bench v1.0 evaluates model-and-harness configurations on practical financial modeling and financial-health research tasks using private investment-research knowledge bases, public search, artifact tools, and task-specific rubrics. Its total score is the arithmetic mean of available 0–100 task evaluations; published averages cover evolving work and do not imply statistical confidence intervals.

Top models (higher is better)

ModelScore
GPT-5.6 Sol88.3
GPT-5.6 Luna86.2
Claude Fable 583.2
Sonnet 582.8
GPT-5.582.0
GPT-5.6 Terra81.3
Grok 4.579.2
Kimi K378.0
Opus 4.877.3
Muse Spark 1.169.3
DeepSeek-V4-Pro65.5
Gemini 3.5 Flash64.5
Kimi K2.657.8
Grok 4.341.5
Loading Atlas data…