Atlas

Benchmarks

← All benchmarks

Diligence Stack Agent — Efficiency

Professional Work · 2026-07-09

Diligence Stack Agent Efficiency is the normalized percentage of available rubric points earned for balancing output quality with production cost, averaged across evaluated work.

Top models (higher is better)

ModelScore
GPT-5.6 Luna91.2
GPT-5.6 Terra82.9
DeepSeek-V4-Pro80.0
Kimi K2.680.0
Grok 4.577.0
Opus 4.866.0
Gemini 3.5 Flash62.0
GPT-5.558.0
Sonnet 551.5
GPT-5.6 Sol41.3
Claude Fable 538.3
Kimi K330.0
Muse Spark 1.130.0
Grok 4.30.0
Loading Atlas data…