Diligence Stack Agent — Efficiency
Professional Work · 2026-07-09
Diligence Stack Agent Efficiency is the normalized percentage of available rubric points earned for balancing output quality with production cost, averaged across evaluated work.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-5.6 Luna | 91.2 |
| GPT-5.6 Terra | 82.9 |
| DeepSeek-V4-Pro | 80.0 |
| Kimi K2.6 | 80.0 |
| Grok 4.5 | 77.0 |
| Opus 4.8 | 66.0 |
| Gemini 3.5 Flash | 62.0 |
| GPT-5.5 | 58.0 |
| Sonnet 5 | 51.5 |
| GPT-5.6 Sol | 41.3 |
| Claude Fable 5 | 38.3 |
| Kimi K3 | 30.0 |
| Muse Spark 1.1 | 30.0 |
| Grok 4.3 | 0.0 |