Diligence Stack Agent Bench
Professional Work · 2026-07-09
Diligence Stack Agent Bench v1.0 evaluates model-and-harness configurations on practical financial modeling and financial-health research tasks using private investment-research knowledge bases, public search, artifact tools, and task-specific rubrics. Its total score is the arithmetic mean of available 0–100 task evaluations; published averages cover evolving work and do not imply statistical confidence intervals.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-5.6 Sol | 88.3 |
| GPT-5.6 Luna | 86.2 |
| Claude Fable 5 | 83.2 |
| Sonnet 5 | 82.8 |
| GPT-5.5 | 82.0 |
| GPT-5.6 Terra | 81.3 |
| Grok 4.5 | 79.2 |
| Kimi K3 | 78.0 |
| Opus 4.8 | 77.3 |
| Muse Spark 1.1 | 69.3 |
| DeepSeek-V4-Pro | 65.5 |
| Gemini 3.5 Flash | 64.5 |
| Kimi K2.6 | 57.8 |
| Grok 4.3 | 41.5 |