Atlas

Benchmarks

← All benchmarks

Diligence Stack Agent — Artifact Quality

Professional Work · 2026-07-09

Diligence Stack Agent Artifact Quality is the normalized percentage of available rubric points earned for readable, usable, and reviewable files, averaged across evaluated work.

Top models (higher is better)

ModelScore
GPT-5.6 Sol94.5
GPT-5.6 Luna83.0
Kimi K383.0
Claude Fable 581.4
Sonnet 580.4
GPT-5.575.0
GPT-5.6 Terra74.4
Opus 4.873.4
Grok 4.571.0
Kimi K2.668.0
Gemini 3.5 Flash67.5
DeepSeek-V4-Pro61.9
Muse Spark 1.148.0
Grok 4.30.0
Loading Atlas data…