Atlas

Benchmarks

← All benchmarks

DECK-Bench

Professional Work · 2026-07-16

DECK-Bench is a Moonshot AI in-house agentic benchmark reported in the Kimi K3 release materials alongside office-productivity evaluations such as OfficeQA Pro and SpreadsheetBench 2. The launch table reports comparative scores for Kimi K3 and competing frontier models, but the task set is not publicly documented.

Top models (higher is better)

ModelScore
GPT-5.6 Sol74.7
Kimi K373.5
Claude Fable 573.0
GLM-5.268.6
GPT-5.568.2
Opus 4.866.9
Loading Atlas data…