Atlas

Benchmarks

← All benchmarks

EnterpriseBench: CoreCraft

Professional Work · 2026-02-18

Score on enterprise agent tasks in a simulated startup world.

Top models (higher is better)

ModelScore
Claude Fable 570.3
GPT-5.552.8
Opus 4.852.3
Gemini 3.5 Flash50.8
GPT-5.242.6
GPT-5.436.4
Opus 4.735.9
Opus 4.630.8
DeepSeek-V4-Flash30.8
GLM-5.227.7
Gemini 3.1 Pro Preview27.2
Qwen3.7-Max26.2
Kimi K2.624.6
DeepSeek-V3.224.1
Grok 4.1 Fast20.5
GPT-5.2-Codex20.1
Gemini 3 Flash Preview20.0
GLM-517.4
Sonnet 4.616.4
Gemini 3 Pro Preview14.4
Qwen3.5 Plus (2026-02-15)11.3
Nova 2 Pro (Preview)8.9
Kimi K2.58.7
Mistral Large 3 675B Instruct 25123.6
Qwen3 Max (rolling alias)3.6
Loading Atlas data…