EnterpriseBench: CoreCraft
Professional Work · 2026-02-18
Score on enterprise agent tasks in a simulated startup world.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Fable 5 | 70.3 |
| GPT-5.5 | 52.8 |
| Opus 4.8 | 52.3 |
| Gemini 3.5 Flash | 50.8 |
| GPT-5.2 | 42.6 |
| GPT-5.4 | 36.4 |
| Opus 4.7 | 35.9 |
| Opus 4.6 | 30.8 |
| DeepSeek-V4-Flash | 30.8 |
| GLM-5.2 | 27.7 |
| Gemini 3.1 Pro Preview | 27.2 |
| Qwen3.7-Max | 26.2 |
| Kimi K2.6 | 24.6 |
| DeepSeek-V3.2 | 24.1 |
| Grok 4.1 Fast | 20.5 |
| GPT-5.2-Codex | 20.1 |
| Gemini 3 Flash Preview | 20.0 |
| GLM-5 | 17.4 |
| Sonnet 4.6 | 16.4 |
| Gemini 3 Pro Preview | 14.4 |
| Qwen3.5 Plus (2026-02-15) | 11.3 |
| Nova 2 Pro (Preview) | 8.9 |
| Kimi K2.5 | 8.7 |
| Mistral Large 3 675B Instruct 2512 | 3.6 |
| Qwen3 Max (rolling alias) | 3.6 |