EnterpriseOps-Gym-AA
Agents · 2026-07-08
Artificial Analysis' independent implementation of ServiceNow's EnterpriseOps-Gym evaluates LLM agents on stateful, multi-step enterprise workflows across eight business domains using live tool use in Stirrup. The headline score is overall task success rate in oracle tool mode, graded from the final state of the underlying databases.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Fable 5 | 51.1 |
| Gemini 3.5 Flash | 50.1 |
| Claude Opus 5 | 47.5 |
| Muse Spark 1.1 | 47.2 |
| GPT-5.5 | 46.6 |
| Kimi K3 | 45.3 |
| Qwen3.7-Max | 45.0 |
| Sonnet 5 | 44.7 |
| Opus 4.8 | 44.0 |
| GPT-5.6 Sol | 42.9 |
| GLM-5.2 | 42.7 |
| Gemini 3.1 Pro Preview | 42.2 |
| Grok 4.5 | 40.8 |
| Qwen3.7-Plus | 40.6 |
| DeepSeek-V4-Pro | 40.4 |
| Kimi K2.7 Code | 40.2 |
| DeepSeek-V4-Flash | 39.6 |
| Kimi K2.6 | 38.5 |
| Inkling | 38.1 |
| GPT-5.4 Mini | 34.6 |
| Mistral Medium 3.5 | 33.7 |
| MiniMax M3 | 32.1 |
| GPT-5.4 Nano | 31.5 |
| Qwen3.5 397B A17B | 30.9 |
| Haiku 4.5 | 29.4 |
| Step 3.7 Flash | 29.1 |
| Nemotron 3 Ultra 550B A55B | 28.9 |
| Gemma 4 31B IT | 28.3 |
| Gemini 3.1 Flash-Lite Preview | 28.0 |
| Grok 4.3 | 26.3 |
| gpt-oss-120b | 25.5 |