Atlas

Benchmarks

← All benchmarks

EnterpriseOps-Gym-AA

Agents · 2026-07-08

Artificial Analysis' independent implementation of ServiceNow's EnterpriseOps-Gym evaluates LLM agents on stateful, multi-step enterprise workflows across eight business domains using live tool use in Stirrup. The headline score is overall task success rate in oracle tool mode, graded from the final state of the underlying databases.

Top models (higher is better)

ModelScore
Claude Fable 551.1
Gemini 3.5 Flash50.1
Claude Opus 547.5
Muse Spark 1.147.2
GPT-5.546.6
Kimi K345.3
Qwen3.7-Max45.0
Sonnet 544.7
Opus 4.844.0
GPT-5.6 Sol42.9
GLM-5.242.7
Gemini 3.1 Pro Preview42.2
Grok 4.540.8
Qwen3.7-Plus40.6
DeepSeek-V4-Pro40.4
Kimi K2.7 Code40.2
DeepSeek-V4-Flash39.6
Kimi K2.638.5
Inkling38.1
GPT-5.4 Mini34.6
Mistral Medium 3.533.7
MiniMax M332.1
GPT-5.4 Nano31.5
Qwen3.5 397B A17B30.9
Haiku 4.529.4
Step 3.7 Flash29.1
Nemotron 3 Ultra 550B A55B28.9
Gemma 4 31B IT28.3
Gemini 3.1 Flash-Lite Preview28.0
Grok 4.326.3
gpt-oss-120b25.5
Loading Atlas data…