Atlas

Benchmarks

← All benchmarks

ITBench-AA

Agents · 2026-05-27

ITBench-AA is Artificial Analysis' implementation of IBM's ITBench for Kubernetes incident root-cause analysis. Models inspect offline incident snapshots and identify the entities causing the failure; the headline score is average precision at full recall.

Top models (higher is better)

ModelScore
GPT-5.6 Sol56.2
GPT-5.6 Terra51.0
Kimi K347.7
Opus 4.746.7
GPT-5.545.8
GLM-5.242.7
Qwen3.7-Max42.5
Gemini 3.5 Flash40.3
GPT-5.6 Luna40.3
GLM-5.140.3
Sonnet 4.639.8
DeepSeek-V4-Pro38.3
MiMo-V2.5-Pro38.2
Gemma 4 31B IT37.3
Qwen3.5-27B35.5
GPT-5.4 Mini35.2
Qwen3.5 397B A17B34.1
Grok 4.332.7
DeepSeek-V4-Flash31.5
Kimi K2.631.2
Gemini 3.1 Pro Preview30.3
Step 3.7 Flash30.3
Haiku 4.527.3
MiniMax M2.726.5
GPT-5.4 Nano24.4
Gemma 4 26B A4B IT23.6
Qwen3.5 35B A3B21.5
Grok 4.1 Fast17.9
gpt-oss-120b5.6
Nemotron 3 Super 120B A12B1.1
Llama 3.3 70B Instruct0.6
Loading Atlas data…