ITBench-AA
Agents · 2026-05-27
ITBench-AA is Artificial Analysis' implementation of IBM's ITBench for Kubernetes incident root-cause analysis. Models inspect offline incident snapshots and identify the entities causing the failure; the headline score is average precision at full recall.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-5.6 Sol | 56.2 |
| GPT-5.6 Terra | 51.0 |
| Kimi K3 | 47.7 |
| Opus 4.7 | 46.7 |
| GPT-5.5 | 45.8 |
| GLM-5.2 | 42.7 |
| Qwen3.7-Max | 42.5 |
| Gemini 3.5 Flash | 40.3 |
| GPT-5.6 Luna | 40.3 |
| GLM-5.1 | 40.3 |
| Sonnet 4.6 | 39.8 |
| DeepSeek-V4-Pro | 38.3 |
| MiMo-V2.5-Pro | 38.2 |
| Gemma 4 31B IT | 37.3 |
| Qwen3.5-27B | 35.5 |
| GPT-5.4 Mini | 35.2 |
| Qwen3.5 397B A17B | 34.1 |
| Grok 4.3 | 32.7 |
| DeepSeek-V4-Flash | 31.5 |
| Kimi K2.6 | 31.2 |
| Gemini 3.1 Pro Preview | 30.3 |
| Step 3.7 Flash | 30.3 |
| Haiku 4.5 | 27.3 |
| MiniMax M2.7 | 26.5 |
| GPT-5.4 Nano | 24.4 |
| Gemma 4 26B A4B IT | 23.6 |
| Qwen3.5 35B A3B | 21.5 |
| Grok 4.1 Fast | 17.9 |
| gpt-oss-120b | 5.6 |
| Nemotron 3 Super 120B A12B | 1.1 |
| Llama 3.3 70B Instruct | 0.6 |