HiL-Bench (Human-in-Loop Benchmark)
Agents · 2026-04-10
Score on help-seeking judgment and selective escalation in agents.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Opus 5 | 57.0 |
| Claude Fable 5 | 56.3 |
| GLM-5.2 | 43.7 |
| Opus 4.7 | 41.7 |
| GPT-5.5 | 39.7 |
| Opus 4.6 | 38.3 |
| Opus 4.8 | 35.3 |
| Gemini 3.1 Pro Preview | 35.3 |
| GPT-5.6 Sol | 32.3 |
| Gemini 3.5 Flash | 27.7 |
| Grok 4.20 | 20.0 |
| Kimi K2.6 | 18.7 |
| GPT-5.4 | 9.7 |
| MiniMax M2.5 | 6.3 |
| GPT-5.3-Codex | 4.3 |