Atlas

Benchmarks

← All benchmarks

HiL-Bench (Human-in-Loop Benchmark)

Agents · 2026-04-10

Score on help-seeking judgment and selective escalation in agents.

Top models (higher is better)

ModelScore
Claude Opus 557.0
Claude Fable 556.3
GLM-5.243.7
Opus 4.741.7
GPT-5.539.7
Opus 4.638.3
Opus 4.835.3
Gemini 3.1 Pro Preview35.3
GPT-5.6 Sol32.3
Gemini 3.5 Flash27.7
Grok 4.2020.0
Kimi K2.618.7
GPT-5.49.7
MiniMax M2.56.3
GPT-5.3-Codex4.3
Loading Atlas data…