Atlas

Benchmarks

← All benchmarks

HANDBOOK.md

Professional Work · 2026-06-24

Long-context agentic instruction-following benchmark in enterprise RL environments with company handbooks, internal tools, and external MCP servers across five domains.

Top models (higher is better)

ModelScore
Claude Fable 536.2
Claude Opus 532.3
GPT-5.6 Sol23.5
Opus 4.821.9
GPT-5.521.5
Grok 4.515.8
Muse Spark 1.113.5
GLM-5.212.7
Kimi K311.9
Gemini 3.5 Flash11.2
Sonnet 4.610.4
Gemini 3.1 Pro Preview10.0
DeepSeek-V4-Pro9.2
Qwen3.7-Max8.5
Hy37.7
DeepSeek-V4-Flash7.3
Kimi K2.66.9
Gemini 3.6 Flash5.0
Gemini 3.5 Flash-Lite3.1
Grok 4.31.9
Inkling1.9
Nemotron 3 Ultra 550B A55B1.5
Loading Atlas data…