HANDBOOK.md
Professional Work · 2026-06-24
Long-context agentic instruction-following benchmark in enterprise RL environments with company handbooks, internal tools, and external MCP servers across five domains.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Fable 5 | 36.2 |
| Claude Opus 5 | 32.3 |
| GPT-5.6 Sol | 23.5 |
| Opus 4.8 | 21.9 |
| GPT-5.5 | 21.5 |
| Grok 4.5 | 15.8 |
| Muse Spark 1.1 | 13.5 |
| GLM-5.2 | 12.7 |
| Kimi K3 | 11.9 |
| Gemini 3.5 Flash | 11.2 |
| Sonnet 4.6 | 10.4 |
| Gemini 3.1 Pro Preview | 10.0 |
| DeepSeek-V4-Pro | 9.2 |
| Qwen3.7-Max | 8.5 |
| Hy3 | 7.7 |
| DeepSeek-V4-Flash | 7.3 |
| Kimi K2.6 | 6.9 |
| Gemini 3.6 Flash | 5.0 |
| Gemini 3.5 Flash-Lite | 3.1 |
| Grok 4.3 | 1.9 |
| Inkling | 1.9 |
| Nemotron 3 Ultra 550B A55B | 1.5 |