Atlas

Benchmarks

← All benchmarks

AHK-Eval Hard

Code · 2026-06-11

AHK-Eval Hard is the task-pass rate across the suite's 12 author-designated hard tasks.

Top models (higher is better)

ModelScore
GPT-5.6 Sol Pro100.0
Claude Fable 5100.0
GPT-5.6 Terra100.0
GPT-5.591.7
GPT-5.6 Luna Pro91.7
Kimi K391.7
GPT-5.6 Sol83.3
GPT-5.6 Luna83.3
Gemini 3.1 Pro Preview83.3
Grok 4.575.0
Kimi K2.675.0
Muse Spark 1.175.0
GPT-5.6 Terra Pro66.7
Opus 4.866.7
GLM-566.7
GLM-5.266.7
Grok 4.358.3
Sonnet 4.658.3
GPT-5.158.3
Aion 3.050.0
MiniMax M350.0
Aion 3.0 Mini41.7
DeepSeek-V4-Pro41.7
Hy333.3
Qwen3-Coder-480B-A35B-Instruct33.3
Laguna XS 2.125.0
Mistral Large 3 675B Instruct 25128.3
Loading Atlas data…