Atlas

Benchmarks

← All benchmarks

AHK-Eval Mid

Code · 2026-06-11

AHK-Eval Mid is the task-pass rate across the suite's 12 author-designated medium-difficulty tasks.

Top models (higher is better)

ModelScore
GPT-5.6 Sol Pro100.0
GPT-5.5100.0
GPT-5.6 Sol100.0
GPT-5.6 Terra Pro100.0
GPT-5.6 Luna Pro91.7
Claude Fable 591.7
Kimi K391.7
GPT-5.6 Luna91.7
Gemini 3.1 Pro Preview91.7
Kimi K2.691.7
Opus 4.891.7
Aion 3.091.7
GPT-5.6 Terra83.3
Grok 4.583.3
Muse Spark 1.183.3
Grok 4.383.3
MiniMax M375.0
Sonnet 4.675.0
GPT-5.175.0
Hy375.0
DeepSeek-V4-Pro75.0
GLM-558.3
GLM-5.258.3
Aion 3.0 Mini58.3
Qwen3-Coder-480B-A35B-Instruct41.7
Mistral Large 3 675B Instruct 251216.7
Laguna XS 2.116.7
Loading Atlas data…