Atlas

Benchmarks

← All benchmarks

AHK-Eval Easy

Code · 2026-06-11

AHK-Eval Easy is the task-pass rate across the suite's 12 author-designated easy tasks.

Top models (higher is better)

ModelScore
GPT-5.5100.0
GPT-5.6 Sol100.0
GPT-5.6 Luna Pro100.0
Kimi K3100.0
GPT-5.6 Luna100.0
Grok 4.5100.0
GPT-5.6 Sol Pro91.7
Claude Fable 591.7
Gemini 3.1 Pro Preview91.7
GPT-5.6 Terra Pro91.7
Muse Spark 1.191.7
Grok 4.391.7
MiniMax M391.7
GPT-5.6 Terra83.3
Kimi K2.683.3
Opus 4.883.3
Aion 3.083.3
GLM-583.3
Aion 3.0 Mini83.3
Hy383.3
Sonnet 4.675.0
GLM-5.275.0
DeepSeek-V4-Pro66.7
Qwen3-Coder-480B-A35B-Instruct58.3
GPT-5.150.0
Mistral Large 3 675B Instruct 251250.0
Laguna XS 2.141.7
Loading Atlas data…