Atlas

Benchmarks

← All benchmarks

AHK-Eval

Code · 2026-06-11

AHK-Eval is a HumanEval-style AutoHotkey v2 benchmark of 36 independent function-generation tasks. Each submission is parsed and executed on the pinned AutoHotkey v2.1-alpha.30+Console fork against 181 hidden functional cases. A task passes only when every hidden case passes; the canonical cold protocol uses one fresh API call per task, no tools or retries, temperature 0.2, and an 8,000-token output cap.

Top models (higher is better)

ModelScore
Claude Fable 5100.0
Opus 4.8100.0
GPT-5.6 Sol Pro97.2
GPT-5.597.2
Gemini 3.1 Pro Preview97.2
Sonnet 4.697.2
GPT-5.6 Sol94.4
GPT-5.6 Luna Pro94.4
Kimi K394.4
GPT-5.6 Luna91.7
GPT-5.6 Terra88.9
Kimi K2.688.9
GPT-5.6 Terra Pro86.1
Grok 4.586.1
Grok 4.386.1
Muse Spark 1.183.3
MiniMax M383.3
GLM-580.6
Aion 3.075.0
DeepSeek-V4-Pro75.0
GPT-5.169.4
GLM-5.266.7
Aion 3.0 Mini61.1
Hy361.1
Qwen3-Coder-480B-A35B-Instruct50.0
Mistral Large 3 675B Instruct 251241.7
Laguna XS 2.125.0
Loading Atlas data…