Atlas

Benchmarks

← All benchmarks

AHK-Eval Numbers

Code · 2026-06-11

AHK-Eval Numbers contains six independent numeric formatting and conversion function tasks, scored by the share for which every hidden case passes.

Top models (higher is better)

ModelScore
GPT-5.6 Sol Pro100.0
GPT-5.5100.0
GPT-5.6 Luna Pro100.0
Kimi K3100.0
GPT-5.6 Luna100.0
GPT-5.6 Terra100.0
Gemini 3.1 Pro Preview100.0
GPT-5.6 Terra Pro100.0
Grok 4.5100.0
Kimi K2.6100.0
Aion 3.0100.0
MiniMax M3100.0
GPT-5.1100.0
GPT-5.6 Sol83.3
Claude Fable 583.3
Muse Spark 1.183.3
Opus 4.883.3
Grok 4.383.3
GLM-583.3
Aion 3.0 Mini83.3
Hy383.3
DeepSeek-V4-Pro83.3
Qwen3-Coder-480B-A35B-Instruct83.3
Sonnet 4.666.7
GLM-5.266.7
Mistral Large 3 675B Instruct 251266.7
Laguna XS 2.150.0
Loading Atlas data…