Atlas

Benchmarks

← All benchmarks

AHK-Eval Algorithms

Code · 2026-06-11

AHK-Eval Algorithms contains six independent algorithmic function tasks, including natural sorting and reverse-Polish-notation evaluation, scored by the share for which every hidden case passes.

Top models (higher is better)

ModelScore
GPT-5.6 Sol Pro100.0
GPT-5.6 Luna Pro100.0
Claude Fable 5100.0
GPT-5.6 Luna100.0
GPT-5.583.3
GPT-5.6 Sol83.3
Kimi K383.3
GPT-5.6 Terra83.3
Muse Spark 1.183.3
Opus 4.883.3
Grok 4.383.3
GLM-583.3
Gemini 3.1 Pro Preview66.7
GPT-5.6 Terra Pro66.7
Grok 4.566.7
Kimi K2.666.7
Aion 3.066.7
MiniMax M366.7
Sonnet 4.666.7
GLM-5.266.7
Aion 3.0 Mini66.7
GPT-5.166.7
Hy366.7
DeepSeek-V4-Pro66.7
Qwen3-Coder-480B-A35B-Instruct50.0
Mistral Large 3 675B Instruct 251216.7
Laguna XS 2.116.7
Loading Atlas data…