Atlas

Benchmarks

← All benchmarks

AHK-Eval Strings

Code · 2026-06-11

AHK-Eval Strings contains six independent string-processing function tasks, scored by the share for which every hidden case passes.

Top models (higher is better)

ModelScore
GPT-5.6 Sol Pro100.0
GPT-5.5100.0
GPT-5.6 Sol100.0
GPT-5.6 Luna Pro100.0
Claude Fable 5100.0
Kimi K3100.0
GPT-5.6 Luna100.0
GPT-5.6 Terra100.0
Gemini 3.1 Pro Preview100.0
GPT-5.6 Terra Pro100.0
Grok 4.5100.0
Kimi K2.6100.0
Muse Spark 1.1100.0
Opus 4.8100.0
Aion 3.083.3
GLM-5.283.3
GPT-5.183.3
Grok 4.366.7
MiniMax M366.7
Sonnet 4.666.7
Hy366.7
GLM-550.0
DeepSeek-V4-Pro50.0
Qwen3-Coder-480B-A35B-Instruct50.0
Aion 3.0 Mini33.3
Mistral Large 3 675B Instruct 251216.7
Laguna XS 2.10.0
Loading Atlas data…