Atlas

Benchmarks

← All benchmarks

AHK-Eval Regex

Code · 2026-06-11

AHK-Eval Regex contains six independent regular-expression and structured-text function tasks, scored by the share for which every hidden case passes.

Top models (higher is better)

ModelScore
GPT-5.5100.0
GPT-5.6 Sol100.0
Kimi K3100.0
GPT-5.6 Luna100.0
Grok 4.5100.0
Muse Spark 1.1100.0
Grok 4.3100.0
MiniMax M3100.0
GPT-5.6 Sol Pro83.3
GPT-5.6 Luna Pro83.3
Claude Fable 583.3
GPT-5.6 Terra83.3
Gemini 3.1 Pro Preview83.3
GPT-5.6 Terra Pro83.3
Aion 3.0 Mini83.3
DeepSeek-V4-Pro83.3
Aion 3.066.7
Sonnet 4.666.7
Hy366.7
Laguna XS 2.166.7
Kimi K2.650.0
Opus 4.850.0
GLM-550.0
GLM-5.250.0
GPT-5.150.0
Qwen3-Coder-480B-A35B-Instruct50.0
Mistral Large 3 675B Instruct 251233.3
Loading Atlas data…