Atlas

Benchmarks

← All benchmarks

IOI (Vals v2)

Code

Vals IOI v2 evaluates eighteen official IOI problems from 2024, 2025 and 2026 using OpenCode in an offline C++20 sandbox. Agents receive no submission or grading feedback during their attempts; only the final solution is graded. The score averages the three six-problem yearly scores. This protocol is explicitly incomparable with the earlier interactive-grading IOI board.

Top models (higher is better)

ModelScore
GPT-6 Astra100.0
Claude Opus 5.595.1
GPT-5.6 Sol91.2
Claude Fable 5.190.8
GPT-5.6 Terra87.6
Claude Opus 584.3
GPT-6 Sol82.6
Qwen3.8-Max68.9
GLM-5.368.4
Gemini 3.7 Flash67.8
GPT-5.6 Luna61.8
Hy4 Preview59.3
Grok 4.757.7
Gemini 3.8 Flash56.9
Muse Spark 1.356.6
GPT-6 Luna55.6
GPT-5.3-Codex53.8
GLM-5.3 Flash52.5
DeepSeek V4 Pro 081351.6
Kimi K348.9
Grok 4.647.6
Sonnet 545.0
Grok 4.540.6
DeepSeek V4.1 Flash40.3
MiMo-V2.6-Pro39.3
Qwen3.8 27B39.1
Gemini 3.6 Flash35.1
DeepSeek-V4-Flash-073132.7
Muse Spark 1.221.8
Inkling14.9
Inkling-Small9.3
Mercury 2.52.4
Loading Atlas data…