Atlas

Benchmarks

← All benchmarks

IOI

Code · 2025-08-11

Competitive-programming benchmark built from International Olympiad in Informatics problems, covering IOI 2024 and 2025 so contamination can be checked. Agents get up to 50 submissions, each graded per subtask, and credit for separate subtasks is combined across submissions -- so the score is partial credit out of 100 points, not the share of problems fully solved.

Top models (higher is better)

ModelScore
Claude Opus 591.7
GPT-5.6 Sol86.7
GPT-5.6 Luna72.9
Claude Fable 572.3
GPT-5.467.8
GPT-5.6 Terra65.3
GPT-5.254.8
Opus 4.747.1
Qwen3.7-Max46.8
GPT-5.3-Codex43.8
Gemini 3 Flash Preview39.1
Gemini 3 Pro Preview38.8
DeepSeek-V4-Pro35.8
Grok 4.2030.2
Gemini 3.5 Flash-Lite26.2
Grok 426.2
Opus 4.523.6
GLM-522.0
GPT-5.121.5
GPT-5.1-Codex-Max21.4
GPT-520.0
Sonnet 4.518.3
Kimi K2.517.7
Gemini 2.5 Pro17.1
Qwen3 Max (rolling alias)15.7
Grok 4.315.3
GPT-5.4 Nano15.3
DeepSeek-V3.214.4
Qwen3 Max Thinking (2026-01-23)13.8
Opus 4.112.5
Grok 4 Fast11.5
GPT-5-Codex9.8
Qwen3 Max Preview7.8
Grok 4.1 Fast7.7
GLM-4.77.6
GPT-5 Mini6.8
MiniMax M2.56.7
Sonnet 46.5
GPT-5.4 Mini6.4
Haiku 4.56.2
Loading Atlas data…