IOI
Code · 2025-08-11
Competitive-programming benchmark built from International Olympiad in Informatics problems, covering IOI 2024 and 2025 so contamination can be checked. Agents get up to 50 submissions, each graded per subtask, and credit for separate subtasks is combined across submissions -- so the score is partial credit out of 100 points, not the share of problems fully solved.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Opus 5 | 91.7 |
| GPT-5.6 Sol | 86.7 |
| GPT-5.6 Luna | 72.9 |
| Claude Fable 5 | 72.3 |
| GPT-5.4 | 67.8 |
| GPT-5.6 Terra | 65.3 |
| GPT-5.2 | 54.8 |
| Opus 4.7 | 47.1 |
| Qwen3.7-Max | 46.8 |
| GPT-5.3-Codex | 43.8 |
| Gemini 3 Flash Preview | 39.1 |
| Gemini 3 Pro Preview | 38.8 |
| DeepSeek-V4-Pro | 35.8 |
| Grok 4.20 | 30.2 |
| Gemini 3.5 Flash-Lite | 26.2 |
| Grok 4 | 26.2 |
| Opus 4.5 | 23.6 |
| GLM-5 | 22.0 |
| GPT-5.1 | 21.5 |
| GPT-5.1-Codex-Max | 21.4 |
| GPT-5 | 20.0 |
| Sonnet 4.5 | 18.3 |
| Kimi K2.5 | 17.7 |
| Gemini 2.5 Pro | 17.1 |
| Qwen3 Max (rolling alias) | 15.7 |
| Grok 4.3 | 15.3 |
| GPT-5.4 Nano | 15.3 |
| DeepSeek-V3.2 | 14.4 |
| Qwen3 Max Thinking (2026-01-23) | 13.8 |
| Opus 4.1 | 12.5 |
| Grok 4 Fast | 11.5 |
| GPT-5-Codex | 9.8 |
| Qwen3 Max Preview | 7.8 |
| Grok 4.1 Fast | 7.7 |
| GLM-4.7 | 7.6 |
| GPT-5 Mini | 6.8 |
| MiniMax M2.5 | 6.7 |
| Sonnet 4 | 6.5 |
| GPT-5.4 Mini | 6.4 |
| Haiku 4.5 | 6.2 |