FrontierCode 1.1 Main - Score
Code · 2026-07-07
FrontierCode 1.1 Main weighted score on the 100 hardest tasks in the 150-task Extended set. The 1.1 methodology adds fair-internet-use instructions and verification, and relaxes overly strict blocker criteria; flagged runs receive zero.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Fable 5 | 53.5 |
| Claude Opus 5 | 53.4 |
| GPT-5.6 Sol | 47.5 |
| Opus 4.8 | 46.5 |
| GPT-5.5 | 43.0 |
| Sonnet 5 | 42.7 |
| Grok 4.5 | 42.4 |
| SWE-1.7 | 42.3 |
| GPT-5.6 Terra | 41.3 |
| GPT-5.6 Luna | 39.8 |
| Opus 4.7 | 38.5 |
| Kimi K2.7 Code | 30.1 |
| Composer 2.5 | 25.6 |
| GLM-5.2 | 24.5 |
| DeepSeek-V4-Pro | 17.6 |
| MiniMax M3 | 14.7 |
| Inkling | 14.0 |
| Qwen3.7-Plus | 10.2 |
| SWE-1.6 | 9.4 |