ARC-AGI-2 Semi-Private (CAISI mean)
General QA · 2026-05-01
CAISI's semi-private ARC-AGI-2 evaluation uses a non-public task set that may have had limited third-party exposure. Scores are the mean score across tasks rather than ARC Prize's official aggregation, so this protocol is kept separate from the public benchmark.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-5.5 | 79.0 |
| Opus 4.6 | 63.0 |
| DeepSeek-V4-Pro | 46.0 |