Atlas

Benchmarks

← All benchmarks

ARC-AGI-2 Semi-Private (CAISI mean)

General QA · 2026-05-01

CAISI's semi-private ARC-AGI-2 evaluation uses a non-public task set that may have had limited third-party exposure. Scores are the mean score across tasks rather than ARC Prize's official aggregation, so this protocol is kept separate from the public benchmark.

Top models (higher is better)

ModelScore
GPT-5.579.0
Opus 4.663.0
DeepSeek-V4-Pro46.0
Loading Atlas data…