Atlas

Benchmarks

← All benchmarks

OTIS-AIME 2025 (CAISI 30-task set)

Math · 2025-09-30

CAISI's OTIS-AIME 2025 set contains 30 advanced high-school mathematics problems with integer answers between 0 and 999. It is a different set from Atlas's 45-task OTIS Mock AIME 2024-2025 benchmark, and scores report answer accuracy.

Top models (higher is better)

ModelScore
GPT-5.5100.0
Claude Mythos Preview99.5
GLM-5.298.6
Opus 4.897.1
DeepSeek-V4-Pro97.1
Opus 4.692.0
GPT-591.9
GPT-5.4 Mini90.0
DeepSeek-V3.177.6
DeepSeek-R1-052873.3
gpt-oss-120b72.9
Opus 466.7
DeepSeek-R158.3
Loading Atlas data…