Atlas

Benchmarks

← All benchmarks

FrontierMath Tier 4 (v1) - Private

Math · 2025-07-01

The Private split of FrontierMath Tier 4 (v1), the original 48-problem research-level private set Epoch AI evaluated before the v2 revision of 2026-06-12. Scores on v1 are not comparable with the v2 rows.

Top models (higher is better)

ModelScore
GPT-5.5 Pro39.6
GPT-5.4 Pro38.0
GPT-5.535.4
GPT-5.2 Pro31.3
GPT-5.427.1
Opus 4.722.9
GPT-5.218.8
Gemini 3.1 Pro Preview16.7
Loading Atlas data…