Atlas

Benchmarks

← All benchmarks

FrontierMath Tiers 1-3 (v1) - Private

Math · 2025-02-28

The Private split of FrontierMath Tiers 1-3 (v1), the original 290-problem private set Epoch AI evaluated before the v2 revision of 2026-06-12 corrected errors in 42% of problems. Scores on v1 are not comparable with the v2 rows.

Top models (higher is better)

ModelScore
GPT-5.5 Pro52.4
GPT-5.551.7
GPT-5.4 Pro50.0
GPT-5.447.6
Opus 4.743.8
GPT-5.240.7
Gemini 3.1 Pro Preview36.9
Loading Atlas data…