Atlas

Benchmarks

← All benchmarks

GDPval

Professional Work · 2025-09-25

OpenAI's GDPval benchmark of economically valuable knowledge-work tasks across 44 occupations, graded by expert human reviewers via blind pairwise comparison. The score is the rate at which a model's deliverables win or tie against human experts.

Top models (higher is better)

ModelScore
GPT-5.584.9
GPT-5.483.0
GPT-5.5 Pro82.3
GPT-5.4 Pro82.0
Opus 4.780.3
GPT-5.2 Pro74.1
GPT-5.270.9
GPT-5.3-Codex70.9
Gemini 3.1 Pro Preview67.3
Opus 4.545.5
Opus 4.143.6
Sonnet 4.542.5
Gemini 3 Pro Preview40.3
GPT-534.8
o330.8
O4 Mini25.3
Gemini 2.5 Pro23.3
GPT-4o (2024-11-20)9.9
Loading Atlas data…