Atlas

Benchmarks

← All benchmarks

DeepResearchBench

Search · 2025-06-13

A benchmark of models' ability to gather information from the internet to answer questions, testing models' ability to find and synthesize information.

Top models (higher is better)

ModelScore
Opus 4.60.6
Sonnet 4.60.5
GPT-5.50.5
Sonnet 4.50.5
GPT-50.5
Opus 4.80.5
Gemini 2.5 Pro0.5
Opus 4.10.5
Opus 40.5
Grok 40.5
Sonnet 40.5
o30.5
Claude 3.7 Sonnet0.4
Gemini 2.5 Pro Preview 06-050.4
GPT-5.10.4
GPT-5.20.4
Sonar Pro0.4
Sonar0.4
Gemini 3.1 Flash-Lite0.4
GPT-5.4 Mini0.4
DeepSeek-R1-05280.4
GPT-5.40.4
Gemini 2.5 Pro Preview 05-060.3
GPT-4.10.3
Gemini 2.5 Flash Preview 04-170.3
Loading Atlas data…