Atlas

Benchmarks

← All benchmarks

Deep Research Bench — Gather Evidence

Search · 2025-05-06

The DRB Gather Evidence category asks an agent to identify key evidence relevant to a query.

Top models (higher is better)

ModelScore
Opus 40.4
GPT-50.4
o30.4
Gemini 3 Pro Preview0.4
Opus 4.10.4
Sonnet 40.3
Grok 40.3
Sonnet 4.60.3
GPT-5.10.3
Opus 4.50.3
Opus 4.60.3
Gemini 3 Flash Preview0.3
Haiku 4.50.3
Gemini 2.5 Pro0.3
Sonnet 4.50.3
GPT-5.20.3
Opus 4.80.3
Gemini 3.1 Pro Preview0.3
GPT-5.50.2
Gemini 3.1 Flash-Lite Preview0.2
GPT-5.4 Mini0.1
GPT-5.40.1
Loading Atlas data…