Atlas

Benchmarks

← All benchmarks

DeepSearchQA (F1 score)

Search · 2026-01-28

DeepSearchQA F1 scores exhaustive answer lists for 900 multi-step web information-seeking prompts across 17 fields, balancing precision and recall.

Top models (higher is better)

ModelScore
Claude Opus 595.0
Kimi K395.0
Claude Mythos Preview94.4
Claude Fable 594.4
Claude Mythos 594.2
Opus 4.893.1
Opus 4.691.3
Sonnet 4.689.2
Opus 4.789.1
GPT-5 Pro79.0
Gemini 3 Pro Preview76.9
GPT-573.2
o3-deep-research66.5
o4-mini-deep-research (2025-06-26)61.8
Gemini 2.5 Flash43.0
Opus 4.540.2
Sonnet 4.527.9
Haiku 4.522.2
Loading Atlas data…