DeepResearchBench
Search · 2025-06-13
A benchmark of models' ability to gather information from the internet to answer questions, testing models' ability to find and synthesize information.
Top models (higher is better)
| Model | Score |
|---|---|
| Opus 4.6 | 0.6 |
| Sonnet 4.6 | 0.5 |
| GPT-5.5 | 0.5 |
| Sonnet 4.5 | 0.5 |
| GPT-5 | 0.5 |
| Opus 4.8 | 0.5 |
| Gemini 2.5 Pro | 0.5 |
| Opus 4.1 | 0.5 |
| Opus 4 | 0.5 |
| Grok 4 | 0.5 |
| Sonnet 4 | 0.5 |
| o3 | 0.5 |
| Claude 3.7 Sonnet | 0.4 |
| Gemini 2.5 Pro Preview 06-05 | 0.4 |
| GPT-5.1 | 0.4 |
| GPT-5.2 | 0.4 |
| Sonar Pro | 0.4 |
| Sonar | 0.4 |
| Gemini 3.1 Flash-Lite | 0.4 |
| GPT-5.4 Mini | 0.4 |
| DeepSeek-R1-0528 | 0.4 |
| GPT-5.4 | 0.4 |
| Gemini 2.5 Pro Preview 05-06 | 0.3 |
| GPT-4.1 | 0.3 |
| Gemini 2.5 Flash Preview 04-17 | 0.3 |