Atlas

Benchmarks

← All benchmarks

BrowseComp

Search · 2025-04-10

Hard web-browsing benchmark for locating obscure facts.

Top models (higher is better)

ModelScore
Claude Mythos 593.3
GPT-5.6 Sol92.2
Kimi K391.2
Claude Opus 590.8
GPT-5.5 Pro90.1
GPT-5.4 Pro89.3
Opus 4.888.5
Claude Fable 588.0
Claude Mythos Preview87.9
GPT-5.6 Terra87.5
Gemini 3.1 Pro Preview85.9
Sonnet 584.7
GPT-5.584.4
GPT-5.6 Luna84.0
Opus 4.684.0
DeepSeek-V4-Pro83.4
Kimi K2.683.2
GPT-5.482.7
Opus 4.779.3
Qwen3.5 397B A17B78.6
GPT-5.2 Pro77.9
Inkling-Small77.4
GPT-5.3-Codex77.3
Inkling77.1
MiniMax M2.776.3
Kimi K2.574.9
Sonnet 4.674.7
DeepSeek-V4-Flash73.2
GPT-5.265.8
Nemotron 3 Ultra 550B A55B63.0
Gemini 3 Pro Preview59.2
O19.9
GPT-4o Search Preview1.9
GPT-4.50.9
GPT-4o (2024-08-06)0.6
Loading Atlas data…