Atlas

Benchmarks

← All benchmarks

Public Benefits Bench v1 - Web Search

Professional Work · 2026-06-09

The Web Search split of Public Benefits Bench v1 allows web search without multi-turn interaction. It is kept under the archived v1 benchmark because v1.1 publishes a different leaderboard without these four condition splits.

Top models (higher is better)

ModelScore
Claude Fable 567.7
Opus 4.858.9
Gemini 3.5 Flash54.1
DeepSeek-V4-Pro54.0
GPT-5.553.9
MiniMax M353.6
Sonnet 4.652.4
GLM-5.150.7
Grok 4.344.0
Kimi K2.643.6
Gemini 3.1 Pro Preview41.8
Haiku 4.536.1
Grok 4.1 Fast33.1
Loading Atlas data…