Atlas

Benchmarks

← All benchmarks

Public Benefits Bench v1 - Both

Professional Work · 2026-06-09

The Both split of Public Benefits Bench v1 allows both web search and multi-turn interaction. It is kept under the archived v1 benchmark because v1.1 publishes a different leaderboard without these four condition splits.

Top models (higher is better)

ModelScore
Claude Fable 571.7
Opus 4.862.1
MiniMax M360.7
Sonnet 4.658.5
Gemini 3.5 Flash58.0
GLM-5.157.9
DeepSeek-V4-Pro57.6
GPT-5.557.2
Kimi K2.653.4
Gemini 3.1 Pro Preview50.9
Grok 4.350.1
Haiku 4.549.5
Grok 4.1 Fast44.3
Loading Atlas data…