Atlas

Benchmarks

← All benchmarks

Public Benefits Bench v1 - Neither

Professional Work · 2026-06-09

The Neither split of Public Benefits Bench v1 allows neither web search nor multi-turn interaction. It is kept under the archived v1 benchmark because v1.1 publishes a different leaderboard without these four condition splits.

Top models (higher is better)

ModelScore
Claude Fable 563.4
Opus 4.853.1
Gemini 3.5 Flash50.3
GPT-5.549.7
Gemini 3.1 Pro Preview47.4
Sonnet 4.645.6
DeepSeek-V4-Pro43.2
MiniMax M340.2
Grok 4.338.0
GLM-5.136.2
Kimi K2.634.4
Grok 4.1 Fast33.0
Haiku 4.522.6
Loading Atlas data…