Atlas

Benchmarks

← All benchmarks

ZeroBench Main Mean pass@1 (five samples)

Multimodal

ZeroBench main-question accuracy estimated from five samplings per question: average the fraction correct for each of the 100 questions, then average over questions. The official filled-circle mean pass@1 is distinct from its one-sample greedy pass@1. Current leaderboard scores include corrections after evaluation red teaming.

Top models (higher is better)

ModelScore
GPT-5.6 Sol22.0
Claude Opus 517.2
GPT-5.517.2
GPT-5.416.4
Claude Fable 515.8
GPT-5.6 Terra14.2
GPT-5.6 Luna13.4
Gemini 3.5 Flash12.8
Gemini 3.6 Flash12.2
Gemini 3.1 Pro Preview11.6
Grok 4.511.4
GPT-5.211.2
Grok 4.611.0
Opus 4.810.4
Gemini 3 Pro Preview10.0
Sonnet 59.4
Gemini 3 Flash Preview9.4
GPT-5.4 Mini9.0
Opus 4.78.8
Sonnet 4.67.8
Grok 4.36.4
Opus 4.66.4
GPT-5 Mini6.2
Gemini 3.5 Flash-Lite6.0
O4 Mini4.6
Opus 4.53.6
Gemini 3.1 Flash-Lite Preview3.4
Opus 4.13.4
Sonnet 42.8
Gemini 2.5 Pro Experimental 03-252.8
Opus 42.4
o32.4
Claude 3.7 Sonnet2.4
GPT-52.2
Sonnet 4.52.0
GPT-4.52.0
GPT-5.11.8
Llama 4 Scout Instruct1.6
Grok 41.4
Gemini 2.5 Flash Preview 04-171.4
Loading Atlas data…