ZeroBench Main Questions pass^5
Multimodal · 2025-02-13
ZeroBench pass^5 is the main-question consistency metric on ZeroBench, a visual reasoning benchmark for large multimodal models. It is computed from five samples and reported alongside pass@5 on the official leaderboard.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-5.5 | 10.0 |
| Claude Fable 5 | 8.0 |
| GPT-5.4 | 8.0 |
| Gemini 3.1 Pro Preview | 7.0 |
| GPT-5.2 | 6.0 |
| Gemini 3.5 Flash | 5.0 |
| GPT-5.4 Mini | 5.0 |
| Gemini 3 Pro Preview | 5.0 |
| Opus 4.8 | 4.0 |
| Opus 4.7 | 4.0 |
| GPT-5 Mini | 3.0 |
| Sonnet 4.6 | 2.0 |
| Opus 4.6 | 2.0 |
| Gemini 3 Flash Preview | 2.0 |
| Gemini 3.1 Flash-Lite Preview | 1.0 |
| Opus 4.5 | 1.0 |
| Opus 4.1 | 1.0 |
| Opus 4 | 1.0 |
| Sonnet 4 | 1.0 |
| Llama 4 Scout Instruct | 1.0 |
| Gemini 2.5 Pro Experimental 03-25 | 1.0 |
| Gemini 2.5 Flash Preview 04-17 | 1.0 |
| GPT-5.1 | 0.0 |
| Sonnet 4.5 | 0.0 |
| GPT-5 | 0.0 |
| GPT-5 Nano | 0.0 |
| Grok 4 | 0.0 |
| o3 | 0.0 |
| O4 Mini | 0.0 |
| Llama 4 Maverick Instruct | 0.0 |
| GPT-4.1 | 0.0 |
| GPT-4.5 | 0.0 |
| Claude 3.7 Sonnet | 0.0 |
| O1 Pro | 0.0 |
| O1 | 0.0 |
| Gemini 2.0 Flash Thinking Experimental 01-21 | 0.0 |
| QVQ 72B Preview | 0.0 |
| GPT-4o | 0.0 |
| GPT-4o Mini | 0.0 |
| Gemini 2.0 Flash | 0.0 |