ZeroBench (Subquestions, pass@1)
Multimodal · 2025-02-13
This ZeroBench split evaluates multimodal visual reasoning on the benchmark's component subquestions. Scores report pass@1 accuracy in percentage points.
Top models (higher is better)
| Model | Score |
|---|---|
| Seed1.5-VL | 30.8 |
| O4 Mini | 29.1 |
| GPT-5 Mini | 27.8 |
| GPT-4.5 | 27.0 |
| GPT-5 | 26.2 |
| Gemini 2.5 Pro Experimental 03-25 | 26.0 |
| o3 | 25.5 |
| Claude 3.5 Sonnet (Oct 2024) | 25.5 |
| Opus 4.1 | 25.3 |
| Opus 4 | 25.1 |
| Sonnet 4 | 24.6 |
| Gemini 2.5 Flash Preview 04-17 | 24.5 |
| Llama 4 Maverick Instruct | 23.7 |
| Gemini 2.0 Flash | 23.2 |
| O1 Pro | 22.4 |
| GPT-5 Nano | 21.6 |
| Grok 4 | 21.6 |
| Gemini 1.5 Pro | 20.9 |
| Claude 3.5 Sonnet (June 2024) | 20.7 |
| GPT-4.1 | 20.7 |
| Gemini 2.0 Flash Thinking Experimental 01-21 | 20.5 |
| QVQ 72B Preview | 20.5 |
| O1 | 20.2 |
| GPT-4o | 19.6 |
| Claude 3.7 Sonnet | 19.0 |
| Pixtral Large | 18.7 |
| Gemini 1.5 Flash | 17.9 |
| GPT-4o Mini | 16.6 |
| Llama 4 Scout Instruct | 16.3 |
| Claude 3 Sonnet | 16.1 |
| Opus 3 | 15.1 |
| NVLM-D-72B | 14.9 |
| Llama 3.2 90B Vision Instruct | 13.3 |
| Qwen2 VL 72B Instruct | 13.0 |
| Gemini 1.0 Pro Vision | 12.4 |
| Claude 3 Haiku | 12.3 |
| Sonnet 4.5 | 9.5 |