ZeroBench (Main questions, pass@5, with tools)
Multimodal · 2025-02-13
Tool-augmented pass@5 accuracy on ZeroBench's 100 main visual-reasoning questions: the percentage answered correctly at least once across five runs with a Python tool available. It combines the pass@5 aggregation of the main-questions track with the with-tools evaluation setting.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Fable 5 | 46.0 |
| Kimi K3 | 41.0 |
| GPT-5.5 | 41.0 |
| GPT-5.6 Sol | 35.0 |
| Opus 4.8 | 34.0 |