MMMU-Pro with Python
Multimodal · 2026-03-17
MMMU-Pro evaluated with access to a Python tool, as reported in OpenAI's GPT-5.4 mini and nano launch table. This protocol is distinct from the launch page's generic with-tools condition.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Fable 5 | 86.5 |
| GPT-5.6 Sol | 84.6 |
| Kimi K3 | 83.4 |
| GPT-5.5 | 83.2 |
| Opus 4.8 | 82.7 |
| GPT-5.4 | 81.5 |
| GPT-5.4 Mini | 78.0 |
| GPT-5 Mini | 74.1 |
| GPT-5.4 Nano | 69.5 |