Atlas

Benchmarks

← All benchmarks

MMMU-Pro with Python

Multimodal · 2026-03-17

MMMU-Pro evaluated with access to a Python tool, as reported in OpenAI's GPT-5.4 mini and nano launch table. This protocol is distinct from the launch page's generic with-tools condition.

Top models (higher is better)

ModelScore
Claude Fable 586.5
GPT-5.6 Sol84.6
Kimi K383.4
GPT-5.583.2
Opus 4.882.7
GPT-5.481.5
GPT-5.4 Mini78.0
GPT-5 Mini74.1
GPT-5.4 Nano69.5
Loading Atlas data…