Atlas

Benchmarks

← All benchmarks

MATH-Vision with Python

Multimodal · 2026-04-20

MATH-Vision with Python is a tool-enabled evaluation setting for MATH-Vision multimodal mathematical reasoning. It records scores from runs where Python execution is allowed, separate from the standard no-tool MATH-Vision score.

Top models (higher is better)

ModelScore
Claude Fable 598.6
Kimi K397.8
GPT-5.6 Sol97.8
Opus 4.897.1
GPT-5.596.8
GPT-5.496.1
Gemini 3.1 Pro Preview95.7
Kimi K2.693.2
Kimi K2.585.0
Opus 4.684.6
Loading Atlas data…