Atlas

Models

← All models

Pixtral 12B

Mistral AI · 2024-09-11 · 12.7B parameters

Pixtral 12B is Mistral AI's first multimodal model, released in September 2024 with a 12B-class multimodal decoder and a 400M-parameter vision encoder. Built on the NeMo 12B text backbone, it accepts an arbitrary number of images at variable resolutions alongside text, with a 128K-token context window, and excels at document understanding, chart interpretation, and multimodal reasoning. It is available under Apache 2.0 on Hugging Face.

Benchmark scores

BenchmarkScore
AI2D (OpenVLM)79.0
Arena Vision — Chinese (No Style Control)945
Arena Vision — Chinese (Style Controlled)964
Arena Vision — English (No Style Control)1028
Arena Vision — English (Style Controlled)1042
Capability115.5
CharXiv Reasoning42.4
CharXiv-D68.1
GeoBench ACW Country Accuracy37.0
HallusionBench (OpenVLM)47.0
MathVista (mini)58.0
MathVista (OpenVLM)56.3
MM-Vet (OpenVLM)58.5
MMMU Validation (OpenVLM)51.1
MMStar (OpenVLM)54.5
SnakeBench Average Apples0.6
SnakeBench Best Apples2.0
SnakeBench Rating17.0
SnakeBench Total Apples12.0
SnakeBench Win Rate38.1
Vision Arena Overall1026
Vision Arena Overall No Style Control1009
VISTA26.0
Loading Atlas data…