Atlas

Models

← All models

Llama 3.2 90B Vision Instruct

Meta · 2024-09-25 · 88.6B parameters

Llama 3.2 90B Vision Instruct is Meta's instruction-tuned multimodal model built on a Llama 3.1 text backbone with a separately trained image encoder and cross-attention adapter. It accepts text and images and generates text for visual recognition, image reasoning, captioning, document understanding, and assistant-style chat; Meta reports a December 2023 knowledge cutoff.

Benchmark scores

BenchmarkScore
AI2D (OpenVLM)69.5
AidanBench538
Artificial Analysis Intelligence Index6.2
Artificial Analysis Openness Index38.9
BALROG27.3
Capability121.8
Coding984
EnigmaEval0.4
Epoch Capabilities Index (ECI)125
EQ-Bench v282.0
EQ-Bench v2 + MAGI-Hard Combined79.9
GeoBench ACW Country Accuracy52.0
HallusionBench (OpenVLM)44.1
Humanity's Last Exam (Preview)5.5
Humanity's Last Exam (Text-Only)4.9
Humanity's Last Exam Text Only (Preview)5.5
Instruction Following82.1
LiveCodeBench (Artificial Analysis)21.4
MAGI-Hard77.8
MASK54.1
MATH Level 539.4
MATH-500 (Artificial Analysis source)62.9
MathVista (OpenVLM)58.2
MM-Vet (OpenVLM)64.1
MMLU80.3
MMLU Pro67.1
MMMU Pro39.5
MMMU Validation (OpenVLM)60.3
MMMU-Pro (Vals 4-Option)48.1
MMStar (OpenVLM)55.3
Loading Atlas data…