Llama 3.2 90B Vision Instruct
Meta · 2024-09-25 · 88.6B parameters
Llama 3.2 90B Vision Instruct is Meta's instruction-tuned multimodal model built on a Llama 3.1 text backbone with a separately trained image encoder and cross-attention adapter. It accepts text and images and generates text for visual recognition, image reasoning, captioning, document understanding, and assistant-style chat; Meta reports a December 2023 knowledge cutoff.