LLaVA v1.6 Mistral 7B
LLaVA Authors · 2024-01-30 · 7.6B parameters
LLaVA v1.6 Mistral 7B is an open-weight multimodal model developed by Haotian Liu and collaborators, released in January 2024, pairing a CLIP ViT-L/14 vision encoder with Mistral 7B via a two-layer MLP projection. The v1.6 update introduced dynamic high-resolution input—splitting images into variable-resolution tiles—enabling improved text-in-image reading, chart interpretation, and visual question answering relative to LLaVA 1.5. It used 558K stage-1 pretraining samples and 760K stage-2 instruction samples (1.318M total) and is available on Hugging Face.
Benchmark scores
| Benchmark | Score |
|---|---|
| CharXiv Reasoning | 13.9 |
| CharXiv-D | 35.4 |
| ScienceQA | 72.8 |