LLaVA v1.5 7B
LLaVA Authors · 2023-10-05 · 7.1B parameters
LLaVA-1.5 7B is a visual instruction-tuned vision-language model developed by Haotian Liu and collaborators, released in October 2023, built on a Vicuna 7B language backbone combined with a CLIP visual encoder. It was a significant milestone among open-weight multimodal LLMs, demonstrating that carefully curated visual instruction tuning data could match larger models on standard vision-language benchmarks. The model is hosted at liuhaotian/llava-v1.5-7b on Hugging Face and supports image-grounded question answering and conversation.
Benchmark scores
| Benchmark | Score |
|---|---|
| AI2D (OpenVLM) | 55.5 |
| Capability | 86.9 |
| HallusionBench (OpenVLM) | 27.6 |
| MATH-Vision | 8.5 |
| MathVista (OpenVLM) | 25.5 |
| MM-Vet (OpenVLM) | 32.9 |
| MMMU Validation (OpenVLM) | 35.7 |
| MMStar (OpenVLM) | 33.1 |
| ScienceQA | 66.8 |