Atlas

Models

← All models

LLaVA-OneVision 72B

LLaVA Authors · 2024-08-06 · 73.2B parameters

LLaVA-OneVision Qwen2-72B OV SFT is the largest final-stage checkpoint in the 2024 LLaVA-OneVision release from the LLaVA research collaboration, whose authors were affiliated with ByteDance, NTU, CUHK, and HKUST. It combines Qwen2-72B-Instruct with a SigLIP SO400M vision encoder and a two-layer MLP projector, and was trained through a final 1.6-million-example OneVision stage covering single-image, multi-image, and video data. The released model supports a 32,768-token context and text generation from text, image, and video inputs; the paper reports that the 72B model performed between GPT-4V and GPT-4o on most evaluated benchmarks.

Benchmark scores

BenchmarkScore
AI2D (OpenVLM)86.2
Capability120.8
CharXiv Reasoning36.4
CharXiv-D71.8
HallusionBench (OpenVLM)47.9
MATH-Vision25.3
MathVista (OpenVLM)68.4
MM-Vet (OpenVLM)60.6
MMMU Validation (OpenVLM)56.6
MMStar (OpenVLM)65.8
OSWorld2.4
OSWorld (self-reported, A11y tree)2.7
OSWorld (self-reported, Screenshot + A11y tree)2.4
OSWorld (self-reported, Set-of-Mark)2.4
Video-MME66.3
Loading Atlas data…