LLaVA-OneVision 72B
LLaVA Authors · 2024-08-06 · 73.2B parameters
LLaVA-OneVision Qwen2-72B OV SFT is the largest final-stage checkpoint in the 2024 LLaVA-OneVision release from the LLaVA research collaboration, whose authors were affiliated with ByteDance, NTU, CUHK, and HKUST. It combines Qwen2-72B-Instruct with a SigLIP SO400M vision encoder and a two-layer MLP projector, and was trained through a final 1.6-million-example OneVision stage covering single-image, multi-image, and video data. The released model supports a 32,768-token context and text generation from text, image, and video inputs; the paper reports that the 72B model performed between GPT-4V and GPT-4o on most evaluated benchmarks.