VideoChat2 Stage 3 Mistral 7B
OpenGVLab · 2024-05-22
VideoChat2 Stage 3 Mistral 7B is OpenGVLab's instruction-tuned multimodal checkpoint built from a UMT-L/16 vision encoder, a Q-Former, and Mistral-7B-Instruct-v0.2 with rank-16 LoRA. Stage 3 trains the vision stack, Q-Former, projector, and LoRA adapters for three epochs on 2 million image-and-video instruction annotations using eight 224-pixel frames and a 512-token maximum training text length. The published 2.07 GB stage-3 checkpoint is loaded with separately supplied UMT/Q-Former, Stage 2, and Mistral base assets; it accepts text with image or video input and generates text.
Benchmark scores
| Benchmark | Score |
|---|
No benchmark scores recorded.