Phi-3.5-vision-instruct
Microsoft Research · 2024-08-22 · 4.1B parameters
Phi-3.5-vision-instruct is Microsoft’s open-weight, 4.15B-class multimodal instruction model, publicly announced on August 22, 2024. It accepts text and one or more images and generates text over a 131,072-token LongRoPE context; the released checkpoint combines a 32-layer Phi-3 Mini language backbone with a CLIP ViT-L/14-336 vision encoder and HD-transform MLP projection. Microsoft reports training on 500B vision-and-text tokens followed by supervised fine-tuning and direct preference optimization for image understanding, OCR, charts and tables, multi-image comparison, and video-frame summarization.