Atlas

Models

← All models

Phi-3.5-vision-instruct

Microsoft Research · 2024-08-22 · 4.1B parameters

Phi-3.5-vision-instruct is Microsoft’s open-weight, 4.15B-class multimodal instruction model, publicly announced on August 22, 2024. It accepts text and one or more images and generates text over a 131,072-token LongRoPE context; the released checkpoint combines a 32-layer Phi-3 Mini language backbone with a CLIP ViT-L/14-336 vision encoder and HD-transform MLP projection. Microsoft reports training on 500B vision-and-text tokens followed by supervised fine-tuning and direct preference optimization for image understanding, OCR, charts and tables, multi-image comparison, and video-frame summarization.

Benchmark scores

BenchmarkScore
Arena Vision — English (No Style Control)876
Arena Vision — English (Style Controlled)936
Capability105.0
CharXiv Reasoning32.7
CharXiv-D55.0
MathVista (mini)43.9
ScienceQA91.3
Vision Arena Overall921
Vision Arena Overall No Style Control852
VISTA15.2
Loading Atlas data…