Mage-VL
Microsoft Research
Mage-VL is Microsoft's open-weight multimodal generative checkpoint for image and video understanding. It combines the Mage-ViT visual encoder, a two-layer projector, and the Qwen3-4B-Instruct-2507 language backbone. A cognition gate within the same checkpoint supports event-triggered streaming responses. The released BF16 weights use Apache 2.0; the language configuration supports 262,144 context positions. The exact complete checkpoint parameter count and release date are not established by the reviewed primary sources.