Atlas

Models

← All models

Mage-VL

Microsoft Research

Mage-VL is Microsoft's open-weight multimodal generative checkpoint for image and video understanding. It combines the Mage-ViT visual encoder, a two-layer projector, and the Qwen3-4B-Instruct-2507 language backbone. A cognition gate within the same checkpoint supports event-triggered streaming responses. The released BF16 weights use Apache 2.0; the language configuration supports 262,144 context positions. The exact complete checkpoint parameter count and release date are not established by the reviewed primary sources.

Benchmark scores

BenchmarkScore
Roboflow Vision Evals — Counting, Judged (September)26.6
Roboflow Vision Evals — Counting, Strict (September)26.6
Roboflow Vision Evals — Detection mAP@50 (September)2.5
Roboflow Vision Evals — Detection mAP@50:95 (September)0.7
Roboflow Vision Evals — Detection mAP@75 (September)0.0
Roboflow Vision Evals — Extraction, Judged (September)55.0
Roboflow Vision Evals — Extraction, Strict (September)47.1
Roboflow Vision Evals — Identification, Judged (September)59.4
Roboflow Vision Evals — Identification, Strict (September)44.8
Roboflow Vision Evals — OCR (September)60.6
Roboflow Vision Evals — Reasoning, Judged (September)24.7
Roboflow Vision Evals — Reasoning, Strict (September)18.8
Roboflow Vision Evals (September 2026 protocol)38.1
Loading Atlas data…