Atlas

Models

← All models

Qwen2.5-VL-72B-Instruct

Alibaba · 2025-01-26 · 73.4B parameters

Qwen2.5-VL-72B-Instruct is Alibaba's instruction-tuned 72B checkpoint in the Qwen2.5-VL family. It accepts text, image, and video inputs and generates text for multimodal chat, OCR, document understanding, visual reasoning, structured outputs, and visual-agent workflows; its published configuration sets 128,000 maximum positions, while the accompanying model-card guidance identifies 32,768 tokens as the original context and prescribes YaRN for longer text inputs.

Benchmark scores

BenchmarkScore
A-Fantasia81.3
A-Fantasia Backwards Spelling100.0
A-Fantasia Chess68.0
A-Fantasia Cube Rotation76.0
Arena Vision — Chinese (No Style Control)1101
Arena Vision — Chinese (Style Controlled)1123
Arena Vision — English (No Style Control)1116
Arena Vision — English (Style Controlled)1125
BuseyBench SVG1.9
Capability130.1
CharXiv Reasoning49.7
CharXiv-D87.4
ClockBench Accuracy6.1
GeoBench ACW Country Accuracy62.0
MATH-Vision38.1
ObviousBench37.5
OmniDocBench v1.587.0
OSWorld8.8
OSWorld-Verified5.0
ScreenSpot-Pro53.3
SnakeBench Average Apples2.6
SnakeBench Best Apples8.0
SnakeBench Rating19.7
SnakeBench Total Apples53.0
SnakeBench Win Rate38.9
SOLO Bench Easy5.2
SpatialViz-Bench33.3
Video-MME73.5
Vision Arena Overall1122
Vision Arena Overall No Style Control1108
Loading Atlas data…