Atlas

Models

← All models

Qwen3-VL-235B-A22B-Instruct

Alibaba · 2025-09-22 · 235.7B parameters

Qwen3-VL-235B-A22B-Instruct is the instruction-tuned variant of Alibaba's flagship open-weight Qwen3-VL vision-language MoE model. It supports text, image, and video understanding, multilingual OCR, spatial grounding, and visual-agent workflows in a non-thinking instruction-following mode. Its native context window is 262,144 tokens and can be extended to 1,000,000 tokens with YaRN.

Benchmark scores

BenchmarkScore
AA-LCR31.7
AA-Omniscience Index-50.7
AIME 202570.7
Artificial Analysis Intelligence Index14.3
Artificial Analysis Omniscience Accuracy20.2
Artificial Analysis Omniscience Hallucination Rate88.7
Artificial Analysis Openness Index50.0
BuseyBench SVG2.0
Capability139.9
ClockBench Accuracy39.4
CritPt0.0
GPQA Diamond71.2
Humanity's Last Exam (Text-Only)6.3
IFBench (Artificial Analysis)42.7
Kangaroo 2025 1-258.3
Kangaroo 2025 11-1285.8
Kangaroo 2025 3-458.3
Kangaroo 2025 5-660.8
Kangaroo 2025 7-882.5
Kangaroo 2025 9-1089.2
LiveCodeBench (Artificial Analysis)59.4
MATH-Vision66.0
MMLU Pro82.3
MMMU Pro67.6
ObviousBench63.9
OmniDocBench v1.589.2
Roboflow Vision Evals68.0
Roboflow Vision Evals - Data Extraction86.6
Roboflow Vision Evals - Object Counting47.3
Roboflow Vision Evals - Object Detection (mAP@50:95)40.3
Loading Atlas data…