GLM-4.1V-9B-Thinking
Zhipu AI (Z.ai) · 2025-07-02 · 10.3B parameters
GLM-4.1V-9B is Zhipu AI's compact vision-language model built around a 9B language backbone and post-trained with Reinforcement Learning with Curriculum Sampling (RLCS) for versatile multimodal reasoning. The Thinking checkpoint was publicly released on July 2, 2025 and supports images, videos, and documents across STEM, GUI-agent, long-document, and code-generation tasks. Despite its smaller size, it achieves strong performance relative to much larger vision-language models.
Benchmark scores
| Benchmark | Score |
|---|---|
| MATH-Vision | 54.4 |
| SnakeBench Average Apples | 0.6 |
| SnakeBench Best Apples | 2.0 |
| SnakeBench Rating | 13.1 |
| SnakeBench Total Apples | 7.0 |
| SnakeBench Win Rate | 33.3 |