GLM 4.7 Flash
Zhipu AI (Z.ai) · 2026-01-19 · 31.2B parameters
GLM-4.7-Flash is Z.ai's January 2026 open-weight lightweight MoE model with a 30B-A3B architecture, designed to balance performance and efficiency for local deployment. The Hugging Face model card positions it as a strong 30B-class model for coding, reasoning, and agentic tasks. It is released under the MIT license and documents serving paths for vLLM, SGLang, Transformers, and Docker Model Runner.