GLM-4.1V-9B-Thinking is an open-source Vision Language Model (VLM) jointly released by Zhipu AI and the KEG Laboratory at Tsinghua University, designed specifically for handling complex multimodal cognitive tasks. Based on the GLM-4-9B-0414 foundation model, it significantly enhances cross-modal reasoning ability and stability by introducing the “Chain-of-Thought” reasoning mechanism and using reinforcement learning strategies. As a lightweight model with 9 billion parameters, it strikes a balance between deployment efficiency and performance. In 28 authoritative benchmark evaluations, it matched or even outperformed the 72-billion-parameter Qwen-2.5-VL-72B model in 18 tasks. The model excels not only in image-text understanding, mathematical and scientific reasoning, and video understanding, but also supports images up to 4K resolution and inputs of arbitrary aspect ratios.
← Models
THUDM/GLM-4.1V-9B-Thinking
THUDM/GLM-4.1V-9B-Thinking Compare
Compare pricing, specifications, performance, and benchmarks for up to four models.
THUDM/GLM-4.1V-9B-Thinking+ Add model · 1/4
THUDM/GLM-4.1V-9B-Thinking
Z.AI · → text
Input$0.10 /M
Output$0.10 /M
Pick a second model to start comparing.
Popular comparisons
Related model match-ups readers also look at.
