GLM 4.5 Vision
Z.AI · text, image, video → text
GLM-4.5V is a vision-language foundational model designed for multimodal agent applications. Based on a mixture-of-experts (MoE) architecture, it has 106 billion parameters and 12 billion active parameters. It delivers outstanding performance in video understanding, image question answering, OCR, and document parsing, and achieves significant improvements in front-end web encoding, basic reasoning, and spatial reasoning.
Input$0.27 /M
Output$0.82 /M

GLM 4.5 Vision