GLM 5 Vision Turbo
Z.AI · text, image, video → text
GLM-5V-Turbo is Zhipu's first multimodal coding foundation model, built for visual programming tasks. It can natively handle multimodal inputs such as images, videos, and text, and is adept at long-horizon planning, complex programming, and action execution; deeply adapted to Agent workflows, it can collaborate closely with agents like Claude Code and OpenClaw to complete the full closed loop of "understand the environment → plan actions → execute tasks."
Input$0.70 /M
Output$3.10 /M

GLM 5 Vision Turbo