by Z.AI
GLM-5.2-Fast-Preview is the high-speed version of Zhipu AI’s flagship model GLM-5.2, supporting a 1M ultra-long context. The model’s capabilities are aligned with the GLM-5.2 standard version, offering logical reasoning, long-text understanding, and code generation. Through inference acceleration optimizations, output TPS can reach 1.5–2× that of the GLM-5.2 standard version, significantly improving output speed. It is suitable for scenarios sensitive to output speed, such as real-time dialogue, multi-turn Agent calls, and streaming code generation.
| Input | $2.25 / 1M tokens |
| Output | $7.89 / 1M tokens |
| Cache read | $0.56 / 1M tokens |
| Context window | 1,000,000 tokens |
| Modalities | text |
| Features | thinking, tools, function_calling, structured_outputs |
| Endpoints | chat_completions, claude_api |
Use glm-5.2-fast-preview via the AIHubMix unified API — one interface for every major LLM.