Pricing
- Input Tokens: $1.200 /M tokens
- Output Tokens: $4.000 /M tokens
- Cache Read: $0.240 /M tokens
Input Modalities
Output Modalities
- Text
Try this model
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["AIHUBMIX_API_KEY"],
base_url="https://aihubmix.com/v1",
)
response = client.chat.completions.create(
model="zai-glm-5-turbo",
messages=[
{
"role": "user",
"content": "Hello, how are you?"
}
],
max_tokens=1024,
stream=False,
)
print(response.choices[0].message.content)Frequently asked questions
How much does Zai Glm 5 Turbo cost?
How do I call Zai Glm 5 Turbo via API?
Who created Zai Glm 5 Turbo?
Compare Zai Glm 5 Turbo
More models from Z.AI
See all Z.AI models →GLM-5.3 is Z.AI’s coding and agentic reasoning model, built for complex software engineering, long-running agent tasks, vulnerability analysis, and other demanding workloads. Building on GLM-5.2, it incorporates further post-training improvements to deliver stronger coding performance, better task execution, and greater token efficiency. We currently offer the production-ready GLM-5.3 API with unlimited concurrency, making it well suited for high-throughput workloads, coding agents, and large-scale automation. For a limited time, GLM-5.3 is available at 10% off.
GLM-5.3 is Z.ai’s reasoning model for coding and agentic workflows, designed for complex software engineering, long-running agents, and vulnerability analysis. It uses the same base model as GLM-5.2, with scaled post-training improving coding, task execution, and token efficiency. This model is a limited-time preview version of GLM-5.3, intended for testing and evaluation only. Service stability is not guaranteed, and we do not recommend using it in production environments. We’re waiting for the official commercial API release and will integrate it as soon as official support becomes available.
GLM 5.2 is a large-scale reasoning model developed by Z Ai that supports text input and output. Featuring a 128,000-token context window, this model is designed to handle complex, data-heavy tasks. It is exceptionally well-suited for long-horizon agent workflows and project-level software engineering.
GLM-5.2 is Z.ai’s flagship model for the era of long-horizon tasks. With a truly usable 1M-token context window, it can handle project-level engineering context, execute long-running tasks more reliably, follow engineering standards more consistently, and complete the full development workflow from requirements to multi-platform deployment in a single task.
GLM-5.2-Fast-Preview is the high-speed version of Zhipu AI’s flagship model GLM-5.2, supporting a 1M ultra-long context. The model’s capabilities are aligned with the GLM-5.2 standard version, offering logical reasoning, long-text understanding, and code generation. Through inference acceleration optimizations, output TPS can reach 1.5–2× that of the GLM-5.2 standard version, significantly improving output speed. It is suitable for scenarios sensitive to output speed, such as real-time dialogue, multi-turn Agent calls, and streaming code generation.
coding glm 5.2 free (coding-glm-5.2-free) is a free coding-focused API route offered by AIHubMix, powered by Z.ai’s open-source flagship GLM-5.2 model. It is suitable for code generation, debugging, project-level development, tool use, and long-running agent tasks, with support for reasoning, function calling, and structured outputs. Compared with GLM-5.1, it offers stronger coding capabilities and a 1-million-token context window. Each account is limited to 5 requests per minute, 500 requests per day, and 1 million free tokens per day.
AIHubMix© 2023 - 2026 AIHubMix, LLC