# GLM 5.2 Fast Preview Model id on AIHubMix: `glm-5.2-fast-preview` Create an API key: https://console.aihubmix.com/?utm_source=llms-agent&utm_medium=model-llms > GLM-5.2-Fast-Preview is the high-speed version of Zhipu AI’s flagship model GLM-5.2, supporting a 1M ultra-long context. The model’s capabilities are aligned with the GLM-5.2 standard version, offering logical reasoning, long-text understanding, and code generation. Through inference acceleration optimizations, output TPS can reach 1.5–2× that of the GLM-5.2 standard version, significantly improving output speed. It is suitable for scenarios sensitive to output speed, such as real-time dialogue, multi-turn Agent calls, and streaming code generation. - Developer: Z.AI - Context window: 1,000,000 tokens - Input modalities: text - Capabilities: Thinking, Streaming, Tool calling, Structured outputs, Prompt caching - Release date: 2026-06-16 - Pricing: $2.254/M input tokens, $7.889/M output tokens, $0.564/M cached input ## Endpoints (base URL: https://aihubmix.com) - `POST /v1/chat/completions` — OpenAI Chat Completions (`Authorization: Bearer $AIHUBMIX_API_KEY`) - `POST /v1/messages` — Anthropic Messages (`x-api-key: $AIHUBMIX_API_KEY` + `anthropic-version: 2023-06-01`) ## Example ```bash curl -s https://aihubmix.com/v1/chat/completions \ -H "Authorization: Bearer $AIHUBMIX_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model":"glm-5.2-fast-preview","messages":[{"role":"user","content":"Hello"}]}' ``` ## Response Without `stream`, `/v1/chat/completions` returns a standard Chat Completions object: ```json {"id":"...","object":"chat.completion","model":"glm-5.2-fast-preview","choices":[{"message":{"role":"assistant","content":"..."}}],"usage":{"prompt_tokens":12,"completion_tokens":24,"total_tokens":36}} ``` With `"stream": true` the response is `text/event-stream`: read each `data:` JSON chunk until `data: [DONE]`. The Messages and Gemini endpoints return their protocols' native response shapes (Anthropic / Google). ## Errors Error responses carry a `tid` (trace id) — include it when contacting support. Reference: https://docs.aihubmix.com/en/FAQs/HTTP-Codes.md - 400 — parameter error; most are passed through from the upstream provider (media: `prompt_missing`, `size_not_supported`, `n_not_within_range`, …) - 401 — missing `Authorization` header, or the key is invalid/expired - 403 — `insufficient_user_quota` (top up at https://console.aihubmix.com/?utm_source=llms-agent&utm_medium=model-llms), account suspended, or this key is not allowed to use this model - 429 — rate limited; back off and retry - 503 — no channel can serve the request (check the model id and your access), or the upstream provider is throttling; retry later ## More - Model page: https://aihubmix.com/model/glm-5.2-fast-preview - Try in browser: https://playground.aihubmix.com/?model=glm-5.2-fast-preview - Compare with another model (human-facing, side-by-side specs and pricing): https://aihubmix.com/compare — pick this model and a peer there; published pairs are listed in https://aihubmix.com/sitemap-compare.xml, unpublished pairs 404 so do not compose the path by hand - Full parameter schema (machine-readable, authoritative): https://aihubmix.com/model-data/models/glm-5.2-fast-preview.27ee8829.json — per-protocol parameters with types, ranges, enums and defaults. Refreshed together with this page; if it ever 404s, re-resolve via `https://aihubmix.com/model-data/index.json` (find this id, fetch its `path`) - Generate runnable code programmatically: npm `@aihubmix/codegen` — the generator behind the Playground's "Get Code" (4 protocols × 7 languages, media endpoints included); the body it builds is the exact wire body the Playground sends, so generated snippets and real requests cannot diverge. `@aihubmix/model-schema` (npm) translates the parameter schema above into codegen input - Site index for agents: https://aihubmix.com/llms.txt · Onboarding: https://aihubmix.com/agents.md --- Canonical version of this document: https://aihubmix.com/model/glm-5.2-fast-preview/llms.txt