Models

GLM 5.3 Flashx Compare

Compare pricing, specifications, performance, and benchmarks for up to four models.

Z.AIGLM 5.3 Flashx
Z.AI logo
GLM 5.3 Flashx
Z.AI · text, image, video → text

GLM-5.3-FlashX is Z.AI’s high-speed inference model, designed for coding agents, real-time interactions, long-running agentic workflows, and applications that require fast token generation. It is heavily optimized at the inference infrastructure level for lower latency and higher execution efficiency, reaching generation speeds of up to 200 tokens per second. Compared with GLM-5.3-Flash, GLM-5.3-FlashX delivers up to 5× faster inference, with a stronger focus on low-latency and high-throughput workloads.

Input$0.37 /M
Output$1.25 /M
Cache read$0.075 /M

Pick a second model to start comparing.

Popular comparisons

Related model match-ups readers also look at.