Z.AI Models

71 modelsGeneral models free to startUp to 1.05M contextOfficial site ↗

Usage

Last 30 days · 2026-09-07 to 2026-10-06

Tokens

809B

Requests

5.2M

Models in use

60 of 71

Tokens per day, stacked by model

022.5B45B09-0709-1409-2109-2810-052026-09-07 — 34,211,324,710 tokens glm-5.3-flash: 19,059,816,875 glm-5.3: 8,506,221,890 coding-glm-5.3-flash: 3,673,599,520 coding-glm-5.3: 1,706,993,930 39 more models: 882,021,835 coding-glm-5.3-free: 203,013,875 glm-5.2: 170,388,190 coding-glm-4.6: 9,268,5952026-09-08 — 40,512,087,435 tokens glm-5.3-flash: 31,111,412,135 glm-5.3: 4,765,267,430 coding-glm-5.3-flash: 1,836,691,175 coding-glm-5.3: 1,186,262,485 39 more models: 1,008,178,900 coding-glm-5.3-free: 489,010,705 glm-5.2: 81,151,885 coding-glm-4.6: 34,112,7202026-09-09 — 37,302,208,485 tokens glm-5.3-flash: 28,819,987,670 glm-5.3: 2,668,476,660 coding-glm-5.3-flash: 2,630,894,580 coding-glm-5.3: 1,787,867,080 39 more models: 670,287,505 coding-glm-5.3-free: 665,144,695 glm-5.2: 33,290,630 coding-glm-4.6: 26,259,6652026-09-10 — 36,761,777,540 tokens glm-5.3-flash: 20,838,310,355 coding-glm-5.3: 6,256,329,915 coding-glm-5.3-flash: 4,554,588,335 glm-5.3: 3,918,748,480 39 more models: 648,165,995 coding-glm-5.3-free: 337,835,180 glm-5.2: 179,967,880 coding-glm-4.6: 27,831,4002026-09-11 — 27,464,306,905 tokens glm-5.3-flash: 13,499,143,900 coding-glm-5.3: 6,302,106,195 coding-glm-5.3-flash: 4,176,616,120 glm-5.3: 2,423,067,985 39 more models: 448,672,990 coding-glm-5.3-free: 387,048,340 glm-5.2: 195,702,255 coding-glm-4.6: 31,949,1202026-09-12 — 28,040,164,260 tokens glm-5.3-flash: 12,697,067,645 coding-glm-5.3: 5,913,778,670 coding-glm-5.3-flash: 5,901,329,840 glm-5.3: 2,562,207,840 39 more models: 509,233,405 coding-glm-5.3-free: 349,691,025 glm-5.2: 55,059,175 coding-glm-4.6: 51,796,6602026-09-13 — 28,973,843,095 tokens glm-5.3-flash: 15,594,889,085 coding-glm-5.3-flash: 7,183,268,420 coding-glm-5.3: 3,550,681,185 glm-5.3: 1,597,157,395 39 more models: 631,708,785 coding-glm-5.3-free: 270,211,240 glm-5.2: 93,276,330 coding-glm-4.6: 52,650,6552026-09-14 — 29,853,050,580 tokens glm-5.3-flash: 13,233,719,365 coding-glm-5.3-flash: 7,729,760,530 coding-glm-5.3: 3,938,101,995 glm-5.3: 3,303,513,370 39 more models: 816,206,435 glm-5.2: 428,183,875 coding-glm-5.3-free: 362,332,145 coding-glm-4.6: 41,232,8652026-09-15 — 31,273,410,215 tokens glm-5.3-flash: 13,130,306,175 coding-glm-5.3-flash: 8,211,169,435 glm-5.3: 5,153,701,615 coding-glm-5.3: 3,269,071,695 39 more models: 683,567,760 coding-glm-5.3-free: 448,137,335 glm-5.2: 336,176,875 coding-glm-4.6: 41,279,3252026-09-16 — 32,518,141,520 tokens glm-5.3-flash: 14,586,531,655 glm-5.3: 9,593,025,950 coding-glm-5.3-flash: 3,440,419,375 coding-glm-5.3: 3,124,363,945 39 more models: 1,049,132,920 coding-glm-5.3-free: 515,564,400 glm-5.2: 159,317,925 coding-glm-4.6: 49,785,3502026-09-17 — 44,993,223,120 tokens glm-5.3-flash: 15,877,690,165 coding-glm-5.3: 11,582,806,730 glm-5.3: 10,737,226,180 coding-glm-5.3-flash: 5,313,812,065 39 more models: 648,208,745 coding-glm-5.3-free: 398,743,550 glm-5.2: 344,222,045 coding-glm-4.6: 90,513,6402026-09-18 — 36,122,929,035 tokens coding-glm-5.3: 13,823,751,070 glm-5.3-flash: 11,001,676,670 glm-5.3: 6,410,825,905 coding-glm-5.3-flash: 3,477,628,505 39 more models: 725,919,010 coding-glm-5.3-free: 396,123,935 glm-5.2: 213,453,700 coding-glm-4.6: 67,991,485 glm-5.3-flashx: 5,558,7552026-09-19 — 28,533,962,090 tokens coding-glm-5.3: 11,623,725,605 glm-5.3-flash: 6,963,528,085 coding-glm-5.3-flash: 4,645,912,290 glm-5.3: 2,312,909,370 glm-5.3-flashx: 1,395,826,035 39 more models: 695,188,980 coding-glm-5.3-free: 680,200,690 glm-5.2: 145,791,305 coding-glm-4.6: 70,879,7302026-09-20 — 19,400,823,010 tokens glm-5.3-flash: 8,460,224,920 glm-5.3: 4,200,871,800 coding-glm-5.3: 2,918,865,425 glm-5.3-flashx: 2,808,218,655 39 more models: 412,065,890 glm-5.2: 277,743,675 coding-glm-5.3-free: 193,201,335 coding-glm-5.3-flash: 99,320,465 coding-glm-4.6: 30,310,8452026-09-21 — 37,804,096,995 tokens glm-5.3-flash: 19,942,791,035 glm-5.3: 13,101,303,460 glm-5.3-flashx: 2,069,376,710 coding-glm-5.3: 1,670,930,525 glm-5.2: 430,262,135 39 more models: 271,481,805 coding-glm-5.3-flash: 144,302,820 coding-glm-4.6: 116,002,130 coding-glm-5.3-free: 57,646,3752026-09-22 — 35,869,308,860 tokens glm-5.3: 10,710,281,705 glm-5.3-flash: 10,339,650,560 glm-5.3-flashx: 5,418,770,705 coding-glm-5.3-flash: 4,192,272,935 coding-glm-5.3: 3,633,681,365 39 more models: 562,076,110 glm-5.2: 498,997,350 coding-glm-5.3-free: 302,025,405 coding-glm-4.6: 211,552,7252026-09-23 — 30,465,635,140 tokens glm-5.3-flash: 8,747,566,540 glm-5.3: 7,301,127,995 coding-glm-5.3: 5,410,572,435 glm-5.3-flashx: 4,369,220,515 coding-glm-5.3-flash: 3,066,424,605 glm-5.2: 745,096,625 39 more models: 420,725,160 coding-glm-5.3-free: 304,845,810 coding-glm-4.6: 100,055,4552026-09-24 — 34,164,094,985 tokens glm-5.3-flash: 10,569,837,200 glm-5.3: 8,364,816,115 coding-glm-5.3: 5,898,677,820 coding-glm-5.3-flash: 4,785,211,620 glm-5.3-flashx: 3,327,104,625 39 more models: 707,203,680 coding-glm-5.3-free: 272,651,730 glm-5.2: 143,812,845 coding-glm-4.6: 94,779,3502026-09-25 — 15,343,110,980 tokens glm-5.3-flash: 4,985,971,445 coding-glm-5.3-flash: 4,051,147,350 coding-glm-5.3: 2,883,232,820 glm-5.3: 967,827,430 glm-5.3-flashx: 840,149,940 glm-5.2: 666,315,440 39 more models: 505,358,035 coding-glm-5.3-free: 312,695,275 coding-glm-4.6: 130,413,2452026-09-26 — 20,651,087,210 tokens coding-glm-5.3: 8,605,799,410 glm-5.3-flash: 6,829,929,810 coding-glm-5.3-flash: 3,320,270,070 glm-5.3: 946,614,890 39 more models: 482,853,450 coding-glm-5.3-free: 324,836,260 coding-glm-4.6: 73,031,300 glm-5.2: 34,994,665 glm-5.3-flashx: 32,757,3552026-09-27 — 13,615,805,890 tokens glm-5.3-flash: 4,872,180,945 coding-glm-5.3: 4,716,002,910 coding-glm-5.3-flash: 2,075,529,115 glm-5.3: 1,004,714,310 39 more models: 344,547,340 coding-glm-5.3-free: 270,785,990 glm-5.3-flashx: 132,078,085 glm-5.2: 114,022,645 coding-glm-4.6: 85,944,5502026-09-28 — 19,556,045,830 tokens coding-glm-5.3: 6,818,499,240 glm-5.3-flash: 5,431,874,900 coding-glm-5.3-flash: 3,167,034,665 glm-5.3: 2,905,327,720 39 more models: 595,064,745 coding-glm-5.3-free: 343,690,100 glm-5.3-flashx: 219,024,280 coding-glm-4.6: 50,961,430 glm-5.2: 24,568,7502026-09-29 — 25,181,650,980 tokens coding-glm-5.3-flash: 7,876,009,835 coding-glm-5.3: 7,106,682,645 glm-5.3-flash: 5,300,363,395 glm-5.3: 2,705,659,930 glm-5.3-flashx: 1,186,217,030 39 more models: 386,499,330 coding-glm-5.3-free: 284,678,940 coding-glm-4.6: 215,554,515 glm-5.2: 119,985,3602026-09-30 — 17,091,670,370 tokens glm-5.3-flash: 5,004,491,970 coding-glm-5.3-flash: 4,938,282,050 coding-glm-5.3: 4,289,874,190 glm-5.3-flashx: 795,160,420 glm-5.3: 637,618,345 39 more models: 519,364,270 coding-glm-4.6: 516,698,335 coding-glm-5.3-free: 271,319,615 glm-5.2: 118,861,1752026-10-01 — 16,228,282,305 tokens glm-5.3-flash: 5,218,315,370 coding-glm-5.3-flash: 5,092,997,610 glm-5.3: 2,673,209,540 coding-glm-5.3: 1,347,471,725 coding-glm-4.6: 700,716,325 39 more models: 488,980,375 glm-5.3-flashx: 402,956,680 coding-glm-5.3-free: 186,515,730 glm-5.2: 117,118,9502026-10-02 — 28,131,681,730 tokens coding-glm-5.3-flash: 16,045,784,060 coding-glm-5.3: 3,919,846,055 glm-5.3-flash: 3,697,162,225 glm-5.3: 2,986,531,205 39 more models: 353,653,935 glm-5.3-flashx: 350,196,290 coding-glm-4.6: 333,044,635 coding-glm-5.3-free: 310,576,505 glm-5.2: 134,886,8202026-10-03 — 20,655,880,165 tokens coding-glm-5.3-flash: 9,792,613,930 coding-glm-5.3: 4,721,071,280 glm-5.3-flash: 4,397,598,080 39 more models: 581,832,470 glm-5.3: 475,975,980 coding-glm-5.3-free: 324,514,250 glm-5.2: 188,391,270 coding-glm-4.6: 106,215,875 glm-5.3-flashx: 67,667,0302026-10-04 — 14,428,007,190 tokens coding-glm-5.3: 4,966,577,780 glm-5.3-flash: 4,004,269,205 coding-glm-5.3-flash: 3,743,535,715 39 more models: 686,390,315 glm-5.3: 346,382,610 coding-glm-5.3-free: 296,558,520 glm-5.2: 179,539,410 glm-5.3-flashx: 124,080,365 coding-glm-4.6: 80,673,2702026-10-05 — 14,700,403,750 tokens coding-glm-5.3: 5,810,931,345 glm-5.3-flash: 5,219,238,295 coding-glm-5.3-flash: 2,461,165,070 glm-5.3: 369,382,260 39 more models: 275,269,845 glm-5.2: 215,417,980 coding-glm-5.3-free: 187,834,870 glm-5.3-flashx: 116,899,020 coding-glm-4.6: 44,265,0652026-10-06 — 9,530,492,620 tokens glm-5.3-flash: 5,387,307,380 coding-glm-5.3: 1,255,473,860 coding-glm-5.3-flash: 1,116,090,780 glm-5.3: 948,055,000 39 more models: 451,593,030 coding-glm-5.3-free: 121,525,420 glm-5.2: 108,812,060 glm-5.3-flashx: 80,380,175 coding-glm-4.6: 61,254,915
  • glm-5.3-flash
  • coding-glm-5.3
  • coding-glm-5.3-flash
  • glm-5.3
  • glm-5.3-flashx
  • coding-glm-5.3-free
  • glm-5.2
  • coding-glm-4.6
  • 39 more models

Which models that traffic went to

  1. GLM 5.3 Flash41.4%335B
  2. Coding GLM 5.318.5%150B
  3. Coding GLM 5.3 Flash17.1%139B
  4. GLM 5.315.4%125B
  5. GLM 5.3 Flashx2.9%23.7B
  6. Coding GLM 5.3 (free)1.2%9.9B
  7. GLM 5.20.8%6.6B
  8. Coding GLM 4.60.4%3.5B
  9. 39 more models2.2%17.5B

Share of 809B tokens. 13 models with traffic report no token counts and cannot be ranked here, including coding-glm-5-turbo-free and ox-alpha — they are in the request view.

The two views disagree on purpose: a model can take a large share of the calls and a small share of the tokens — many short requests — or the reverse. Which one matters depends on whether your cost is driven by call volume or by prompt length. Measured on AIHubMix over the last 30 days, counting the 71 model IDs listed on this page; traffic routed through upstream-specific IDs that are not in the public catalog is not included.

All 71 Z.AI Models

Open in model list
Z.AI models on AIHubMix with input and output modalities, context length, maximum output, price per million tokens including cache read and cache write rates, and measured throughput and latency.
Modalities
coding-glm-5.3-freeTakes text, returns text.1.05M131KFreeFree/M—49 tok/s4.07 s
ox-alphaTakes text, vision, video, returns text.1.05M131KFreeFree/M———
coding-glm-5.3Takes text, returns text.1.05M131K$0.06$0.22/M$0.015/M41 tok/s4.07 s
glm-5.3-flashTakes text, vision, video, returns text.1.05M131K$0.1127$0.3944/M$0.0282/M53 tok/s1.45 s
glm-5.3Takes text, returns text.1.05M131K$1.1268$3.9438/M$0.2817/M47 tok/s1.07 s
coding-glm-5.2-freeTakes text, returns text.1M131KFreeFree/M—47 tok/s4.48 s
coding-glm-5.3-flash-freeTakes text, vision, video. Output modality not published.1M131KFreeFree/M—25 tok/s5.71 s
coding-glm-5.3-flashTakes text, vision, video. Output modality not published.1M131K$0.0282$0.0986/M$0.007/M34 tok/s5.56 s
coding-glm-5.2Takes text, returns text.1M131K$0.06$0.22/M—47 tok/s4.21 s
glm-5.3-flashxTakes text, vision, video, returns text.1M—$0.37$1.25/M$0.075/M79 tok/s2.74 s
glm-5.2Takes text, returns text.1M131K$1.1268$3.9438/M$0.2817/M32 tok/s0.74 s
cloudflare-glm-5.2Takes text, returns text.1M131K$1.4$4.4002/M$0.2604/M——
glm-5.2-fast-previewTakes text, returns text.1M131K$2.254$7.889/M$0.5635/M68 tok/s1.00 s
coding-glm-5-turbo-freeTakes text, returns text.205K131KFreeFree/M———
cc-glm-5-turboTakes text, returns text.205K131K$0.06$0.22/M—8 tok/s1.86 s
coding-glm-5-turboTakes text, returns text.205K131K$0.06$0.22/M———
glm-5-turboTakes text, returns text.205K131K$1.2$3.9996/M$0.24/M15 tok/s3.38 s
zai-glm-5-turboTakes text, returns text.205K131K$1.2$3.9996/M$0.24/M15 tok/s3.38 s
coding-glm-4.6-freeTakes text, returns text.200K131KFreeFree/M—36 tok/s5.36 s
coding-glm-4.7-freeTakes text, returns text.200K131KFreeFree/M—33 tok/s4.50 s
coding-glm-5-freeTakes text, returns text.200K131KFreeFree/M—45 tok/s2.40 s
coding-glm-5.1-freeTakes text, returns text.200K131KFreeFree/M—48 tok/s3.28 s
glm-4.6Takes text, returns text.200K131KFreeFree/MFree/M20 tok/s4.65 s
glm-4.7-flash-freeTakes text, returns text.200K131KFreeFree/M—44 tok/s20.17 s
cc-glm-5Takes text, returns text.200K131K$0.06$0.22/M—19 tok/s3.43 s
cc-glm-5.1Takes text, returns text.200K131K$0.06$0.22/M———
coding-glm-4.6Takes text, returns text.200K131K$0.06$0.22/M$0.011/M37 tok/s1.89 s
coding-glm-4.7Takes text, returns text.200K131K$0.06$0.22/M$0.011/M36 tok/s1.82 s
coding-glm-5Takes text, returns text.200K131K$0.06$0.22/M—56 tok/s0.95 s
coding-glm-5.1Takes text, returns text.200K131K$0.06$0.22/M—49 tok/s6.75 s
glm-4.7Takes text, returns text.200K131K$0.274$1.0959/M$0.0548/M24 tok/s6.37 s
glm-5v-turboTakes text, vision, video, returns text.200K131K$0.7042$3.0985/M$0.169/M24 tok/s2.45 s
glm-5.1Takes text, returns text.200K131K$0.845$3.38/M$0.1831/M22 tok/s1.59 s
coding-glm-4.5-airTakes text. Output modality not published.131K—$0.014$0.084/M—21 tok/s6.01 s
glm-4.6vTakes text, vision, video, returns text.131K33K$0.137$0.411/M$0.0274/M55 tok/s2.20 s
glm-4.5-airTakes text. Output modality not published.131K98K$0.14$0.84/M—91 tok/s1.42 s
glm-4.5Takes text. Output modality not published.131K98K$0.4$1.6/M—107 tok/s0.46 s
glm-4.5vTakes text, vision, video, returns text.66K16K$0.274$0.822/M—37 tok/s4.10 s
glm-ocrTakes vision, returns text.32K—$0.0282$0.0282/M———
embedding-2Takes text. Output modality not published.8K—$0.0686$0.0686/M———
embedding-3Takes text. Output modality not published.8K—$0.0686$0.0686/M———
glm-imageTakes text, returns vision.——FreeFree/M———
Pro/THUDM/GLM-4.1V-9B-Thinking——$0.04$0.16/M———
THUDM/GLM-4-9B-0414——$0.05$0.05/M———
THUDM/GLM-Z1-9B-0414——$0.05$0.05/M———
cc-glm-4.6——$0.06$0.22/M———
cc-glm-4.7——$0.06$0.22/M———
THUDM/GLM-4-32B-0414——$0.08$0.08/M———
THUDM/GLM-Z1-32B-0414——$0.08$0.08/M———
glm-4-flash——$0.1$0.1/M———
THUDM/GLM-4.1V-9B-Thinking——$0.1$0.1/M———
doubao-1-5-pro-32k-250115——$0.108$0.27/M———
chatglm_lite——$0.2858$0.2858/M———
alicloud-glm-4.7——$0.411$1.9178/M$0.411/M44 tok/s1.01 s
alicloud-glm-5——$0.5634$2.5353/M$0.1127/M37 tok/s9.99 s
doubao-1-5-pro-256k-250115——$0.684$1.2312/M———
glm-3-turbo——$0.71$0.71/M———
chatglm_std——$0.7144$0.7144/M———
chatglm_turbo——$0.7144$0.7144/M———
glm-4.5-airxTakes text. Output modality not published.——$1.1$4.51/M$0.22/M——
chatglm_pro——$1.4286$1.4286/M———
glm-4v-plus——$2$2/M———
glm-zero-preview——$2$2/M———
glm-4.5-xTakes text. Output modality not published.——$2.2$8.91/M$0.44/M1 tok/s0.59 s
cbs-glm-4.7——$2.25$2.75/M———
glm-4-plus——$8$8/M———
cogview-3-plus——$10$10/M———
glm-4——$14.2$14.2/M———
glm-4v——$14.2$14.2/M———
code-davinci-edit-001——$20$20/M———
cogview-3——$35.5$35.5/M———

Prices are USD per million tokens; cache read and cache write are the rates for prompt-cache hits and for writing a prompt into the cache. Throughput and latency are measured on AIHubMix — the same figures the model detail page shows — not vendor claims. A dash means the catalog does not publish that field for that model, which is not the same as the model not supporting it.

Z.AI on AIHubMix

Which Z.AI model should I start with?

coding-glm-4.6-free is free on input — the cheapest entry here that declares tool calling, and it carries a 200K context. Move up to cogview-3 when answer quality matters more than cost, or to coding-glm-5.3-free for long-form reasoning.

Which of these models reason before answering?

29 of the 71 models here declare a reasoning phase — they work through the problem before producing an answer, which helps on multi-step problems at the cost of extra output tokens. Use the Reasoning filter above the table to see them. The catalog does not record anything further about how they differ, so this page does not sort them into families.

Why are there several entries for the same model?

Because each row is a route you can call, not a model release. Some IDs name an upstream (azure-, alicloud-, cc-), some are the open-weight repository form (THUDM/…), and some differ only in capitalisation, kept so older integrations keep working.

The catalog does not carry a field saying which of those a given row is, so this page does not sort them into buckets it would have to invent. Every row shows that route’s own price, context and speed — compare those directly, and open a model to see the upstreams that serve it.

How is cached input billed?

The Cache read column is the rate for input tokens served from the prompt cache — for example coding-glm-4.6 bills cache hits at 18.33% of the input rate and coding-glm-4.7 bills cache hits at 18.33% of the input rate. Cache write is the surcharge for putting a prompt into the cache in the first place, and only a few upstreams bill it separately. A dash in either column means the catalog carries no cache rate for that model, so plan on paying the full input rate.

Do I need a separate Z.AI account?

No. One AIHubMix key covers every model on this page, and switching between them is a change to the model string — billing, rate limits, and logs stay in one place.

Start calling Z.AI in one line

One key, one endpoint, 913 models across 42 model authors.