Z.AI Models

66 modelsGeneral models free to startUp to 1.05M contextOfficial site

Usage

Last 30 days · 2026-07-22 to 2026-08-20

Tokens

294B

Requests

2.7M

Models in use

56 of 66

Tokens per day, stacked by model

014.1B28.3B07-2207-2908-0508-1208-192026-07-22 — 11,124,628,340 tokens glm-5.2: 6,005,081,095 coding-glm-5: 2,382,155,695 glm-5.1: 1,409,570,620 coding-glm-4.6: 576,161,170 coding-glm-5.1: 300,904,550 38 more models: 254,320,930 coding-glm-5.2: 180,348,905 coding-glm-4.7: 16,085,3752026-07-23 — 15,843,694,810 tokens glm-5.2: 6,861,096,065 coding-glm-5: 6,354,510,535 glm-5.1: 736,011,655 coding-glm-4.6: 702,173,625 coding-glm-5.1: 624,440,500 coding-glm-5.2: 307,838,850 38 more models: 239,023,985 coding-glm-4.7: 18,599,5952026-07-24 — 8,883,641,515 tokens glm-5.2: 4,879,494,005 coding-glm-5: 1,960,650,500 coding-glm-4.6: 657,090,715 coding-glm-5.2: 592,579,890 coding-glm-5.1: 341,731,500 38 more models: 331,395,035 glm-5.1: 97,143,135 coding-glm-4.7: 23,556,7352026-07-25 — 3,615,730,225 tokens glm-5.2: 1,801,703,600 coding-glm-4.6: 591,196,795 coding-glm-5: 412,246,310 38 more models: 266,544,040 coding-glm-5.2: 223,028,315 coding-glm-5.1: 179,173,810 glm-5.1: 136,147,045 coding-glm-4.7: 5,690,3102026-07-26 — 2,993,448,495 tokens glm-5.2: 1,246,434,250 coding-glm-4.6: 723,500,005 coding-glm-5.2: 355,446,270 coding-glm-5: 331,818,805 coding-glm-5.1: 147,428,075 38 more models: 109,146,180 glm-5.1: 79,321,815 coding-glm-4.7: 353,0952026-07-27 — 3,184,329,180 tokens glm-5.2: 1,928,084,420 coding-glm-4.6: 650,552,975 38 more models: 265,660,925 coding-glm-5.1: 200,362,615 glm-5.1: 87,349,305 coding-glm-5.2: 46,839,225 coding-glm-4.7: 5,479,7152026-07-28 — 4,743,023,160 tokens glm-5.2: 3,019,814,460 38 more models: 749,842,555 coding-glm-4.6: 665,739,050 coding-glm-5.1: 186,280,565 glm-5.1: 76,261,890 coding-glm-5.2: 42,484,945 coding-glm-4.7: 2,599,6952026-07-29 — 6,356,994,015 tokens coding-glm-5: 3,383,685,565 glm-5.2: 1,673,889,220 coding-glm-4.6: 700,346,260 coding-glm-5.1: 270,573,970 38 more models: 255,992,275 glm-5.1: 68,407,110 coding-glm-5.2: 4,099,6152026-07-30 — 12,792,717,625 tokens coding-glm-5: 7,019,128,480 glm-5.2: 3,826,474,430 coding-glm-4.6: 728,535,980 coding-glm-5.1: 646,015,185 38 more models: 477,635,640 glm-5.1: 61,496,795 coding-glm-5.2: 33,349,095 coding-glm-4.7: 82,0202026-07-31 — 5,725,622,850 tokens glm-5.2: 4,339,203,630 coding-glm-4.6: 650,897,510 38 more models: 244,220,475 coding-glm-5: 229,979,190 coding-glm-5.1: 104,535,965 coding-glm-5.2: 99,472,420 glm-5.1: 46,589,015 coding-glm-4.7: 10,724,6452026-08-01 — 3,673,283,425 tokens glm-5.2: 1,653,512,820 coding-glm-4.6: 799,564,875 coding-glm-5: 589,923,110 coding-glm-5.1: 334,767,795 38 more models: 222,649,780 glm-5.1: 47,515,420 coding-glm-5.2: 25,349,6252026-08-02 — 2,507,995,155 tokens coding-glm-4.6: 708,474,105 coding-glm-5: 598,037,920 glm-5.2: 572,625,540 coding-glm-5.1: 283,911,985 38 more models: 218,097,290 glm-5.1: 125,624,540 coding-glm-5.2: 1,028,820 coding-glm-4.7: 194,9552026-08-03 — 4,463,338,270 tokens glm-5.2: 1,269,027,325 glm-5.1: 1,064,460,725 coding-glm-5: 884,499,220 coding-glm-4.6: 632,725,970 coding-glm-5.1: 259,694,260 38 more models: 243,263,140 coding-glm-5.2: 109,086,900 coding-glm-4.7: 580,7302026-08-04 — 10,181,814,485 tokens coding-glm-5: 4,371,633,950 glm-5.2: 2,870,253,770 glm-5.1: 943,627,780 coding-glm-4.6: 784,551,105 coding-glm-5.1: 645,673,130 38 more models: 561,473,175 coding-glm-5.2: 4,550,020 coding-glm-4.7: 51,5552026-08-05 — 8,392,866,115 tokens coding-glm-5: 4,825,857,940 glm-5.1: 968,459,515 coding-glm-4.6: 877,101,420 glm-5.2: 858,483,960 coding-glm-5.1: 509,630,370 38 more models: 324,932,285 coding-glm-5.2: 28,285,440 coding-glm-4.7: 115,1852026-08-06 — 10,657,976,760 tokens coding-glm-5: 4,551,619,675 coding-glm-5.2: 3,123,752,490 glm-5.1: 890,584,475 glm-5.2: 681,881,285 coding-glm-4.6: 642,975,700 coding-glm-5.1: 414,884,975 38 more models: 352,278,1602026-08-07 — 13,338,325,155 tokens coding-glm-5: 5,565,066,830 coding-glm-5.2: 4,663,565,115 glm-5.2: 883,979,050 coding-glm-4.6: 783,953,435 coding-glm-5.1: 610,045,745 38 more models: 448,921,815 glm-5.1: 278,299,390 coding-glm-4.7: 104,493,7752026-08-08 — 7,003,480,570 tokens coding-glm-5.2: 2,466,588,140 coding-glm-5: 2,252,876,180 glm-5.2: 598,526,110 coding-glm-4.7: 524,405,245 coding-glm-4.6: 382,489,870 coding-glm-5.1: 357,971,280 38 more models: 338,318,015 glm-5.1: 82,305,7302026-08-09 — 5,851,429,285 tokens glm-5.2: 3,833,910,675 coding-glm-5.2: 504,978,785 coding-glm-4.7: 498,967,700 coding-glm-4.6: 390,876,415 coding-glm-5: 266,812,400 38 more models: 184,540,250 coding-glm-5.1: 102,790,105 glm-5.1: 68,552,9552026-08-10 — 4,734,214,155 tokens glm-5.2: 2,408,502,230 coding-glm-5.2: 650,831,420 coding-glm-4.7: 459,578,135 coding-glm-5: 449,062,685 coding-glm-4.6: 393,649,350 38 more models: 172,890,005 coding-glm-5.1: 123,934,720 glm-5.1: 75,765,6102026-08-11 — 9,145,389,365 tokens coding-glm-5.2: 3,336,317,865 coding-glm-5: 2,428,312,195 glm-5.2: 1,621,252,820 coding-glm-4.7: 457,408,950 coding-glm-5.1: 448,238,685 coding-glm-4.6: 369,573,025 38 more models: 277,470,630 glm-5.1: 206,815,1952026-08-12 — 28,288,549,480 tokens glm-5.2: 12,373,654,255 coding-glm-5.2: 10,343,943,815 coding-glm-5: 3,735,401,105 coding-glm-5.1: 496,575,430 coding-glm-4.7: 489,608,025 coding-glm-4.6: 414,077,520 38 more models: 363,281,200 glm-5.1: 72,008,1302026-08-13 — 16,976,099,525 tokens coding-glm-5.2: 8,971,879,860 coding-glm-5: 3,790,038,285 glm-5.2: 2,498,606,080 coding-glm-4.7: 608,350,400 coding-glm-5.1: 400,529,045 38 more models: 331,357,345 coding-glm-4.6: 309,100,995 glm-5.1: 66,237,5152026-08-14 — 13,979,625,255 tokens coding-glm-5.2: 5,237,336,550 glm-5.2: 2,869,854,170 coding-glm-5: 2,754,609,075 coding-glm-5.3: 1,314,751,205 coding-glm-4.7: 645,696,585 coding-glm-4.6: 377,225,780 coding-glm-5.1: 347,590,535 38 more models: 302,179,570 glm-5.1: 130,381,7852026-08-15 — 8,927,698,295 tokens coding-glm-5.2: 2,894,599,215 coding-glm-5.3: 2,712,119,810 coding-glm-4.7: 971,683,360 glm-5.2: 736,207,660 coding-glm-5: 646,401,965 coding-glm-4.6: 313,378,320 38 more models: 300,696,970 coding-glm-5.1: 259,488,705 glm-5.1: 93,122,2902026-08-16 — 10,858,625,150 tokens coding-glm-5.3: 6,240,755,890 glm-5.2: 1,277,709,835 coding-glm-5.1: 1,083,042,280 coding-glm-5: 968,461,655 coding-glm-4.7: 550,792,815 coding-glm-4.6: 354,536,760 38 more models: 209,996,940 glm-5.1: 161,730,400 coding-glm-5.2: 11,598,5752026-08-17 — 10,885,006,880 tokens coding-glm-5.3: 4,856,441,140 glm-5.2: 3,542,775,825 coding-glm-5: 846,574,650 coding-glm-4.7: 534,609,745 coding-glm-4.6: 336,002,060 coding-glm-5.1: 265,979,190 38 more models: 254,960,365 glm-5.1: 138,537,970 coding-glm-5.2: 109,125,9352026-08-18 — 12,680,941,755 tokens coding-glm-5.3: 5,588,652,325 glm-5.2: 3,295,448,645 38 more models: 911,376,825 coding-glm-4.7: 801,399,540 coding-glm-5: 755,660,520 coding-glm-5.2: 559,263,310 coding-glm-4.6: 379,723,500 coding-glm-5.1: 322,857,955 glm-5.1: 66,559,1352026-08-19 — 16,957,620,515 tokens coding-glm-5.3: 10,813,449,925 38 more models: 2,374,312,175 glm-5.2: 1,311,421,615 coding-glm-4.7: 774,367,860 coding-glm-5: 669,518,985 coding-glm-5.1: 396,205,620 coding-glm-4.6: 345,056,485 coding-glm-5.2: 171,080,425 glm-5.1: 102,207,4252026-08-20 — 18,954,357,385 tokens coding-glm-5.3: 10,854,316,955 38 more models: 3,550,068,430 glm-5.2: 1,984,767,335 coding-glm-5: 700,718,890 coding-glm-4.7: 601,639,450 coding-glm-5.1: 522,209,770 coding-glm-4.6: 379,215,900 glm-5.1: 250,345,535 coding-glm-5.2: 111,075,120
  • glm-5.2
  • coding-glm-5
  • coding-glm-5.2
  • coding-glm-5.3
  • coding-glm-4.6
  • coding-glm-5.1
  • glm-5.1
  • coding-glm-4.7
  • 38 more models

Which models that traffic went to

  1. GLM 5.228.2%82.7B
  2. Coding GLM 521.7%63.7B
  3. Coding GLM 5.215.4%45.2B
  4. Coding GLM 5.314.4%42.4B
  5. Coding GLM 4.65.7%16.6B
  6. Coding GLM 5.13.8%11.2B
  7. GLM 5.12.9%8.6B
  8. Coding GLM 4.72.8%8.1B
  9. 38 more models5.2%15.1B

Share of 294B tokens. 10 models with traffic report no token counts and cannot be ranked here, including coding-glm-5-turbo-free and glm-4-flash — they are in the request view.

The two views disagree on purpose: a model can take a large share of the calls and a small share of the tokens — many short requests — or the reverse. Which one matters depends on whether your cost is driven by call volume or by prompt length. Measured on AIHubMix over the last 30 days, counting the 66 model IDs listed on this page; traffic routed through upstream-specific IDs that are not in the public catalog is not included.

All 66 Z.AI Models

Open in model list
Z.AI models on AIHubMix with input and output modalities, context length, maximum output, price per million tokens including cache read and cache write rates, and measured throughput and latency.
Modalities
coding-glm-5.3Takes text, returns text.1.05M$0.06$0.22/M
glm-5.3Takes text, returns text.1.05M128K$1.13$3.94/M$0.28/M32 tok/s4.76 s
coding-glm-5.2-freeTakes text, returns text.1MFreeFree/M
glm-5.2Takes text, returns text.1M128K$1.13$3.94/M$0.28/M46 tok/s1.09 s
glm-5.2-fast-previewTakes text, returns text.1M128K$2.25$7.89/M$0.56/M38 tok/s2.00 s
glm-4.6Takes text, returns text.205K131KFreeFree/MFree/M27 tok/s4.65 s
glm-5Takes text, returns text.203KFreeFree/MFree/M82 tok/s1.18 s
glm-5-turboTakes text, returns text.203K$1.20$4.00/M$0.24/M8 tok/s5.94 s
coding-glm-4.6-freeTakes text, returns text.200K128KFreeFree/M
glm-4.7Takes text, returns text.200K128K$0.27$1.10/M$0.05/M35 tok/s15.14 s
glm-5v-turboTakes text, vision, video, returns text.200K128K$0.70$3.10/M$0.17/M31 tok/s2.58 s
glm-5.1Takes text, returns text.200K128K$0.84$3.38/M$0.18/M20 tok/s1.64 s
coding-glm-4.5-airTakes text. Output modality not published.131K$0.01$0.08/M
glm-4.5-airTakes text. Output modality not published.131K98K$0.14$0.84/M91 tok/s1.42 s
glm-4.5Takes text. Output modality not published.131K98K$0.40$1.60/M107 tok/s0.46 s
glm-4.6vTakes text, vision, video, returns text.128K$0.14$0.41/M$0.03/M6 tok/s0.91 s
glm-4.5vTakes text, vision, video, returns text.64K16K$0.27$0.82/M80 tok/s1.25 s
glm-ocrTakes vision, returns text.32K$0.03$0.03/M
embedding-2Takes text. Output modality not published.8K$0.07$0.07/M
embedding-3Takes text. Output modality not published.8K$0.07$0.07/M
coding-glm-4.7-freeTakes text, returns text.FreeFree/M
coding-glm-5-freeTakes text, returns text.FreeFree/M
coding-glm-5-turbo-freeTakes text, returns text.FreeFree/M
coding-glm-5.1-freeTakes text, returns text.FreeFree/M
glm-4.7-flash-freeTakes text, returns text.FreeFree/M
glm-imageTakes text, returns vision.FreeFree/M
Pro/THUDM/GLM-4.1V-9B-Thinking$0.04$0.16/M
THUDM/GLM-4-9B-0414$0.05$0.05/M
THUDM/GLM-Z1-9B-0414$0.05$0.05/M
cc-glm-4.6$0.06$0.22/M
cc-glm-4.7$0.06$0.22/M
cc-glm-5Takes text, returns text.$0.06$0.22/M
cc-glm-5-turboTakes text, returns text.$0.06$0.22/M
cc-glm-5.1Takes text, returns text.$0.06$0.22/M
coding-glm-4.6Takes text, returns text.$0.06$0.22/M$0.01/M
coding-glm-4.7Takes text, returns text.$0.06$0.22/M$0.01/M
coding-glm-5Takes text, returns text.$0.06$0.22/M
coding-glm-5-turboTakes text, returns text.$0.06$0.22/M
coding-glm-5.1Takes text, returns text.$0.06$0.22/M
coding-glm-5.2Takes text, returns text.$0.06$0.22/M
THUDM/GLM-4-32B-0414$0.08$0.08/M
THUDM/GLM-Z1-32B-0414$0.08$0.08/M
glm-4-flash$0.10$0.10/M
THUDM/GLM-4.1V-9B-Thinking$0.10$0.10/M
doubao-1-5-pro-32k-250115$0.11$0.27/M
chatglm_lite$0.29$0.29/M
alicloud-glm-4.7$0.41$1.92/M$0.41/M
alicloud-glm-5$0.56$2.54/M$0.11/M
doubao-1-5-pro-256k-250115$0.68$1.23/M
glm-3-turbo$0.71$0.71/M
chatglm_std$0.71$0.71/M
chatglm_turbo$0.71$0.71/M
glm-4.5-airxTakes text. Output modality not published.$1.10$4.51/M$0.22/M
zai-glm-5-turboTakes , returns text.$1.20$4.00/M$0.24/M
cloudflare-glm-5.2Takes , returns text.$1.40$4.40/M$0.26/M
chatglm_pro$1.43$1.43/M
glm-4v-plus$2.00$2.00/M
glm-zero-preview$2.00$2.00/M
glm-4.5-xTakes text. Output modality not published.$2.20$8.91/M$0.44/M1 tok/s0.59 s
cbs-glm-4.7$2.25$2.75/M
glm-4-plus$8.00$8.00/M
cogview-3-plus$10.00$10.00/M
glm-4$14.20$14.20/M
glm-4v$14.20$14.20/M
code-davinci-edit-001$20.00$20.00/M
cogview-3$35.50$35.50/M

Prices are USD per million tokens; cache read and cache write are the rates for prompt-cache hits and for writing a prompt into the cache. Throughput and latency are measured on AIHubMix — the same figures the model detail page shows — not vendor claims. A dash means the catalog does not publish that field for that model, which is not the same as the model not supporting it.

Z.AI on AIHubMix

Which Z.AI model should I start with?

coding-glm-4.6-free is free on input — the cheapest entry here that declares tool calling, and it carries a 200K context. Move up to cogview-3 when answer quality matters more than cost, or to coding-glm-5.3 for long-form reasoning.

Which of these models reason before answering?

25 of the 66 models here declare a reasoning phase — they work through the problem before producing an answer, which helps on multi-step problems at the cost of extra output tokens. Use the Reasoning filter above the table to see them. The catalog does not record anything further about how they differ, so this page does not sort them into families.

Why are there several entries for the same model?

Because each row is a route you can call, not a model release. Some IDs name an upstream (azure-, alicloud-, cc-), some are the open-weight repository form (THUDM/…), and some differ only in capitalisation, kept so older integrations keep working.

The catalog does not carry a field saying which of those a given row is, so this page does not sort them into buckets it would have to invent. Every row shows that route’s own price, context and speed — compare those directly, and open a model to see the upstreams that serve it.

How is cached input billed?

The Cache read column is the rate for input tokens served from the prompt cache — for example coding-glm-4.6 bills cache hits at 18.33% of the input rate and coding-glm-4.7 bills cache hits at 18.33% of the input rate. Cache write is the surcharge for putting a prompt into the cache in the first place, and only a few upstreams bill it separately. A dash in either column means the catalog carries no cache rate for that model, so plan on paying the full input rate.

Do I need a separate Z.AI account?

No. One AIHubMix key covers every model on this page, and switching between them is a change to the model string — billing, rate limits, and logs stay in one place.

Start calling Z.AI in one line

One key, one endpoint, 859 models across 37 model authors.