The short answer
Yes, GLM is in the Whizi catalogue. Z.ai rows are priced across the GLM 4 and GLM 5 generations, one row opens on the Pro plan at 2 credits per message, and every other GLM row needs Powerhouse.
| Model | Context window | Credits per message | Plan required |
|---|---|---|---|
| GLM 4.6 | 205K | 1 | Powerhouse |
| GLM 4.7 | 205K | 2 | Powerhouse |
| GLM 5 | 205K | 2 | Pro |
| GLM 5.1 | not in the price index | 3 | Powerhouse |
| GLM 5.2 | 1M | 3 | Powerhouse |
| GLM 5.3 | 1M | 5 | Powerhouse |
| GLM 5 Turbo | not in the price index | 5 | Powerhouse |
| GLM 5V Turbo | not in the price index | 5 | Powerhouse |
A GLM identifier not listed here is still charged a rung and still requires Powerhouse.
Free accounts are not model-gated. The chat path returns for the free tier before the model gate runs, so a prospect can try any text model in the catalogue, GLM included, inside a lifetime allowance of 7 messages that never resets.
The one row on Pro is not the cheapest row
Pro resolves to exactly one Z.ai identifier, GLM 5, at 2 credits per message. The curation rule the Pro list is built on is set out in the model reference, and on this family it lands on a single row, which is the whole reason the rest of the generation sits a tier up.
That single row is not the cheap one. The cheapest GLM row in credits charges 1 credit per message and requires Powerhouse, while the Pro row charges twice that. Credit cost and plan tier are two separate decisions in Whizi, and neither one predicts the other.
Starter does not include any GLM model. Starter reaches exactly four picker entries: Auto, the house model, one OpenAI model and one Google model. Every one of them is a 1x or 2x row, which is why the Starter credit allowance and the Starter message allowance are the same number.
Powerhouse covers the rest, by name in one case and by fallback in every other. The frontier reasoning row in the GLM 5 generation is named explicitly among the models held back for Powerhouse, alongside a Grok row, the Qwen thinking row, the Kimi thinking row and others. Every remaining GLM row lands on the top tier because any identifier that is not in the Starter set and not in the explicit Pro list resolves there.
Gating is enforced on the server, not in the interface. Selecting a model above your tier returns HTTP 403 with the code tier_upgrade_required and the message "Upgrade your plan to use this model." If you hit that on a GLM row, see why a model shows as unavailable.
The GLM 4 pair, where the rung and the running cost disagree
In the GLM 4 pair the two columns disagree. The row that is cheaper per standard answer carries the higher rung, and its sibling carries the lower one. A rung is measured on a reference turn of 3,000 input and 800 output tokens, the index figure on a 1,000 plus 500 standard answer, and two rows whose input and output rates are weighted differently can swap places between those two mixes. The rung method is in the credits reference.
| Model | Cost per standard answer | Cost per thousand answers | Credits per message |
|---|---|---|---|
| GLM 4.7 | $0.001275 | $1.275 | 2 |
| GLM 4.6 | $0.0015 | $1.50 | 1 |
| GLM 5 | $0.00156 | $1.56 | 2 |
| GLM 5.2 | $0.002484 | $2.484 | 3 |
| GLM 5.3 | $0.0036 | $3.60 | 5 |
The Whizi cost index prices 100 OpenRouter rows on one identical unit so providers can be compared: a standard answer of 1,000 input tokens plus 500 output tokens. Source prices were fetched from OpenRouter on 2026-08-20, across 29 providers.
For scale: the median priced row in the index is $0.00185 per standard answer, the cheapest is $0.000053 and the dearest is $0.105, a spread of roughly 2000x. Three of the five priced GLM rows sit below that median, and two sit above it.
Rows marked "not in the price index" in the table at the top carry a credit charge in the product but are not among the 100 rows the cost index prices, so there is no per answer figure to publish for them.
What a GLM message spends from your allowance
A credit is one message on the house model, and credit cost is fixed per message by the row you picked rather than by length. The method is in the credits reference. What it means on this family: the GLM rows listed at the top of this page span a single band from 1 credit to 5, so the dearest of them costs exactly five of the cheapest, and the whole family sits inside the bottom five rungs of an eighteen rung ladder.
| Plan | Price | Credits per month | Messages on the Pro GLM row (2 credits) | Messages on the dearest GLM row (5 credits) |
|---|---|---|---|---|
| Starter | $15.99/month, or $10.99/mo billed annually at $131.88 | 400 | no GLM row on this plan | no GLM row on this plan |
| Pro | $29.99/month, or $19.99/mo billed annually at $239.88 | 2,000 | 1,000 | not on this plan |
| Powerhouse | $49.99/month, or $34.99/mo billed annually at $419.88 | 8,000 | 4,000 | 1,600 |
On Powerhouse the 1 credit GLM row buys 8,000 messages from an 8,000 credit allowance, and the 3 credit rows buy 2,666. Credits do not roll over: usage is summed against a period key, so the arrival of a new period is itself the reset.
A GLM turn that will not fit the remaining balance is rejected outright instead of being charged in part. That is deliberate, because charging 3 credits for a 5 credit row would make one model cost different amounts depending on when it was sent.
Credit pricing is gated per platform by an allowlist whose production value today is the web client. Off that allowlist a GLM row shows a flat 1x in the picker, and every turn costs one against the older message allowance rather than a rung against the credit allowance.
Why the 205K rows and the 1M rows are sized the same
A GLM row with a 1M window and a GLM row with a 205K window get identical room inside Whizi. Both sit above the 93,000 token threshold, and every catalogue model above that threshold gets one flat budget of 40,000 input tokens and 20,000 output tokens per turn, which is true of every GLM row the cost index prices. Models below the threshold get a smaller proportional budget instead, described in when a conversation is too long.
So moving up to the larger window does not buy a longer prompt. What it buys is headroom on the provider side, since OpenRouter counts input plus requested maximum output against the window rather than input alone.
Moving a thread onto GLM and off again
Auto will not take you into GLM. Auto has no rung of its own, and its ladder is six fixed routes across five model rows: the house model for short questions, a Google fast row for rewrites and for attachments, an OpenAI row for long form, an Anthropic row for code, and an OpenAI reasoning row for multi step reasoning. No GLM row is on that ladder, so a turn only lands on GLM when you pick it yourself.
Every model row served to the web client carries its cost signal, so the credit price of a GLM row is visible before you send rather than after.
A GLM thread you opened months ago still opens today, because a conversation keeps the identifier it was created with even after the picker has moved on to newer rows.
Whichever way you move, the charge follows the model that actually answered the turn, not the model the conversation started on.
Where GLM does not appear
Media generation does not use GLM. The image generation section is four fixed rows, video generation is a single row, and audio generation is a single row, all of them house or partner generators, so no GLM row is involved in generating an image, a video or music.
There is no per model vision flag to quote. Whizi does not track which models can read images. Attachments are handled per request rather than per model, and on the web client the text of a PDF, Word or spreadsheet file is extracted in the browser before anything is sent. See supported file types for what that means in practice.
Web search is a request level toggle, not a model capability. Search runs in one of three modes, and turning it on affects the request rather than selecting a different model, so there is no published list of GLM rows that can or cannot browse.
Agent personas are not GLM. The four agent personas are a Pro feature rather than a model tier, and every one of them runs on the same base model, so choosing a persona never routes a turn to Z.ai.
- Whizi prices GLM models from Z.ai across the GLM 4 and GLM 5 generations
- GLM 5 is the only GLM row reachable on Pro, at 2 credits per message
- Every other GLM row requires Powerhouse
- Starter includes no GLM model at all
- Free accounts can try any text model, GLM included, within a lifetime allowance of 7 messages
- Credit cost is fixed per message and does not vary with length
- Every GLM row the cost index prices runs on the flat 40,000 input and 20,000 output token budget
- Auto never routes a turn to a GLM row
Frequently asked questions
Is GLM included in Whizi?
Yes. GLM rows sit in the All models section of the picker, under the Z.ai vendor logo rather than the generic fallback icon that most smaller labs get. They are not in the Recommended section, which is five fixed rows and carries no GLM entry, so browse All models rather than the top of the picker to find them.
Which Whizi plan do I need to use GLM?
Pro reaches one GLM row, GLM 5. Everything else in the family is Powerhouse. If you send on a row your tier does not include, the worker answers HTTP 403 with a JSON body shaped {"error": {"code": "tier_upgrade_required", "message": "Upgrade your plan to use this model."}}, and the website re-wraps that into its older code and message pair before the chat page reads it, so the sentence you see on screen is the backend string verbatim rather than a client rewrite of it.
How many credits does a GLM message cost?
Between 1 and 5 credits depending on the row. The GLM 4 generation is 1 or 2 credits, GLM 5 is 2, two further rows in the GLM 5 generation sit at 3 credits, and the dearest GLM rows are 5.
Can I use GLM on the free tier?
Yes, within a hard cap. Free accounts get a credit allowance of 0 and a message allowance of 0, but they do get a lifetime allowance of 7 messages that never resets, and the free path deliberately skips the model gate so a prospect can try any text model inside that cap. When it is used up the app returns the code free_limit_reached and the message "Your free messages are used up. Start a subscription to keep chatting."
What is the context window on GLM in Whizi?
Before any budget is checked, the measured input is scaled by 1.8x, because the token estimator can undercount a real tokenizer by as much as 1.66x on content such as JSON. A GLM prompt heavy with structured data is therefore sized larger than it literally measures. The priced GLM rows carry 205K or 1M token windows depending on the row, and neither number is what one turn sends: the per turn budget is in the section on why the two window sizes come out the same.