Z.AI Review: A Cheap, Capable API With Unpublished Data Terms
Z.AI publishes its token prices to the cent and nothing we could find about zero data retention.
Z.AI is worth testing in a paid pilot, but it's difficult to approve for regulated data. At Layer3Labs, we build and operate AI systems inside other companies, so choosing a China-hosted API is a decision we make for clients, not a thought experiment.
Search for a Z.AI review, and the first page is mostly Reddit threads, Trustpilot ratings, and Product Hunt comments. No established publisher appears in the results.
Those discussions might tell you whether the coding plan felt fast last Tuesday. They won't tell you whether Z.AI retains your prompts or which plan gives you a written retention guarantee.
For the basics, what Z.AI is and who owns it, see our Z.AI explainer.
Our Z.AI Review Verdict on Cost and Capability
Z.AI's published prices are the strongest part of the product. Its pricing page lists GLM-5.3 at $1.40 per million input tokens, $0.26 per million cached input tokens, and $4.40 per million output tokens. At those rates a small team can run a high-volume coding or document workload without a procurement conversation about the bill.
GLM-5.3-Flash sits far below that at $0.075 per million input, $0.015 cached, and $0.25 output. That is cheap enough to put a model inside a background job that runs thousands of times a day. Most small teams never automate that kind of work because the per-call cost scares them off.
Three models cost nothing at all: GLM-4.7-Flash, GLM-4.5-Flash, and GLM-4.6V-Flash. You can build and test a classifier or a document-extraction step on those before you spend a dollar. GLM-5.3-Flash is not one of them. Cheap and free are two different things, and people confuse those two more often than anything else in the GLM lineup.
The flagship itself is recent. Z.AI released GLM-5.3 on 2026-08-17 with a stated focus on coding, long-horizon tasks, and cybersecurity, and it carries 753 billion parameters with a 1-million-token maximum context. A million tokens means a whole repository or a long contract goes into one request instead of being chopped into chunks you then have to reconcile.
One detail that surprises teams on their first invoice: Z.AI's built-in web-search tool bills at $0.01 per call on top of token cost. An agent that searches on every turn adds a line item that token math alone will not predict.
- GLM-5.3 API: $1.40 input, $0.26 cached input, $4.40 output, per million tokens, so a heavy coding workload stays inside a small monthly budget
- GLM-5.3-Flash: $0.075 input, $0.015 cached, $0.25 output, cheap enough for high-frequency background jobs
- Genuinely free ($0) models: GLM-4.7-Flash, GLM-4.5-Flash, GLM-4.6V-Flash, so start prototyping there
- GLM-5.3: released 2026-08-17, 753B parameters, 1M-token maximum context, so a full repository fits in one request
- Built-in web search: $0.01 per call on top of tokens, a line item agent builders forget

First Month Free
Get one month of Starlink free when you sign up through this link. Fast, reliable internet at home and on the go.
What We Verified About Z.AI Ourselves
We checked Z.AI's own pages on 2026-09-02 and split what we found into confirmed and unconfirmed. The confirmed half is short. The unconfirmed half decides a vendor review.
Four things held up against first-party documentation. The per-million token prices above come straight from the pricing page. The GLM Coding Plan credit tiers are published in the devpack overview. Lite gets 2,000 credits per 5 hours and 10,000 per week, Pro gets 12,000 and 60,000, and Max gets 28,000 and 140,000. Off-peak usage is discounted 50%. The same page states the plan serves GLM-5.3 and GLM-5.3-Flash only, and silently reroutes requests for GLM-5.2 and GLM-5.1 to GLM-5.3, and requests for GLM-4.7 to GLM-5.3-Flash. GLM-5.3 weights are published on Hugging Face under the zai-org account.
Two of those carry a consequence worth naming. The rerouting means a coding-plan subscription cannot pin an older model. Any evaluation you ran on GLM-5.2 through the plan is no longer the model you are buying. The plan's dual quota window, 5 hours and weekly at once, is also a real throughput ceiling. Z.AI's documentation pages we checked publish no per-tier request or token rate limit at all.
Three things we could not confirm. Every GLM-5.3 benchmark figure in circulation is Z.AI's own claim, not a reproduced result. The reported licence carve-out for large companies is a third-party report, not a line we can quote from the model card. And the data-retention question has no published answer at all, which the next section takes on its own.
For a read on GLM-5.3 that Z.AI did not produce, Artificial Analysis runs its own evaluations across models. Check there before you take a lab's own chart at face value.
- Confirmed: the per-million token prices, from Z.AI's own pricing page
- Confirmed: Coding Plan credit tiers, at 2,000 credits per 5 hours and 10,000 per week on Lite, 12,000 and 60,000 on Pro, 28,000 and 140,000 on Max, plus a 50% off-peak discount
- Confirmed: the plan serves only GLM-5.3 and GLM-5.3-Flash, and reroutes GLM-5.2, GLM-5.1 and GLM-4.7 requests automatically
- Confirmed: GLM-5.3 weights are published on Hugging Face under zai-org
- Not confirmed: any benchmark score, the licence revenue carve-out, or any data-retention or deletion commitment
Does Z.AI Train on Your Data?
Z.AI's public documentation does not answer whether prompts sent to its hosted API are used to train future GLM models. We read the pricing page, the quick-start guide and the devpack overview on 2026-09-02. Those pages document tokens, endpoints, model ids, and quotas. They do not document what happens to your text after the response comes back.
The same gap covers zero data retention. We found no published zero-data-retention option for Z.AI, and no statement of which plan or contract tier one would attach to. That is an absence in the public documentation, not a policy we can quote against you or for you.
Treat an absence as a question you have not asked yet. Before any regulated or client-confidential text goes near the hosted API, get these six things in writing from Z.AI and attach them to the order form.
- A written statement of whether API inputs and outputs are used to train or fine-tune any model, with the answer applying to your account specifically
- The retention period for prompts, completions, and logs, stated in days, plus what triggers deletion
- Whether a zero-data-retention configuration exists, what it costs, and which plan carries it
- Which legal entity contracts with you and which jurisdiction governs the agreement
- Where inference runs and where logs are stored, named by region rather than by company
- Which subprocessors touch the data, and how you are told when that list changes
What Zhipu Ownership Means for a US Vendor Review
Zhipu ownership changes who your security questionnaire is aimed at, not whether GLM-5.3 does the work. Z.AI is the international brand of the Chinese AI lab Zhipu, and a US buyer's review process reacts to that fact long before anyone opens a benchmark chart.
Across the rollouts we run for small and mid-market teams, the failure mode we hit most often is not a model that underperforms. It is a security lead asking for a retention clause and finding no published page that contains one. The pilot then sits frozen for weeks while nobody escalates. A cheap model that stalls in review for a quarter is more expensive than a model that costs four times as much and clears in a week.
Two practical moves shorten that. Ask for the six items above before the pilot rather than after, because a security team that gets them on day one will usually clear a non-regulated pilot quickly. Then split the review into two questions. Where the data goes is a contract question you can negotiate with Z.AI. Who built the model is a procurement-policy question you cannot.
Country of origin is also a live policy variable rather than a fixed one. US federal and state procurement rules on Chinese-owned technology change, so the answer your legal team gave last year is not a standing approval. Our comparison of what each major lab publishes on compliance is at AI model compliance comparison, and the broader risk picture is at Chinese AI models and security risks.
- Run the contract question (retention, training, deletion, jurisdiction) and the policy question (Chinese-owned vendor allowed at all) as two separate tracks
- Get written retention terms before the pilot starts, so a security review does not freeze work that is already underway
- Re-check procurement policy at renewal rather than treating a past approval as permanent
- If the policy track fails, self-hosted open weights keeps GLM-5.3 and removes the hosted-service question entirely
Where Z.AI Falls Short
The GLM-5.3 licence is the biggest change since GLM-5.2, and it moves away from open. GLM-5.2 shipped under MIT. GLM-5.3 ships under a bespoke licence named "glm-5.3" on its Hugging Face model card, while GLM-5.3-Flash shipped under plain MIT seven days apart. If your legal team approved GLM-5.2 on the strength of MIT, that approval does not carry forward. Someone has to read the new LICENSE file before you host GLM-5.3 commercially.
The New Stack has reported that the new licence requires companies above a large aggregate-revenue threshold to pass a Z.AI security review before hosting the model commercially. That threshold figure is not on the Hugging Face model card, so we are not stating a dollar number as fact. Read the LICENSE file on the GLM-5.3 model card yourself and have counsel confirm whether your revenue puts you inside it.
Every GLM-5.3 performance number is Z.AI's own claim. Z.AI reports a 50% gain over GLM-5.2 on its internal Z.AI Code Bench, and a Terminal-Bench 3.0 score of 28.3 up from 4.6. It also reports a CyberGym result of 84.5% against 83.8% for Anthropic's Mythos 5 and 83.6% for OpenAI's GPT-5.6 Sol. Z.AI also reports 2,436 confirmed vulnerabilities found across 269 open-source projects, 1,097 of them critical or high severity. None of that has been independently reproduced, and a self-scored comparison against two competitors is a marketing artefact until someone else runs it.
There is no verified official Z.AI mobile or desktop client. As of 2026-09-02, checked against z.ai and chat.z.ai, the web app is the product and its page links to no iOS, Android, or desktop app. Two lookalikes do sit in the stores. "Z.ai : App Directions" is an unofficial guide app that states it is not affiliated with or endorsed by Zhipu AI, and "Z-AI Multi Translator" is a different product entirely. Check z.ai before you install anything that carries the name.
Two prices are not published. Z.AI's coding-plan documentation publishes only "Starting at just 18 USD per month" and no per-tier dollar price for Pro or Max. Third-party sites disagree with each other on what those tiers cost. Confirm the current figure on Z.AI's own subscribe page rather than on a comparison blog.
- GLM-5.3 is not MIT licensed, so a prior GLM-5.2 legal approval has to be redone
- A reported revenue-threshold security review sits in the LICENSE file, and the threshold is not on the model card
- All headline benchmark scores are Z.AI-reported, including the comparisons against Mythos 5 and GPT-5.6 Sol
- No official native app exists as of 2026-09-02, and two unaffiliated lookalikes use the name in app stores
- Coding Plan Pro and Max dollar prices are unpublished, so budget from the subscribe page and not a third-party chart
Which Z.AI Plan Fits Which Job
Three ways to buy Z.AI exist, and they suit different jobs. The table compares them on the facts Z.AI publishes.
| GLM-5.3 API | GLM-5.3-Flash API | GLM Coding Plan | |
|---|---|---|---|
| Published price | $1.40 in / $4.40 out per million tokens | $0.075 in / $0.25 out per million tokens | "Starting at just 18 USD per month", with Pro and Max prices unpublished |
| Models served | Any model id you call | Any model id you call | GLM-5.3 and GLM-5.3-Flash only, with older ids rerouted |
| Main limit | No per-tier rate limit published | No per-tier rate limit published | Credits capped in a 5-hour and a weekly window at once |
| Best fit | Long-context repository and contract work | High-frequency background jobs | Daily coding inside Claude Code and similar clients |
| Verdict | Buy for the 1M-token window | Buy for volume, but do not mistake it for free | Buy only if you can live with GLM-5.3 as the sole model |
For a full breakdown of the subscription mechanics, including credits and the off-peak discount, see the GLM Coding Plan guide. For the token-by-token cost picture across the whole lineup, see Z.AI pricing.
Who Should Not Buy Z.AI
Three groups should not put the Z.AI hosted API into production today, and each has a different alternative.
Teams handling regulated data are the clearest case. If you carry protected health information, criminal-justice data, export-controlled material, or client data under a confidentiality obligation, do not send it to a hosted API whose retention and training terms are not published. Run GLM-5.3 from the open weights inside your own cloud region instead, or use a provider that publishes its data-processing terms in full.
Teams that need a fixed model for reproducibility are the second. If your evaluation harness, your audit trail, or a client contract depends on the exact model version staying constant, the Coding Plan's automatic rerouting of GLM-5.2 and GLM-5.1 requests to GLM-5.3 will break that. Use the direct API with an explicit model id, or self-host a pinned checkpoint.
Teams that sell to buyers with a stated policy against Chinese-owned vendors are the third. The engineering merits do not matter if a customer security questionnaire disqualifies you at the supplier stage. Price a Western-hosted model into the deal and treat the cost difference as sales expense.
- Regulated-data teams: self-host the open weights or pick a provider that publishes data-processing terms
- Reproducibility-bound teams: call the direct API with an explicit model id, or pin a self-hosted checkpoint
- Teams whose customers ban Chinese-owned vendors: budget for a Western-hosted model instead of arguing the point
- Anyone expecting a native mobile app: use the web app at chat.z.ai and install nothing from an app store
What Would Change This Z.AI Review
Four specific changes would move Z.AI from a pilot recommendation to a production one. Each is something Z.AI controls and could publish tomorrow.
A published data-processing page stating retention periods, a training opt-out, and deletion timelines would remove the largest objection on this page. A named zero-data-retention tier with a price attached would remove it entirely for most business buyers.
Independent benchmark reproduction would fix the second problem. If Artificial Analysis or another evaluator confirms the Terminal-Bench and CyberGym results, the capability claim stops resting on Z.AI's own scoring.
The other two are smaller. Publishing Pro and Max dollar prices would let a buyer budget without guessing, and letting Coding Plan subscribers pin a model id would make the plan usable for teams with an audit trail. Absent those, our position holds: pilot it on work you would not mind seeing in public, and keep regulated workloads on infrastructure you control.
A Z.AI review that matters to you is the one you run on your own repositories. Start a two-week pilot at the published GLM-5.3 rate on non-sensitive code, and ask for the retention clause in writing before anything regulated goes near it.
- If Z.AI publishes retention, training opt-out, and deletion terms, the main objection on this review goes away
- If a named zero-data-retention tier ships with a price, regulated pilots become possible on the hosted API
- If another evaluator reproduces the Terminal-Bench 3.0 and CyberGym scores, the capability claim stands on its own
- If Z.AI publishes Pro and Max prices and allows model pinning, the Coding Plan becomes budgetable and auditable
Frequently Asked Questions
- Z.AI does not publish per-tier request or token rate limits on the documentation pages we checked on 2026-09-02, so throughput on the direct API is an unknown you should test rather than assume. The GLM Coding Plan is more predictable because its ceilings are published: Lite gets 2,000 credits per 5 hours and 10,000 per week, Pro gets 12,000 and 60,000, and Max gets 28,000 and 140,000, with both windows applying at the same time. Build retry-with-backoff logic and pilot at your expected load before you depend on it.
- No. Three GLM models cost $0 per token on Z.AI's published price list: GLM-4.7-Flash, GLM-4.5-Flash, and GLM-4.6V-Flash. GLM-5.3-Flash is often mistaken for one of them and is not free, at $0.075 per million input tokens and $0.25 per million output tokens. The flagship GLM-5.3 is paid, and the GLM Coding Plan is a subscription starting at 18 USD per month.
- Z.AI's published API rates are $1.40 per million input tokens, $0.26 cached input, and $4.40 per million output tokens for GLM-5.3, and $0.075, $0.015, and $0.25 for GLM-5.3-Flash. The built-in web-search tool adds $0.01 per call on top of token cost. The GLM Coding Plan documentation publishes only a starting price of 18 USD per month and no per-tier dollar figure for Pro or Max, so confirm those on Z.AI's own subscribe page at https://z.ai/subscribe.
- Z.AI's public documentation does not say. We read the pricing, quick-start, and coding-plan pages on 2026-09-02 and found no statement of how long prompts and completions are retained, whether they can be used for training, or how deletion works. That is an unanswered question rather than a stated policy either way. Ask Z.AI for retention and training terms in writing before sending anything confidential, and self-host the open GLM-5.3 weights if you cannot get them.
- Z.AI does not publish an answer on the documentation pages we checked on 2026-09-02. There is no public training opt-out statement covering hosted API inputs and outputs. Request a written statement that applies to your specific account, along with a retention period in days, before a pilot handles anything you would not publish.
- We found no published zero-data-retention option for Z.AI as of 2026-09-02, and no statement of which plan or contract tier would carry one. If zero data retention is a hard requirement for your workload, ask Z.AI to confirm in writing whether it can be configured and at what price. The alternative that needs no vendor answer is running the published GLM-5.3 weights from Hugging Face inside your own cloud region.
- They win on different things, and the comparison is not settled by a benchmark chart. Z.AI reports a CyberGym score of 84.5% against 83.8% for Anthropic's Mythos 5, but that number is Z.AI's own and has not been independently reproduced. Price and terms are verifiable. GLM-5.3 costs $1.40 per million input tokens and $4.40 output, and Z.AI publishes no data-retention policy we could find, while Anthropic publishes its data-handling terms openly. For non-sensitive high-volume coding, Z.AI is the cheaper option. For regulated work, published terms matter more than a score.
Deciding Whether Z.AI Clears Your Vendor Review?
Book a free 30-minute AI workflow audit with Layer3Labs. We will map which of your workloads can run on the Z.AI hosted API today, which need self-hosted GLM-5.3 weights, and exactly what to ask Z.AI for in writing before a pilot starts.
Book Now