Z.AI Pricing: What GLM Models Actually Cost
A buyer-focused breakdown of Z.AI's GLM API pricing, Coding Plan tiers, free access, enterprise options, and how the numbers compare to Claude, OpenAI, and DeepSeek.
Z.AI is the international brand of Chinese AI lab Zhipu AI. Its GLM models have emerged as a frequently cited, lower-cost alternative to Claude and GPT. But comparing costs isn't straightforward: Z.AI's pricing page combines per-token API rates, a subscription-based Coding Plan, and a limited free tier.
This page uses figures directly from Z.AI's pricing documentation, verified on 2026-07-21. Every published price is sourced, and where public pricing isn't available, such as Coding Plan Pro and Max tiers or enterprise contracts, we make that clear and direct you to the live page rather than filling in the gaps.
Below, you'll find API rates for every available GLM tier, details on what the free tier includes, a high-level overview of the GLM Coding Plan, and what enterprise buyers can expect. The full subscription breakdown is available on our GLM Coding Plan page. We also compare Z.AI's pricing with Anthropic Claude, OpenAI, and DeepSeek.
Model prices change frequently, so always check the vendor's live pricing page before signing a contract or building a cost model.
Z.AI Pricing at a Glance
Z.AI sells access to its GLM models three ways. Understanding which one you need is the first cost decision, because the price-per-token math only applies to one of them.
- Pay-as-you-go API: per-1M-token pricing for input, output, and cached input. This is the standard developer path and the mode most cost comparisons refer to.
- GLM Coding Plan: a flat monthly subscription with prompt quotas, built around the GLM Coding tier. Starts at $18/month for the Lite plan. Suited to individual developers using the model inside coding IDEs.
- Free tier: GLM-4.5-Flash is priced at zero for input, output, and cached input on Z.AI's public API pricing page — a genuine free model, not a trial credit.
- Enterprise / custom: volume commits and dedicated deployments are not publicly listed — contact Z.AI sales.

First Month Free
Get one month of Starlink free when you sign up through this link. Fast, reliable internet at home and on the go.
API Pricing by GLM Model
These are Z.AI's published per-1M-token API rates as listed on the official pricing page on 2026-09-02. All values are USD. Cached-input pricing applies when a prefix has been seen recently and Z.AI can reuse the KV cache.
- GLM-5.3 (current flagship, released 2026-08-17): $1.40 per 1M input, $4.40 per 1M output, $0.26 per 1M cached input.
- GLM-5.3-Flash: $0.075 per 1M input, $0.25 per 1M output, $0.015 per 1M cached input. Cheap, but not free.
- GLM-5.2 and GLM-5.1: $1.40 per 1M input, $4.40 per 1M output, $0.26 per 1M cached input, the same rates as GLM-5.3.
- GLM-5: $1.00 per 1M input, $3.20 per 1M output, $0.20 per 1M cached input.
- GLM-4.7: $0.60 per 1M input, $2.20 per 1M output, $0.11 per 1M cached input.
- GLM-4.6: $0.60 per 1M input tokens, $2.20 per 1M output tokens, $0.11 per 1M cached-input tokens.
- GLM-4.5: $0.60 per 1M input, $2.20 per 1M output, $0.11 per 1M cached input. Same public rates as GLM-4.6.
- GLM-4.5-Air (lightweight): $0.20 per 1M input, $1.10 per 1M output, $0.03 per 1M cached input.
- GLM-4.5-Flash: free — $0 for input, output, and cached input on the public pricing page.
- GLM-4.5V (vision): $0.60 per 1M input, $1.80 per 1M output, $0.11 per 1M cached input.
- GLM-4-32B-0414-128K: $0.10 per 1M input, $0.10 per 1M output.
- GLM-4.7-Flash and GLM-4.6V-Flash are also listed at $0, alongside GLM-4.5-Flash.
- The GLM-5 family now appears on the same public per-token table as the 4.x line. That was not true when this page was first written, so an older cost model built on "GLM-5 is Coding Plan only" needs redoing against the rates above.
GLM Coding Plan (Subscription)
The GLM Coding Plan is a flat monthly subscription that gives coding-tool users a prompt quota rather than per-token billing. It is aimed at developers using GLM inside IDEs and CLI coding assistants.
Z.AI's docs list three tiers — Lite, Pro, and Max — each with its own 5-hour and weekly prompt cap. The Lite plan starts at $18/month. Pro and Max monthly prices are not published on the docs overview page verified here; check z.ai/subscribe for the current tier prices.
For the full breakdown of quotas, model access, and how the Coding Plan compares to per-token API billing, see our dedicated /guides/glm-coding-plan-explained page.
- Lite: ~80 prompts per 5 hours, ~400 per week. Starts at $18/month.
- Pro: ~400 prompts per 5 hours, ~2,000 per week. Monthly price not publicly listed on the docs overview page — check z.ai/subscribe.
- Max: ~1,600 prompts per 5 hours, ~8,000 per week. Monthly price not publicly listed on the docs overview page — check z.ai/subscribe.
- Z.AI states the monthly quota is roughly equivalent to 15–30x the subscription fee in API-billing terms, so the plan is designed to be cheaper than per-token API for steady coding workloads.
- Advanced models (GLM-5.2, GLM-5-Turbo) consume quota at higher multiples during peak Beijing hours.
Free-Tier Limits
Three Z.AI models are free on the public API: GLM-4.7-Flash, GLM-4.5-Flash, and GLM-4.6V-Flash. GLM-5.3-Flash is not one of them, which is the single most common mix-up on Z.AI pricing, and the flagship GLM-5.3 is not free either. Z.AI gives free access by pricing whole models at zero rather than by handing out trial credits. What it does not publish is the rate limit that applies to them, so the list below separates what Z.AI states from what only outside trackers report.
- GLM-4.7-Flash, GLM-4.5-Flash, and GLM-4.6V-Flash are each listed at $0 for input, output, and cached input on the Z.AI pricing page as of 2026-09-02.
- GLM-5.3-Flash costs $0.075 per 1M input and $0.25 per 1M output. It is the cheapest current-generation model, not a free one.
- The free models are older generations. If you want GLM-5.3 quality at no marginal cost, the only route is self-hosting the open weights, and running GLM locally covers the hardware that takes.
- Z.AI's public pricing and docs pages do not publish a numeric rate limit for free-tier access, whether requests per minute, tokens per minute, or a daily cap. The page states only that standard API rate limits and Terms of Service apply, with no number attached. Check your own quota in the Z.AI console.
- Third-party API trackers report free-tier throughput of roughly 1 request per second, about 60 a minute, with a daily cap near 1,000 requests. That figure comes from outside observation rather than Z.AI's documentation, so treat it as directional and not contractual.
- Z.AI does not commit to keeping these models free. They are priced at zero today, and that can change without notice.
- For the paid models, Z.AI's docs mention a limited-time free benefit on cached-input storage. Storing cached prefixes is free during that promotional window, but the cached-input read price above still applies.
- There is no bundled free-credit grant on the public pricing page. If you need trial credits, higher free-tier throughput, or a written rate-limit commitment for a proof of concept, ask Z.AI directly through the subscribe page.
- For how to reach these models rather than what they cost, the Z.AI API guide covers keys, base URLs and endpoints, and our Z.AI review covers the data-retention terms a buyer has to ask for in writing. If you are looking for a phone or desktop client rather than an API, which Z.AI apps exist has the dated answer.
Enterprise / Custom Pricing
Z.AI does not publish enterprise-tier pricing. Volume discounts, dedicated deployments, private endpoints, data-residency options, and SLA-backed availability are all negotiated case-by-case.
- Volume discounts: not publicly listed — contact Z.AI sales for token-volume commitments.
- Dedicated / private deployment: not publicly listed — enterprise-only. Buyers focused on data isolation often self-host the open-weight GLM releases instead (see the self-hosted section below).
- Data-processing agreements and regional hosting: needed for GDPR-scope data. Not documented on the public pricing page — request through sales.
- Support SLAs: not publicly listed. Standard API access has documented rate limits but no uptime SLA on the public page.
How Z.AI Pricing Compares
The chart below puts Z.AI's flagship per-token API pricing next to the equivalent tier from Anthropic, OpenAI, and DeepSeek. All figures are verified against each vendor's public pricing page on 2026-07-21 and are USD per 1M tokens.
- Z.AI GLM-4.6: $0.60 input / $2.20 output. Source: docs.z.ai/guides/overview/pricing.
- Z.AI GLM-4.5-Air: $0.20 input / $1.10 output. Source: docs.z.ai/guides/overview/pricing.
- Anthropic Claude Sonnet 4.5: $3.00 input / $15.00 output. Source: anthropic.com/pricing. GLM-4.6 lists roughly 5x lower on input, ~7x lower on output.
- Anthropic Claude Opus 4.5: $15.00 input / $75.00 output. Source: anthropic.com/pricing. Opus targets a different tier — this is a ceiling reference, not a like-for-like swap.
- OpenAI GPT-5 (standard tier): pricing shifts often — verify at openai.com/api/pricing before modeling. GLM-4.6 sits well below the GPT-5 standard tier on the pricing page verified today.
- DeepSeek V3: $0.27 input / $1.10 output at standard rates. Source: deepseek.com. DeepSeek undercuts GLM-4.6 on input; GLM-4.5-Air is closer to DeepSeek territory.
Self-Hosted Cost (Open Weights)
Several GLM releases (including GLM-4.5, GLM-4.5-Air, and GLM-5.2) ship as open weights under permissive licenses. Self-hosting means no per-token API cost, but you pay for GPUs, engineers, and operational overhead.
- Per-token cost: $0. There is no license fee for the open-weight GLM releases.
- Hardware class: flagship GLM models require a multi-GPU node — typically 8x H100/H200-class accelerators for full-precision serving. Quantized deployments can run on smaller footprints with quality trade-offs.
- GPU capital or rental: exact list prices are not published across cloud vendors — check the live pricing on your target cloud (AWS, GCP, Azure, Lambda, CoreWeave) before modeling. A single H100 rental typically runs several dollars per hour on-demand at public cloud vendors.
- Operational cost: engineering time to stand up vLLM/SGLang, monitoring, autoscaling, model updates, and security patching. These are the costs that surprise most teams.
- Break-even: self-hosting only wins on cost when steady-state token volume is very high, or when the driver is data isolation rather than pure economics.
Total Cost of Ownership
The sticker price on a pricing page is only one line of the real cost of running GLM in production. A realistic TCO model has to include the operating cost around the token bill.
- Token spend: per-1M rates above, multiplied by realistic input and output volumes for your workload — not the marketing example.
- Prompt caching: Z.AI's cached-input rate is roughly 5x cheaper than standard input for GLM-4.6. Structured prompts that reuse a long system prefix can meaningfully lower the effective per-request cost.
- Guardrails and moderation: teams routing production traffic to any Chinese-lab model typically add a moderation layer, log review, and prompt-filtering — this is engineering time and, often, a second API bill.
- Data-processing and compliance: GDPR-scope traffic requires a DPA or self-hosted deployment. Compliance review is a real line item, not a rounding error.
- Fallback provider: production stacks usually keep a second provider on standby for outages or rate-limit events. Budget for the second provider's minimums.
- Engineering time to swap: switching from Claude or GPT to GLM is not free — evaluation, prompt-tuning, and eval-set rebuilds all cost engineering hours.
When Z.AI's Pricing Wins
Z.AI's pricing is genuinely competitive, but not for every workload. The scenarios below are where GLM's per-token math translates into real savings.
- High-volume batch coding: overnight refactors, vulnerability scans, or agentic code review where per-token cost dominates and latency is not the constraint.
- Long-context document processing: legal review, contract analysis, or research summarization where GLM's long-context tiers avoid chunking overhead.
- Individual developers on the Coding Plan: at $18/month for Lite, the effective per-prompt cost is well below equivalent Claude Sonnet or GPT-5 usage inside coding IDEs.
- Free-tier experimentation: GLM-4.5-Flash at $0 is a real free model for prototyping, evaluation harnesses, and internal tooling.
- Cost-driven agent workloads: multi-step agents that burn tokens quickly benefit most from the lower per-1M rates.
- Scenarios where Z.AI does NOT win: real-time customer-facing chat where latency and moderation matter more than token price; regulated workloads that require a US-based provider with an existing BAA/DPA; teams already deeply invested in Claude or OpenAI tooling where switching cost exceeds the savings.
What you need to run GLM-5.2 yourself
GLM-5.2 is a frontier-scale Mixture-of-Experts model, so "running it yourself" is a real infrastructure decision — not something a single laptop or gaming GPU can do. Match the path below to how seriously you need to self-host. For most teams the API or rented GPUs are the right answer; buying hardware only pays off at steady, high volume or when your data can never leave your walls.
| Path | What it is | Best for | Get started |
|---|---|---|---|
| Call the hosted API | Use GLM-5.2 as a pay-per-token API — zero hardware | Most teams; evaluating before committing | OpenRouter |
| Rent GPUs by the hour | Spin up H100 / A100 nodes on demand, tear them down after | Self-hosting without capital outlay; bursty workloads | RunPod |
| Local on unified memory | A single workstation with enough unified memory to hold a 4-bit quant | One powerful on-prem box; privacy-first solo/SMB use | Apple Mac Studio (M3 Ultra, 512GB) |
| Local on workstation GPUs | Multiple 48GB professional cards for MoE offload / tensor parallelism | Power users and small clusters that want cards they own | NVIDIA RTX 6000 Ada (48GB) |
Once GLM-5.2 is running, the fastest way to put it to work day to day is inside Cursor — point it at the model through OpenRouter as a custom model. And if you would rather run a model on one affordable box, see Best mini PCs for local AI and Local AI hardware calculator.

Frequently Asked Questions
- It depends which product you buy. On the pay-as-you-go API, GLM-4.6 lists at $0.60 per 1M input tokens and $2.20 per 1M output tokens on Z.AI's docs pricing page as of 2026-07-21. The GLM Coding Plan starts at $18/month for the Lite tier. GLM-4.5-Flash is priced at $0 on the same public pricing page. Enterprise pricing is not publicly listed — contact Z.AI sales.
- The platform itself is not free, but one shipping model — GLM-4.5-Flash — is listed at $0 for input, output, and cached input on Z.AI's public API pricing page as of 2026-07-21. That is a genuine free-to-call model, not a trial credit, though standard rate limits and Terms of Service apply and pricing can change without notice.
- Yes, in the form of a free model rather than a credit grant. GLM-4.5-Flash is priced at zero on the public API pricing page. There is no advertised trial-credit grant on the pricing page verified here — if you need larger free access for a proof-of-concept, request it from Z.AI directly.
- On the pricing pages verified on 2026-07-21, yes — significantly. GLM-4.6 lists at $0.60 input / $2.20 output per 1M tokens on docs.z.ai. Anthropic Claude Sonnet 4.5 lists at $3.00 input / $15.00 output per 1M on anthropic.com/pricing. That is roughly 5x cheaper on input and about 7x cheaper on output. Whether that math translates into a lower TCO for your workload depends on caching, moderation, fallbacks, and how well GLM performs on your evals — not on the sticker price alone.
- For flagship-to-flagship, DeepSeek V3 is cheaper on the pricing pages verified on 2026-07-21 — DeepSeek V3 lists at roughly $0.27 input / $1.10 output per 1M tokens on deepseek.com, versus $0.60 input / $2.20 output for Z.AI GLM-4.6. Z.AI's GLM-4.5-Air ($0.20 input / $1.10 output) is closer to DeepSeek's per-token math. The right choice depends on eval quality on your task, not sticker price.
- Volume discounts are not publicly listed on Z.AI's pricing page. Enterprise buyers with meaningful token commitments should contact Z.AI sales for a custom quote — volume commits and dedicated deployments are handled through direct negotiation rather than a published discount ladder.
- The GLM Coding Plan starts at $18/month for the Lite tier, which allows roughly 80 prompts per 5 hours and 400 per week per Z.AI's public docs on 2026-07-21. The Pro and Max tier monthly prices are not published on the docs overview page verified here — check z.ai/subscribe for the current tier prices. For a full breakdown of quotas, model access, and how the plan compares to per-token API billing, see our dedicated GLM Coding Plan page.
- Yes. Several GLM releases ship as open weights under permissive licenses, and self-hosting eliminates per-token API fees. You still pay for multi-GPU hardware (flagship tiers typically need an 8x H100/H200-class node), engineering time, and operations. Self-hosting only wins on pure cost at very high steady-state token volume, or when the driver is data isolation rather than economics.
- Z.AI's public pricing page lists GLM-4.5-Flash at $0, but it does not publish a specific rate limit for that free access. The docs say standard API rate limits and Terms of Service apply, without stating a number. Third-party API trackers report roughly 1 request per second (~60/minute) with a daily cap around 1,000 requests, but that is an outside estimate, not a Z.AI-published figure — confirm current limits in your Z.AI console, and contact Z.AI directly if you need a written commitment for a production workload.
Not Sure Which GLM Tier Actually Saves You Money?
Per-token math looks simple until you factor in caching, guardrails, fallbacks, and the cost of switching. Book a workflow audit and we will model the real cost of moving a specific workload to GLM against your current provider.
Book a Free AI Cost Review