Z.AI Pricing: What GLM Models Actually Cost
A buyer-focused breakdown of Z.AI's GLM API pricing, Coding Plan tiers, free access, enterprise options, and how the numbers compare to Claude, OpenAI, and DeepSeek.
Z.AI is the international brand of Chinese AI lab Zhipu AI, and its GLM family of models has become one of the most-quoted low-cost alternatives to Claude and GPT. But the pricing page mixes per-token API rates, a subscription-style Coding Plan, and a limited free tier, and that makes real cost comparisons harder than they should be.
This page pulls the numbers straight from Z.AI's own pricing documentation as of 2026-07-21, keeps every figure sourced, and stays out of the gap where a page invents rates that were never published. Where a number is not on the public page (Coding Plan Pro and Max tiers, enterprise contracts), we say so and point you to the live page.
We cover the per-model API rates for every shipping GLM tier, what the free tier actually includes, how the GLM Coding Plan works at a high level (the full breakdown lives on our GLM Coding Plan page), what enterprise buyers should expect, and side-by-side comparisons against Anthropic Claude, OpenAI, and DeepSeek.
All dollar figures are verified against vendor pricing pages on 2026-07-21. Model prices change often — always confirm on the vendor's live page before signing a contract or building a cost model.
Z.AI Pricing at a Glance
Z.AI sells access to its GLM models three ways. Understanding which one you need is the first cost decision, because the price-per-token math only applies to one of them.
- Pay-as-you-go API: per-1M-token pricing for input, output, and cached input. This is the standard developer path and the mode most cost comparisons refer to.
- GLM Coding Plan: a flat monthly subscription with prompt quotas, built around the GLM Coding tier. Starts at $18/month for the Lite plan. Suited to individual developers using the model inside coding IDEs.
- Free tier: GLM-4.5-Flash is priced at zero for input, output, and cached input on Z.AI's public API pricing page — a genuine free model, not a trial credit.
- Enterprise / custom: volume commits and dedicated deployments are not publicly listed — contact Z.AI sales.
Want a real cost model for switching a workload to GLM — including caching, fallback, and the engineering time to migrate? Book a workflow audit and we will build it against your current provider.
Book a ConsultationAPI Pricing by GLM Model
These are Z.AI's published per-1M-token API rates as listed on the official pricing page on 2026-07-21. All values are USD. Cached-input pricing applies when a prefix has been seen recently and Z.AI can reuse the KV cache.
- GLM-4.6 (flagship): $0.60 per 1M input tokens, $2.20 per 1M output tokens, $0.11 per 1M cached-input tokens.
- GLM-4.5: $0.60 per 1M input, $2.20 per 1M output, $0.11 per 1M cached input. Same public rates as GLM-4.6.
- GLM-4.5-Air (lightweight): $0.20 per 1M input, $1.10 per 1M output, $0.03 per 1M cached input.
- GLM-4.5-Flash: free — $0 for input, output, and cached input on the public pricing page.
- GLM-4.5V (vision): $0.60 per 1M input, $1.80 per 1M output, $0.11 per 1M cached input.
- GLM-4-32B-0414-128K: $0.10 per 1M input, $0.10 per 1M output.
- GLM-5 and GLM-5.2 tiers are quoted through the GLM Coding Plan and enterprise channels — per-token rates for the newest GLM-5 family are not on the same public pricing table at the time of verification; check the live page before modeling costs.
GLM Coding Plan (Subscription)
The GLM Coding Plan is a flat monthly subscription that gives coding-tool users a prompt quota rather than per-token billing. It is aimed at developers using GLM inside IDEs and CLI coding assistants.
Z.AI's docs list three tiers — Lite, Pro, and Max — each with its own 5-hour and weekly prompt cap. The Lite plan starts at $18/month. Pro and Max monthly prices are not published on the docs overview page verified here; check z.ai/subscribe for the current tier prices.
For the full breakdown of quotas, model access, and how the Coding Plan compares to per-token API billing, see our dedicated /guides/glm-coding-plan-explained page.
- Lite: ~80 prompts per 5 hours, ~400 per week. Starts at $18/month.
- Pro: ~400 prompts per 5 hours, ~2,000 per week. Monthly price not publicly listed on the docs overview page — check z.ai/subscribe.
- Max: ~1,600 prompts per 5 hours, ~8,000 per week. Monthly price not publicly listed on the docs overview page — check z.ai/subscribe.
- Z.AI states the monthly quota is roughly equivalent to 15–30x the subscription fee in API-billing terms, so the plan is designed to be cheaper than per-token API for steady coding workloads.
- Advanced models (GLM-5.2, GLM-5-Turbo) consume quota at higher multiples during peak Beijing hours.
Free-Tier Limits
Z.AI's free-tier story is unusual: instead of trial credits, one shipping model is priced at zero on the public API pricing page.
- GLM-4.5-Flash is listed at $0 for input, output, and cached input tokens on the Z.AI pricing page as of 2026-07-21.
- Standard API rate limits and Terms of Service still apply. Z.AI does not commit on the pricing page to a permanent-forever free status — the line item is priced Free, but pricing can change without notice.
- For the paid models, Z.AI's docs mention a 'Limited-time Free' benefit on cached-input storage for the paid tiers — the storage of cached prefixes is free during the promotional window, but the cached-input read price still applies.
- There is no bundled free-credit grant advertised on the public pricing page verified here — if you need trial credits for a proof-of-concept, contact Z.AI directly.
Enterprise / Custom Pricing
Z.AI does not publish enterprise-tier pricing. Volume discounts, dedicated deployments, private endpoints, data-residency options, and SLA-backed availability are all negotiated case-by-case.
- Volume discounts: not publicly listed — contact Z.AI sales for token-volume commitments.
- Dedicated / private deployment: not publicly listed — enterprise-only. Buyers focused on data isolation often self-host the open-weight GLM releases instead (see the self-hosted section below).
- Data-processing agreements and regional hosting: needed for GDPR-scope data. Not documented on the public pricing page — request through sales.
- Support SLAs: not publicly listed. Standard API access has documented rate limits but no uptime SLA on the public page.
How Z.AI Pricing Compares
The chart below puts Z.AI's flagship per-token API pricing next to the equivalent tier from Anthropic, OpenAI, and DeepSeek. All figures are verified against each vendor's public pricing page on 2026-07-21 and are USD per 1M tokens.
- Z.AI GLM-4.6: $0.60 input / $2.20 output. Source: docs.z.ai/guides/overview/pricing.
- Z.AI GLM-4.5-Air: $0.20 input / $1.10 output. Source: docs.z.ai/guides/overview/pricing.
- Anthropic Claude Sonnet 4.5: $3.00 input / $15.00 output. Source: anthropic.com/pricing. GLM-4.6 lists roughly 5x lower on input, ~7x lower on output.
- Anthropic Claude Opus 4.5: $15.00 input / $75.00 output. Source: anthropic.com/pricing. Opus targets a different tier — this is a ceiling reference, not a like-for-like swap.
- OpenAI GPT-5 (standard tier): pricing shifts often — verify at openai.com/api/pricing before modeling. GLM-4.6 sits well below the GPT-5 standard tier on the pricing page verified today.
- DeepSeek V3: $0.27 input / $1.10 output at standard rates. Source: deepseek.com. DeepSeek undercuts GLM-4.6 on input; GLM-4.5-Air is closer to DeepSeek territory.
Self-Hosted Cost (Open Weights)
Several GLM releases (including GLM-4.5, GLM-4.5-Air, and GLM-5.2) ship as open weights under permissive licenses. Self-hosting means no per-token API cost, but you pay for GPUs, engineers, and operational overhead.
- Per-token cost: $0. There is no license fee for the open-weight GLM releases.
- Hardware class: flagship GLM models require a multi-GPU node — typically 8x H100/H200-class accelerators for full-precision serving. Quantized deployments can run on smaller footprints with quality trade-offs.
- GPU capital or rental: exact list prices are not published across cloud vendors — check the live pricing on your target cloud (AWS, GCP, Azure, Lambda, CoreWeave) before modeling. A single H100 rental typically runs several dollars per hour on-demand at public cloud vendors.
- Operational cost: engineering time to stand up vLLM/SGLang, monitoring, autoscaling, model updates, and security patching. These are the costs that surprise most teams.
- Break-even: self-hosting only wins on cost when steady-state token volume is very high, or when the driver is data isolation rather than pure economics.
Total Cost of Ownership
The sticker price on a pricing page is only one line of the real cost of running GLM in production. A realistic TCO model has to include the operating cost around the token bill.
- Token spend: per-1M rates above, multiplied by realistic input and output volumes for your workload — not the marketing example.
- Prompt caching: Z.AI's cached-input rate is roughly 5x cheaper than standard input for GLM-4.6. Structured prompts that reuse a long system prefix can meaningfully lower the effective per-request cost.
- Guardrails and moderation: teams routing production traffic to any Chinese-lab model typically add a moderation layer, log review, and prompt-filtering — this is engineering time and, often, a second API bill.
- Data-processing and compliance: GDPR-scope traffic requires a DPA or self-hosted deployment. Compliance review is a real line item, not a rounding error.
- Fallback provider: production stacks usually keep a second provider on standby for outages or rate-limit events. Budget for the second provider's minimums.
- Engineering time to swap: switching from Claude or GPT to GLM is not free — evaluation, prompt-tuning, and eval-set rebuilds all cost engineering hours.
When Z.AI's Pricing Wins
Z.AI's pricing is genuinely competitive, but not for every workload. The scenarios below are where GLM's per-token math translates into real savings.
- High-volume batch coding: overnight refactors, vulnerability scans, or agentic code review where per-token cost dominates and latency is not the constraint.
- Long-context document processing: legal review, contract analysis, or research summarization where GLM's long-context tiers avoid chunking overhead.
- Individual developers on the Coding Plan: at $18/month for Lite, the effective per-prompt cost is well below equivalent Claude Sonnet or GPT-5 usage inside coding IDEs.
- Free-tier experimentation: GLM-4.5-Flash at $0 is a real free model for prototyping, evaluation harnesses, and internal tooling.
- Cost-driven agent workloads: multi-step agents that burn tokens quickly benefit most from the lower per-1M rates.
- Scenarios where Z.AI does NOT win: real-time customer-facing chat where latency and moderation matter more than token price; regulated workloads that require a US-based provider with an existing BAA/DPA; teams already deeply invested in Claude or OpenAI tooling where switching cost exceeds the savings.
Frequently Asked Questions
- It depends which product you buy. On the pay-as-you-go API, GLM-4.6 lists at $0.60 per 1M input tokens and $2.20 per 1M output tokens on Z.AI's docs pricing page as of 2026-07-21. The GLM Coding Plan starts at $18/month for the Lite tier. GLM-4.5-Flash is priced at $0 on the same public pricing page. Enterprise pricing is not publicly listed — contact Z.AI sales.
- The platform itself is not free, but one shipping model — GLM-4.5-Flash — is listed at $0 for input, output, and cached input on Z.AI's public API pricing page as of 2026-07-21. That is a genuine free-to-call model, not a trial credit, though standard rate limits and Terms of Service apply and pricing can change without notice.
- Yes, in the form of a free model rather than a credit grant. GLM-4.5-Flash is priced at zero on the public API pricing page. There is no advertised trial-credit grant on the pricing page verified here — if you need larger free access for a proof-of-concept, request it from Z.AI directly.
- On the pricing pages verified on 2026-07-21, yes — significantly. GLM-4.6 lists at $0.60 input / $2.20 output per 1M tokens on docs.z.ai. Anthropic Claude Sonnet 4.5 lists at $3.00 input / $15.00 output per 1M on anthropic.com/pricing. That is roughly 5x cheaper on input and about 7x cheaper on output. Whether that math translates into a lower TCO for your workload depends on caching, moderation, fallbacks, and how well GLM performs on your evals — not on the sticker price alone.
- For flagship-to-flagship, DeepSeek V3 is cheaper on the pricing pages verified on 2026-07-21 — DeepSeek V3 lists at roughly $0.27 input / $1.10 output per 1M tokens on deepseek.com, versus $0.60 input / $2.20 output for Z.AI GLM-4.6. Z.AI's GLM-4.5-Air ($0.20 input / $1.10 output) is closer to DeepSeek's per-token math. The right choice depends on eval quality on your task, not sticker price.
- Volume discounts are not publicly listed on Z.AI's pricing page. Enterprise buyers with meaningful token commitments should contact Z.AI sales for a custom quote — volume commits and dedicated deployments are handled through direct negotiation rather than a published discount ladder.
- The GLM Coding Plan starts at $18/month for the Lite tier, which allows roughly 80 prompts per 5 hours and 400 per week per Z.AI's public docs on 2026-07-21. The Pro and Max tier monthly prices are not published on the docs overview page verified here — check z.ai/subscribe for the current tier prices. For a full breakdown of quotas, model access, and how the plan compares to per-token API billing, see our dedicated GLM Coding Plan page.
- Yes. Several GLM releases ship as open weights under permissive licenses, and self-hosting eliminates per-token API fees. You still pay for multi-GPU hardware (flagship tiers typically need an 8x H100/H200-class node), engineering time, and operations. Self-hosting only wins on pure cost at very high steady-state token volume, or when the driver is data isolation rather than economics.
Not Sure Which GLM Tier Actually Saves You Money?
Per-token math looks simple until you factor in caching, guardrails, fallbacks, and the cost of switching. Book a workflow audit and we will model the real cost of moving a specific workload to GLM against your current provider.
/ai-workflow-audit