Reviewed by Jonathan West · Updated Jul 21, 2026

Z.AI Pricing: What GLM Models Actually Cost

A buyer-focused breakdown of Z.AI's GLM API pricing, Coding Plan tiers, free access, enterprise options, and how the numbers compare to Claude, OpenAI, and DeepSeek.

Reviewed by Jonathan West · Updated Jul 21, 2026

Z.AI is the international brand of Chinese AI lab Zhipu AI, and its GLM family of models has become one of the most-quoted low-cost alternatives to Claude and GPT. But the pricing page mixes per-token API rates, a subscription-style Coding Plan, and a limited free tier, and that makes real cost comparisons harder than they should be.

This page pulls the numbers straight from Z.AI's own pricing documentation as of 2026-07-21, keeps every figure sourced, and stays out of the gap where a page invents rates that were never published. Where a number is not on the public page (Coding Plan Pro and Max tiers, enterprise contracts), we say so and point you to the live page.

We cover the per-model API rates for every shipping GLM tier, what the free tier actually includes, how the GLM Coding Plan works at a high level (the full breakdown lives on our GLM Coding Plan page), what enterprise buyers should expect, and side-by-side comparisons against Anthropic Claude, OpenAI, and DeepSeek.

All dollar figures are verified against vendor pricing pages on 2026-07-21. Model prices change often — always confirm on the vendor's live page before signing a contract or building a cost model.


Z.AI Pricing at a Glance

Z.AI sells access to its GLM models three ways. Understanding which one you need is the first cost decision, because the price-per-token math only applies to one of them.

  • Pay-as-you-go API: per-1M-token pricing for input, output, and cached input. This is the standard developer path and the mode most cost comparisons refer to.
  • GLM Coding Plan: a flat monthly subscription with prompt quotas, built around the GLM Coding tier. Starts at $18/month for the Lite plan. Suited to individual developers using the model inside coding IDEs.
  • Free tier: GLM-4.5-Flash is priced at zero for input, output, and cached input on Z.AI's public API pricing page — a genuine free model, not a trial credit.
  • Enterprise / custom: volume commits and dedicated deployments are not publicly listed — contact Z.AI sales.

Want a real cost model for switching a workload to GLM — including caching, fallback, and the engineering time to migrate? Book a workflow audit and we will build it against your current provider.

Book a Consultation

API Pricing by GLM Model

These are Z.AI's published per-1M-token API rates as listed on the official pricing page on 2026-07-21. All values are USD. Cached-input pricing applies when a prefix has been seen recently and Z.AI can reuse the KV cache.

  • GLM-4.6 (flagship): $0.60 per 1M input tokens, $2.20 per 1M output tokens, $0.11 per 1M cached-input tokens.
  • GLM-4.5: $0.60 per 1M input, $2.20 per 1M output, $0.11 per 1M cached input. Same public rates as GLM-4.6.
  • GLM-4.5-Air (lightweight): $0.20 per 1M input, $1.10 per 1M output, $0.03 per 1M cached input.
  • GLM-4.5-Flash: free — $0 for input, output, and cached input on the public pricing page.
  • GLM-4.5V (vision): $0.60 per 1M input, $1.80 per 1M output, $0.11 per 1M cached input.
  • GLM-4-32B-0414-128K: $0.10 per 1M input, $0.10 per 1M output.
  • GLM-5 and GLM-5.2 tiers are quoted through the GLM Coding Plan and enterprise channels — per-token rates for the newest GLM-5 family are not on the same public pricing table at the time of verification; check the live page before modeling costs.
Prices verified against docs.z.ai on 2026-07-21. Vendor pricing changes frequently — treat these figures as a snapshot and re-check the live page before contract or budget commitments.

GLM Coding Plan (Subscription)

The GLM Coding Plan is a flat monthly subscription that gives coding-tool users a prompt quota rather than per-token billing. It is aimed at developers using GLM inside IDEs and CLI coding assistants.

Z.AI's docs list three tiers — Lite, Pro, and Max — each with its own 5-hour and weekly prompt cap. The Lite plan starts at $18/month. Pro and Max monthly prices are not published on the docs overview page verified here; check z.ai/subscribe for the current tier prices.

For the full breakdown of quotas, model access, and how the Coding Plan compares to per-token API billing, see our dedicated /guides/glm-coding-plan-explained page.

  • Lite: ~80 prompts per 5 hours, ~400 per week. Starts at $18/month.
  • Pro: ~400 prompts per 5 hours, ~2,000 per week. Monthly price not publicly listed on the docs overview page — check z.ai/subscribe.
  • Max: ~1,600 prompts per 5 hours, ~8,000 per week. Monthly price not publicly listed on the docs overview page — check z.ai/subscribe.
  • Z.AI states the monthly quota is roughly equivalent to 15–30x the subscription fee in API-billing terms, so the plan is designed to be cheaper than per-token API for steady coding workloads.
  • Advanced models (GLM-5.2, GLM-5-Turbo) consume quota at higher multiples during peak Beijing hours.

Free-Tier Limits

Z.AI's free-tier story is unusual: instead of trial credits, one shipping model is priced at zero on the public API pricing page.

  • GLM-4.5-Flash is listed at $0 for input, output, and cached input tokens on the Z.AI pricing page as of 2026-07-21.
  • Standard API rate limits and Terms of Service still apply. Z.AI does not commit on the pricing page to a permanent-forever free status — the line item is priced Free, but pricing can change without notice.
  • For the paid models, Z.AI's docs mention a 'Limited-time Free' benefit on cached-input storage for the paid tiers — the storage of cached prefixes is free during the promotional window, but the cached-input read price still applies.
  • There is no bundled free-credit grant advertised on the public pricing page verified here — if you need trial credits for a proof-of-concept, contact Z.AI directly.

Enterprise / Custom Pricing

Z.AI does not publish enterprise-tier pricing. Volume discounts, dedicated deployments, private endpoints, data-residency options, and SLA-backed availability are all negotiated case-by-case.

  • Volume discounts: not publicly listed — contact Z.AI sales for token-volume commitments.
  • Dedicated / private deployment: not publicly listed — enterprise-only. Buyers focused on data isolation often self-host the open-weight GLM releases instead (see the self-hosted section below).
  • Data-processing agreements and regional hosting: needed for GDPR-scope data. Not documented on the public pricing page — request through sales.
  • Support SLAs: not publicly listed. Standard API access has documented rate limits but no uptime SLA on the public page.

How Z.AI Pricing Compares

The chart below puts Z.AI's flagship per-token API pricing next to the equivalent tier from Anthropic, OpenAI, and DeepSeek. All figures are verified against each vendor's public pricing page on 2026-07-21 and are USD per 1M tokens.

  • Z.AI GLM-4.6: $0.60 input / $2.20 output. Source: docs.z.ai/guides/overview/pricing.
  • Z.AI GLM-4.5-Air: $0.20 input / $1.10 output. Source: docs.z.ai/guides/overview/pricing.
  • Anthropic Claude Sonnet 4.5: $3.00 input / $15.00 output. Source: anthropic.com/pricing. GLM-4.6 lists roughly 5x lower on input, ~7x lower on output.
  • Anthropic Claude Opus 4.5: $15.00 input / $75.00 output. Source: anthropic.com/pricing. Opus targets a different tier — this is a ceiling reference, not a like-for-like swap.
  • OpenAI GPT-5 (standard tier): pricing shifts often — verify at openai.com/api/pricing before modeling. GLM-4.6 sits well below the GPT-5 standard tier on the pricing page verified today.
  • DeepSeek V3: $0.27 input / $1.10 output at standard rates. Source: deepseek.com. DeepSeek undercuts GLM-4.6 on input; GLM-4.5-Air is closer to DeepSeek territory.
Every price above is a snapshot from each vendor's public pricing page on 2026-07-21. Model pricing changes on the order of months — before you build a business case, click through and re-verify each row on the vendor's live page.

Self-Hosted Cost (Open Weights)

Several GLM releases (including GLM-4.5, GLM-4.5-Air, and GLM-5.2) ship as open weights under permissive licenses. Self-hosting means no per-token API cost, but you pay for GPUs, engineers, and operational overhead.

  • Per-token cost: $0. There is no license fee for the open-weight GLM releases.
  • Hardware class: flagship GLM models require a multi-GPU node — typically 8x H100/H200-class accelerators for full-precision serving. Quantized deployments can run on smaller footprints with quality trade-offs.
  • GPU capital or rental: exact list prices are not published across cloud vendors — check the live pricing on your target cloud (AWS, GCP, Azure, Lambda, CoreWeave) before modeling. A single H100 rental typically runs several dollars per hour on-demand at public cloud vendors.
  • Operational cost: engineering time to stand up vLLM/SGLang, monitoring, autoscaling, model updates, and security patching. These are the costs that surprise most teams.
  • Break-even: self-hosting only wins on cost when steady-state token volume is very high, or when the driver is data isolation rather than pure economics.

Total Cost of Ownership

The sticker price on a pricing page is only one line of the real cost of running GLM in production. A realistic TCO model has to include the operating cost around the token bill.

  • Token spend: per-1M rates above, multiplied by realistic input and output volumes for your workload — not the marketing example.
  • Prompt caching: Z.AI's cached-input rate is roughly 5x cheaper than standard input for GLM-4.6. Structured prompts that reuse a long system prefix can meaningfully lower the effective per-request cost.
  • Guardrails and moderation: teams routing production traffic to any Chinese-lab model typically add a moderation layer, log review, and prompt-filtering — this is engineering time and, often, a second API bill.
  • Data-processing and compliance: GDPR-scope traffic requires a DPA or self-hosted deployment. Compliance review is a real line item, not a rounding error.
  • Fallback provider: production stacks usually keep a second provider on standby for outages or rate-limit events. Budget for the second provider's minimums.
  • Engineering time to swap: switching from Claude or GPT to GLM is not free — evaluation, prompt-tuning, and eval-set rebuilds all cost engineering hours.
Sources

When Z.AI's Pricing Wins

Z.AI's pricing is genuinely competitive, but not for every workload. The scenarios below are where GLM's per-token math translates into real savings.

  • High-volume batch coding: overnight refactors, vulnerability scans, or agentic code review where per-token cost dominates and latency is not the constraint.
  • Long-context document processing: legal review, contract analysis, or research summarization where GLM's long-context tiers avoid chunking overhead.
  • Individual developers on the Coding Plan: at $18/month for Lite, the effective per-prompt cost is well below equivalent Claude Sonnet or GPT-5 usage inside coding IDEs.
  • Free-tier experimentation: GLM-4.5-Flash at $0 is a real free model for prototyping, evaluation harnesses, and internal tooling.
  • Cost-driven agent workloads: multi-step agents that burn tokens quickly benefit most from the lower per-1M rates.
  • Scenarios where Z.AI does NOT win: real-time customer-facing chat where latency and moderation matter more than token price; regulated workloads that require a US-based provider with an existing BAA/DPA; teams already deeply invested in Claude or OpenAI tooling where switching cost exceeds the savings.
Sources

Frequently Asked Questions

  • It depends which product you buy. On the pay-as-you-go API, GLM-4.6 lists at $0.60 per 1M input tokens and $2.20 per 1M output tokens on Z.AI's docs pricing page as of 2026-07-21. The GLM Coding Plan starts at $18/month for the Lite tier. GLM-4.5-Flash is priced at $0 on the same public pricing page. Enterprise pricing is not publicly listed — contact Z.AI sales.
  • The platform itself is not free, but one shipping model — GLM-4.5-Flash — is listed at $0 for input, output, and cached input on Z.AI's public API pricing page as of 2026-07-21. That is a genuine free-to-call model, not a trial credit, though standard rate limits and Terms of Service apply and pricing can change without notice.
  • Yes, in the form of a free model rather than a credit grant. GLM-4.5-Flash is priced at zero on the public API pricing page. There is no advertised trial-credit grant on the pricing page verified here — if you need larger free access for a proof-of-concept, request it from Z.AI directly.
  • On the pricing pages verified on 2026-07-21, yes — significantly. GLM-4.6 lists at $0.60 input / $2.20 output per 1M tokens on docs.z.ai. Anthropic Claude Sonnet 4.5 lists at $3.00 input / $15.00 output per 1M on anthropic.com/pricing. That is roughly 5x cheaper on input and about 7x cheaper on output. Whether that math translates into a lower TCO for your workload depends on caching, moderation, fallbacks, and how well GLM performs on your evals — not on the sticker price alone.
  • For flagship-to-flagship, DeepSeek V3 is cheaper on the pricing pages verified on 2026-07-21 — DeepSeek V3 lists at roughly $0.27 input / $1.10 output per 1M tokens on deepseek.com, versus $0.60 input / $2.20 output for Z.AI GLM-4.6. Z.AI's GLM-4.5-Air ($0.20 input / $1.10 output) is closer to DeepSeek's per-token math. The right choice depends on eval quality on your task, not sticker price.
  • Volume discounts are not publicly listed on Z.AI's pricing page. Enterprise buyers with meaningful token commitments should contact Z.AI sales for a custom quote — volume commits and dedicated deployments are handled through direct negotiation rather than a published discount ladder.
  • The GLM Coding Plan starts at $18/month for the Lite tier, which allows roughly 80 prompts per 5 hours and 400 per week per Z.AI's public docs on 2026-07-21. The Pro and Max tier monthly prices are not published on the docs overview page verified here — check z.ai/subscribe for the current tier prices. For a full breakdown of quotas, model access, and how the plan compares to per-token API billing, see our dedicated GLM Coding Plan page.
  • Yes. Several GLM releases ship as open weights under permissive licenses, and self-hosting eliminates per-token API fees. You still pay for multi-GPU hardware (flagship tiers typically need an 8x H100/H200-class node), engineering time, and operations. Self-hosting only wins on pure cost at very high steady-state token volume, or when the driver is data isolation rather than economics.

Not Sure Which GLM Tier Actually Saves You Money?

Per-token math looks simple until you factor in caching, guardrails, fallbacks, and the cost of switching. Book a workflow audit and we will model the real cost of moving a specific workload to GLM against your current provider.

/ai-workflow-audit