Reviewed by Jonathan West · Updated Aug 6, 2026

Grok 4.5 API Pricing: Per-Token Rates and Cost Modeling

A developer-focused breakdown of what xAI charges for Grok 4.5, how prompt caching and long-context surcharges shift the math, and how the bill compares to Claude and GPT-5.6 for the same workload.

Reviewed by Jonathan West · Updated Aug 6, 2026

Grok 4.5 is priced at $2.00 per million input tokens and $6.00 per million output tokens on the xAI API, with a 75% prompt-caching discount that drops cached input to $0.50 per million, per xAI's own pricing page at docs.x.ai. Numbers here should be re-verified before you build a budget on them.

For most agent and RAG workloads, that puts Grok 4.5 in the same tier as Claude Sonnet 5 on input and cheaper than every flagship on output. It is not the cheapest model on the market, but the effective rate after caching a long system prompt is competitive with mid-tier frontier options.

This page walks through the current rate card, the two surcharges that trip up first-time budgets, a worked example for 1,000 chat requests, and a side-by-side comparison against Claude Sonnet 5, Claude Opus 5, and the GPT-5.6 family.


Headline Grok 4.5 API Rates

Grok 4.5 costs $2.00 per million input tokens and $6.00 per million output tokens on the standard xAI API tier. Cached input tokens are billed at $0.50 per million, a 75% discount versus fresh input.

The context window is 500,000 tokens on the API. Web and X search tool calls bill separately at roughly $5.00 per 1,000 calls when the model chooses to invoke them.

These are the current numbers surfaced by xAI's pricing documentation and mirrored by third-party trackers. Always confirm on docs.x.ai before you commit a monthly budget, because xAI has repriced Grok models mid-quarter before.

There is no published batch-processing discount for Grok 4.5, unlike the 20% batch discount xAI still lists for Grok 4.3.

Trying to decide whether Grok 4.5 at $2 / $6 per million tokens beats Claude Sonnet 5 or GPT-5.6 Terra for your specific stack? We can model the caching, surcharge, and quality tradeoffs against your real prompts.

Get a Grok 4.5 cost model

The Two Surcharges Every Buyer Misses

Two line items catch teams off guard when the first invoice arrives. The first is the long-context surcharge, and the second is priority processing.

xAI doubles per-token rates on prompts that exceed 200,000 tokens. Grok 4.5 at high context bills at $4.00 input and $12.00 output per million tokens rather than the headline $2 / $6.

Priority processing, which trades cost for a faster place in the queue, also doubles the base rate to $4 / $12 per million tokens. Stack it with long context and you pay 4x the sticker price.

Neither surcharge is optional in the client SDK by default, but both are visible in the console.x.ai billing view once they trigger. Alert on them before they surprise finance.


Prompt Caching Changes the Math

Prompt caching is where the effective cost of Grok 4.5 actually lands for most production workloads. xAI charges $0.50 per million cached input tokens versus $2.00 for fresh input, a 75% cut.

A typical agent design keeps a large, stable system prompt plus a growing conversation. If 80% of your input tokens per call are cache hits, blended input drops to roughly $0.80 per million.

Cache hits require the same prefix across calls, so caching pays off in RAG pipelines, tool-using agents, and long-lived chat threads. It rarely helps one-shot generation.

In our own work on Layer3Labs client stacks, we treat the caching-eligible portion of the prompt as a hard design constraint from day one. Retrofitting caching later is painful, and the price gap versus non-cached calls is large enough to be worth planning around.


A Simple Cost Formula

Use this formula to estimate a monthly Grok 4.5 API bill before you run traffic through it. It covers the four levers that actually move the number.

Monthly cost in USD = (fresh input tokens / 1,000,000 x $2.00) + (cached input tokens / 1,000,000 x $0.50) + (output tokens / 1,000,000 x $6.00) + (search tool calls / 1,000 x $5.00).

If any single request crosses 200,000 tokens of context, double the input and output rates for that request's tokens. If you enable priority processing on a request, double them again on top of that.

Track fresh vs cached input separately in your logging from day one. Blended reporting hides the caching lever, which is the largest single knob you can turn on Grok 4.5 spend.


Worked Example: 1,000 Chat Requests

Assume 1,000 typical support-agent requests. Each request sends a 4,000-token cached system prompt, 1,000 tokens of fresh user context, and produces a 500-token answer. No long context, no priority tier, no search calls.

Cached input: 4,000,000 tokens at $0.50 per million = $2.00. Fresh input: 1,000,000 tokens at $2.00 per million = $2.00. Output: 500,000 tokens at $6.00 per million = $3.00.

Total for 1,000 requests: $7.00, or roughly $0.007 per request. Without caching, the same volume runs about $13, so the cache is doing real work.

Now scale to 100,000 requests per day and the picture shifts. Daily spend lands near $700, monthly near $21,000, and the surcharge cliff at 200,000 tokens per request becomes worth engineering around.


Grok 4.5 vs Claude and GPT-5.6 on Cost

Grok 4.5's $2 / $6 rate card sits below every flagship model on output pricing. Claude Opus 5 lists at $5 / $25 per million tokens, and Claude Sonnet 5 lists at an introductory $2 / $10 through August 31, 2026, then $3 / $15 standard.

GPT-5.6 splits into three tiers. Sol runs $5 / $30, Terra runs $2 / $12, and Luna runs $0.20 / $1.20 per million input and output tokens.

On the same 1,000-request example above, Grok 4.5 lands at $7.00, Claude Sonnet 5 at introductory pricing lands near $10, Claude Opus 5 near $27, GPT-5.6 Terra near $11, and GPT-5.6 Sol near $32. Luna at $0.20 / $1.20 is cheapest by a wide margin, but it is not a like-for-like intelligence tier.

Price alone does not decide the model. Evaluate on quality against your actual prompts, then let the cost delta break ties.

In our own work running content and agent workloads across the Layer3Labs portfolio, we see teams pick Grok 4.5 for high-volume drafting and tool-use loops where output tokens dominate the bill, and pick Claude Sonnet 5 or GPT-5.6 Terra where reasoning quality matters more than the per-call rate.


Buying Checklist Before You Commit

Run a live evaluation on 100 real prompts from your workload. Log fresh input, cached input, output, and any tool calls separately, then multiply against the rate card.

Confirm the surcharge triggers. If any prompt in your evaluation approaches 200,000 tokens, price that request at the doubled rate, not the headline rate.

Check console.x.ai for organization-level rate limits and any negotiated enterprise discount. Public pricing is the starting point, not the ending point, for teams above roughly $10,000 in monthly spend.

Model the cache hit rate honestly. A design that assumes 90% cache hits and delivers 30% will run 2-3x over budget within the first month.

Frequently Asked Questions

  • Grok 4.5 is $2.00 per million input tokens and $6.00 per million output tokens on the standard tier, with cached input at $0.50 per million. These rates are published at docs.x.ai and should be reverified before budgeting.
  • Yes. Cached input tokens are billed at $0.50 per million, which is 75% below the $2.00 fresh-input rate. Cache hits require the same prompt prefix across calls, so agents and RAG pipelines with stable system prompts benefit most.
  • The Grok 4.5 API supports a 500,000-token context window. Prompts above 200,000 tokens trigger a long-context surcharge that doubles the per-token rate to $4 input and $12 output per million tokens.
  • Grok 4.5 is cheaper than Claude Sonnet 5 on output tokens and matches it on input at introductory pricing. It is cheaper than Claude Opus 5 and GPT-5.6 Sol across the board, but GPT-5.6 Luna undercuts it at a lower intelligence tier.
  • Two line items surprise buyers. The long-context surcharge doubles rates above 200,000 tokens per request, and priority processing doubles rates again for faster queue placement. Search tool calls bill separately at roughly $5.00 per 1,000 calls.
  • Split traffic into fresh input, cached input, output tokens, and tool calls. Multiply each by its rate, then double any request that crosses the 200,000-token threshold or uses priority mode. Track cache hit rate separately, since it is the largest lever on effective cost.

Turn a Grok 4.5 API bill into a workflow that pays for itself

Layer3Labs helps SMB and regulated-industry teams pick the right model tier, wire up prompt caching, and set surcharge alarms before the first invoice lands. We map cost to actual output, not sticker price.

Book an AI Workflow Audit