AI Model Cost Calculator
Enter your monthly token volume and see the estimated cost across every major AI model, ranked cheapest to most expensive.
AI model pricing is split into an input price and an output price per million tokens, and the gap between the cheapest and most expensive models is often 100x or more. This tool multiplies your monthly token volume against every model in our pricing table so you can see the real cost gap for your own workload, not just a generic per-token number.
Estimate Your Monthly Cost
Not sure about your volume? A busy support agent replying to ~2,000 tickets a month with a 1,000-token context each is roughly 2M input / 1M output tokens — scale from there.
| Rank | Model | Vendor | Est. monthly cost |
|---|---|---|---|
| 1 | Nova Micro | Amazon | $3.15 |
| 2 | Nova Lite | Amazon | $5.40 |
| 3 | DeepSeek V4 Flash | DeepSeek | $9.80 |
| 4 | Mistral Small 4 | Mistral | $13.50 |
| 5 | Llama 4 Scout | Meta | $15.10 |
| 6 | Llama 4 Maverick | Meta | $21.70 |
| 7 | Codestral | Mistral | $24.00 |
| 8 | Gemini 3.1 Flash-Lite | $27.50 | |
| 9 | DeepSeek V4 Pro | DeepSeek | $30.45 |
| 10 | Gemini 2.5 Flash | $40.00 | |
| 11 | Mistral Large 3 | Mistral | $40.00 |
| 12 | Command R | Cohere | $40.00 |
| 13 | Sonar | Perplexity | $60.00 |
| 14 | Grok Build | xAI | $70.00 |
| 15 | Nova Pro | Amazon | $72.00 |
| 16 | Grok 4.3 | xAI | $87.50 |
| 17 | Claude Haiku 4.5 | Anthropic | $100 |
| 18 | GPT-5.6 Luna | OpenAI | $110 |
| 19 | Mistral Medium 3.5 | Mistral | $150 |
| 20 | Magistral Medium | Mistral | $150 |
| 21 | Gemini 2.5 Pro | $163 | |
| 22 | Gemini 3.5 Flash | $165 | |
| 23 | Sonar Reasoning Pro | Perplexity | $180 |
| 24 | Sonar Deep Research | Perplexity | $180 |
| 25 | Claude Sonnet 5 | Anthropic | $200 |
| 26 | Gemini 3.1 Pro | $220 | |
| 27 | Command R+ | Cohere | $225 |
| 28 | Nova Premier | Amazon | $250 |
| 29 | GPT-5.6 Terra | OpenAI | $275 |
| 30 | Sonar Pro | Perplexity | $300 |
| 31 | Claude Opus 4.8 | Anthropic | $500 |
| 32 | GPT-5.6 Sol | OpenAI | $550 |
| 33 | Claude Fable 5 | Anthropic | $1,000 |
| 34 | Claude Mythos 5 | Anthropic | $1,000 |
Estimate only — treats every token as list-priced with no prompt caching. Cached input tokens (where a vendor offers them) cost far less on repeated context; check the full pricing table for cached rates and source links before you commit spend.
How the Estimate Works
The calculator takes two numbers — your monthly input tokens and monthly output tokens, both in millions — and multiplies each against every model's published per-million-token price. Input and output are priced separately because output almost always costs several times more than input.
- Monthly cost = (input tokens in millions × input price) + (output tokens in millions × output price)
- Every model is ranked from cheapest to most expensive at your specific volume
- Prices come from the same ledger as our full AI model pricing page, sourced from each vendor's official pricing page
Why Cost Varies So Much Between Models
The cheapest models on this list, such as Amazon Nova Micro or DeepSeek V4 Flash, cost a small fraction of a cent per million tokens. The most capable flagship models, such as Claude Fable 5 or GPT-5.6 Sol, cost tens of dollars per million tokens because they are built for the hardest reasoning, coding, and research work.
What This Calculator Does Not Cover
This is a token-price estimate, not a full bill. It does not include prompt-caching discounts (several vendors price repeated context far cheaper on a cache hit), per-request fees some vendors add on top of token price, or seat-based subscription plans like ChatGPT Plus or Claude Pro, which are priced flat per user rather than per token.
Use the result to compare models against each other for your workload, then confirm the current price and any caching terms on the vendor's official pricing page before you commit spend.
Frequently Asked Questions
- Start from a real workload. Count roughly how many requests you send per month and the average input and output length per request, then multiply. A support agent answering 2,000 tickets a month with a 1,000-token context is close to 2 million input and 1 million output tokens.
- Yes, and that is the reason to run the numbers instead of trusting a single per-token price. Because input and output are priced separately, a model that looks cheap on input can slip down the ranking once you weight it by a heavy-output workload (long generated answers, not long prompts). Enter your own volume above to see the ranking for your actual mix, not a generic list.
- No. It prices every token at the standard list rate. Several vendors, including Anthropic and Google, offer a much cheaper cached-input rate for repeated context. If your workload reuses a large system prompt or document, your real cost will be lower than this estimate.
- No. This tool estimates API (pay-per-token) costs only. Seat-based subscription plans are a flat monthly price regardless of usage — see the full pricing table for those.
- Real bills depend on your exact token counts, caching hit rate, and any per-request fees a vendor charges (Perplexity Sonar, for example, adds a search fee per request). Treat this as a starting comparison, then confirm the current price on the vendor's official pricing page.
Not Sure Which Model Fits Your Workload?
Token price is only one input. Our free AI workflow audit maps the right model, plan, and routing strategy to your exact tasks, budget, and compliance needs.
Book a Consultation