LLM Routing Cost Calculator

Compare what AI model routing costs you versus going direct. See the markup, and the compliance risk your routing layer can hide.

An AI routing layer like OpenRouter, LiteLLM, or Portkey sits between your app and the model providers. It buys you convenience — one API for many models — but that convenience has a price, and the price is not always the sticker rate. This tool estimates your monthly cost across each routing option for your own token volume and model mix, then flags the compliance exposure most teams never price in. For the full picture, start with AI model routing explained.

Compare Your Routing Cost

Monthly input tokens1M
Monthly output tokens500K

Not sure about your volume? A support agent answering ~2,000 tickets a month with a 1,000-token context each is roughly 2M input / 1M output tokens — scale from there. With several models selected, your volume is split evenly across them.

OpenAI
Anthropic
Google
DeepSeek
Meta (via providers)
Mistral
Alibaba (via providers)
OpenRouter adds an estimated 5.5% credit fee on spend.
Direct API$8.25/mo$99.00/yrCheapest
OpenRouter$8.70/mo$104/yr
LiteLLM (self-hosted)$58.25/mo$699/yr
Portkey$57.25/mo$687/yr
You could save $50.00/mo ($600/yr) by routing through Direct API versus the most expensive option.
ModelDirect APIOpenRouterLiteLLMPortkeyCheapest
GPT-4.1$3.00$3.17 +5.5%$28.00 +833%$27.50 +817%Direct API
Claude Sonnet 5$5.25$5.54 +5.5%$30.25 +476%$29.75 +467%Direct API
Compliance check: none of your selected models are Chinese-hosted. Verify your routing tool does not auto-select them on cost — audit your model allowlist.

Estimate only. Token prices are list rates last verified July 2026 and exclude prompt-caching discounts; OpenRouter, LiteLLM infra, and Portkey figures are configurable assumptions, not quotes. Models marked * are provider-dependent and approximate. Confirm current rates on each vendor's pricing page — see the OpenRouter pricing and LiteLLM pricing breakdowns — before you commit spend.

How the Estimate Works

You set two numbers — your monthly input and output tokens — and pick the models you use. The calculator prices each model at its published per-million-token rate, splits your volume evenly across the models you selected, and totals the monthly cost under four routing options.

  • Direct API — provider list rates, no markup. The baseline every other column is measured against.
  • OpenRouter — the same per-token rates plus an estimated credit fee on spend. See the OpenRouter pricing breakdown.
  • LiteLLM — provider rates plus the monthly infrastructure cost you enter to self-host the proxy. See LiteLLM pricing.
  • Portkey — provider rates plus a managed-gateway platform fee.

Where the Markup Hides

At low volume, a hosted gateway's per-token pass-through pricing is hard to beat — the fixed cost of self-hosting LiteLLM is not worth it for a few dollars of traffic. As volume grows, that flips: the marketplace fee scales with spend while a self-hosted proxy's infrastructure cost stays flat, so LiteLLM pulls ahead. Enter your real numbers above to find your own crossover point rather than trusting a generic rule of thumb. The OpenRouter vs LiteLLM comparison walks through the trade-off in detail.

The Hidden Compliance Cost

The cheapest model at a given quality tier is often a Chinese-developed one — DeepSeek or Qwen. A cost-optimizing router will reach for it unless you explicitly block it, which means a single API call can send regulated data to infrastructure governed by Chinese data law. That is not a line item on any pricing page, but it is the most expensive mistake in the whole stack.

If the calculator flags a Chinese-hosted model in your selection, treat it as a prompt to audit your routing tool's model allowlist — see the AI routing vendor audit checklist and the risks detailed in DeepSeek data privacy and security risks.

Frequently Asked Questions

  • OpenRouter charges the same per-token rate as the underlying provider for most models, then takes a fee on the credits you buy (roughly 5% plus payment processing). So the markup is on your total spend, not each token. Use the calculator above to see the effect on your model mix, and verify current terms on openrouter.ai/pricing.
  • The open-source proxy is free to run. You still pay the underlying model provider token costs, plus your own infrastructure — a server to host it, monitoring, and engineer time. The calculator lets you set that infrastructure cost. LiteLLM also sells a managed enterprise tier for a fee.
  • Direct API access is cheapest per token because there is no middleman. The trade-off is managing a separate integration for each provider. Self-hosted LiteLLM is usually the cheapest unified option once your volume covers the fixed infrastructure cost.
  • Yes. If your routing tool can reach Chinese-developed models such as DeepSeek or Qwen by default, a cost-optimized or misconfigured request can send sensitive data to servers governed by Chinese data laws (PIPL). The calculator flags any Chinese-hosted models you select. Audit your model allowlist before you route production traffic.
  • Frequently — major providers adjust pricing several times a year, usually downward. This calculator uses list rates last verified in July 2026 and is updated quarterly. Always confirm the current price on the vendor site before budgeting.

Not Sure Which Routing Setup Fits Your Compliance Requirements?

Cost is only half the decision. Our free AI workflow audit maps the right routing layer, model allowlist, and data-residency controls to your exact stack and regulatory obligations.

Book Your Free AI Routing Audit