LLM Routing Cost Calculator
Compare what AI model routing costs you versus going direct. See the markup, and the compliance risk your routing layer can hide.
An AI routing layer like OpenRouter, LiteLLM, or Portkey sits between your app and the model providers. It buys you convenience — one API for many models — but that convenience has a price, and the price is not always the sticker rate. This tool estimates your monthly cost across each routing option for your own token volume and model mix, then flags the compliance exposure most teams never price in. For the full picture, start with AI model routing explained.
Compare Your Routing Cost
Not sure about your volume? A support agent answering ~2,000 tickets a month with a 1,000-token context each is roughly 2M input / 1M output tokens — scale from there. With several models selected, your volume is split evenly across them.
| Model | Direct API | OpenRouter | LiteLLM | Portkey | Cheapest |
|---|---|---|---|---|---|
| GPT-4.1 | $3.00 | $3.17 +5.5% | $28.00 +833% | $27.50 +817% | Direct API |
| Claude Sonnet 5 | $5.25 | $5.54 +5.5% | $30.25 +476% | $29.75 +467% | Direct API |
Estimate only. Token prices are list rates last verified July 2026 and exclude prompt-caching discounts; OpenRouter, LiteLLM infra, and Portkey figures are configurable assumptions, not quotes. Models marked * are provider-dependent and approximate. Confirm current rates on each vendor's pricing page — see the OpenRouter pricing and LiteLLM pricing breakdowns — before you commit spend.
How the Estimate Works
You set two numbers — your monthly input and output tokens — and pick the models you use. The calculator prices each model at its published per-million-token rate, splits your volume evenly across the models you selected, and totals the monthly cost under four routing options.
- Direct API — provider list rates, no markup. The baseline every other column is measured against.
- OpenRouter — the same per-token rates plus an estimated credit fee on spend. See the OpenRouter pricing breakdown.
- LiteLLM — provider rates plus the monthly infrastructure cost you enter to self-host the proxy. See LiteLLM pricing.
- Portkey — provider rates plus a managed-gateway platform fee.
Where the Markup Hides
At low volume, a hosted gateway's per-token pass-through pricing is hard to beat — the fixed cost of self-hosting LiteLLM is not worth it for a few dollars of traffic. As volume grows, that flips: the marketplace fee scales with spend while a self-hosted proxy's infrastructure cost stays flat, so LiteLLM pulls ahead. Enter your real numbers above to find your own crossover point rather than trusting a generic rule of thumb. The OpenRouter vs LiteLLM comparison walks through the trade-off in detail.
The Hidden Compliance Cost
The cheapest model at a given quality tier is often a Chinese-developed one — DeepSeek or Qwen. A cost-optimizing router will reach for it unless you explicitly block it, which means a single API call can send regulated data to infrastructure governed by Chinese data law. That is not a line item on any pricing page, but it is the most expensive mistake in the whole stack.
Frequently Asked Questions
- OpenRouter charges the same per-token rate as the underlying provider for most models, then takes a fee on the credits you buy (roughly 5% plus payment processing). So the markup is on your total spend, not each token. Use the calculator above to see the effect on your model mix, and verify current terms on openrouter.ai/pricing.
- The open-source proxy is free to run. You still pay the underlying model provider token costs, plus your own infrastructure — a server to host it, monitoring, and engineer time. The calculator lets you set that infrastructure cost. LiteLLM also sells a managed enterprise tier for a fee.
- Direct API access is cheapest per token because there is no middleman. The trade-off is managing a separate integration for each provider. Self-hosted LiteLLM is usually the cheapest unified option once your volume covers the fixed infrastructure cost.
- Yes. If your routing tool can reach Chinese-developed models such as DeepSeek or Qwen by default, a cost-optimized or misconfigured request can send sensitive data to servers governed by Chinese data laws (PIPL). The calculator flags any Chinese-hosted models you select. Audit your model allowlist before you route production traffic.
- Frequently — major providers adjust pricing several times a year, usually downward. This calculator uses list rates last verified in July 2026 and is updated quarterly. Always confirm the current price on the vendor site before budgeting.
Not Sure Which Routing Setup Fits Your Compliance Requirements?
Cost is only half the decision. Our free AI workflow audit maps the right routing layer, model allowlist, and data-residency controls to your exact stack and regulatory obligations.
Book Your Free AI Routing Audit