Reviewed by Jonathan West · Updated Aug 14, 2026

Nemotron 3.5 Lightning Pricing

What it costs to run NVIDIA's open agent model, two ways.

Reviewed by Jonathan West · Updated Aug 14, 2026

Nemotron 3.5 Lightning has no license fee, because it ships with open weights. Self-hosting costs $0 to license; you pay only for hardware and operations.

If you use a hosted endpoint instead, you pay per token. That covers the compute someone else runs for you.

This guide breaks down both cost models. It explains who each one fits, and where the real spending hides. We will not quote a per-token rate, because those change; check the live rate before you budget.


The Two Cost Models

Nemotron 3.5 Lightning has two cost paths: self-host or hosted API. The license is free either way, because the weights are open under OpenMDW-1.1.

Self-hosting means you download the weights and run them on your own GPUs. There is no per-token license charge. Your cost is hardware, power, and staff time.

A hosted endpoint means a provider runs the model and bills you per token. You trade a running fee for zero setup and instant scale.

Both paths use the same open weights, so quality does not change between them. What changes is who owns the hardware and the ops work.

That makes the choice a pure cost and control question. You are not picking a better or worse model, just a better fit for your team.

Trying to decide if Nemotron 3.5 Lightning is cheaper self-hosted or hosted? We build cost models for open-weight deployments.

Book a Consultation

What Self-hosting Really Costs

Self-hosting is free to license but not free to run. Your bill is the hardware and the people who keep it healthy.

You buy or rent GPUs, then pay to power and cool them. Those costs run whether the model is busy or idle.

The model has 30 billion total parameters. You need enough GPU memory to load it, though only 3 billion activate per token at runtime.

NVIDIA ships an NVFP4 4-bit checkpoint alongside the BF16 one. The 4-bit version needs far less memory, so it can run on smaller or cheaper hardware.

Add power, cooling, and engineering time to the total. At high, steady volume, owning the hardware often beats per-token fees.


What Hosted Endpoints Cost

Hosted endpoints charge per token, split between input and output. You pay for what you send and what the model returns.

NVIDIA offers a hosted endpoint on build.nvidia.com. Third-party providers may also host the open weights and set their own rates.

Do not trust any fixed price you read in a guide, including this one. Rates move. Confirm the current number on build.nvidia.com or your provider before you plan a budget.

For hosted use, your monthly cost tracks your token volume. Long prompts and long outputs cost more, so a 1 million token context can add up fast.


Which Cost Model Fits You

Choose self-host if you run high, steady volume or need data to stay in house. The fixed hardware cost gets cheaper per call as usage climbs.

Choose a hosted endpoint if your volume is low, spiky, or just starting. You skip setup and pay only for what you use.

Hybrid setups also work. Some teams self-host their steady base load and burst to a hosted endpoint at peak times.

Many teams start hosted to test, then move to self-host once volume justifies the hardware. The open license makes that switch easy, since the weights are the same.

Across the model launches we track, teams most often overpay by staying on hosted per-token billing well past the point where owned hardware would be cheaper. Re-run that math as your volume grows.


How the Checkpoint Choice Changes Cost

Your checkpoint choice moves your self-hosting cost. NVIDIA ships two: NVFP4 at 4-bit and BF16 at 16-bit.

The NVFP4 checkpoint uses less GPU memory, so it fits on smaller and cheaper cards. That lowers the hardware you must buy or rent.

The BF16 checkpoint keeps full precision and drives the published benchmark scores. It costs more memory, so it needs bigger hardware.

For many agent tasks, the 4-bit version runs well at a lower hardware bill. Test both on your workload to see if the quality gap matters for you.


Costs People Forget to Count

The token or hardware bill is not your whole cost. A few line items surprise teams later.

Long context is a real expense. Filling the 1 million token window on a hosted endpoint means paying for every one of those input tokens.

Self-host adds ops work: updates, monitoring, and failover. That is staff time, not a software fee.

Speed can save money. NVIDIA claims up to 4x output speed versus similar-sized models, which can cut GPU hours per task. Verify that claim against your own workload.


A Simple Way to Budget

Budget in two steps: estimate your token volume, then price both paths against it. That gives you a fair side-by-side.

First, estimate tokens per task and tasks per month. Multiply to get your monthly token load, including long-context prompts.

Second, price the hosted path at the live rate on build.nvidia.com. Then price the self-host path as hardware plus power plus staff time.

Compare the two totals at your real volume. Re-run the math each quarter, since both your usage and vendor rates will change.

Do not forget a ramp period. Early on, hosted billing avoids upfront hardware spend while you learn your true volume.

Once volume is steady and large, the self-host total often wins. The open license lets you switch without changing the model.

Frequently Asked Questions

  • Self-hosting Nemotron 3.5 Lightning costs $0 in license fees, because the weights are open under OpenMDW-1.1; you pay only for hardware and operations. Hosted endpoints on build.nvidia.com or third-party providers charge per token, so check the live rate there before budgeting.
  • The weights are free to download and self-host under the OpenMDW-1.1 license. It is not free to run, because you still pay for the GPUs, power, and staff time, or for per-token usage on a hosted endpoint.
  • Per-token prices are set by each hosted provider and change over time, so no fixed number here would stay accurate. Check the current input and output rates on build.nvidia.com or your chosen provider before you plan spend.
  • Self-hosting is usually cheaper at high, steady volume, because the fixed hardware cost spreads across many calls. Hosted APIs are usually cheaper for low or spiky usage, since you pay only for the tokens you use.
  • Yes, the NVFP4 4-bit checkpoint needs far less GPU memory than the BF16 version. That lets you run the model on smaller or cheaper hardware, which lowers self-hosting cost.

Model the true cost before you commit

Self-host or hosted is a math problem, and the wrong call gets expensive. We build the cost model for your real volume so you pick with confidence. Book a call to run the numbers.

Book a Call