Is Nemotron 3.5 Lightning Free?
Open weights, commercial use, and the real bills you still pay.
Yes, Nemotron 3.5 Lightning is free to download and self-host, including for commercial use. NVIDIA released it as an open-weight model under the OpenMDW-1.1 license on August 11, 2026.
Free here means the weights cost nothing and the license permits business use. It does not mean running the model is cost-free, because you still pay for hardware and operations.
This guide explains what is free, what costs money, and what the OpenMDW-1.1 license actually allows in plain English.
The Short Answer
Nemotron 3.5 Lightning is free to use if you host it yourself. The weights are open and the license permits commercial use.
NVIDIA published two checkpoints, NVFP4 and BF16, on Hugging Face. Anyone can download them at no charge.
The company went further than weights alone. It also released training data and recipes under the same open framing.
So the model itself carries no license fee. Your only costs come from the compute you run it on and the people who run it.
The catch is small but real. Free to license does not mean free to operate, and the next sections separate the two clearly.
Deciding whether free self-hosted Nemotron 3.5 Lightning beats a paid endpoint for you? We build the cost model on your real volume.
Book a ConsultationWhat Is Actually Free
Three things are free: the model weights, the training data, and the training recipes. That is an unusually open release.
The weights are the model itself. You download them and run inference without paying NVIDIA per token.
The data and recipes let researchers study and reproduce the work. That transparency helps teams trust and adapt the model.
Free also covers commercial products under the license terms. You can build a paid app on top of the self-hosted model.
Both the NVFP4 and BF16 checkpoints are free to download. You choose the format that fits your hardware without any change in price.
Fine-tuning is free too. You can adapt the weights to your own data and keep the specialized model private, all without a usage fee.
What Still Costs Money
Two things cost money: the compute to run the model and the hosted API if you skip self-hosting. Neither is a license fee.
Self-hosting needs NVIDIA GPUs, power, and engineers to keep the service healthy. Those are real, ongoing costs.
The hosted endpoint on build.nvidia.com charges per token. You trade GPU ownership for a usage bill.
Third-party providers may also host the model for a per-token rate. Free weights do not make hosted tokens free.
There are softer costs too. Setting up serving, monitoring, and updates takes engineering time you should budget for.
The BF16 checkpoint needs more memory than the NVFP4 4-bit version. Your checkpoint choice changes the hardware you must buy or rent.
OpenMDW-1.1 in Plain English
OpenMDW-1.1 is an open model license that lets you use, modify, and redistribute the model, including for commercial work. It is designed to be permissive.
In practice, you can download the weights, fine-tune them, and ship a product. The license aims to keep those rights broad.
As with any license, you should read the exact terms before you build a business on it. Terms can carry conditions on attribution or redistribution.
If your use is large or regulated, have counsel review the full text. A five-minute read now prevents a costly surprise later.
OpenMDW-1.1 is also a step past the NVIDIA Open Model License used for the earlier Nemotron 3 family. The direction of travel is toward broader, simpler openness.
Free to License Is Not Free to Run
A free license and a free deployment are different things. Confusing them is the most common budgeting mistake we see.
Across the model launches we track, teams often assume open weights mean zero cost, then meet a large GPU bill. The license is free; the electrons are not.
The real comparison weighs self-host compute against hosted per-token fees at your real volume. One wins for steady heavy use, the other for bursty light use.
Model the numbers for your traffic before you commit. The right answer depends on volume, privacy needs, and team capacity.
A quick way to sanity-check is to price a month of expected tokens at the hosted rate, then compare it to the monthly cost of a GPU that could serve the same load. The cheaper path is your answer.
Can You Build a Paid Product on It?
Yes, you can build a paid product on Nemotron 3.5 Lightning under the OpenMDW-1.1 license. Commercial use is permitted, including inside a self-hosted app.
Many businesses want an open model precisely for this reason. They can ship a product without a per-seat or per-token fee flowing to the model vendor.
You can also fine-tune the weights on your own data. That lets you specialize the model for your domain and keep the result private.
Before launch, confirm any attribution or redistribution conditions in the license text. Meeting a simple attribution rule is easy once you know it exists.
For a regulated product, loop in legal review early. The license is permissive, but your industry may add its own rules on model use.
How Open This Release Really Is
This release is more open than most, because NVIDIA shared weights, data, and recipes together. Many open-weight models share only the weights.
Open weights let you run and fine-tune the model. Open data and recipes let you study how it was built and reproduce parts of the work.
That transparency matters for trust. Teams in careful fields can inspect the ingredients instead of taking a black box on faith.
It also compares favorably to the prior Nemotron 3 line, which used the NVIDIA Open Model License. The move to OpenMDW-1.1 signals a broader openness push.
Openness is not the same as free support. You own the operations, so plan for the engineering time any self-hosted model requires.
For most buyers, the practical takeaway is simple. You get real freedom to use, change, and ship the model, in exchange for owning the cost of running it.
When Free Self-Hosting Wins
Self-hosting wins when your volume is steady, your data is sensitive, or you already own GPUs. Then the free license turns into real savings.
High steady throughput spreads hardware cost across many requests. Per-token fees would add up faster.
Strict privacy rules favor keeping data on your own machines. Self-hosting keeps prompts off a third-party endpoint.
If none of those apply, the hosted API may be cheaper and simpler. Free weights are a tool, not an obligation.
The decision is rarely permanent. Teams often start on the hosted API, then move to self-hosting once volume and confidence grow.
You can also run both at once. Serve steady baseline traffic on your own GPUs and send overflow spikes to the hosted API, so you pay for peak capacity only when you use it.
Frequently Asked Questions
- Yes. Nemotron 3.5 Lightning is free to download and self-host under the OpenMDW-1.1 license, including for commercial use. You pay only for hardware, operations, or a hosted endpoint.
- Yes, the OpenMDW-1.1 license permits commercial use of the self-hosted model. Read the full license terms before building a product, since conditions can apply to attribution or redistribution.
- Nemotron 3.5 Lightning uses OpenMDW-1.1, an open model license. NVIDIA also released the training data and recipes, making this an unusually open release.
- The license is free, but running the model is not. Self-hosting needs GPUs and engineers, and the hosted API on build.nvidia.com charges per token.
- The NVFP4 and BF16 checkpoints are on Hugging Face under the nvidia organization. Both are free to download and run on NVIDIA hardware.
Model the True Cost Before You Commit
Open weights are free, but running them is not. At Layer3Labs, we build the self-host versus hosted math around your real traffic so the choice is clear.
Book a Consultation