Model card
Nemotron-3.5-Lightning is NVIDIA’s high-efficiency Mixture-of-Experts (MoE) model designed specifically for low-latency, high-throughput environments. While the total parameter count sits at 30B, it only utilizes 3B active parameters per token, offering a massive performance leap for developers needing rapid inference without the computational overhead of dense large-scale models. For engineers building agentic workflows, tool-calling loops, or real-time RAG pipelines, this model strikes a pragmatic balance between reasoning depth and execution speed. Unlike general-purpose heavyweights, Lightning is optimized for specialized task execution and high-frequency API calls. It integrates seamlessly into existing NVIDIA-optimized stacks and is particularly effective when you need to scale agentic reasoning across thousands of concurrent sessions where traditional LLMs would become a cost or latency bottleneck.
Model files and versions
Download this model
How to use
- 01Step 1
Read the model card and source information.
- 02Step 2
Start with a small, non-sensitive evaluation.
- 03Step 3
Review quality, licensing and usage limits.
- 04Step 4
Adopt it only after validation.
Discussions
Use this space to keep checking source information, usage experience and maintenance status.
Open source page