Model card
Nemotron-3-Super-120B is a high-efficiency hybrid architecture designed specifically for complex, multi-agent workflows. Unlike standard dense models, it utilizes a Mixture-of-Experts (MoE) approach that activates only 12B parameters per token, offering a massive 120B parameter knowledge base with the latency and compute footprint of a much smaller model. What sets this apart for developers is its hybrid Mamba-Transformer backbone, which addresses the quadratic scaling issues of traditional attention mechanisms, making it highly effective for processing long-context reasoning tasks. For those building autonomous agent loops or RAG pipelines, this model provides a sweet spot between high-fidelity instruction following and rapid inference speeds. It is particularly well-suited for integration into orchestration frameworks where low-latency decision-making and high reasoning accuracy are non-negotiable.
Model files and versions
Download this model
How to use
- 01Step 1
Read the model card and source information.
- 02Step 2
Start with a small, non-sensitive evaluation.
- 03Step 3
Review quality, licensing and usage limits.
- 04Step 4
Adopt it only after validation.
Discussions
Use this space to keep checking source information, usage experience and maintenance status.
Open source page