Model card
Nemotron-3-Super-120B-A12B is a high-efficiency Mixture-of-Experts (MoE) model designed for developers building complex, multi-agent workflows. While it maintains a 120B parameter footprint, its hybrid Mamba-Transformer architecture ensures that only 12B parameters are active during inference. This design provides a critical sweet spot: the reasoning depth of a large-scale model with the low latency and compute cost typically associated with much smaller models. For developers, this means you can deploy sophisticated agentic logic and long-context reasoning without the prohibitive hardware overhead. The architecture is particularly optimized for tasks requiring high precision in instruction following and structured data generation. Whether you are integrating it via API for real-time applications or fine-tuning it for specialized reasoning tasks, the model offers a scalable path for moving from simple chatbots to autonomous multi-agent systems.
Model files and versions
Download this model
How to use
- 01Step 1
Read the model card and source information.
- 02Step 2
Start with a small, non-sensitive evaluation.
- 03Step 3
Review quality, licensing and usage limits.
- 04Step 4
Adopt it only after validation.
Discussions
Use this space to keep checking source information, usage experience and maintenance status.
Open source page