Model card
The Qwen3.6 35B A3B NVFP4 is a specialized iteration of the Qwen series, optimized specifically for NVIDIA hardware using the NVFP4 quantization format. For developers, the primary draw here is the efficiency gain; by leveraging 4-bit floating point precision, this model significantly reduces VRAM overhead without the drastic perplexity loss typically seen in integer quantization. It is designed for high-throughput text generation and complex reasoning tasks where latency is critical. Integration is streamlined for NVIDIA TensorRT-LLM environments, making it an ideal candidate for production-grade RAG pipelines or agentic workflows where you need the intelligence of a mid-sized model but the speed of a much smaller one.
Model files and versions
Download this model
We recommend using the ModelScope CLI or SDK. Install ModelScope first, then choose a full snapshot, single file, SDK or Git LFS workflow.
nvidia/Qwen3.6-35B-A3B-NVFP4Install the CLI and SDK dependency before downloading.
pip install modelscopeDownload the complete weights, configuration and model card.
modelscope download --model nvidia/Qwen3.6-35B-A3B-NVFP4README.md is used as an example; replace it with another repository file when needed.
modelscope download --model nvidia/Qwen3.6-35B-A3B-NVFP4 README.md --local_dir ./dirUseful in Python projects and automation scripts.
from modelscope import snapshot_download
model_dir = snapshot_download('nvidia/Qwen3.6-35B-A3B-NVFP4')Make sure Git LFS is installed correctly.
git lfs install
git clone https://www.modelscope.cn/nvidia/Qwen3.6-35B-A3B-NVFP4.gitFetch the repository structure first, then pull large files when needed.
GIT_LFS_SKIP_SMUDGE=1 git clone https://www.modelscope.cn/nvidia/Qwen3.6-35B-A3B-NVFP4.gitHow to use
- 01Step 1
Read the model card and source information.
- 02Step 2
Start with a small, non-sensitive evaluation.
- 03Step 3
Review quality, licensing and usage limits.
- 04Step 4
Adopt it only after validation.
Discussions
Use this space to keep checking source information, usage experience and maintenance status.
Open source page