Model card
Llama-4-Scout 17B is a specialized Mixture-of-Experts (MoE) model designed to balance high-density knowledge with efficient inference. While the total parameter count sits at 109B, the architecture only activates 17B parameters per token, making it significantly faster and more cost-effective for real-time applications than dense models of a similar scale. For developers, the standout feature is its native multimodality; you can pass image and text inputs through a single unified interface without needing separate vision encoders. With a massive 1.3M token context window, it is purpose-built for long-form document analysis, complex codebase reasoning, and large-scale data extraction. Unlike previous iterations that required heavy fine-tuning for specific tasks, Scout's instruction-tuned weights are optimized for direct integration into RAG pipelines and agentic workflows where low latency and high reasoning accuracy are critical.
Model files and versions
Download this model
How to use
- 01Step 1
Read the model card and source information.
- 02Step 2
Start with a small, non-sensitive evaluation.
- 03Step 3
Review quality, licensing and usage limits.
- 04Step 4
Adopt it only after validation.
Discussions
Use this space to keep checking source information, usage experience and maintenance status.
Open source page