Model card
gpt-oss-120b is a high-density Mixture-of-Experts (MoE) model designed to bridge the gap between massive parameter counts and production-grade inference efficiency. By activating only 5.1B parameters per token, it offers a unique value proposition: the reasoning capabilities of a large-scale model with the low latency typically associated with much smaller architectures. For developers, this means improved throughput for agentic workflows and complex multi-step reasoning tasks without the massive compute overhead of dense 100B+ models. With a 131k context window, it is well-suited for long-form document analysis, RAG pipelines, and maintaining state in complex autonomous agents. Unlike standard dense models, its MoE structure allows for specialized knowledge retrieval during the forward pass, making it particularly effective for coding, mathematical reasoning, and structured data extraction. It is built for integration into existing API-driven stacks where high reliability and reasoning depth are non-negotiable.
Model files and versions
Download this model
How to use
- 01Step 1
Read the model card and source information.
- 02Step 2
Start with a small, non-sensitive evaluation.
- 03Step 3
Review quality, licensing and usage limits.
- 04Step 4
Adopt it only after validation.
Discussions
Use this space to keep checking source information, usage experience and maintenance status.
Open source page