Model card
Inkling-small is a specialized multimodal Mixture-of-Experts (MoE) model designed for developers who need high-reasoning capabilities without the massive compute overhead of dense trillion-parameter models. While it sits within a 276B total parameter architecture, it only activates 12B parameters per token, offering a highly efficient inference profile that balances throughput with intelligence. For international engineering teams, this means you can deploy sophisticated multimodal workflows—combining vision and text—on more modest hardware compared to traditional dense models. It excels in complex reasoning tasks and long-context applications, supported by a massive 1M token context window. Whether you are integrating it via API for rapid prototyping or leveraging its open-weight nature for fine-tuning, inkling-small provides a scalable middle ground between lightweight edge models and heavy-duty frontier LLMs.
Model files and versions
Download this model
How to use
- 01Step 1
Read the model card and source information.
- 02Step 2
Start with a small, non-sensitive evaluation.
- 03Step 3
Review quality, licensing and usage limits.
- 04Step 4
Adopt it only after validation.
Discussions
Use this space to keep checking source information, usage experience and maintenance status.
Open source page