Model card
Inkling-small is a high-efficiency multimodal Mixture-of-Experts (MoE) model designed for developers who need large-scale reasoning capabilities without the massive compute overhead of dense models. While it sits on a 276B total parameter architecture, it utilizes only 12B active parameters per token, making it highly optimized for low-latency inference and high-throughput API integration. For developers working with complex multimodal datasets, its standout feature is the massive 1M token context window, which allows for deep document analysis and long-form reasoning that standard small models cannot handle. Compared to dense 7B or 13B models, Inkling-small offers a significantly higher intelligence ceiling by leveraging its MoE structure, making it an ideal candidate for RAG pipelines, long-context summarization, and multimodal agentic workflows where cost-to-performance ratios are critical.
Model files and versions
Download this model
How to use
- 01Step 1
Read the model card and source information.
- 02Step 2
Start with a small, non-sensitive evaluation.
- 03Step 3
Review quality, licensing and usage limits.
- 04Step 4
Adopt it only after validation.
Discussions
Use this space to keep checking source information, usage experience and maintenance status.
Open source page