Model card
Inkling:free is a high-efficiency multimodal Mixture-of-Experts (MoE) model designed for developers building complex, agentic workflows. While the total parameter count sits at 975B, the architecture optimizes performance by utilizing only 41B active parameters per token, offering a massive knowledge base with the inference speed typically associated with much smaller models. For engineers, the standout feature is the massive 1M+ token context window, which makes it ideal for deep codebase analysis, long-form document reasoning, and maintaining state in multi-turn agentic loops. Unlike dense models that scale latency linearly with parameter count, Inkling provides a pragmatic middle ground for tool-use and automated reasoning tasks. It is particularly well-suited for integration into RAG pipelines and autonomous coding assistants where both high-level reasoning and rapid response times are critical requirements.
Model files and versions
Download this model
How to use
- 01Step 1
Read the model card and source information.
- 02Step 2
Start with a small, non-sensitive evaluation.
- 03Step 3
Review quality, licensing and usage limits.
- 04Step 4
Adopt it only after validation.
Discussions
Use this space to keep checking source information, usage experience and maintenance status.
Open source page