Model card
inkling:batch is a massive-scale multimodal Mixture-of-Experts (MoE) model from Thinking Machines Lab, engineered for high-throughput reasoning and complex agentic workflows. While the total parameter count sits at 975B, the architecture optimizes efficiency by activating only 41B parameters per token, making it a competitive choice for developers needing deep logic without the latency of dense models. It excels in code generation, multi-step tool use, and long-context reasoning, supported by a substantial 524k context window. For teams building autonomous agents or RAG pipelines, inkling:batch offers a robust middle ground: the intelligence of a frontier-class model with the specialized efficiency of an MoE structure. It is primarily accessible via API, making it easy to integrate into existing production environments that require reliable, scalable reasoning capabilities.
Model files and versions
Download this model
How to use
- 01Step 1
Read the model card and source information.
- 02Step 2
Start with a small, non-sensitive evaluation.
- 03Step 3
Review quality, licensing and usage limits.
- 04Step 4
Adopt it only after validation.
Discussions
Use this space to keep checking source information, usage experience and maintenance status.
Open source page