Model card
Mercury-2 introduces a fundamental shift in inference architecture by moving away from traditional autoregressive token generation. As the first reasoning diffusion LLM (dLLM), it utilizes a parallel refinement process rather than sequential prediction. For developers, this means a significant reduction in time-to-first-token and overall latency, particularly for complex reasoning tasks that typically bottleneck standard models. The 128k context window makes it highly capable for long-form document analysis, code refactoring, and multi-step logic chains. Unlike standard LLMs that struggle with the 'thinking' overhead, Mercury-2's diffusion-based approach allows it to refine its internal logic mid-generation. It is designed for high-throughput environments where speed and reasoning depth must coexist, making it an ideal candidate for real-time agentic workflows and complex automated debugging pipelines via API integration.
Model files and versions
Download this model
How to use
- 01Step 1
Read the model card and source information.
- 02Step 2
Start with a small, non-sensitive evaluation.
- 03Step 3
Review quality, licensing and usage limits.
- 04Step 4
Adopt it only after validation.
Discussions
Use this space to keep checking source information, usage experience and maintenance status.
Open source page