Global AI chat room · 14 online now Join now
DIRECTORY / 02

AI Models | Open-Source LLM Directory

Discover and compare open-source LLMs, language models and multimodal models by capability, scale, license, downloads and provenance.

Compare modelsFind the right building block for your next workflow
Directory overview
591
curated entries
24 topic groupsLive
02 / MODEL INDEX

Find the right model for the job

Context first, better decisions. Every entry keeps the signal that matters.

CURATED DIRECTORY591 results

relace-apply-3

relace
256000 ctx

Relace-apply-3 is a specialized utility model designed to solve the 'last mile' problem in AI-assisted development: the manual effort of applying LLM-generated diffs to local source code. Unlike general-purpose models that simply output code blocks, this model is optimized for precise code-patching. It acts as a bridge between high-level reasoning models—like GPT-4o or Claude—and your actual filesystem. By ingesting suggested edits and merging them directly into existing files, it minimizes the risk of syntax errors and manual copy-paste mistakes. For developers building automated refactoring tools or IDE extensions, it provides a reliable mechanism to transform abstract suggestions into concrete, file-level updates. It is particularly useful in CI/CD pipelines or agentic workflows where autonomous code modification is required without human intervention.

text generationAPI

cydonia-24b-v4.1

thedrummer
131072 ctx

Cydonia-24b-v4.1 is a specialized fine-tune of the Mistral Small 3.2 architecture, optimized specifically for high-fidelity creative writing and complex instruction following. For developers building agentic workflows or narrative-driven applications, this model bridges the gap between lightweight 7B models and heavy-duty 70B+ parameters. It excels in scenarios where strict adherence to stylistic constraints and long-context recall are critical, making it a strong candidate for roleplay engines, automated storytelling, and nuanced content generation. Unlike standard safety-tuned models that often trigger false positives during creative tasks, Cydonia is designed to maintain narrative momentum without excessive filtering. With a 131k context window, it is well-suited for processing large document sets or maintaining deep continuity in long-form dialogue. It offers a high intelligence-to-latency ratio, providing a performant middle ground for developers who need sophisticated reasoning without the massive compute overhead of larger frontier models.

text generationAPI

deepseek-v3.2-exp

deepseek
163840 ctx

DeepSeek-V3.2-Exp is an experimental bridge model designed to test architectural refinements before the next major release. For developers, the primary technical differentiator is the introduction of DeepSeek Sparse Attention (DSA). This fine-grained mechanism aims to optimize computational efficiency and context handling, potentially offering a better performance-to-latency ratio than previous iterations. While it serves as an intermediate step, it remains a robust tool for complex text generation and reasoning tasks. Integration is straightforward via API, making it suitable for testing high-throughput applications where attention mechanism efficiency is critical. Compared to its predecessors, expect more nuanced long-context management and improved throughput, though as an 'experimental' release, it is best utilized for benchmarking and iterative development rather than mission-critical production stability.

text generationAPI

glm-4.6

z-ai
204800 ctx

GLM-4.6 is the latest iteration in the GLM series, specifically engineered to address the scaling limitations of its predecessors. For developers working with massive datasets or complex codebase analysis, the most significant upgrade is the expansion of the context window to 200K tokens. This allows for much deeper reasoning over long-form documentation and larger repository structures without the typical loss of coherence seen in smaller windows. While previous versions established a strong baseline for multilingual tasks, 4.6 focuses on improving instruction following and reducing latency in high-throughput production environments. It is designed as an API-first model, making it easy to integrate into existing RAG (Retrieval-Augmented Generation) pipelines or agentic workflows where long-term memory and precise context retrieval are critical. If your use case involves summarizing extensive legal documents or maintaining state across lengthy multi-turn dialogues, this model offers a more robust architectural foundation than the 128K-limited 4.5 version.

text generationAPI

qwen3-vl-30b-a3b-instruct

qwen
262144 ctx

For developers building vision-centric applications, qwen3-vl-30b-a3b-instruct represents a significant step forward in multimodal reasoning. Unlike standard LLMs that rely on external vision encoders, this model unifies visual perception with text generation, allowing for more nuanced understanding of both static images and temporal video sequences. The 30B parameter scale strikes a balance between high-level reasoning capabilities and deployment efficiency, making it suitable for complex tasks like visual question answering (VQA), document parsing, and automated video captioning. Its instruction-tuned architecture is specifically optimized for following multi-step prompts, which is critical for integrating the model into agentic workflows where visual input drives decision-making. Whether you are building sophisticated OCR pipelines or interactive visual assistants, this model offers the low-latency responsiveness and high contextual accuracy required for production-grade multimodal integrations.

text generationAPI

qwen3-vl-30b-a3b-thinking

qwen
262144 ctx

For developers building vision-centric applications, Qwen3-VL-30B-A3B-Thinking represents a significant step forward in multimodal reasoning. Unlike standard vision-language models that often struggle with spatial logic or multi-step visual deduction, this 'Thinking' variant integrates a specialized reasoning trace to handle complex STEM problems and intricate video analysis. It is designed for high-precision tasks where the model must not only describe an image but interpret the underlying logic—such as solving mathematical equations from a whiteboard or debugging code from a screen recording. With a 262k context window, it is well-suited for long-form video understanding and large-scale document processing. For integration, the API-first approach allows for seamless deployment into existing pipelines, offering a competitive middle ground between lightweight vision models and massive, computationally expensive frontier models.

text generationAPI

qwen3-vl-8b-instruct

qwen
262144 ctx

Qwen3-VL-8B-Instruct is a lightweight yet highly capable multimodal model designed for developers needing efficient vision-language reasoning. Unlike standard LLMs, this model utilizes Interleaved-MRoPE to maintain spatial and temporal coherence, making it particularly effective for tasks involving long-form video analysis and complex document understanding. For developers, the 8B parameter footprint offers a sweet spot: it provides enough reasoning depth for high-fidelity image captioning and visual question answering (VQA) while remaining computationally accessible for low-latency applications or edge deployments. Compared to previous iterations, the improved multimodal fusion allows for better integration of interleaved text and visual data, reducing the 'hallucination' effect in complex spatial reasoning tasks. It is an ideal candidate for building intelligent visual agents, automated video indexing tools, or advanced OCR pipelines where context across frames or dense visual layouts is critical.

text generationAPI

qwen3-vl-8b-thinking

qwen
131072 ctx

For developers building vision-centric applications, Qwen3-VL-8B-Thinking represents a significant shift toward multimodal reasoning rather than simple pattern recognition. While standard VL models excel at captioning, this variant is specifically tuned to handle complex visual logic, such as interpreting intricate document layouts, analyzing temporal changes in video sequences, and solving spatial reasoning tasks. It bridges the gap between 'seeing' and 'understanding' by integrating a dedicated thinking process that allows the model to decompose visual queries before generating a response. At 8B parameters, it offers a high performance-to-latency ratio, making it an ideal candidate for edge-integrated workflows or high-throughput agentic loops where reasoning depth is required without the overhead of a massive parameter count. Whether you are automating document extraction or building sophisticated visual agents, this model provides the granular logical framework necessary for high-accuracy multimodal deployments.

text generationAPI

granite-4.0-h-micro

ibm-granite
131000 ctx

Granite-4.0-H-Micro is a specialized 3B parameter model from IBM's latest Granite family, engineered specifically for high-efficiency text generation tasks. For developers working in resource-constrained environments or building low-latency pipelines, this model offers a strategic balance between a small memory footprint and robust reasoning capabilities. Unlike larger general-purpose models, the 'Micro' architecture is optimized for speed and throughput without sacrificing the contextual depth required for enterprise workflows. It supports a substantial 131,000 token context window, making it an ideal candidate for long-form document analysis, complex RAG (Retrieval-Augmented Generation) implementations, and automated summarization. Integration is straightforward via API, allowing you to deploy sophisticated NLP features into edge computing or microservice architectures where minimizing inference costs and latency is critical. If your roadmap requires a lightweight, scalable model that handles extended context better than typical small-scale LLMs, this is a highly competitive option.

text generationAPI

qwen3-vl-32b-instruct

qwen
131072 ctx

Qwen3-VL-32B-Instruct is a high-density multimodal model engineered for developers needing a balance between reasoning depth and inference efficiency. Moving beyond simple OCR, this 32B parameter model excels at complex visual reasoning, temporal understanding in video streams, and high-resolution document parsing. For engineers building agentic workflows, its ability to map visual spatial coordinates to text makes it a strong candidate for UI automation and robotic vision tasks. Compared to larger flagship models, it offers a significantly better performance-to-latency ratio, making it suitable for real-time applications like visual QA or automated content moderation. The model supports a massive 131k context window, allowing for long-form video analysis and multi-image document processing within a single prompt. Integration is straightforward via API, making it a viable drop-in upgrade for existing vision-language pipelines requiring higher precision in structured data extraction.

text generationAPI

minimax-m2

minimax
204800 ctx

MiniMax-M2 is a high-efficiency MoE (Mixture-of-Experts) model designed specifically for developers building autonomous agents and complex coding pipelines. While it boasts a massive 230B total parameter architecture, it only activates 10B parameters per token, striking a strategic balance between frontier-level reasoning and low-latency execution. For engineers, this means you get the cognitive depth required for multi-step logic and code generation without the prohibitive inference costs typically associated with massive dense models. Its 204k context window makes it a viable candidate for RAG-heavy applications and analyzing entire codebases. Unlike general-purpose chat models that prioritize conversational fluff, M2 is tuned for the structured, iterative reasoning required in agentic workflows, making it a strong competitor for developers looking to deploy reliable, task-oriented AI agents via API.

text generationAPI

voxtral-small-24b-2507

mistralai
32768 ctx

Voxtral-small-24b-2507 is a specialized multimodal evolution of the Mistral Small 3 architecture, specifically engineered for developers building audio-native applications. While it maintains the high-reasoning text performance expected from the Mistral lineage, its core differentiator is the integrated native audio input layer. Unlike traditional pipelines that rely on a separate Whisper-style STT model followed by a text LLM, Voxtral processes raw audio signals directly. This reduces latency and preserves prosodic nuances—like tone and emotion—that are often lost in standard transcription. For developers, this means more seamless integration for real-time translation, complex audio summarization, and voice-driven agentic workflows. It operates within a 32k context window, making it suitable for long-form speech analysis. If your roadmap involves moving beyond simple text prompts into sophisticated voice interfaces or automated meeting intelligence, this model offers a more cohesive architectural approach than decoupled speech-to-text systems.

text generationAPI

sonar-pro-search

perplexity
200000 ctx

Sonar Pro Search represents a shift from simple retrieval to agentic reasoning. Unlike standard LLMs that rely on static training data, this model functions as an autonomous research agent designed to navigate the live web, synthesize multi-step queries, and verify information through iterative searching. For developers, this means moving beyond basic RAG pipelines toward true autonomous research workflows. It is particularly effective for use cases requiring high factual accuracy, such as real-time market analysis, technical documentation lookup, or complex fact-checking. While standard models often hallucinate when faced with recent events, Sonar Pro utilizes a deep reasoning loop to cross-reference sources before generating a response. Integration via the OpenRouter API allows for seamless implementation into existing agentic frameworks, offering a 200k context window to handle extensive retrieved data without losing coherence.

text generationAPI

nova-premier-v1

amazon
1000000 ctx

Nova Premier v1 is Amazon's high-reasoning multimodal model designed for developers tackling complex logic and high-density data processing. Unlike lightweight models optimized for latency, this version prioritizes depth, making it an ideal engine for sophisticated agentic workflows and multi-step reasoning tasks. A standout feature for engineering teams is its utility in model distillation; its high-quality outputs serve as a robust gold standard for training smaller, specialized local models. With a massive 1-million-token context window, it handles massive codebases or extensive documentation without losing structural coherence. Whether you are building complex RAG pipelines or need a high-fidelity 'teacher' model to optimize your production inference costs through distillation, Nova Premier provides the architectural depth required for enterprise-grade reasoning.

text generationAPI

kimi-k2-thinking

moonshotai
262144 ctx

For developers building complex, multi-step workflows, kimi-k2-thinking represents a significant shift from standard chat completion to agentic reasoning. Built on a large-scale Mixture-of-Experts (MoE) architecture, this model is specifically optimized for long-horizon tasks that require deep logical decomposition rather than just pattern matching. Unlike traditional LLMs that may struggle with cascading errors in complex prompts, the K2 series utilizes an enhanced reasoning trace to navigate intricate problem sets. This makes it particularly effective for autonomous coding agents, mathematical verification, and complex data synthesis where precision is non-negotiable. With a substantial 262k context window, it handles massive technical documentation or large codebases without losing the thread of logic. For integration, it functions via API, allowing you to plug high-level cognitive capabilities into existing agentic frameworks or RAG pipelines that require more than just simple retrieval.

text generationAPI

deepseek-v3.2

deepseek
163840 ctx

DeepSeek-V3.2 represents a significant step forward for developers building autonomous agents and complex reasoning pipelines. Unlike general-purpose models that struggle with long-context coherence, this iteration leverages DeepSeek Sparse Attention (DSA) to maintain high computational efficiency without sacrificing the granular precision required for multi-step logic. For engineers, the primary value proposition lies in its optimized tool-use capabilities; it is architected to act as a reliable reasoning engine within agentic workflows, making it a strong candidate for replacing heavier, more expensive models in production environments. While many models focus on raw parameter count, V3.2 prioritizes the throughput-to-intelligence ratio, offering a streamlined integration path via API for those needing high-performance text generation and structured data extraction. If your stack requires a model that can handle complex instruction following and external tool calls with minimal latency, this is a highly competitive alternative to existing frontier models.

text generationAPI

mistral-large-2512:batch

mistralai
262144 ctx

Mistral Large 2512 represents a significant architectural shift for developers requiring high-performance reasoning at scale. Built on a sparse Mixture-of-Experts (MoE) framework, it utilizes 41B active parameters out of a 675B total parameter pool, optimizing the trade-off between computational efficiency and deep cognitive capability. For engineers, the standout feature is the Apache 2.0 licensing, which provides unprecedented flexibility for commercial deployment and fine-tuning compared to many closed-source competitors. The model excels in complex multilingual tasks, advanced code generation, and sophisticated logical reasoning. With a massive 262k context window, it is particularly well-suited for processing extensive documentation or large-scale codebase analysis. While it competes directly with top-tier proprietary models, its combination of high parameter density and open-license accessibility makes it a premier choice for developers building production-grade, sovereign AI applications that require both intelligence and legal predictability.

text generationAPI

ministral-3b-2512

mistralai
131072 ctx

Ministral 3B is Mistral AI's high-efficiency edge solution, designed specifically for developers needing low-latency performance without sacrificing reasoning depth. Unlike standard small language models (SLMs) that focus solely on text, this 3B parameter variant integrates native vision capabilities, allowing for multimodal workflows in resource-constrained environments. It features a substantial 128k context window, making it surprisingly capable of processing long-form documents or complex instruction sets that typically require much larger models. For developers, this means you can deploy it on local hardware, mobile devices, or lightweight edge instances to handle tasks like visual document parsing, real-time chat, and structured data extraction. While it lacks the massive world knowledge of its larger siblings, its architectural optimization makes it a superior choice for specialized, high-throughput pipelines where latency and compute costs are the primary constraints.

text generationAPI

ministral-8b-2512:batch

mistralai
262144 ctx

For developers building latency-sensitive applications, Ministral 8B represents a strategic middle ground between ultra-lightweight edge models and heavy-duty frontier LLMs. Part of the Ministral 3 family, this 8B parameter model is optimized for high-throughput batch processing and efficient inference without sacrificing reasoning depth. Unlike standard text-only small models, it features native multimodal vision capabilities, allowing you to integrate visual reasoning directly into your workflows. Whether you are implementing real-time agentic loops, automated data extraction from documents, or local RAG pipelines, the model offers a high performance-to-compute ratio. Its massive 262k context window is a significant technical advantage, enabling the processing of extensive codebase documentation or long-form visual sequences that typically choke smaller architectures. It is designed to be integrated via API for scalable production environments where cost-per-token and response speed are critical KPIs.

text generationAPI

ministral-8b-2512

mistralai
262144 ctx

Ministral-8b-2512 is a specialized small language model (SLM) designed for developers who need to balance low-latency performance with sophisticated reasoning. Unlike standard text-only small models, this iteration integrates native vision capabilities, allowing for multimodal workflows like document parsing, visual reasoning, and UI automation within a compact footprint. It is optimized for edge deployment and high-throughput API environments where memory constraints are a primary concern. For engineers building agentic workflows or RAG pipelines, the model offers a significant upgrade in intelligence-per-parameter, providing a more robust alternative to basic 7B-class models while maintaining a manageable computational overhead. Whether you are integrating it into mobile applications or scaling microservices, it serves as a high-efficiency backbone for real-time, multimodal task execution.

text generationAPI

ministral-14b-2512

mistralai
262144 ctx

Ministral-14b-2512 represents a strategic middle ground for developers needing frontier-level reasoning without the latency or cost overhead of massive parameter models. While it sits at 14B parameters, its architecture is tuned to punch significantly above its weight class, delivering performance benchmarks that rival the 24B-class Mistral Small 3.2. For engineers building agentic workflows, RAG pipelines, or complex tool-use applications, this model offers a high intelligence-to-compute ratio. It is designed for low-latency deployment in production environments where throughput is critical but logical depth cannot be sacrificed. Unlike general-purpose giants, Ministral focuses on efficient instruction following and high-context reasoning, making it an ideal candidate for integration into edge-heavy or cost-sensitive microservices. If your stack requires a model that balances sophisticated multi-step reasoning with rapid inference speeds, this is a highly competitive option for your deployment lifecycle.

text generationAPI

nova-2-lite-v1

amazon
1000000 ctx

For developers balancing latency requirements with complex reasoning needs, nova-2-lite-v1 offers a pragmatic middle ground. Unlike massive frontier models that incur high inference costs, this model is optimized for high-throughput, everyday workloads that demand multimodal comprehension. It handles text, image, and video inputs natively, making it particularly effective for automated visual inspection, video captioning, or extracting structured data from visual documentation. While it isn't designed to compete with heavy-duty reasoning giants on deep logic, its strength lies in its efficiency and the massive 1M context window, which allows for processing extensive document sets or long-form video sequences in a single pass. Integration via API is straightforward, making it a viable candidate for production-grade agents and real-time multimodal applications where cost-per-token and response speed are critical KPIs.

text generationAPI

bodybuilder

openrouter
128000 ctx

Bodybuilder is a specialized utility model designed to bridge the gap between high-level intent and low-level API implementation. Instead of manually constructing complex JSON payloads for OpenRouter, developers can provide natural language descriptions of their desired AI workflows, and the model will output valid, structured API request objects. This significantly reduces the friction of prototyping multi-model pipelines or experimenting with different provider configurations. It is particularly useful for building autonomous agents or dynamic orchestration layers where the system needs to programmatically decide which model parameters, headers, or provider settings to use based on a user's prompt. Unlike general-purpose chat models, Bodybuilder is optimized for schema adherence and technical precision, making it a reliable middleware component for developers looking to automate the boilerplate of LLM integration.

text generationAPI

glm-4.6v

z-ai
131072 ctx

GLM-4.6V is a high-capacity multimodal model engineered for developers building vision-centric applications that require deep document intelligence and long-context reasoning. Unlike standard vision models that struggle with dense information, this model excels at parsing complex page layouts, intricate charts, and mixed-media documents. With a 128K token context window, it allows you to feed entire technical manuals or multi-page reports into a single prompt for holistic analysis. For engineers, the primary value lies in its ability to bridge the gap between raw visual data and structured reasoning, making it a strong candidate for automated OCR pipelines, intelligent document processing (IDP), and complex visual QA systems. While many models handle simple image captioning, GLM-4.6V is optimized for the structural nuances of professional documentation, offering a more robust alternative for enterprise-grade automation workflows.

text generationAPI
Email