Global AI chat room · 14 online now Join now
DIRECTORY / 02

AI Models | Open-Source LLM Directory

Discover and compare open-source LLMs, language models and multimodal models by capability, scale, license, downloads and provenance.

Compare modelsFind the right building block for your next workflow
Directory overview
591
curated entries
24 topic groupsLive
02 / MODEL INDEX

Find the right model for the job

Context first, better decisions. Every entry keeps the signal that matters.

CURATED DIRECTORY591 results

qwen3-coder

qwen
262144 ctx

Qwen3-Coder-480B-A35B-Instruct is a high-parameter Mixture-of-Experts (MoE) model specifically engineered for complex software engineering workflows. Unlike standard LLMs that focus on simple autocomplete, this model is architected for agentic autonomy. It excels in high-reasoning tasks such as multi-step function calling, tool orchestration, and navigating massive codebases via its extensive 262k context window. For developers building autonomous coding agents or sophisticated IDE extensions, the MoE architecture offers a strategic balance: the intelligence of a massive parameter set with the inference efficiency required for real-time development cycles. It bridges the gap between simple snippet generation and full-scale repository reasoning, making it a strong contender for integration into CI/CD pipelines and automated debugging environments where long-context coherence is non-negotiable.

text generationAPI

qwen3-235b-a22b-thinking-2507

qwen
131072 ctx

For developers building high-logic applications, qwen3-235b-a22b-thinking-2507 represents a significant shift toward efficient, large-scale reasoning. Built on a Mixture-of-Experts (MoE) architecture, this model optimizes compute by activating only 22B parameters per token, despite having a massive 235B parameter footprint. This provides the intelligence of a dense flagship model with the latency benefits of a much smaller one. The standout feature is its massive 262k context window, making it a viable candidate for long-form document analysis, codebase auditing, and complex multi-turn agentic workflows. Unlike standard chat models, this iteration is specifically tuned for 'thinking'—meaning it excels at chain-of-thought processes required for mathematical proofs, advanced coding, and structured logical deduction. If you are moving beyond simple RAG into autonomous reasoning agents, this model offers the scale and context depth necessary to maintain coherence over extended operations.

text generationAPI

glm-4.5-air

z-ai
131072 ctx

GLM-4.5-Air is a high-efficiency Mixture-of-Experts (MoE) model designed specifically for developers building autonomous agent workflows. While it maintains the reasoning capabilities of the flagship GLM-4.5 series, this 'Air' variant is optimized for lower latency and reduced computational overhead, making it ideal for high-throughput production environments. For developers, the primary value proposition lies in its agent-centric architecture, which excels at tool calling, multi-step planning, and following complex instructions within long-context windows. Unlike heavy-parameter dense models, its MoE structure allows for faster inference speeds without a significant drop in logical reasoning. It is best suited for integration into RAG pipelines, automated coding assistants, and complex task-oriented bots where response time and cost-per-token are critical constraints. If you are transitioning from smaller models to a more capable reasoning engine, this provides a balanced middle ground between raw power and operational efficiency.

text generationAPI

glm-4.5

z-ai
131072 ctx

GLM-4.5 represents a significant shift toward agentic workflows, moving beyond simple chat completion to focus on complex, multi-step reasoning. Built on a Mixture-of-Experts (MoE) architecture, the model optimizes computational efficiency while maintaining high performance across diverse reasoning tasks. For developers, the standout feature is its native optimization for agent-based applications, making it a strong candidate for autonomous tool-use, complex planning, and long-context retrieval. With a 128k token context window, it handles large-scale documentation and codebase analysis effectively. Compared to previous iterations, GLM-4.5 offers improved instruction following and more reliable structured output, which is critical when integrating LLMs into automated pipelines. Whether you are building sophisticated RAG systems or autonomous software agents, this model provides the stability and reasoning depth required for production-grade deployment via API.

text generationAPI

qwen3-30b-a3b-instruct-2507

qwen
262144 ctx

For developers looking to balance high-performance reasoning with low-latency inference, the Qwen3-30B-A3B-Instruct-2507 offers a compelling Mixture-of-Experts (MoE) architecture. While the total parameter count sits at 30.5B, the model only activates approximately 3.3B parameters per token. This design allows for sophisticated instruction following and deep multilingual comprehension without the massive computational overhead typically associated with dense 30B-class models. It is optimized for standard instruction-following tasks rather than extended 'thinking' or chain-of-thought reasoning, making it an ideal candidate for real-time applications like chat interfaces, automated coding assistants, and complex data extraction. With a massive 262,144 context window, it handles long-form document analysis and large codebase ingestion far more effectively than most mid-sized models. Integration is straightforward via API, providing a scalable solution for production environments where throughput and cost-per-token are critical KPIs.

text generationAPI

qwen3-coder-30b-a3b-instruct

qwen
262144 ctx

For developers building autonomous agents or complex software tooling, qwen3-coder-30b-a3b-instruct represents a significant shift toward efficient, high-density reasoning. Unlike dense models of similar scale, this 30.5B parameter Mixture-of-Experts (MoE) architecture utilizes only 8 active experts per forward pass, offering a high performance-to-latency ratio that is critical for real-time IDE integrations. The model is specifically tuned for repository-scale context, moving beyond simple snippet completion to handle deep dependency logic and multi-file architectural understanding. It excels in agentic workflows, showing improved reliability in structured tool calling and function execution compared to previous iterations. Whether you are integrating it via API for automated code reviews or deploying it as a local reasoning engine, its ability to navigate massive context windows makes it a viable alternative to much larger, more expensive proprietary models.

text generationAPI

codestral-2508:batch

mistralai
256000 ctx

Codestral-2508:batch is Mistral's latest specialized iteration optimized for high-throughput coding workflows. Unlike general-purpose LLMs that prioritize conversational fluidity, this model is architected specifically for the technical demands of the SDLC. It excels in low-latency, high-frequency operations, making it an ideal engine for IDE integrations where speed is critical. Developers can leverage its advanced Fill-In-the-Middle (FIM) capabilities for seamless code completion, automated unit test generation, and rapid debugging. With a massive 256,000 token context window, it handles large-scale repository analysis and complex refactoring tasks that would typically exceed standard model limits. While many models struggle with long-range dependencies in large codebases, Codestral provides the structural awareness necessary for maintaining consistency across multiple files. For teams looking to automate CI/CD linting or build sophisticated autonomous coding agents, this model offers a highly efficient, API-driven solution that balances performance with significant operational scale.

text generationAPI

codestral-2508

mistralai
256000 ctx

Codestral-2508 is Mistral AI's latest specialized release designed specifically for the high-velocity requirements of modern software engineering workflows. Unlike general-purpose LLMs that struggle with the nuances of syntax-heavy tasks, this model is optimized for low-latency execution, making it an ideal candidate for IDE integrations and real-time autocomplete engines. A key technical differentiator is its refined proficiency in Fill-In-the-Middle (FIM) patterns, which allows for much more accurate code insertion and context-aware completions compared to standard causal models. Beyond simple generation, it excels at structural debugging, automated test suite construction, and complex refactoring tasks. With a massive 256k context window, developers can feed entire repository structures into the prompt to maintain architectural consistency. For teams looking to build custom coding assistants or automated CI/CD agents, Codestral-2508 provides a highly efficient, API-driven backbone that prioritizes speed and precision over broad, conversational fluff.

text generationAPI

glm-4.5v

z-ai
65536 ctx

GLM-4.5V is a high-capacity multimodal foundation model designed specifically for developers building vision-centric agentic workflows. Moving beyond simple image captioning, this model leverages a Mixture-of-Experts (MoE) architecture—utilizing 106B total parameters with a highly efficient 12B active parameter count—to balance deep reasoning with inference speed. For developers, the primary value lies in its sophisticated video understanding and spatial reasoning capabilities, making it a strong candidate for complex automation tasks like UI navigation, video analysis, and real-time visual monitoring. Compared to monolithic dense models, its MoE structure offers a more granular approach to processing diverse visual inputs, providing a scalable backbone for applications requiring high-fidelity multimodal integration via API. It is particularly suited for developers looking to bridge the gap between raw visual data and actionable logic in autonomous agent systems.

text generationAPI

mistral-medium-3.1:batch

mistralai
131072 ctx

Mistral Medium 3.1:batch is an optimized, high-throughput iteration of Mistral's enterprise-grade architecture, specifically engineered for large-scale asynchronous processing. For developers managing high-volume workloads, this batch-optimized version offers a strategic middle ground between lightweight models and heavy-duty frontier models. It maintains high reasoning capabilities and complex instruction following while significantly lowering the cost-per-token compared to real-time inference. The 128k context window makes it ideal for processing massive datasets, long-form document summarization, or large-scale data extraction tasks where immediate latency is less critical than cost efficiency and throughput. Integrating this via API allows for seamless scaling of background jobs, such as offline content moderation, batch translation, or bulk analytical labeling, without the overhead of maintaining real-time connection stability for massive payloads.

text generationAPI

mistral-medium-3.1

mistralai
131072 ctx

Mistral Medium 3.1 is a strategic update to the previous Medium iteration, specifically tuned to bridge the gap between cost-efficiency and frontier-level reasoning. For developers building production-grade applications, this model offers a sweet spot: it delivers the high-order logic required for complex instruction following and structured data extraction without the prohibitive latency or inference costs of massive-scale models. With a 128k context window, it is well-suited for long-form document analysis, RAG pipelines, and multi-turn conversational agents. Unlike many general-purpose models that prioritize broad chat capabilities, Medium 3.1 is optimized for enterprise workflows where reliability and predictable output formats are paramount. It integrates seamlessly via API, making it a viable backbone for developers looking to scale sophisticated agentic workflows or automated reasoning tasks while maintaining tight control over operational overhead.

text generationAPI

deepseek-chat-v3.1

deepseek
163840 ctx

DeepSeek-V3.1 is a high-density hybrid reasoning model designed for developers who need to toggle between rapid inference and deep logical processing. Built on a massive 671B parameter architecture with only 37B active parameters per token, it offers a highly efficient MoE (Mixture-of-Experts) structure that balances throughput with sophisticated reasoning capabilities. The standout feature is its dual-mode execution: you can trigger a 'thinking' mode for complex algorithmic tasks and multi-step logic, or use standard non-thinking modes for low-latency text generation and chat applications. With a 164k context window, it is well-suited for large-scale codebase analysis, long-document summarization, and complex RAG pipelines. For integration, the model provides a predictable API surface that allows you to control reasoning depth via specific prompt templates, making it a versatile alternative to closed-source frontier models when optimizing for both cost and intelligence.

text generationAPI

hermes-4-405b

nousresearch
131072 ctx

Hermes-4-405B is a high-parameter reasoning model engineered by Nous Research, leveraging the Meta-Llama-3.1-405B architecture as its foundation. Unlike standard LLMs that provide immediate token streams, this model implements a hybrid reasoning mode. It can autonomously decide when to engage in internal deliberation before delivering a final response, making it particularly effective for complex logic, multi-step mathematical problems, and deep code synthesis. For developers, this means a significant reduction in hallucination rates for high-stakes reasoning tasks. The model supports a massive 131k context window, allowing for the ingestion of extensive documentation or entire codebases. While it carries the raw power of a 405B parameter model, the integrated reasoning capability offers a more efficient alternative to manual chain-of-thought prompting. It is designed for seamless API integration into agentic workflows where decision-making accuracy is more critical than raw throughput.

text generationAPI

qwen3-30b-a3b-thinking-2507

qwen
81920 ctx

For developers building agentic workflows or complex reasoning pipelines, qwen3-30b-a3b-thinking-2507 represents a significant step in specialized MoE architectures. Unlike standard dense models, this 30B parameter Mixture-of-Experts model is purpose-built for deep reasoning tasks where accuracy in multi-step logic is more critical than raw token throughput. The standout feature is its dedicated 'thinking mode,' which isolates internal reasoning traces from the final output. This architectural choice is a game-changer for debugging and observability, allowing you to inspect the model's chain-of-thought without polluting your application's primary response stream. While it may not match the sheer speed of smaller, general-purpose models, its ability to handle intricate instruction following and mathematical or logical decomposition makes it a superior choice for RAG-based reasoning, code generation, and automated problem-solving agents. It integrates seamlessly via API, providing a high-intelligence backbone for developers who need verifiable logic rather than just probabilistic text completion.

text generationAPI

kimi-k2-0905

moonshotai
262144 ctx

Kimi-k2-0905 is the latest iteration of Moonshot AI’s large-scale Mixture-of-Experts (MoE) architecture, specifically optimized for high-throughput reasoning and complex instruction following. For developers, the core value lies in its massive 1-trillion parameter scale, which enables sophisticated pattern recognition and deep logical reasoning across diverse datasets. Unlike dense models, its MoE structure offers a more efficient compute-to-performance ratio, making it a strong candidate for agentic workflows and long-context reasoning tasks. The model supports a significant context window of 262,144 tokens, allowing for the processing of entire codebases or extensive documentation in a single pass. Whether you are building autonomous agents, complex RAG pipelines, or advanced coding assistants, Kimi-k2-0905 provides the architectural depth required for production-grade applications where precision and context retention are non-negotiable.

text generationAPI

qwen-plus-2025-07-28

qwen
1000000 ctx

Qwen-plus-2025-07-28 is a high-throughput reasoning model built on the Qwen3 architecture, specifically designed for developers needing a middle ground between lightweight chat models and heavy-duty reasoning engines. The standout feature is its 1-million-token context window, which makes it highly effective for long-form document analysis, large-scale codebase ingestion, and complex multi-turn retrieval tasks. Unlike ultra-large models that trade latency for intelligence, this model optimizes for a 'balanced' profile—providing significant reasoning depth while maintaining the speed and cost-efficiency required for production-scale RAG pipelines and automated agent workflows. For developers integrating via API, it offers a predictable performance-to-cost ratio, making it an ideal candidate for scaling enterprise applications that require both deep context handling and rapid inference cycles.

text generationAPI

qwen3-next-80b-a3b-instruct

qwen
262144 ctx

Qwen3-Next-80B-A3B-Instruct is a high-throughput, instruction-tuned model designed for production environments where low latency is as critical as reasoning depth. Unlike models that output lengthy chain-of-thought traces, this iteration is optimized for direct, stable responses, making it ideal for real-time chat interfaces and automated agentic workflows. With an expansive 262k context window, developers can process massive documentation sets or long-form codebase analysis without hitting immediate token limits. It bridges the gap between heavy-duty reasoning models and lightweight edge models, offering a balanced profile for complex code generation, multilingual QA, and structured data extraction. For teams integrating via API, the focus here is on predictable output patterns and reduced time-to-first-token, providing a robust backbone for applications that require intelligence without the overhead of verbose reasoning steps.

text generationAPI

qwen3-next-80b-a3b-thinking

qwen
262144 ctx

For developers building complex autonomous agents or heavy-duty reasoning pipelines, qwen3-next-80b-a3b-thinking represents a shift toward transparent, chain-of-thought computation. Unlike standard LLMs that jump straight to an answer, this model prioritizes a structured 'thinking' trace, allowing you to inspect its internal logic before the final output is generated. This makes it particularly potent for high-stakes tasks like debugging intricate codebases, verifying mathematical proofs, or managing multi-step agentic workflows where error propagation is a risk. With a massive 262k context window, it handles deep document analysis and large-scale codebase ingestion without losing the thread. While standard models excel at quick chat interactions, this model is engineered for accuracy in logic-dense environments, offering a more predictable way to integrate complex reasoning into your existing API-driven infrastructure.

text generationAPI

qwen3-coder-flash

qwen
1000000 ctx

For developers building autonomous software agents, Qwen3-Coder-Flash offers a strategic balance between low latency and high-reasoning capabilities. While the 'Plus' variant serves heavy-duty architecture tasks, the Flash model is specifically optimized for high-throughput environments where speed and cost-efficiency are critical. Its primary strength lies in its refined tool-calling proficiency, making it an ideal engine for agentic workflows that require frequent interaction with compilers, debuggers, and file systems. Unlike general-purpose models that often struggle with precise syntax in long-context loops, this model is fine-tuned for the iterative nature of programming. Whether you are integrating it into a CI/CD pipeline for automated code reviews or deploying it as the backbone of a real-time coding assistant, it provides the reliability of a specialized coding model without the heavy compute overhead of larger parameter sets.

text generationAPI

deepseek-v3.1-terminus

deepseek
163840 ctx

DeepSeek-V3.1 Terminus is a refined iteration of the V3.1 architecture, specifically engineered to resolve common friction points in multi-turn reasoning and cross-lingual stability. For developers building complex agentic workflows, this update is significant; it directly addresses previous inconsistencies in instruction following and language switching, making it a more reliable backbone for autonomous agents. While the core intelligence remains consistent with the base V3.1 model, the 'Terminus' tuning optimizes the model's ability to maintain context across long-form interactions. It is particularly well-suited for developers integrating LLMs into production-grade tool-use environments where predictable output formats and linguistic precision are non-negotiable. Compared to its predecessor, you can expect fewer hallucinations during complex reasoning steps and more robust performance in non-English language tasks, all while maintaining the high throughput efficiency characteristic of the DeepSeek series.

text generationAPI

qwen3-coder-plus

qwen
1000000 ctx

Qwen3-Coder-Plus is a proprietary, high-parameter evolution of the open-source Qwen3 Coder architecture, specifically optimized for autonomous software engineering workflows. Unlike standard LLMs that merely suggest snippets, this model is architected as a dedicated coding agent. It excels in complex, multi-step reasoning tasks through advanced tool-calling capabilities, allowing it to interact directly with compilers, debuggers, and file systems. For developers, this means moving beyond simple autocomplete toward true agentic workflows where the model can plan, execute, and verify code autonomously. With a massive 1M token context window, it can ingest entire repositories to maintain global state awareness, making it a viable replacement for human-in-the-loop code reviews and large-scale refactoring projects. It bridges the gap between a chat interface and a fully integrated development environment (IDE) agent.

text generationAPI

qwen3-max

qwen
262144 ctx

Qwen3-Max represents a significant architectural leap in the Qwen series, specifically targeting developers who require high-fidelity reasoning and precise instruction following. Unlike previous iterations, this model shows marked improvements in handling complex, multi-step logic and expanding its long-tail knowledge base, making it more reliable for niche domain tasks. For developers building autonomous agents or sophisticated RAG pipelines, the updated multilingual capabilities and enhanced context handling provide a more stable foundation for global applications. While earlier versions were competitive in general chat, Qwen3-Max is engineered for production-grade accuracy in code generation and structured data extraction. It integrates seamlessly via API, offering a scalable way to deploy state-of-the-art intelligence without the overhead of self-hosting massive parameter sets. If your workflow demands a model that minimizes hallucination during complex reasoning tasks, this is a substantial upgrade over the January 2025 baseline.

text generationAPI

qwen3-vl-235b-a22b-instruct

qwen
262144 ctx

Qwen3-VL-235B-A22B-Instruct is a heavy-duty multimodal model designed for developers needing high-fidelity visual reasoning and text generation in a single pipeline. Unlike smaller vision models that struggle with dense data, this model excels at complex document parsing, intricate chart analysis, and long-form video understanding. For engineers building automated inspection tools, data extraction pipelines, or advanced VQA interfaces, the 235B scale provides a significant reasoning advantage over lightweight alternatives. It bridges the gap between simple image captioning and deep semantic understanding of temporal video data. Integration is straightforward via API, making it a viable backbone for production-grade agents that require a unified vision-language architecture without the overhead of managing massive local weights.

text generationAPI

qwen3-vl-235b-a22b-thinking

qwen
131072 ctx

Qwen3-VL-235B-A22B-Thinking is a massive-scale multimodal model designed for developers who need more than just simple image captioning. Unlike standard vision-language models, this architecture integrates a specialized 'thinking' process to handle complex reasoning tasks involving both static images and temporal video data. For developers working in STEM, automated mathematics, or technical documentation, the model excels at interpreting intricate diagrams, handwritten equations, and multi-step visual logic. It bridges the gap between raw perception and logical deduction, making it a powerful engine for agentic workflows that require visual grounding. While it carries a significant parameter footprint, its ability to perform deep reasoning over high-resolution visual inputs sets it apart from smaller, faster models that often struggle with spatial accuracy or complex mathematical reasoning in visual contexts. Integration via API allows for scaling these high-reasoning capabilities into production-ready visual agents.

text generationAPI
Email