Global AI chat room · 18 online now Join now
DIRECTORY / 02

AI Models | Open-Source LLM Directory

Discover and compare open-source LLMs, language models and multimodal models by capability, scale, license, downloads and provenance.

Compare modelsFind the right building block for your next workflow
Directory overview
839
curated entries
24 topic groupsLive
02 / MODEL INDEX

Find the right model for the job

Context first, better decisions. Every entry keeps the signal that matters.

CURATED DIRECTORY839 results

grok-4.20

x-ai
2000000 ctx

Grok-4.20 is a high-performance reasoning model engineered for developers building complex, autonomous agent workflows. Unlike standard LLMs that struggle with multi-step logic, this model is optimized for agentic tool calling and strict instruction following, making it a reliable backbone for automated pipelines. It features a massive 2-million token context window, allowing you to ingest entire codebases or massive documentation sets without losing coherence. While many models trade speed for reasoning depth, Grok-4.20 maintains industry-leading latency, making it suitable for real-time applications. For developers, the primary value proposition lies in its significantly reduced hallucination rate, which minimizes the need for manual output verification in production environments. Whether you are integrating it via API for RAG-heavy architectures or deploying it for complex decision-making agents, it offers a robust balance of throughput and logical precision.

text generationAPI

grok-4.20-multi-agent

x-ai
2000000 ctx

Grok-4.20-multi-agent represents a shift from single-prompt reasoning to orchestrated, collaborative intelligence. Unlike standard LLMs that process tasks in a linear fashion, this variant is architected specifically for agentic workflows where multiple specialized instances operate in parallel. For developers, this means moving beyond simple text generation into complex task decomposition, autonomous research, and coordinated tool execution. The model is optimized for high-context environments, leveraging a massive 2M token window to maintain state across multi-turn agent interactions. While standard models often struggle with 'agent drift' or loss of coherence during long-running tasks, this multi-agent framework is designed to synthesize disparate data points into a unified output. It is particularly suited for building autonomous DevOps pipelines, complex data analysis engines, or automated research agents that require real-time tool calling and cross-verification between sub-agents.

text generationAPI

trinity-large-thinking

arcee-ai
262144 ctx

Trinity Large Thinking is a specialized reasoning model from Arcee AI designed to bridge the gap between standard LLMs and complex agentic workflows. Unlike general-purpose chat models, this architecture is optimized for high-density reasoning tasks and multi-step logic, as evidenced by its performance on the PinchBench benchmark. For developers, this means a significant reduction in logical hallucinations when building autonomous agents or complex tool-use pipelines. It supports a substantial 262k context window, making it viable for deep document analysis and long-form codebase reasoning. While many models struggle with the 'chain-of-thought' overhead, Trinity is fine-tuned to handle agentic workloads where precision in decision-making is more critical than mere conversational fluency. It offers a robust alternative for teams looking to integrate sophisticated reasoning capabilities into their existing RAG or agentic frameworks via API.

text generationAPI

glm-5v-turbo

z-ai
202752 ctx

GLM-5V-Turbo is a native multimodal foundation model designed specifically for developers building autonomous agents and vision-centric applications. Unlike models that rely on separate vision encoders, this architecture treats image, video, and text as unified inputs, which significantly reduces latency and improves reasoning consistency across modalities. For engineers, the primary value lies in its long-horizon planning capabilities and its specialized proficiency in vision-based coding tasks—making it a strong candidate for automated UI testing, visual debugging, and complex workflow orchestration. While many multimodal models struggle with temporal consistency in video or precise spatial reasoning in code generation, GLM-5V-Turbo is optimized for these high-stakes agentic loops. It is accessible via API, making it easy to integrate into existing RAG pipelines or agent frameworks that require a model to 'see' and 'act' within a digital environment.

text generationAPI

qwen3.6-plus

qwen
1000000 ctx

Qwen 3.6 Plus introduces a significant architectural shift by merging linear attention mechanisms with a sparse Mixture-of-Experts (MoE) routing system. For developers, this means a more efficient scaling path where model capacity increases without a proportional surge in computational latency. Unlike the previous 3.5 series, this iteration is optimized for high-throughput inference, making it particularly suitable for real-time applications and complex agentic workflows. The model excels in reasoning-heavy tasks and long-context processing, maintaining stability across its 1M token window. Whether you are integrating via API for RAG pipelines or building autonomous tool-use agents, the 3.6 Plus offers a more granular balance between parameter density and inference speed, positioning it as a highly competitive alternative to existing large-scale MoE models in the production environment.

text generationAPI

gemma-4-31b-it:free

google
262144 ctx

Gemma 4 31B Instruct is a high-density multimodal model designed for developers requiring a balance between sophisticated reasoning and efficient deployment. Unlike standard text-only LLMs, this model natively processes both text and image inputs, making it ideal for visual reasoning, document analysis, and complex multimodal workflows. A standout technical feature is its massive 256K token context window, which allows for the ingestion of entire codebases or extensive technical documentation in a single prompt. For developers building autonomous agents, the model supports configurable reasoning modes and native function calling, enabling seamless integration into existing software ecosystems. While it maintains the accessibility of the Gemma open-weights lineage, its 30.7B parameter architecture provides a significant leap in logical depth compared to smaller edge models, positioning it as a versatile middle-weight powerhouse for production-grade applications.

text generationAPI

gemma-4-31b-it

google
262144 ctx

Gemma 4 31B Instruct is a high-density multimodal model designed for developers requiring a balance between sophisticated reasoning and efficient deployment. Unlike smaller parameter models, this 30.7B dense architecture handles complex instruction following with a significantly expanded 256K token context window, making it ideal for large-scale document analysis and long-form codebase reasoning. A standout feature is the configurable reasoning mode, which allows you to toggle deep 'thinking' processes for logic-heavy tasks or prioritize low-latency responses for standard chat applications. It natively supports multimodal inputs, enabling seamless integration of visual data into your text-based workflows. For engineers building agentic systems, its native function calling capabilities provide a reliable bridge between LLM reasoning and external tool execution. Whether you are fine-tuning for specific domain expertise or integrating via API for production-grade RAG pipelines, Gemma 4 offers a robust middle-ground between lightweight edge models and massive, high-latency frontier models.

text generationAPI

gemma-4-26b-a4b-it:free

google
262144 ctx

Gemma 4 26B A4B IT is a specialized Mixture-of-Experts (MoE) model engineered to bridge the gap between lightweight inference and high-parameter reasoning. While the architecture carries 25.2B total parameters, its sparse activation strategy only engages roughly 3.8B parameters per token. For developers, this means you get the intelligence profile of a ~30B parameter dense model but with the latency and throughput characteristics of a much smaller footprint. This makes it an ideal candidate for real-time applications, complex instruction following, and RAG pipelines where response speed is critical. It excels in reasoning-heavy tasks and nuanced text generation while maintaining a significantly lower compute cost per request compared to traditional dense models. If you are looking to optimize your inference budget without sacrificing logic or linguistic precision, this model offers a highly efficient middle ground for production-grade deployments.

text generationAPI

gemma-4-26b-a4b-it

google
262144 ctx

Gemma 4 26B A4B IT is a specialized Mixture-of-Experts (MoE) model designed to bridge the gap between lightweight efficiency and high-parameter reasoning. For developers, the standout feature is its architecture: while the model holds 25.2B total parameters, it only activates approximately 3.8B per token. This allows you to deploy a model with the intelligence profile of a 30B+ parameter dense model while maintaining the low latency and reduced compute costs of a much smaller footprint. It is instruction-tuned for high-precision following, making it ideal for complex RAG pipelines, agentic workflows, and real-time conversational interfaces. Whether you are optimizing for throughput in a production environment or seeking deep reasoning capabilities without the massive VRAM overhead, this model provides a highly efficient scaling path for sophisticated text-based applications.

text generationAPI

glm-5.1

z-ai
204800 ctx

GLM-5.1 marks a strategic shift from chat-based interaction toward autonomous agentic workflows. While previous iterations focused on short-form instruction following, this model is architected to handle long-horizon reasoning and complex, multi-step coding tasks. For developers, this means moving beyond simple snippet generation to delegating entire development cycles, such as debugging large codebases or implementing multi-file features. With a 204,800 token context window, it maintains high retrieval accuracy across extensive documentation and repository structures, making it a viable backbone for autonomous coding agents. Compared to standard LLMs that struggle with task drift during long sessions, GLM-5.1 is optimized for continuity and independent execution, offering a robust API for integrating sophisticated reasoning into existing CI/CD pipelines and developer tools.

text generationAPI

kimi-k2.6

moonshotai
262144 ctx

Kimi K2.6 is Moonshot AI’s latest multimodal powerhouse, specifically architected to move beyond simple chat and into autonomous engineering workflows. For developers, the real value lies in its high-context reasoning capabilities and its specialized training for long-horizon coding tasks. Unlike standard LLMs that struggle with architectural consistency, K2.6 is optimized for end-to-end development cycles, spanning languages like Python, Rust, and Go. It excels in scenarios where you need to bridge the gap between high-level UI/UX design and functional code, or when orchestrating complex multi-agent systems to solve multi-step logic problems. With a massive 262k context window, it is built for deep repository analysis and maintaining state across extended debugging sessions. If you are building automated DevOps pipelines, complex software agents, or rapid prototyping tools, K2.6 provides the structural stability and multimodal understanding required for production-grade integration.

text generationAPI

pareto-code

openrouter
2000000 ctx

Pareto-code is a dynamic routing layer designed to optimize the trade-off between coding performance and latency/cost. Instead of hitting a single static model, this router maintains a tiered shortlist of high-performing LLMs, ranked specifically by their coding percentiles from Artificial Analysis. For developers, this means you aren't overpaying for GPT-4o when a smaller, faster model can handle a simple refactor, yet you aren't sacrificing quality on complex algorithmic tasks. The core mechanism relies on the 'min_coding_score' parameter, allowing you to set a quality threshold between 0 and 1. If a task requires high reasoning, the router escalates to top-tier models; if the task is trivial, it routes to more efficient engines. It integrates seamlessly via the OpenRouter API, making it an ideal middle-layer for building autonomous coding agents or IDE extensions where reliability and cost-efficiency are equally critical.

text generationAPI

mimo-v2.5

xiaomi
1050000 ctx

MiMo-V2.5 is Xiaomi's latest native omnimodal model, engineered specifically to bridge the gap between high-end agentic reasoning and production-scale cost efficiency. For developers building autonomous agents or complex multimodal workflows, this model offers a significant shift in the performance-to-cost ratio, delivering Pro-level capabilities at approximately 50% of the typical inference overhead. Unlike models that rely on bolted-on vision encoders, MiMo-V2.5 features a native architecture that enhances its perception of both static images and temporal video data. This makes it particularly effective for real-time visual reasoning, video analysis, and complex tool-use scenarios where context retention is critical. With a massive 1,050,000 token context window, it is designed to handle extensive documentation or long-form video streams without losing coherence. Whether you are integrating via API for mobile ecosystem automation or building sophisticated vision-language applications, MiMo-V2.5 provides a highly scalable alternative to more expensive, heavyweight multimodal models.

text generationAPI

mimo-v2.5-pro

xiaomi
1050000 ctx

mimo-v2.5-pro is Xiaomi’s high-performance flagship model, specifically engineered for developers working on autonomous agents and complex software engineering workflows. Unlike standard chat models, this iteration is optimized for long-horizon reasoning and multi-step task execution, making it a viable backbone for agentic frameworks. It demonstrates significant proficiency in software development lifecycles, as evidenced by its performance on SWE-bench Pro, and handles extended context requirements with a 1.05M token window. For engineers, this means the ability to ingest entire codebases or massive documentation sets to maintain state across complex debugging or refactoring sessions. While many models struggle with the 'drift' seen in long-form reasoning, mimo-v2.5-pro is tuned to maintain logical consistency through deep-reasoning tasks. It is accessible via API, making it a plug-and-play option for integrating advanced reasoning into existing CI/CD pipelines or automated developer tools.

text generationAPI

hy3-preview

tencent
262144 ctx

Hy3-preview is a specialized Mixture-of-Experts (MoE) model from Tencent, engineered specifically for agentic workflows and high-throughput production environments. Unlike standard monolithic LLMs, Hy3 offers a unique architectural advantage through configurable reasoning modes. Developers can toggle between disabled, low, and high reasoning levels, providing a granular way to balance inference latency against complex problem-solving capabilities. This makes it particularly effective for multi-step agent loops where cost and speed are as critical as logic. With a substantial 262k context window, it handles long-form documentation and extensive conversation histories without significant degradation. For teams integrating via API, Hy3 serves as a scalable middle ground between lightweight chat models and heavy-duty reasoning engines, offering a predictable way to scale compute resources based on the specific complexity of the incoming task.

text generationAPI

deepseek-v4-flash

deepseek
1048576 ctx

DeepSeek-V4-Flash is a high-throughput Mixture-of-Experts (MoE) model engineered specifically for developers prioritizing low-latency inference without sacrificing reasoning depth. With a massive 284B total parameter architecture, it leverages a sparse activation strategy—utilizing only 13B parameters per token—to deliver a performance-to-cost ratio that challenges much larger, dense models. The standout feature for production environments is the 1M-token context window, making it an ideal candidate for long-form document analysis, massive codebase ingestion, and complex multi-turn agentic workflows. Unlike standard lightweight models that struggle with nuance, V4-Flash maintains high instruction-following accuracy, making it a versatile drop-in replacement for RAG pipelines and automated coding assistants where speed and context density are the primary bottlenecks.

text generationAPI

deepseek-v4-pro

deepseek
1048576 ctx

DeepSeek-V4-Pro is a high-performance Mixture-of-Experts (MoE) model engineered for developers who require massive scale without the typical latency overhead of dense architectures. With 1.6 trillion total parameters and 49 billion activated per token, it strikes a sophisticated balance between deep reasoning capabilities and computational efficiency. The standout feature for production environments is the expansive 1M-token context window, making it a viable backbone for complex RAG pipelines, long-form codebase analysis, and multi-document synthesis. Unlike many general-purpose models, V4 Pro shows significant strength in structured logic and advanced programming tasks, positioning it as a direct competitor to top-tier frontier models. For integration, its API-first approach allows for seamless deployment into existing workflows, offering a scalable solution for applications requiring high-density information processing and complex instruction following.

text generationAPI

qwen3.6-27b

qwen
262144 ctx

Qwen3.6-27B represents a significant step forward for developers seeking a versatile mid-sized model that balances computational efficiency with multimodal intelligence. Unlike purely text-based LLMs, this dense 27B architecture natively handles text, image, and video inputs, making it a robust engine for complex reasoning tasks that require visual context. For engineering teams, the 262k context window is a standout feature, enabling the processing of massive codebases or long-form video documentation without frequent truncation. While larger models offer higher reasoning ceilings, the 27B parameter count is optimized for high-throughput production environments where latency and cost-per-token are critical KPIs. It is particularly well-suited for building agentic workflows, automated video analysis tools, and sophisticated RAG pipelines that incorporate non-textual data. Integration is straightforward via API, offering a scalable alternative to heavier frontier models when deployment speed and multimodal versatility are the primary requirements.

text generationAPI

qwen3.6-max-preview

qwen
262144 ctx

Qwen3.6-Max-Preview is a frontier-class sparse Mixture-of-Experts (MoE) model designed to bridge the gap between general reasoning and specialized agentic workflows. With a massive parameter scale and a 262k context window, it is specifically tuned for high-density tasks like complex codebase navigation, multi-step tool orchestration, and autonomous software engineering. For developers, the primary value lies in its improved reliability during function calling and its ability to maintain coherence across large-scale documentation ingestion. Unlike dense models that may struggle with latency-to-intelligence ratios, this MoE architecture optimizes for high-throughput reasoning, making it a viable backbone for production-grade AI agents and automated DevOps pipelines. If your stack requires deep integration with external APIs or sophisticated code generation within large repositories, this model offers a significant step up in instruction following and structural accuracy.

text generationAPI

qwen3.6-35b-a3b

qwen
262144 ctx

For developers building high-throughput applications, Qwen3.6-35B-A3B offers a compelling middle ground between lightweight edge models and massive dense architectures. By utilizing a Sparse Mixture-of-Experts (SMoE) design, it delivers the reasoning capabilities of a much larger model while only activating 3 billion parameters per token. This significantly reduces inference latency and compute costs without sacrificing the nuanced understanding required for complex tasks. The model is natively multimodal, making it a versatile choice for pipelines involving both vision and text. With a massive 262k context window, it excels at long-document processing, codebase analysis, and complex RAG workflows. Unlike monolithic models, this architecture is optimized for efficient scaling, allowing you to maintain high performance in production environments where tokens-per-second and cost-efficiency are critical KPIs. Whether you are integrating via API or fine-tuning for specific domain logic, the efficiency-to-intelligence ratio here is highly competitive for modern AI orchestration.

text generationAPI

qwen3.6-flash

qwen
1000000 ctx

Qwen3.6 Flash is Alibaba's latest efficient language model optimized for speed and cost-effectiveness. It supports text, image, and video inputs with a massive 1M token context window, making it suitable for complex multimodal applications. The model offers tiered pricing, allowing developers to balance performance and budget based on their specific needs. Compared to previous versions, it delivers faster inference times while maintaining strong reasoning capabilities across multiple languages. Integration is straightforward through standard APIs, and it's particularly well-suited for real-time applications like chatbots, content generation, and data analysis where latency matters. Its multilingual support and long-context handling make it a solid choice for global development teams working on scalable AI products.

text generationAPI

qwen3.5-plus-20260420

qwen
1000000 ctx

Qwen3.5-Plus (April 2026) represents a significant leap in multimodal reasoning for developers building complex, data-heavy applications. Unlike previous iterations that focused primarily on text, this model natively processes text, high-resolution images, and video streams within a single inference pass. For engineers, the standout feature is the 1M token context window, which effectively moves long-form video analysis and massive codebase auditing from a retrieval-augmented generation (RAG) problem to a direct context problem. While many models struggle with temporal consistency in video, Qwen3.5-Plus is optimized for long-sequence multimodal understanding. Integration is handled via standard API protocols, making it a drop-in replacement for developers looking to upgrade from text-only LLMs to sophisticated vision-language agents. It is particularly well-suited for automated visual QA, complex video summarization, and multimodal reasoning tasks where high-fidelity spatial understanding is required.

text generationAPI

nemotron-3-nano-omni-30b-a3b-reasoning:free

nvidia
256000 ctx

NVIDIA's Nemotron-3-Nano-Omni is a specialized 30B parameter multimodal model engineered specifically for high-efficiency agentic workflows. Unlike massive general-purpose LLMs, this model is architected as a 'sub-agent' designed to handle perception and context management within larger enterprise systems. It processes text, images, and video, making it an ideal candidate for tasks requiring visual reasoning or multi-modal context extraction before passing structured data to a primary orchestrator. For developers, the standout feature is its ability to act as a lightweight, high-speed sensory layer, reducing latency and token costs in complex RAG or agentic pipelines. While it lacks the broad creative depth of much larger models, its optimization for multimodal input and enterprise-grade reliability makes it a powerful tool for building autonomous systems that need to 'see' and 'understand' environment data in real-time.

text generationAPI

mistral-medium-3-5:batch

mistralai
262144 ctx

Mistral Medium 3.5 is a high-density 128B parameter model engineered for developers who require a balance between reasoning depth and operational efficiency. Unlike smaller, specialized models, this iteration excels in complex instruction-following and multimodal processing, allowing you to feed both text and visual data into a single pipeline. It is specifically optimized for agentic workflows, where the model must maintain logical consistency across multi-step reasoning tasks or complex coding environments. For international teams, the model’s strength lies in its ability to handle high-context tasks within a large 262k window, making it ideal for analyzing extensive documentation or long-form codebases. While it competes with top-tier frontier models, its architectural focus on dense instruction following makes it a highly predictable choice for production-grade automation and sophisticated RAG implementations.

text generationAPI
Email