Global AI chat room · 16 online now Join now
DIRECTORY / 02

AI Models | Open-Source LLM Directory

Discover and compare open-source LLMs, language models and multimodal models by capability, scale, license, downloads and provenance.

Compare modelsFind the right building block for your next workflow
Directory overview
591
curated entries
24 topic groupsLive
02 / MODEL INDEX

Find the right model for the job

Context first, better decisions. Every entry keeps the signal that matters.

CURATED DIRECTORY591 results

minimax-m3

minimax
1048576 ctx

MiniMax-M3 is a multimodal foundation model designed specifically for high-complexity, long-context workflows. Unlike standard LLMs that struggle with information retrieval over large datasets, M3 features a massive 1M-token context window, making it a viable backbone for long-horizon agentic tasks and deep codebase analysis. It processes text, image, and video inputs natively, allowing developers to build sophisticated multimodal agents that can 'see' and 'read' simultaneously. For engineers building autonomous agents or complex RAG pipelines, M3 offers the architectural depth needed to maintain coherence across extended reasoning chains. While many models focus on quick chat interactions, M3 is optimized for structural tasks like heavy-duty coding, multi-step reasoning, and processing massive document sets via its API.

text generationAPI

qwen3.7-plus

qwen
1000000 ctx

Qwen3.7-Plus is a high-efficiency multimodal model designed for developers requiring a balance between reasoning depth and operational cost. While part of a larger series, the 'Plus' iteration focuses on optimizing the performance-to-latency ratio, making it ideal for high-throughput production environments. It natively supports interleaved text and image inputs, allowing for complex visual reasoning tasks such as document parsing, UI automation, and visual Q&A. For engineers, the standout feature is its massive 1,000,000 token context window, which significantly lowers the barrier for processing entire codebases or massive technical datasets without complex RAG architectures. Compared to larger flagship models, Qwen3.7-Plus offers a more streamlined integration path for real-time applications where cost-per-token and response speed are critical constraints. It is built for seamless API integration, targeting developers who need robust multimodal intelligence without the overhead of massive parameter counts.

text generationAPI

nemotron-3-ultra-550b-a55b:free

nvidia
1000000 ctx

Nemotron-3-Ultra is a high-performance Mixture-of-Experts (MoE) model designed for complex reasoning and orchestration tasks. Unlike dense models, it utilizes a hybrid Transformer-Mamba architecture, activating only 55B parameters out of a 550B total pool. For developers, this means you get frontier-level intelligence with significantly lower latency and higher throughput during inference. The model is particularly effective for long-context workflows, supporting up to a 1M token window, making it a strong candidate for massive document analysis, codebase reasoning, and multi-step agentic orchestration. While many models struggle with context decay, the Mamba integration provides a more efficient way to handle linear scaling in long sequences. If your stack requires an orchestrator that can manage complex tool-calling or synthesize information from vast datasets without the typical overhead of massive dense models, this is a highly competitive option.

text generationAPI

nemotron-3-ultra-550b-a55b

nvidia
262144 ctx

Nemotron-3-Ultra-550B is a high-density Mixture-of-Experts (MoE) model designed specifically for complex reasoning and multi-step orchestration tasks. While the total parameter count sits at 550B, its architecture utilizes only 55B active parameters per token, offering a strategic balance between massive knowledge retrieval and inference efficiency. What sets this model apart for developers is its hybrid Transformer-Mamba backbone, which aims to optimize long-context processing and sequential data modeling. Unlike standard dense models, this architecture is built to handle sophisticated agentic workflows where logical consistency and instruction following are paramount. For teams integrating AI into production pipelines, it serves as a robust backbone for autonomous agents, complex code generation, and structured data extraction. It bridges the gap between lightweight specialized models and massive, computationally expensive frontier models, providing a scalable middle ground for high-throughput reasoning applications.

text generationAPI

nemotron-3.5-content-safety:free

nvidia
128000 ctx

For developers building production-grade LLM applications, managing toxicity and safety without sacrificing latency is a constant struggle. Nemotron-3.5-Content-Safety addresses this by providing a specialized, high-efficiency guardrail layer. Built on the Google Gemma-3-4B architecture, this 4B-parameter model is optimized specifically for moderation rather than general reasoning. Unlike massive general-purpose models that can be overkill for safety checks, this compact model is designed to intercept problematic inputs and outputs in both text and vision modalities. It functions as a multimodal filter, making it ideal for developers integrating VLMs where visual context must be scrutinized alongside text. Because it is fine-tuned for safety-specific classification, it offers a more surgical approach to content moderation than zero-shot prompting on larger models, allowing for tighter integration into your inference pipelines with minimal overhead.

text generationAPI

nemotron-3.5-content-safety

nvidia
131072 ctx

NVIDIA Nemotron-3.5-Content-Safety is a specialized 4B-parameter guardrail model designed to sit in the inference pipeline between users and large-scale models. Built upon the Gemma-3 architecture, this multimodal model addresses a critical gap in production AI: the need for low-latency, high-accuracy content moderation for both text and vision inputs. Unlike massive general-purpose models that are too slow for real-time filtering, this compact model is optimized for high-throughput deployment. It functions as a bidirectional safety layer, inspecting incoming user prompts for malicious intent and auditing outgoing model responses to ensure compliance with safety guidelines. For developers building agentic workflows or customer-facing VLMs, it provides a scalable way to mitigate toxicity and policy violations without the massive computational overhead of larger models. Its multimodal capability makes it particularly useful for applications where image-to-text or visual reasoning is central to the user experience.

text generationAPI

kimi-k2.7-code

moonshotai
262144 ctx

kimi-k2.7-code is a specialized MoE (Mixture-of-Experts) model engineered specifically for the software development lifecycle. Unlike general-purpose LLMs, this iteration is optimized for high-precision code generation and complex architectural reasoning across extended contexts. For developers working on large-scale repositories, the standout feature is its ability to maintain coherence over massive codebases, making it highly effective for end-to-end refactoring, debugging, and feature implementation. It moves beyond simple snippet completion, offering deep logical understanding suitable for integrating into automated CI/CD pipelines or custom IDE extensions. While many models struggle with context window degradation, K2.7 is built to handle long-range dependencies, making it a competitive alternative for teams requiring robust, production-ready code assistance within a multimodal framework.

text generationAPI

fusion

openrouter
1000000 ctx

Fusion is a sophisticated orchestration layer designed for complex reasoning tasks that single-model architectures often struggle to resolve. Instead of a linear inference path, Fusion implements a multi-model deliberation pattern. When a prompt is received, it triggers a parallel execution of specialized expert models, augmented by real-time web search and data fetching capabilities. For developers, this means moving away from 'one-shot' prompting toward a structured ensemble approach. It is particularly effective for high-stakes research, multi-step technical troubleshooting, and tasks requiring up-to-the-minute factual accuracy. While standard LLMs rely on static training data, Fusion’s integration of live web retrieval and cross-model verification significantly reduces hallucinations. Integrating via OpenRouter, it provides a scalable way to implement agentic workflows without managing the underlying complexity of multi-model consensus logic yourself.

text generationAPI

glm-5.2:free

z-ai
32768 ctx

GLM-5.2 is a high-reasoning model designed specifically for complex, multi-step technical workflows. Unlike standard chat models optimized for short interactions, this iteration prioritizes long-horizon planning and deep logical reasoning. For developers, the standout feature is the massive 1M-token context window, which makes it a viable engine for project-level software engineering, entire codebase analysis, and massive documentation retrieval. It is built to act as a core reasoning engine for autonomous agents, capable of maintaining coherence across extended task sequences. While many models struggle with context drift in large-scale tasks, GLM-5.2 is engineered to handle the high density of information required for advanced RAG pipelines and automated debugging. If your stack requires an LLM that functions more like a collaborative engineer than a simple text predictor, this model provides the necessary architectural depth for integration into sophisticated agentic frameworks.

text generationAPI

glm-5.2:batch

z-ai
1048576 ctx

GLM-5.2:batch is a high-capacity reasoning model engineered specifically for intensive, long-context workloads. Unlike standard chat models, this iteration is optimized for stability and throughput in complex, multi-step environments. With a massive 1M-token context window, it is purpose-built for developers tackling project-level software engineering, deep codebase analysis, and long-horizon agentic workflows where maintaining state over thousands of lines of code is critical. While many models struggle with 'lost in the middle' phenomena during long sequences, GLM-5.2 is designed to maintain high retrieval accuracy and logical consistency across its entire window. For integration, it serves as a robust backbone for autonomous agents that require deep reasoning rather than just quick responses. If your use case involves processing entire documentation sets or managing complex repository refactoring, this model provides the necessary architectural depth to handle those dependencies without losing context.

text generationAPI

glm-5.2

z-ai
1048576 ctx

GLM-5.2 is a specialized reasoning model designed to handle high-complexity, long-context workflows that typically break standard LLMs. While many models struggle with coherence as context grows, GLM-5.2 leverages a massive 1M-token window specifically optimized for project-level software engineering and autonomous agent orchestration. For developers, this means you can feed entire codebases, massive documentation sets, or multi-step execution logs into a single prompt without losing structural integrity. Unlike general-purpose chat models, its architecture prioritizes long-horizon planning, making it a strong candidate for building agents that require multi-step reasoning and memory retention. Integration is handled via API, making it a drop-in replacement for developers moving from smaller models to more robust, agentic frameworks where deep logical consistency is the primary requirement.

text generationAPI

north-mini-code:free

cohere
256000 ctx

North Mini Code marks Cohere's strategic pivot into agentic coding workflows. Built on a sparse Mixture-of-Experts (MoE) architecture, the model manages 30B total parameters but maintains high efficiency by utilizing only 3B active parameters per token. For developers, this translates to lower latency and improved throughput without sacrificing the reasoning depth required for complex software engineering tasks. Unlike general-purpose LLMs, this model is specifically tuned for code generation, refactoring, and autonomous debugging within agentic frameworks. With a massive 256k context window, it excels at ingestion of entire repositories or extensive documentation, allowing for high-fidelity codebase awareness. While larger monolithic models might offer slightly broader knowledge, North Mini Code is optimized for the specific speed-to-intelligence ratio required for real-time IDE integration and automated CI/CD pipeline agents.

text generationAPI

fugu-ultra

sakana
1000000 ctx

Fugu-Ultra represents a shift away from traditional monolithic architectures toward a learned multi-agent orchestration system. Instead of relying on a single massive parameter set to handle every task, this model functions as an intelligent router trained to decompose complex queries and delegate them to specialized sub-agents. For developers, this means higher precision in reasoning and a significant reduction in 'hallucination noise' typically seen when a single model attempts to context-switch between vastly different domains. Integration is handled via API, making it a drop-in replacement for standard LLM workflows where high-level logic and multi-step planning are required. While standard models often struggle with long-context coherence, Fugu-Ultra’s orchestration layer manages information flow more effectively across its 1M token window. It is particularly suited for complex software engineering pipelines, multi-document legal analysis, and autonomous agentic workflows where the quality of the decision-making process is more critical than raw throughput.

text generationAPI

laguna-xs-2.1:free

poolside
262144 ctx

Laguna XS 2.1 is a specialized coding agent designed for high-efficiency software engineering workflows. Built on a 33B-parameter architecture, it is optimized for the specific logic requirements of code generation, refactoring, and complex debugging. Unlike general-purpose LLMs, this model focuses on maintaining structural integrity and following strict syntax patterns across multiple languages. For developers, the standout feature is the massive 262,144 context window, which allows you to feed entire repositories or extensive documentation into a single prompt without losing coherence. This makes it particularly effective for large-scale refactoring tasks or navigating legacy codebases where cross-file dependency awareness is critical. While it may not match the broad creative writing capabilities of massive frontier models, its performance-to-latency ratio makes it a superior choice for integration into IDE extensions and automated CI/CD pipelines where rapid, accurate code completions are required.

text generationAPI

laguna-xs-2.1

poolside
262144 ctx

Laguna-XS-2.1 is a specialized Mixture-of-Experts (MoE) model designed specifically for high-performance coding tasks. Operating in the 33B-A3B parameter class, it optimizes the trade-off between inference speed and reasoning depth, making it an ideal candidate for real-time IDE integrations and autonomous coding agents. Unlike massive general-purpose models, this iteration focuses on refined logic for complex software engineering workflows, such as codebase refactoring, unit test generation, and multi-file debugging. With a massive 262k context window, it can ingest entire repositories or extensive documentation sets to maintain high architectural awareness during long-form development sessions. For developers, this means lower latency in autocomplete loops and more reliable output when working across deep dependency trees compared to its predecessor, XS.2.

text generationAPI

hy3

tencent
262144 ctx

Hy3 is a high-capacity Mixture-of-Experts (MoE) model from Tencent, engineered specifically for complex reasoning and autonomous agentic workflows. With a total parameter count of 295B and a highly granular architecture featuring 192 experts, it utilizes top-8 routing to maintain efficiency while delivering deep cognitive capabilities. Unlike dense models that struggle with scaling costs, Hy3’s 21B active parameter design allows for sophisticated multi-step logic and tool-use integration without the typical latency overhead. For developers, the standout feature is the configurable reasoning effort, which lets you tune the model's computational intensity based on the task complexity—ranging from rapid text generation to deep, iterative problem-solving. This makes it a versatile backbone for production-grade RAG pipelines, automated coding assistants, and long-context agentic loops. It is optimized for stability in real-world deployment scenarios where predictable reasoning depth is critical.

text generationAPI

aion-3.0

aion-labs
131072 ctx

Aion-3.0 is a specialized multi-model architecture designed specifically for complex narrative generation and roleplaying workflows. Unlike monolithic LLMs that often struggle with character consistency or long-term plot coherence, Aion-3.0 leverages a collaborative ensemble approach based on the GLM family. It utilizes multiple specialized sub-models that work in tandem to manage different layers of storytelling, such as environmental description, character dialogue, and internal monologue. For developers, this means higher fidelity in persona maintenance and reduced logic drift during extended sessions. The model supports a substantial 131k context window, making it viable for deep-lore integration and long-form interactive fiction. While general-purpose models excel at instruction following, Aion-3.0 is optimized for the nuance and creative unpredictability required in high-end RPG engines and automated storytelling tools. Integration is handled via API, allowing for seamless deployment into existing game loops or creative writing platforms.

text generationAPI

aion-3.0-mini

aion-labs
131072 ctx

Aion-3.0-mini is a specialized multi-agent orchestration layer built atop the DeepSeek architecture, specifically optimized for complex narrative generation and roleplaying workflows. Unlike standard monolithic LLMs, this system employs a collaborative generation process where multiple specialized sub-models interact to manage character consistency, world-building, and dialogue logic. For developers, this means a significant reduction in 'character drift' and improved adherence to long-form narrative constraints. With a 128k context window, it is engineered to handle massive story arcs and deep lore repositories without losing structural integrity. While it shares a foundational lineage with DeepSeek, the multi-model coordination makes it a distinct choice for developers building interactive fiction, NPC engines, or sophisticated storytelling agents where nuanced persona maintenance is more critical than raw instruction following.

text generationAPI

grok-4.5

x-ai
500000 ctx

Grok-4.5 represents a significant leap in frontier reasoning, specifically optimized for high-density technical workflows. For developers, the most compelling aspect is its refined performance in complex coding tasks, STEM-based problem solving, and deep knowledge synthesis. Unlike general-purpose models that often struggle with logical consistency in long-form technical documentation, Grok-4.5 maintains high precision across its 500,000-token context window. This makes it an ideal engine for building sophisticated RAG pipelines, automated codebase analysis tools, and advanced mathematical reasoning agents. Integration is handled via a standard API, allowing for seamless deployment into existing CI/CD pipelines or specialized IDE extensions. While many models prioritize conversational fluidity, Grok-4.5 leans into structural accuracy and logical depth, positioning it as a specialized tool for engineers who require more than just a chat interface.

text generationAPI

kat-coder-pro-v2.5

kwaipilot
262144 ctx

KAT-Coder-Pro-v2.5 is a specialized agentic model designed for high-level autonomy in software engineering workflows. Unlike standard chat-based LLMs that require step-by-step prompting, this version is architected to ingest entire business requirements and execute multi-step resolution cycles independently. It excels at navigating complex codebases to locate bugs, refactor legacy modules, and implement end-to-end features with minimal human intervention. For developers, the primary value lies in its ability to function as a virtual teammate rather than a simple autocomplete tool. With a massive 262k context window, it can maintain state across large repositories, making it ideal for integrating into CI/CD pipelines or autonomous agent frameworks. While many models struggle with long-range dependencies in large projects, v2.5 is optimized for the deep reasoning required to manage full-scale issue resolution and architectural consistency.

text generationAPI

muse-spark-1.1

meta
1048576 ctx

Muse Spark 1.1 is Meta’s latest multimodal reasoning engine, specifically architected to power autonomous agentic workflows. Unlike standard LLMs that primarily process text, this model handles a diverse input stream including video, audio, images, and complex PDF documents. For developers, the standout feature is the 1-million-token context window, which allows for deep reasoning over massive datasets or long-form video content without losing coherence. While many models struggle with temporal reasoning in video or structured data extraction from large documents, Muse Spark is optimized for these high-density tasks. It is designed to act as the 'brain' for agents that need to observe an environment through multiple senses and execute logic-driven text outputs. Whether you are building automated document auditors, visual QA systems, or complex multi-step reasoning agents, this model provides the high-capacity context required for production-grade autonomy.

text generationAPI

kimi-k3:batch

moonshotai
1048576 ctx

Kimi-k3:batch is a massive 2.8T parameter multimodal reasoning model designed for high-throughput, complex logic tasks. Unlike standard chat models, this architecture is optimized for long-horizon agentic workflows and deep reasoning, making it a strong candidate for developers building autonomous agents or complex software engineering pipelines. With a massive 1M context window, it excels at ingesting entire codebases or extensive technical documentation to maintain coherence over long-running processes. For developers, the primary value lies in its ability to handle multi-step planning and multimodal inputs without losing the thread of logic. While many open-weight models struggle with deep reasoning consistency, Kimi-k3 bridges the gap between general-purpose LLMs and specialized reasoning engines, offering a scalable solution for knowledge-intensive automation and advanced code synthesis.

text generationAPI

kimi-k3

moonshotai
1048576 ctx

Kimi K3 is a 2.8T parameter open-weight model from Moonshot AI designed specifically for high-reasoning workloads. Unlike standard LLMs that prioritize quick chat responses, K3 is architected for long-horizon agentic workflows and complex logic chains. For developers, the standout feature is its ability to maintain coherence across massive contexts, making it a viable backbone for autonomous agents and sophisticated coding assistants. It bridges the gap between massive proprietary models and open-weight accessibility, offering a specialized focus on deep reasoning and multi-step problem solving. Whether you are building RAG pipelines that require dense information retrieval or complex software engineering agents, K3 provides the parameter scale necessary to handle non-trivial instruction following and nuanced knowledge work without the overhead of much larger closed-source alternatives.

text generationAPI

auto-beta

openrouter
2000000 ctx

Auto-beta is an experimental iteration of our proprietary routing engine, designed to dynamically direct queries to the most efficient underlying model based on task complexity. Unlike static API endpoints, this model functions as an intelligent orchestration layer, optimizing for the trade-off between latency and reasoning depth. For developers, this means you can send generalized prompts without manually selecting a model for every specific sub-task. It is particularly useful for multi-stage pipelines where cost-efficiency and response speed are critical. While it offers a massive 2M token context window, keep in mind that this is a beta release; you should implement it with robust error handling and fallback logic in your production environments. It serves as a high-performance sandbox for testing the latest improvements in automated model selection before they stabilize in our general-purpose routing production tier.

text generationAPI
Email