Global AI chat room · 5 online now Join now
DIRECTORY / 02

AI Models | Open-Source LLM Directory

Discover and compare open-source LLMs, language models and multimodal models by capability, scale, license, downloads and provenance.

Compare modelsFind the right building block for your next workflow
Directory overview
839
curated entries
24 topic groupsLive
02 / MODEL INDEX

Find the right model for the job

Context first, better decisions. Every entry keeps the signal that matters.

CURATED DIRECTORY839 results

mistral-medium-3-5

mistralai
262144 ctx

Mistral Medium 3.5 is a high-density 128B parameter model engineered specifically for developers building autonomous agentic workflows and sophisticated reasoning pipelines. Unlike smaller, faster models that struggle with long-range logic, this model strikes a balance between high-level cognitive reasoning and practical latency, making it ideal for complex coding tasks and multi-step instruction following. It features native multimodal capabilities, allowing you to process image inputs alongside text, which expands its utility in visual reasoning and document analysis. With a massive 262,144 token context window, it is built to ingest entire codebases or massive datasets without losing coherence. For teams moving beyond simple chat interfaces into structured tool-use and automated decision-making, this model provides the stability and instruction adherence required for production-grade integration via API.

text generationAPI

grok-4.3:batch

x-ai
1000000 ctx

Grok 4.3:batch is a high-throughput reasoning model designed for developers building autonomous agentic workflows and complex instruction-following pipelines. Unlike standard chat models, this iteration prioritizes logical consistency and factual precision, making it ideal for RAG-heavy applications and multi-step task execution. It features native multimodal capabilities, allowing you to process both text and visual data within a single context window of up to 1,000,000 tokens. For teams scaling production environments, the 'batch' designation indicates an optimized endpoint for non-latency-sensitive, high-volume processing, offering a cost-effective way to handle large-scale data extraction, document analysis, and automated reasoning tasks. If your stack requires a model that can maintain deep coherence across massive datasets without the overhead of real-time conversational latency, this is a highly competitive option for your deployment pipeline.

text generationAPI

grok-4.3

x-ai
1000000 ctx

Grok 4.3 marks a significant shift toward high-fidelity reasoning and autonomous agentic workflows. Unlike standard LLMs that prioritize conversational fluency, this model is architected for precision in instruction-following and complex multi-step logic. For developers building autonomous agents, the model's ability to process multimodal inputs—specifically text and imagery—allows for more sophisticated environmental perception in software automation. The 1M token context window is a standout feature, enabling the ingestion of massive codebases or extensive technical documentation for RAG-based applications without losing structural coherence. While many models struggle with factual drift during long-form reasoning, Grok 4.3 is optimized for high-density information retrieval. If your stack requires a reliable engine for tool-calling, automated debugging, or complex data extraction, this model provides the stability and reasoning depth necessary to move beyond simple chat interfaces into production-grade agentic systems.

text generationAPI

perceptron-mk1

perceptron
32768 ctx

Perceptron-mk1 is a high-fidelity vision-language model engineered specifically for temporal reasoning and embodied AI applications. Unlike standard multimodal models that treat video as a sequence of static frames, mk1 is optimized to parse complex motion dynamics and spatial relationships over time. For developers building autonomous agents, robotics controllers, or advanced video analytics pipelines, this model provides the granular visual grounding necessary to translate raw video streams into actionable logic. It excels in tasks requiring long-context visual understanding, such as describing multi-step physical actions or diagnosing causal events in a video sequence. Integration is handled via a standard API, supporting a 32k context window to accommodate extended temporal data. While many VLMs struggle with the 'temporal drift' seen in long video clips, mk1 is architected to maintain high-resolution semantic consistency throughout the input stream, making it a robust choice for real-world deployment in vision-centric workflows.

text generationAPI

grok-build-0.1

x-ai
256000 ctx

Grok-build-0.1 is a specialized model engineered specifically for agentic software engineering workflows. Unlike general-purpose LLMs that struggle with long-term architectural consistency, this model is optimized for the high-frequency, iterative loops required by autonomous coding agents. It features a massive 256k context window, making it highly effective for ingesting entire codebases or complex documentation sets to maintain state across multi-step tasks. The model supports multimodal inputs, allowing developers to feed in UI screenshots or architectural diagrams to drive implementation. For integration, it is designed for low-latency interaction, which is critical when deploying agents that must perform rapid reasoning and tool-calling cycles. While general models excel at snippet generation, Grok-build-0.1 is purpose-built for the more demanding lifecycle of autonomous debugging, refactoring, and system-level development.

text generationAPI

qwen3.7-max

qwen
1000000 ctx

Qwen3.7-Max represents a significant shift toward agentic intelligence, moving beyond simple chat interfaces to handle complex, multi-step reasoning workflows. For developers building autonomous agents, this model is optimized for high-reliability tool use and structured output, making it a strong candidate for integration into automated DevOps pipelines or sophisticated software engineering assistants. Unlike general-purpose models that struggle with long-context consistency, Qwen3.7-Max leverages its massive 1M token window to maintain coherence across entire codebases or extensive technical documentation. While many models focus on creative prose, this flagship is engineered for precision in coding, logical reasoning, and productivity automation. If your stack requires a model that can act as a reasoning engine rather than just a text predictor, Qwen3.7-Max provides the architectural depth needed for production-grade agentic tasks.

text generationAPI

step-3.7-flash

stepfun
262144 ctx

Step-3.7-Flash is a high-efficiency multimodal MoE model designed for developers requiring low-latency reasoning and native vision processing. Unlike traditional dense models, it utilizes a 196B-parameter backbone but only activates approximately 11B parameters per token, striking a balance between high-level intelligence and rapid inference speeds. For developers building real-time applications, this architecture minimizes time-to-first-token without sacrificing the depth needed for complex instruction following. The model excels in multimodal workflows, offering integrated image and video understanding that goes beyond simple captioning to include spatial reasoning and temporal analysis. With a massive 262k context window, it is particularly well-suited for long-document processing, large-scale codebase analysis, and multi-frame video reasoning. If your stack requires a scalable API-driven solution that handles vision-heavy tasks with the efficiency of a smaller model, Step-3.7-Flash provides a competitive alternative to standard lightweight LLMs.

text generationAPI

minimax-m3:batch

minimax
524288 ctx

MiniMax-M3:batch is a high-throughput multimodal foundation model designed for developers building complex, long-context applications. Unlike standard chat models, M3 is architected to handle massive context windows—up to 1M tokens—making it a viable engine for deep codebase analysis, extensive document processing, and long-horizon agentic workflows. It natively processes text, image, and video inputs, allowing for sophisticated cross-modal reasoning within a single pipeline. For engineering teams, the 'batch' designation suggests an optimization for non-real-time, high-volume processing tasks where cost-efficiency and throughput are prioritized over instantaneous latency. Whether you are automating complex coding tasks or building autonomous agents that require continuous environmental observation via video, M3 provides the structural depth needed to maintain coherence across extended sequences that typically cause smaller models to lose track.

text generationAPI

minimax-m3

minimax
1048576 ctx

MiniMax-M3 is a multimodal foundation model designed specifically for high-complexity, long-context workflows. Unlike standard LLMs that struggle with information retrieval over large datasets, M3 features a massive 1M-token context window, making it a viable backbone for long-horizon agentic tasks and deep codebase analysis. It processes text, image, and video inputs natively, allowing developers to build sophisticated multimodal agents that can 'see' and 'read' simultaneously. For engineers building autonomous agents or complex RAG pipelines, M3 offers the architectural depth needed to maintain coherence across extended reasoning chains. While many models focus on quick chat interactions, M3 is optimized for structural tasks like heavy-duty coding, multi-step reasoning, and processing massive document sets via its API.

text generationAPI

qwen3.7-plus

qwen
1000000 ctx

Qwen3.7-Plus is a high-efficiency multimodal model designed for developers requiring a balance between reasoning depth and operational cost. While part of a larger series, the 'Plus' iteration focuses on optimizing the performance-to-latency ratio, making it ideal for high-throughput production environments. It natively supports interleaved text and image inputs, allowing for complex visual reasoning tasks such as document parsing, UI automation, and visual Q&A. For engineers, the standout feature is its massive 1,000,000 token context window, which significantly lowers the barrier for processing entire codebases or massive technical datasets without complex RAG architectures. Compared to larger flagship models, Qwen3.7-Plus offers a more streamlined integration path for real-time applications where cost-per-token and response speed are critical constraints. It is built for seamless API integration, targeting developers who need robust multimodal intelligence without the overhead of massive parameter counts.

text generationAPI

nemotron-3-ultra-550b-a55b:free

nvidia
1000000 ctx

Nemotron-3-Ultra is a high-performance Mixture-of-Experts (MoE) model designed for complex reasoning and orchestration tasks. Unlike dense models, it utilizes a hybrid Transformer-Mamba architecture, activating only 55B parameters out of a 550B total pool. For developers, this means you get frontier-level intelligence with significantly lower latency and higher throughput during inference. The model is particularly effective for long-context workflows, supporting up to a 1M token window, making it a strong candidate for massive document analysis, codebase reasoning, and multi-step agentic orchestration. While many models struggle with context decay, the Mamba integration provides a more efficient way to handle linear scaling in long sequences. If your stack requires an orchestrator that can manage complex tool-calling or synthesize information from vast datasets without the typical overhead of massive dense models, this is a highly competitive option.

text generationAPI

nemotron-3-ultra-550b-a55b

nvidia
262144 ctx

Nemotron-3-Ultra-550B is a high-density Mixture-of-Experts (MoE) model designed specifically for complex reasoning and multi-step orchestration tasks. While the total parameter count sits at 550B, its architecture utilizes only 55B active parameters per token, offering a strategic balance between massive knowledge retrieval and inference efficiency. What sets this model apart for developers is its hybrid Transformer-Mamba backbone, which aims to optimize long-context processing and sequential data modeling. Unlike standard dense models, this architecture is built to handle sophisticated agentic workflows where logical consistency and instruction following are paramount. For teams integrating AI into production pipelines, it serves as a robust backbone for autonomous agents, complex code generation, and structured data extraction. It bridges the gap between lightweight specialized models and massive, computationally expensive frontier models, providing a scalable middle ground for high-throughput reasoning applications.

text generationAPI

nemotron-3.5-content-safety:free

nvidia
128000 ctx

For developers building production-grade LLM applications, managing toxicity and safety without sacrificing latency is a constant struggle. Nemotron-3.5-Content-Safety addresses this by providing a specialized, high-efficiency guardrail layer. Built on the Google Gemma-3-4B architecture, this 4B-parameter model is optimized specifically for moderation rather than general reasoning. Unlike massive general-purpose models that can be overkill for safety checks, this compact model is designed to intercept problematic inputs and outputs in both text and vision modalities. It functions as a multimodal filter, making it ideal for developers integrating VLMs where visual context must be scrutinized alongside text. Because it is fine-tuned for safety-specific classification, it offers a more surgical approach to content moderation than zero-shot prompting on larger models, allowing for tighter integration into your inference pipelines with minimal overhead.

text generationAPI

nemotron-3.5-content-safety

nvidia
131072 ctx

NVIDIA Nemotron-3.5-Content-Safety is a specialized 4B-parameter guardrail model designed to sit in the inference pipeline between users and large-scale models. Built upon the Gemma-3 architecture, this multimodal model addresses a critical gap in production AI: the need for low-latency, high-accuracy content moderation for both text and vision inputs. Unlike massive general-purpose models that are too slow for real-time filtering, this compact model is optimized for high-throughput deployment. It functions as a bidirectional safety layer, inspecting incoming user prompts for malicious intent and auditing outgoing model responses to ensure compliance with safety guidelines. For developers building agentic workflows or customer-facing VLMs, it provides a scalable way to mitigate toxicity and policy violations without the massive computational overhead of larger models. Its multimodal capability makes it particularly useful for applications where image-to-text or visual reasoning is central to the user experience.

text generationAPI

kimi-k2.7-code

moonshotai
262144 ctx

kimi-k2.7-code is a specialized MoE (Mixture-of-Experts) model engineered specifically for the software development lifecycle. Unlike general-purpose LLMs, this iteration is optimized for high-precision code generation and complex architectural reasoning across extended contexts. For developers working on large-scale repositories, the standout feature is its ability to maintain coherence over massive codebases, making it highly effective for end-to-end refactoring, debugging, and feature implementation. It moves beyond simple snippet completion, offering deep logical understanding suitable for integrating into automated CI/CD pipelines or custom IDE extensions. While many models struggle with context window degradation, K2.7 is built to handle long-range dependencies, making it a competitive alternative for teams requiring robust, production-ready code assistance within a multimodal framework.

text generationAPI

fusion

openrouter
1000000 ctx

Fusion is a sophisticated orchestration layer designed for complex reasoning tasks that single-model architectures often struggle to resolve. Instead of a linear inference path, Fusion implements a multi-model deliberation pattern. When a prompt is received, it triggers a parallel execution of specialized expert models, augmented by real-time web search and data fetching capabilities. For developers, this means moving away from 'one-shot' prompting toward a structured ensemble approach. It is particularly effective for high-stakes research, multi-step technical troubleshooting, and tasks requiring up-to-the-minute factual accuracy. While standard LLMs rely on static training data, Fusion’s integration of live web retrieval and cross-model verification significantly reduces hallucinations. Integrating via OpenRouter, it provides a scalable way to implement agentic workflows without managing the underlying complexity of multi-model consensus logic yourself.

text generationAPI

glm-5.2:free

z-ai
32768 ctx

GLM-5.2 is a high-reasoning model designed specifically for complex, multi-step technical workflows. Unlike standard chat models optimized for short interactions, this iteration prioritizes long-horizon planning and deep logical reasoning. For developers, the standout feature is the massive 1M-token context window, which makes it a viable engine for project-level software engineering, entire codebase analysis, and massive documentation retrieval. It is built to act as a core reasoning engine for autonomous agents, capable of maintaining coherence across extended task sequences. While many models struggle with context drift in large-scale tasks, GLM-5.2 is engineered to handle the high density of information required for advanced RAG pipelines and automated debugging. If your stack requires an LLM that functions more like a collaborative engineer than a simple text predictor, this model provides the necessary architectural depth for integration into sophisticated agentic frameworks.

text generationAPI

glm-5.2:batch

z-ai
1048576 ctx

GLM-5.2:batch is a high-capacity reasoning model engineered specifically for intensive, long-context workloads. Unlike standard chat models, this iteration is optimized for stability and throughput in complex, multi-step environments. With a massive 1M-token context window, it is purpose-built for developers tackling project-level software engineering, deep codebase analysis, and long-horizon agentic workflows where maintaining state over thousands of lines of code is critical. While many models struggle with 'lost in the middle' phenomena during long sequences, GLM-5.2 is designed to maintain high retrieval accuracy and logical consistency across its entire window. For integration, it serves as a robust backbone for autonomous agents that require deep reasoning rather than just quick responses. If your use case involves processing entire documentation sets or managing complex repository refactoring, this model provides the necessary architectural depth to handle those dependencies without losing context.

text generationAPI

glm-5.2

z-ai
1048576 ctx

GLM-5.2 is a specialized reasoning model designed to handle high-complexity, long-context workflows that typically break standard LLMs. While many models struggle with coherence as context grows, GLM-5.2 leverages a massive 1M-token window specifically optimized for project-level software engineering and autonomous agent orchestration. For developers, this means you can feed entire codebases, massive documentation sets, or multi-step execution logs into a single prompt without losing structural integrity. Unlike general-purpose chat models, its architecture prioritizes long-horizon planning, making it a strong candidate for building agents that require multi-step reasoning and memory retention. Integration is handled via API, making it a drop-in replacement for developers moving from smaller models to more robust, agentic frameworks where deep logical consistency is the primary requirement.

text generationAPI

north-mini-code:free

cohere
256000 ctx

North Mini Code marks Cohere's strategic pivot into agentic coding workflows. Built on a sparse Mixture-of-Experts (MoE) architecture, the model manages 30B total parameters but maintains high efficiency by utilizing only 3B active parameters per token. For developers, this translates to lower latency and improved throughput without sacrificing the reasoning depth required for complex software engineering tasks. Unlike general-purpose LLMs, this model is specifically tuned for code generation, refactoring, and autonomous debugging within agentic frameworks. With a massive 256k context window, it excels at ingestion of entire repositories or extensive documentation, allowing for high-fidelity codebase awareness. While larger monolithic models might offer slightly broader knowledge, North Mini Code is optimized for the specific speed-to-intelligence ratio required for real-time IDE integration and automated CI/CD pipeline agents.

text generationAPI

fugu-ultra

sakana
1000000 ctx

Fugu-Ultra represents a shift away from traditional monolithic architectures toward a learned multi-agent orchestration system. Instead of relying on a single massive parameter set to handle every task, this model functions as an intelligent router trained to decompose complex queries and delegate them to specialized sub-agents. For developers, this means higher precision in reasoning and a significant reduction in 'hallucination noise' typically seen when a single model attempts to context-switch between vastly different domains. Integration is handled via API, making it a drop-in replacement for standard LLM workflows where high-level logic and multi-step planning are required. While standard models often struggle with long-context coherence, Fugu-Ultra’s orchestration layer manages information flow more effectively across its 1M token window. It is particularly suited for complex software engineering pipelines, multi-document legal analysis, and autonomous agentic workflows where the quality of the decision-making process is more critical than raw throughput.

text generationAPI

laguna-xs-2.1:free

poolside
262144 ctx

Laguna XS 2.1 is a specialized coding agent designed for high-efficiency software engineering workflows. Built on a 33B-parameter architecture, it is optimized for the specific logic requirements of code generation, refactoring, and complex debugging. Unlike general-purpose LLMs, this model focuses on maintaining structural integrity and following strict syntax patterns across multiple languages. For developers, the standout feature is the massive 262,144 context window, which allows you to feed entire repositories or extensive documentation into a single prompt without losing coherence. This makes it particularly effective for large-scale refactoring tasks or navigating legacy codebases where cross-file dependency awareness is critical. While it may not match the broad creative writing capabilities of massive frontier models, its performance-to-latency ratio makes it a superior choice for integration into IDE extensions and automated CI/CD pipelines where rapid, accurate code completions are required.

text generationAPI

laguna-xs-2.1

poolside
262144 ctx

Laguna-XS-2.1 is a specialized Mixture-of-Experts (MoE) model designed specifically for high-performance coding tasks. Operating in the 33B-A3B parameter class, it optimizes the trade-off between inference speed and reasoning depth, making it an ideal candidate for real-time IDE integrations and autonomous coding agents. Unlike massive general-purpose models, this iteration focuses on refined logic for complex software engineering workflows, such as codebase refactoring, unit test generation, and multi-file debugging. With a massive 262k context window, it can ingest entire repositories or extensive documentation sets to maintain high architectural awareness during long-form development sessions. For developers, this means lower latency in autocomplete loops and more reliable output when working across deep dependency trees compared to its predecessor, XS.2.

text generationAPI

hy3

tencent
262144 ctx

Hy3 is a high-capacity Mixture-of-Experts (MoE) model from Tencent, engineered specifically for complex reasoning and autonomous agentic workflows. With a total parameter count of 295B and a highly granular architecture featuring 192 experts, it utilizes top-8 routing to maintain efficiency while delivering deep cognitive capabilities. Unlike dense models that struggle with scaling costs, Hy3’s 21B active parameter design allows for sophisticated multi-step logic and tool-use integration without the typical latency overhead. For developers, the standout feature is the configurable reasoning effort, which lets you tune the model's computational intensity based on the task complexity—ranging from rapid text generation to deep, iterative problem-solving. This makes it a versatile backbone for production-grade RAG pipelines, automated coding assistants, and long-context agentic loops. It is optimized for stability in real-world deployment scenarios where predictable reasoning depth is critical.

text generationAPI
Email