Global AI chat room · 16 online now Join now
DIRECTORY / 02

AI Models | Open-Source LLM Directory

Discover and compare open-source LLMs, language models and multimodal models by capability, scale, license, downloads and provenance.

Compare modelsFind the right building block for your next workflow
Directory overview
591
curated entries
24 topic groupsLive
02 / MODEL INDEX

Find the right model for the job

Context first, better decisions. Every entry keeps the signal that matters.

CURATED DIRECTORY591 results

inkling:free

thinkingmachines
1048576 ctx

Inkling:free is a high-efficiency multimodal Mixture-of-Experts (MoE) model designed for developers building complex, agentic workflows. While the total parameter count sits at 975B, the architecture optimizes performance by utilizing only 41B active parameters per token, offering a massive knowledge base with the inference speed typically associated with much smaller models. For engineers, the standout feature is the massive 1M+ token context window, which makes it ideal for deep codebase analysis, long-form document reasoning, and maintaining state in multi-turn agentic loops. Unlike dense models that scale latency linearly with parameter count, Inkling provides a pragmatic middle ground for tool-use and automated reasoning tasks. It is particularly well-suited for integration into RAG pipelines and autonomous coding assistants where both high-level reasoning and rapid response times are critical requirements.

text generationAPI

inkling:batch

thinkingmachines
524288 ctx

inkling:batch is a massive-scale multimodal Mixture-of-Experts (MoE) model from Thinking Machines Lab, engineered for high-throughput reasoning and complex agentic workflows. While the total parameter count sits at 975B, the architecture optimizes efficiency by activating only 41B parameters per token, making it a competitive choice for developers needing deep logic without the latency of dense models. It excels in code generation, multi-step tool use, and long-context reasoning, supported by a substantial 524k context window. For teams building autonomous agents or RAG pipelines, inkling:batch offers a robust middle ground: the intelligence of a frontier-class model with the specialized efficiency of an MoE structure. It is primarily accessible via API, making it easy to integrate into existing production environments that require reliable, scalable reasoning capabilities.

text generationAPI

inkling

thinkingmachines
524288 ctx

Inkling is a high-efficiency multimodal MoE model engineered for developers building complex, agentic workflows. While its total parameter count reaches 975B, the sparse architecture utilizes only 41B active parameters per token, offering a massive knowledge base with the inference latency typically associated with much smaller models. For engineers, the primary value proposition lies in its specialized training for tool-use and multi-step reasoning, making it a strong candidate for autonomous agents and automated coding assistants. Unlike dense models that struggle with scaling reasoning capabilities without massive compute overhead, Inkling’s mixture-of-experts approach provides a high performance-to-cost ratio. It supports a massive 1M context window, allowing for the ingestion of entire codebases or extensive documentation in a single prompt. Whether you are integrating it via API for scalable production apps or fine-tuning its open weights for domain-specific logic, Inkling is built to handle high-reasoning density tasks that standard LLMs often fail.

text generationAPI

longcat-2.0

meituan
1048756 ctx

LongCat 2.0 is a sparse Mixture-of-Experts (MoE) model designed specifically for high-complexity engineering workflows. While it boasts a massive 1.6T total parameter scale, its architectural efficiency shines through 48B active parameters, balancing high-reasoning capabilities with manageable compute requirements. What sets this model apart for developers is its massive 1M+ token context window, which moves beyond simple chat interactions into true repository-level intelligence. It is engineered for tasks that demand long-horizon planning, such as executing multi-step agentic workflows, refactoring entire codebases, and maintaining coherence across massive documentation sets. For teams building autonomous coding agents or deep-context RAG pipelines, LongCat 2.0 provides the structural depth needed to handle dependencies and logic that standard dense models often lose in long-sequence processing.

text generationAPI

laguna-s-2.1:free

poolside
262144 ctx

Laguna S 2.1 is a specialized Mixture-of-Experts (MoE) model engineered specifically for high-performance software engineering workflows. Built by Poolside, the architecture utilizes a 118B total parameter structure with 8B active parameters per token, striking a balance between deep reasoning capabilities and low-latency inference. Unlike general-purpose LLMs, this model is fine-tuned for terminal interaction, complex debugging, and codebase navigation, as evidenced by its strong performance on Terminal-Bench 2.1. For developers, this means more reliable command-line execution and better context awareness during automated refactoring tasks. It supports a massive 262k context window, making it suitable for ingesting entire repositories or extensive documentation for RAG-based development tools. Whether you are integrating it into a CLI agent or a custom IDE extension via API, Laguna S 2.1 offers a highly efficient alternative to heavier, non-specialized models for autonomous coding agents.

text generationAPI

laguna-s-2.1

poolside
1048576 ctx

Laguna S 2.1 is a specialized Mixture-of-Experts (MoE) model engineered specifically for software engineering workflows. While it boasts a 118B total parameter architecture, its 8B active parameter design ensures low-latency inference, making it highly efficient for real-time IDE integrations and automated agentic tasks. Unlike general-purpose LLMs, this model is optimized for terminal interaction and complex codebase reasoning, evidenced by its high performance on the Terminal-Bench 2.1 benchmark. For developers building autonomous coding agents or sophisticated CI/CD automation tools, Laguna S 2.1 offers a high-context (1M tokens) solution that balances massive architectural knowledge with the speed required for iterative development cycles. It is best utilized via API for tasks ranging from automated bug fixing to large-scale refactoring across distributed repositories.

text generationAPI

ling-3.0-flash

inclusionai
262144 ctx

For developers building high-throughput agentic workflows, Ling-3.0-flash offers a strategic balance between intelligence and latency. Built on a 124B Mixture-of-Experts (MoE) architecture, it optimizes compute by activating only 5.1B parameters per token. This design makes it particularly effective for production environments where cost-per-token and inference speed are critical bottlenecks. Unlike dense models that struggle with scaling costs, Ling-3.0-flash is engineered for complex reasoning tasks and long-context orchestration, supporting a massive 262,144 token window. This makes it a strong candidate for RAG pipelines, multi-step agentic reasoning, and large-scale data processing. If your stack requires a model that can handle deep contextual memory without the typical latency penalties of larger dense models, this is a highly efficient integration choice.

text generationAPI

qwen3.7-flash

qwen
1000000 ctx

Qwen3.7-Flash is a high-speed, multimodal model engineered for developers building agentic workflows that require tight integration between visual perception and logical reasoning. Unlike standard LLMs, this model is optimized for low-latency tasks involving spatial intelligence and computer interaction, making it a strong candidate for UI automation and visual coding assistants. It excels at decomposing complex visual scenes into actionable data, which is critical for multimodal agents navigating real-world environments or digital interfaces. For engineering teams, the primary value lies in its balance of high-throughput processing and sophisticated object recognition. While larger models might offer deeper nuance, Flash is designed to minimize inference costs and latency in production pipelines where rapid visual feedback loops are essential. It fits well into existing API-driven architectures, particularly for applications requiring real-time visual search or automated GUI testing.

text generationAPI

inkling-small:free

thinkingmachines
1048576 ctx

Inkling-small is a high-efficiency multimodal Mixture-of-Experts (MoE) model designed for developers who need large-scale reasoning capabilities without the massive compute overhead of dense models. While it sits on a 276B total parameter architecture, it utilizes only 12B active parameters per token, making it highly optimized for low-latency inference and high-throughput API integration. For developers working with complex multimodal datasets, its standout feature is the massive 1M token context window, which allows for deep document analysis and long-form reasoning that standard small models cannot handle. Compared to dense 7B or 13B models, Inkling-small offers a significantly higher intelligence ceiling by leveraging its MoE structure, making it an ideal candidate for RAG pipelines, long-context summarization, and multimodal agentic workflows where cost-to-performance ratios are critical.

text generationAPI

inkling-small

thinkingmachines
524288 ctx

Inkling-small is a specialized multimodal Mixture-of-Experts (MoE) model designed for developers who need high-reasoning capabilities without the massive compute overhead of dense trillion-parameter models. While it sits within a 276B total parameter architecture, it only activates 12B parameters per token, offering a highly efficient inference profile that balances throughput with intelligence. For international engineering teams, this means you can deploy sophisticated multimodal workflows—combining vision and text—on more modest hardware compared to traditional dense models. It excels in complex reasoning tasks and long-context applications, supported by a massive 1M token context window. Whether you are integrating it via API for rapid prototyping or leveraging its open-weight nature for fine-tuning, inkling-small provides a scalable middle ground between lightweight edge models and heavy-duty frontier LLMs.

text generationAPI

deepseek-v4-flash-0731:free

deepseek
1048576 ctx

DeepSeek-V4-Flash-0731 is a high-efficiency sparse Mixture-of-Experts (MoE) model engineered for low-latency, high-throughput applications. While the total parameter count sits at 284B, the architecture only activates 13B parameters per token, offering a pragmatic middle ground between massive dense models and lightweight edge models. For developers, this translates to significantly reduced inference costs and faster time-to-first-token without sacrificing the reasoning depth required for complex logic. The model is specifically tuned for high-density workloads such as automated code generation, multi-step agentic reasoning, and long-context data extraction. With a massive 1M token context window, it is particularly well-suited for analyzing entire codebases or massive document repositories. If your workflow requires balancing sophisticated instruction following with the speed necessary for real-time agent loops, this model serves as a highly competitive alternative to larger, more expensive proprietary models.

text generationAPI

deepseek-v4-flash-0731:batch

deepseek
1048576 ctx

DeepSeek-V4-Flash-0731 is a high-throughput Mixture-of-Experts (MoE) model designed specifically for latency-sensitive applications. While it boasts a massive 284B total parameter architecture, it utilizes a sparse routing mechanism that activates only 13B parameters per token, offering a significant efficiency advantage for high-volume batch processing. For developers, this means a superior performance-to-cost ratio compared to dense models of similar scale. The model is optimized for complex reasoning, code generation, and multi-step agentic workflows, supported by an expansive 1M token context window. Unlike general-purpose LLMs that struggle with long-form document analysis or deep codebase navigation, this revision focuses on maintaining logical coherence across extended sequences. It is an ideal candidate for integrating into automated DevOps pipelines, large-scale data extraction tasks, or autonomous agent frameworks where speed and context depth are non-negotiable requirements.

text generationAPI

deepseek-v4-flash-0731

deepseek
1310720 ctx

DeepSeek-V4-Flash-0731 is a high-efficiency sparse Mixture-of-Experts (MoE) model designed to balance massive scale with low-latency execution. While the total parameter count sits at 284B, the architecture only activates 13B parameters per token, making it an ideal candidate for developers building real-time agentic workflows or complex reasoning loops where cost-per-token and speed are critical constraints. Unlike monolithic dense models, this version is specifically optimized through post-training to excel in structured tasks like code generation and multi-step logical reasoning. For engineers integrating via API, the model offers a massive 131k context window, allowing for deep document analysis and large-scale codebase ingestion without the typical memory overhead seen in traditional large models. It positions itself as a high-performance alternative to larger proprietary models, offering competitive reasoning capabilities at a fraction of the inference latency.

text generationAPI

muse-spark-1.2

meta
1048576 ctx

Muse Spark 1.2 is Meta's latest reasoning-focused model, specifically architected to power autonomous agentic workflows. Unlike standard LLMs that struggle with long-range dependencies, this model leverages a massive 1M-token context window, making it highly effective for analyzing entire codebases, lengthy technical documentation, or multi-hour video files. It is natively multimodal, processing text, images, audio, and video inputs to generate structured text outputs. For developers, the primary value lies in its ability to handle complex, multi-step reasoning tasks that require cross-modal understanding—such as debugging via video screen recordings or extracting insights from massive PDF archives. While many models require specialized vision or audio wrappers, Muse Spark 1.2 integrates these modalities into a single reasoning engine, simplifying the stack for developers building sophisticated AI agents and complex RAG pipelines.

text generationAPI

muse-glimmer-30b:batch

meta
131072 ctx

Muse Glimmer 30B is a dense, open-weight multimodal model designed specifically for developers building autonomous agents. Unlike larger, resource-heavy models, Glimmer is distilled from the Muse Spark architecture to run efficiently on consumer-grade hardware without sacrificing reasoning depth. It excels in long-horizon planning and multi-step task execution, making it a practical choice for agentic workflows where latency and cost-per-token are critical constraints. With a massive 131k context window, it handles large datasets and extended conversation histories with ease. For developers, the primary advantage lies in its balance: it offers the multimodal capabilities required for complex environments while remaining small enough to facilitate local deployment or low-latency API integration. If your roadmap involves autonomous tool-use or complex reasoning loops on edge devices, this model provides a highly optimized middle ground between lightweight SLMs and massive frontier models.

text generationAPI

muse-glimmer-30b

meta
131072 ctx

Muse Glimmer 30B is a specialized open-weight multimodal model designed specifically for developers building autonomous agents. While many models struggle with task drift during complex workflows, Glimmer is distilled from the larger Muse Spark architecture to maintain high reasoning density within a 30B parameter footprint. This makes it uniquely capable of running on high-end consumer hardware without sacrificing the long-horizon planning required for agentic loops. With a massive 131k context window, it handles extensive documentation and multi-turn history with ease. For developers, the primary value lies in its balance: it offers the multimodal reasoning of much larger proprietary models but remains accessible for local deployment and fine-tuning. Whether you are building autonomous web navigators, complex coding assistants, or long-context RAG pipelines, Glimmer provides a predictable, low-latency backbone that bridges the gap between massive cloud models and efficient edge deployment.

text generationAPI

solar-pro4

upstage
524288 ctx

Solar Pro 4 is a specialized LLM designed for developers building high-throughput, long-context applications. Unlike general-purpose models that struggle with information retrieval in massive datasets, Solar Pro 4 features a 524K context window, making it a robust choice for RAG (Retrieval-Augmented Generation) pipelines and complex document analysis. It is specifically optimized for agentic workflows, where the model must maintain state and follow multi-step reasoning across extended instruction sets. For engineering teams, the primary value proposition lies in its balance of cost-efficiency and performance in office productivity automation and heavy-duty text processing. While larger frontier models offer raw reasoning power, Solar Pro 4 is engineered to be the reliable engine for production-grade agents that require deep context without the prohibitive latency or expense of massive parameter models.

text generationAPI

sakana-namazu

sakana
262144 ctx

Sakana Namazu is a specialized reasoning model engineered specifically for high-fidelity Japanese language processing and localized business logic. Built upon the Kimi K2.6 architecture, it moves beyond simple translation by incorporating deep training in Japanese-specific instruction following and professional context awareness. For developers building localized applications, Namazu addresses the common pitfalls of general-purpose LLMs, such as unnatural syntax or cultural misalignment in formal settings. It features a substantial 262k context window, making it highly effective for long-form document analysis, complex legal summarization, and multi-turn reasoning within Japanese enterprise workflows. While general models excel at broad multilingual tasks, Namazu is optimized for developers who require high precision in Japanese semantic nuances and structured business reasoning via API integration.

text generationAPI

nemotron-3.5-lightning:free

nvidia
1000000 ctx

For developers building high-concurrency applications, Nemotron-3.5-Lightning offers a compelling balance between latency and intelligence. Built on a Mixture-of-Experts (MoE) architecture, it utilizes only 3B active parameters out of a 30B total, which significantly optimizes inference speed without the typical performance degradation seen in smaller dense models. This makes it an ideal candidate for agentic workflows where rapid-fire reasoning and tool-calling are required. Unlike general-purpose monolithic models, this version is specifically tuned for high-throughput environments. If your stack requires low-latency text generation, complex instruction following, or real-time data processing within an API-driven architecture, Nemotron-3.5-Lightning provides a specialized alternative to larger, more expensive models. It bridges the gap between lightweight edge models and heavy-duty LLMs, focusing on efficiency for specialized, task-oriented deployments.

text generationAPI

nemotron-3.5-lightning

nvidia
262144 ctx

Nemotron-3.5-Lightning is NVIDIA’s high-efficiency Mixture-of-Experts (MoE) model designed specifically for low-latency, high-throughput environments. While the total parameter count sits at 30B, it only utilizes 3B active parameters per token, offering a massive performance leap for developers needing rapid inference without the computational overhead of dense large-scale models. For engineers building agentic workflows, tool-calling loops, or real-time RAG pipelines, this model strikes a pragmatic balance between reasoning depth and execution speed. Unlike general-purpose heavyweights, Lightning is optimized for specialized task execution and high-frequency API calls. It integrates seamlessly into existing NVIDIA-optimized stacks and is particularly effective when you need to scale agentic reasoning across thousands of concurrent sessions where traditional LLMs would become a cost or latency bottleneck.

text generationAPI

lfm-2.5-2.6b:free

liquid
65536 ctx

LFM-2.5-2.6B is a specialized small language model (SLM) from Liquid AI designed for high-efficiency reasoning within a compact parameter footprint. Unlike general-purpose massive models, this architecture is optimized for structured workflows where latency and cost-per-token are critical constraints. Developers should look to this model for high-density tasks such as complex data extraction, RAG-based retrieval pipelines, and long-context information synthesis. While its 65k context window makes it a strong candidate for processing large document sets, it is important to note its specific design intent: it excels at analytical reasoning and pattern recognition but is not optimized for autonomous code generation. For engineering teams building agentic loops or automated data processing pipelines, LFM-2.5 provides a lightweight alternative to larger LLMs, offering a better balance of throughput and reasoning depth for specialized, non-coding logic tasks.

text generationAPI

grok-4.6

x-ai
500000 ctx

Grok 4.6 represents a significant step forward in high-reasoning LLMs, specifically optimized for complex engineering and mathematical workflows. For developers, the primary value proposition lies in its specialized performance across STEM disciplines and sophisticated code generation tasks. Unlike general-purpose models that prioritize conversational fluidity, Grok 4.6 is architected to handle deep logic and technical knowledge retrieval with higher precision. The model supports a massive 500,000 token context window, making it highly effective for analyzing entire codebases, massive documentation sets, or long-form technical specifications without losing coherence. Integration is streamlined via API, allowing for seamless deployment into automated CI/CD pipelines, technical RAG (Retrieval-Augmented Generation) systems, and advanced coding assistants. While many models struggle with multi-step logical reasoning in niche scientific domains, Grok 4.6 is positioned as a frontier tool for developers building high-stakes, logic-heavy applications.

text generationAPI

deepseek-v4-pro-0813:batch

deepseek
1048576 ctx

DeepSeek-V4-Pro-0813 is a high-capacity Mixture-of-Experts (MoE) model designed for developers requiring massive throughput and high-reasoning capabilities. Unlike dense models, this MoE architecture optimizes compute efficiency, making it ideal for complex logic tasks, large-scale data synthesis, and sophisticated code generation. With a massive 1,048,576 token context window, it is purpose-built for analyzing entire codebases, long-form documentation, or massive unstructured datasets in a single pass. For engineers building agentic workflows or RAG pipelines, this model offers a significant advantage in handling long-range dependencies that typically cause context fragmentation in smaller models. While it excels in general text generation, its true strength lies in high-volume batch processing where reasoning depth and context retention are non-negotiable. Integration is straightforward via API, making it a viable alternative to larger closed-source models for enterprise-grade automation.

text generationAPI

deepseek-v4-pro-0813

deepseek
1048576 ctx

DeepSeek-V4-Pro-0813 is a high-performance Mixture-of-Experts (MoE) model designed for developers requiring a balance between massive scale and inference efficiency. Unlike dense architectures, this MoE implementation optimizes compute routing, allowing it to handle complex reasoning and high-throughput text generation tasks without the typical latency overhead. With a massive 1,048,576 token context window, it is purpose-built for long-form document analysis, codebase auditing, and processing extensive multi-turn dialogues. For integration, the model is accessible via API, making it a viable drop-in replacement for developers migrating from other large-scale providers who need to maintain high logical reasoning capabilities while managing long-context retrieval. It excels in scenarios where precision in instruction following and deep semantic understanding are non-negotiable, particularly in RAG pipelines and automated software engineering workflows.

text generationAPI
Email