Global AI chat room · 5 online now Join now
DIRECTORY / 02

AI Models | Open-Source LLM Directory

Discover and compare open-source LLMs, language models and multimodal models by capability, scale, license, downloads and provenance.

Compare modelsFind the right building block for your next workflow
Directory overview
839
curated entries
24 topic groupsLive
02 / MODEL INDEX

Find the right model for the job

Context first, better decisions. Every entry keeps the signal that matters.

CURATED DIRECTORY839 results

solar-pro4

upstage
524288 ctx

Solar Pro 4 is a specialized LLM designed for developers building high-throughput, long-context applications. Unlike general-purpose models that struggle with information retrieval in massive datasets, Solar Pro 4 features a 524K context window, making it a robust choice for RAG (Retrieval-Augmented Generation) pipelines and complex document analysis. It is specifically optimized for agentic workflows, where the model must maintain state and follow multi-step reasoning across extended instruction sets. For engineering teams, the primary value proposition lies in its balance of cost-efficiency and performance in office productivity automation and heavy-duty text processing. While larger frontier models offer raw reasoning power, Solar Pro 4 is engineered to be the reliable engine for production-grade agents that require deep context without the prohibitive latency or expense of massive parameter models.

text generationAPI

sakana-namazu

sakana
262144 ctx

Sakana Namazu is a specialized reasoning model engineered specifically for high-fidelity Japanese language processing and localized business logic. Built upon the Kimi K2.6 architecture, it moves beyond simple translation by incorporating deep training in Japanese-specific instruction following and professional context awareness. For developers building localized applications, Namazu addresses the common pitfalls of general-purpose LLMs, such as unnatural syntax or cultural misalignment in formal settings. It features a substantial 262k context window, making it highly effective for long-form document analysis, complex legal summarization, and multi-turn reasoning within Japanese enterprise workflows. While general models excel at broad multilingual tasks, Namazu is optimized for developers who require high precision in Japanese semantic nuances and structured business reasoning via API integration.

text generationAPI

nemotron-3.5-lightning:free

nvidia
1000000 ctx

For developers building high-concurrency applications, Nemotron-3.5-Lightning offers a compelling balance between latency and intelligence. Built on a Mixture-of-Experts (MoE) architecture, it utilizes only 3B active parameters out of a 30B total, which significantly optimizes inference speed without the typical performance degradation seen in smaller dense models. This makes it an ideal candidate for agentic workflows where rapid-fire reasoning and tool-calling are required. Unlike general-purpose monolithic models, this version is specifically tuned for high-throughput environments. If your stack requires low-latency text generation, complex instruction following, or real-time data processing within an API-driven architecture, Nemotron-3.5-Lightning provides a specialized alternative to larger, more expensive models. It bridges the gap between lightweight edge models and heavy-duty LLMs, focusing on efficiency for specialized, task-oriented deployments.

text generationAPI

nemotron-3.5-lightning

nvidia
262144 ctx

Nemotron-3.5-Lightning is NVIDIA’s high-efficiency Mixture-of-Experts (MoE) model designed specifically for low-latency, high-throughput environments. While the total parameter count sits at 30B, it only utilizes 3B active parameters per token, offering a massive performance leap for developers needing rapid inference without the computational overhead of dense large-scale models. For engineers building agentic workflows, tool-calling loops, or real-time RAG pipelines, this model strikes a pragmatic balance between reasoning depth and execution speed. Unlike general-purpose heavyweights, Lightning is optimized for specialized task execution and high-frequency API calls. It integrates seamlessly into existing NVIDIA-optimized stacks and is particularly effective when you need to scale agentic reasoning across thousands of concurrent sessions where traditional LLMs would become a cost or latency bottleneck.

text generationAPI

lfm-2.5-2.6b:free

liquid
65536 ctx

LFM-2.5-2.6B is a specialized small language model (SLM) from Liquid AI designed for high-efficiency reasoning within a compact parameter footprint. Unlike general-purpose massive models, this architecture is optimized for structured workflows where latency and cost-per-token are critical constraints. Developers should look to this model for high-density tasks such as complex data extraction, RAG-based retrieval pipelines, and long-context information synthesis. While its 65k context window makes it a strong candidate for processing large document sets, it is important to note its specific design intent: it excels at analytical reasoning and pattern recognition but is not optimized for autonomous code generation. For engineering teams building agentic loops or automated data processing pipelines, LFM-2.5 provides a lightweight alternative to larger LLMs, offering a better balance of throughput and reasoning depth for specialized, non-coding logic tasks.

text generationAPI

grok-4.6

x-ai
500000 ctx

Grok 4.6 represents a significant step forward in high-reasoning LLMs, specifically optimized for complex engineering and mathematical workflows. For developers, the primary value proposition lies in its specialized performance across STEM disciplines and sophisticated code generation tasks. Unlike general-purpose models that prioritize conversational fluidity, Grok 4.6 is architected to handle deep logic and technical knowledge retrieval with higher precision. The model supports a massive 500,000 token context window, making it highly effective for analyzing entire codebases, massive documentation sets, or long-form technical specifications without losing coherence. Integration is streamlined via API, allowing for seamless deployment into automated CI/CD pipelines, technical RAG (Retrieval-Augmented Generation) systems, and advanced coding assistants. While many models struggle with multi-step logical reasoning in niche scientific domains, Grok 4.6 is positioned as a frontier tool for developers building high-stakes, logic-heavy applications.

text generationAPI

deepseek-v4-pro-0813:batch

deepseek
1048576 ctx

DeepSeek-V4-Pro-0813 is a high-capacity Mixture-of-Experts (MoE) model designed for developers requiring massive throughput and high-reasoning capabilities. Unlike dense models, this MoE architecture optimizes compute efficiency, making it ideal for complex logic tasks, large-scale data synthesis, and sophisticated code generation. With a massive 1,048,576 token context window, it is purpose-built for analyzing entire codebases, long-form documentation, or massive unstructured datasets in a single pass. For engineers building agentic workflows or RAG pipelines, this model offers a significant advantage in handling long-range dependencies that typically cause context fragmentation in smaller models. While it excels in general text generation, its true strength lies in high-volume batch processing where reasoning depth and context retention are non-negotiable. Integration is straightforward via API, making it a viable alternative to larger closed-source models for enterprise-grade automation.

text generationAPI

deepseek-v4-pro-0813

deepseek
1048576 ctx

DeepSeek-V4-Pro-0813 is a high-performance Mixture-of-Experts (MoE) model designed for developers requiring a balance between massive scale and inference efficiency. Unlike dense architectures, this MoE implementation optimizes compute routing, allowing it to handle complex reasoning and high-throughput text generation tasks without the typical latency overhead. With a massive 1,048,576 token context window, it is purpose-built for long-form document analysis, codebase auditing, and processing extensive multi-turn dialogues. For integration, the model is accessible via API, making it a viable drop-in replacement for developers migrating from other large-scale providers who need to maintain high logical reasoning capabilities while managing long-context retrieval. It excels in scenarios where precision in instruction following and deep semantic understanding are non-negotiable, particularly in RAG pipelines and automated software engineering workflows.

text generationAPI

seed-2.0-code

bytedance-seed
262144 ctx

Seed 2.0 Code is a specialized model from ByteDance optimized specifically for agentic workflows and high-autonomy coding tasks. Unlike general-purpose LLMs that act primarily as chat interfaces, this model is architected to function as a reasoning engine within autonomous coding agents. It excels in frontend development and complex, multilingual programming environments where context awareness is critical. With a significant 262k context window, it is designed to ingest large codebases, allowing developers to integrate it into IDE-based agents or CI/CD pipelines for automated debugging, refactoring, and feature implementation. For engineers building the next generation of AI software engineers, Seed 2.0 provides the structural reliability needed for multi-step tool use and long-context reasoning that standard models often struggle to maintain.

text generationAPI

qwen3.8-2.4t-a95b:batch

qwen
1010000 ctx

For developers working with massive-scale reasoning tasks, Qwen3.8-2.4T-A95B represents a significant leap in sparse Mixture-of-Experts (MoE) architecture. While the total parameter count sits at 2.4 trillion, the model optimizes inference by activating only 95 billion parameters per token. This design strikes a balance between the deep knowledge density of a frontier-class model and the latency requirements of production environments. It is essentially the open-weight distillation of the Qwen3.8 Max series, making it ideal for complex agentic workflows, sophisticated code generation, and long-context retrieval tasks. Unlike monolithic dense models of similar scale, this MoE approach provides high-throughput capabilities without sacrificing the nuanced reasoning required for multi-step logical deduction. Integration via API allows you to leverage trillion-parameter intelligence for RAG pipelines or autonomous tool-use without the prohibitive hardware overhead of hosting a dense 2T model locally.

text generationAPI

qwen3.8-2.4t-a95b

qwen
1048576 ctx

For developers working with large-scale reasoning tasks, Qwen3.8 2.4T A95B represents a significant step in sparse Mixture-of-Experts (MoE) architecture. While the total parameter count sits at 2.4 trillion, the model only activates 95 billion parameters per token, offering a high-performance profile that balances massive knowledge density with much lower inference latency than a dense model of equivalent scale. This makes it particularly effective for complex multi-step reasoning, sophisticated code generation, and high-context information retrieval. With a massive 1M token context window, it is designed to ingest entire codebases or lengthy documentation sets without losing coherence. Compared to standard dense models, you get the intelligence of a massive frontier model with the throughput efficiency required for production-grade RAG pipelines and autonomous agent workflows. It is an ideal choice for integrating high-level cognitive capabilities into existing software stacks via API.

text generationAPI

seed-2-1-turbo

bytedance-seed
262144 ctx

Seed 2.1 Turbo is a specialized multimodal model from ByteDance designed to bridge the gap between high-level reasoning and practical execution. Unlike standard LLMs that struggle with complex, multi-step processes, this model is architected specifically for long-horizon agent workflows and end-to-end software delivery. For developers, this means moving beyond simple code completion toward autonomous task execution where the model can interpret visual UI elements alongside textual requirements. It excels in environments requiring high context retention, supporting a 262k token window which is critical for navigating large codebases or lengthy documentation. Whether you are building autonomous coding agents or complex automation pipelines that require visual grounding, Seed 2.1 Turbo offers a streamlined API integration for workflows that demand both logical precision and multimodal understanding.

text generationAPI

dots-3-note-preview:free

dots-studio
512000 ctx

Dots-3-Note-Preview is a specialized Mixture-of-Experts (MoE) model designed for high-efficiency text generation. While it sits within a massive 280B parameter architecture, it operates with only 16B active parameters per token, offering a pragmatic balance between reasoning depth and inference speed. For developers, this means you get the sophisticated pattern recognition of a large-scale model without the prohibitive latency typically associated with dense models of this magnitude. With a massive 512,000 context window, it is purpose-built for long-form document analysis, complex codebase summarization, and large-scale data extraction tasks where maintaining long-range dependencies is critical. Unlike standard dense models, its MoE structure allows for more granular task specialization, making it an excellent candidate for RAG pipelines and multi-step reasoning workflows. It serves as an accessible entry point into the Dots 3 ecosystem, optimized for developers who need high-throughput performance on complex, context-heavy workloads.

text generationAPI

qwen3.8-27b:free

qwen
262144 ctx

Qwen3.8-27B is a high-density, open-weight vision-language model engineered for developers who require a balance between reasoning depth and computational efficiency. Unlike standard LLMs, this model integrates multimodal capabilities directly into its architecture, making it highly effective for tasks involving visual reasoning, complex document parsing, and spatial understanding. For engineers building autonomous agents, the model's ability to manage long-running workflows and extended context windows is a significant advantage. It excels in specialized domains such as automated code generation, technical research, and professional-grade data extraction. While larger models offer more raw knowledge, the 27B parameter scale provides a sweet spot for production environments where low-latency inference and high throughput are critical. Whether you are integrating it via API for agentic tasks or deploying it for multimodal RAG pipelines, Qwen3.8 offers a robust, flexible backbone for sophisticated AI applications.

text generationAPI

qwen3.8-27b

qwen
1000000 ctx

Qwen3.8-27B is a dense, open-weight vision-language model designed for developers who need a balance between high-reasoning capabilities and efficient deployment. Unlike purely text-based LLMs, this model integrates multimodal perception, making it capable of processing visual data alongside complex text instructions. It is specifically architected to handle professional-grade workflows, including sophisticated coding tasks, scientific research, and long-context agentic reasoning. For engineers building autonomous agents, the model's ability to maintain coherence over extended task sequences is a significant advantage. While larger models offer higher raw intelligence, the 27B parameter count provides a sweet spot for low-latency integration and cost-effective scaling in production environments. Whether you are implementing visual document parsing or building multi-step reasoning pipelines, Qwen3.8 offers a robust, flexible foundation that competes closely with much larger proprietary systems.

text generationAPI

glm-5.3:batch

z-ai
1048576 ctx

GLM-5.3:batch is a specialized reasoning model engineered specifically for high-complexity software engineering workflows and long-horizon agentic orchestration. Unlike standard chat models optimized for quick interactions, this iteration focuses on maintaining logical consistency across massive datasets, supported by a robust 1-million-token context window. For developers, this means you can ingest entire codebases, extensive documentation, or multi-file logs into a single prompt without losing structural coherence. The 'batch' designation suggests an optimization for high-throughput processing, making it an ideal candidate for automated code reviews, large-scale refactoring tasks, or complex debugging agents that require deep reasoning over extended sequences. While many models struggle with 'lost in the middle' phenomena during long-context retrieval, GLM-5.3 is architected to handle the dependencies inherent in large-scale software architecture, providing a reliable backend for autonomous engineering agents.

text generationAPI

glm-5.3

z-ai
1310720 ctx

GLM-5.3 is a high-reasoning model specifically architected for developers working on complex software engineering workflows and autonomous agent orchestration. Unlike general-purpose chat models, this iteration focuses on long-horizon task execution, making it suitable for multi-step debugging, codebase analysis, and automated system design. The standout technical feature is the 1M-token context window, which allows you to ingest entire repositories or massive documentation sets without losing structural coherence. For integration, it functions via API, providing a scalable backend for building agentic tools that require deep logical consistency over extended sessions. If your roadmap includes building self-correcting coding assistants or sophisticated RAG pipelines that demand high precision in long-context retrieval, GLM-5.3 offers a robust alternative to existing industry standards by prioritizing reasoning depth over simple pattern matching.

text generationAPI

hy-mt2-7b

tencent
8192 ctx

Hy-MT2-7B is a specialized 7B-parameter model from Tencent designed specifically for high-precision machine translation. Unlike general-purpose LLMs that treat translation as a standard text-generation task, this model is architected to handle complex linguistic workflows. It supports 33 major language pairs alongside five specific Chinese dialects and minority languages, making it a robust choice for localized applications in the APAC region. For developers, the standout feature is its granular control over translation logic; it supports structured, delimiter-based, and glossary-based workflows, allowing you to enforce terminology consistency and maintain specific stylistic tones via prompt-driven constraints. With an 8k context window, it can manage longer segments while maintaining contextual coherence. It is ideal for integrating into localization pipelines, technical documentation workflows, or any application requiring high-fidelity cross-lingual mapping with strict adherence to predefined glossaries.

text generationAPI

hy-mt2-30b-a3b

tencent
8192 ctx

For developers building localization pipelines or multilingual interfaces, hy-mt2-30b-a3b represents a specialized shift from general-purpose LLMs toward high-precision translation. Unlike standard models that often struggle with linguistic nuance or structural consistency, this model is architected specifically for translation workflows. It supports 33 language pairs alongside niche Chinese dialects and minority languages, filling a critical gap for localized products in the APAC region. What sets it apart is the built-in support for structured data handling; you can integrate it into existing pipelines using delimiter-based or glossary-based workflows to ensure brand consistency and technical accuracy. Whether you are automating software localization or building real-time translation tools, the model's ability to respect context and specific terminology makes it a more reliable choice than generic chat models for production-grade NMT (Neural Machine Translation) tasks.

text generationAPI

hy-mt2-1.8b

tencent
8192 ctx

For developers building localized applications or real-time communication tools, hy-mt2-1.8b offers a highly efficient, specialized solution for multilingual translation. Unlike general-purpose LLMs that struggle with linguistic nuance at small scales, this 1.8B-parameter model is purpose-built for high-fidelity translation across 33 language pairs, including specific support for Chinese dialects and minority languages. What sets it apart from standard NMT models is its architectural support for complex workflows: you can pass structured delimiters, enforce specific glossaries to maintain brand terminology, and apply style-guided constraints to match a target persona. With an 8k context window, it handles more than just sentence-level snippets, allowing for better contextual awareness in longer passages. It is an ideal candidate for edge deployment or high-throughput microservices where low latency and strict adherence to technical terminology are more critical than broad generative reasoning.

text generationAPI

deepseek-v4-flash-vision-exp:batch

deepseek
1048576 ctx

DeepSeek-V4-Flash-Vision-Exp is an experimental multimodal iteration of the V4 Flash architecture, specifically designed to bridge high-speed text processing with robust visual reasoning. For developers building agentic workflows, this model provides a significant upgrade by allowing the same engine that handles complex logic and tool-calling to also interpret visual inputs. Unlike many vision models that sacrifice reasoning depth for speed, this version maintains the core performance characteristics of the V4 Flash series, making it ideal for real-time applications like automated UI testing, visual document parsing, and multi-modal agent orchestration. It is optimized for high-throughput batch processing, offering a cost-effective way to integrate vision into existing text-based pipelines without the latency overhead typically associated with larger vision-language models. If your stack requires low-latency multimodal feedback loops, this model offers a highly competitive integration path.

text generationAPI

deepseek-v4-flash-vision-exp

deepseek
1048576 ctx

DeepSeek-V4-Flash-Vision-Exp is an experimental multimodal iteration of the V4 Flash architecture, designed to bridge high-speed text processing with robust visual reasoning. For developers, this model offers a significant upgrade over the text-only base by enabling native image understanding within a low-latency framework. It is specifically optimized for agentic workflows where visual context—such as UI screenshots, diagrams, or document scans—must be parsed alongside complex text instructions. Unlike larger, heavier vision models that sacrifice throughput for depth, this 'Flash' variant prioritizes rapid inference and high context window utilization, making it ideal for real-time applications like automated visual QA, accessibility tools, or multimodal RAG pipelines. While it remains in an experimental phase, it maintains the core text capabilities and agentic instruction-following performance of the standard V4 Flash, providing a cost-effective way to integrate vision into existing text-based automation loops.

text generationAPI

muse-spark-1.2-contributor

meta
1048576 ctx

Muse Spark 1.2 Contributor is a specialized reasoning model from Meta designed to bridge the gap between high-performance logic and production-scale cost efficiency. While the flagship Muse Spark models excel at complex architectural reasoning, the Contributor tier is optimized for developers who need consistent logical throughput without the premium price tag. It maintains a massive 1M token context window, making it highly effective for deep codebase analysis, long-form documentation parsing, and multi-file repository reasoning. For developers building agentic workflows or automated debugging tools, this model offers a strategic middle ground: it provides the structured reasoning required for complex task decomposition while significantly lowering the inference overhead. If your application requires high-volume logical processing where latency and cost-per-token are critical constraints, this model serves as a highly scalable alternative to larger, more expensive reasoning engines.

text generationAPI

glm-5.3-flash:batch

z-ai
1048576 ctx

GLM-5.3-Flash:batch is a high-throughput, multimodal model engineered specifically for developers building autonomous agents and complex coding workflows. Unlike standard LLMs that struggle with context drift, this model utilizes a hybrid sparse and linear attention architecture. This design allows it to maintain high precision across its massive 1M+ token window, making it ideal for analyzing entire codebases or processing lengthy document sets without the typical quadratic computational cost. For developers, the 'batch' designation signifies an optimization for asynchronous, large-scale processing tasks where latency is secondary to cost-efficiency and volume. It bridges the gap between lightweight models and heavy-duty reasoning engines, offering a specialized middle ground for long-horizon reasoning and multimodal data extraction. If your roadmap involves agentic loops that require consistent state tracking over long sequences, this model provides a scalable architecture to support those requirements via API.

text generationAPI
Email