Global AI chat room · 17 online now Join now
DIRECTORY / 02

AI Models | Open-Source LLM Directory

Discover and compare open-source LLMs, language models and multimodal models by capability, scale, license, downloads and provenance.

Compare modelsFind the right building block for your next workflow
Directory overview
591
curated entries
24 topic groupsLive
02 / MODEL INDEX

Find the right model for the job

Context first, better decisions. Every entry keeps the signal that matters.

CURATED DIRECTORY591 results

seed-2.0-code

bytedance-seed
262144 ctx

Seed 2.0 Code is a specialized model from ByteDance optimized specifically for agentic workflows and high-autonomy coding tasks. Unlike general-purpose LLMs that act primarily as chat interfaces, this model is architected to function as a reasoning engine within autonomous coding agents. It excels in frontend development and complex, multilingual programming environments where context awareness is critical. With a significant 262k context window, it is designed to ingest large codebases, allowing developers to integrate it into IDE-based agents or CI/CD pipelines for automated debugging, refactoring, and feature implementation. For engineers building the next generation of AI software engineers, Seed 2.0 provides the structural reliability needed for multi-step tool use and long-context reasoning that standard models often struggle to maintain.

text generationAPI

qwen3.8-2.4t-a95b:batch

qwen
1010000 ctx

For developers working with massive-scale reasoning tasks, Qwen3.8-2.4T-A95B represents a significant leap in sparse Mixture-of-Experts (MoE) architecture. While the total parameter count sits at 2.4 trillion, the model optimizes inference by activating only 95 billion parameters per token. This design strikes a balance between the deep knowledge density of a frontier-class model and the latency requirements of production environments. It is essentially the open-weight distillation of the Qwen3.8 Max series, making it ideal for complex agentic workflows, sophisticated code generation, and long-context retrieval tasks. Unlike monolithic dense models of similar scale, this MoE approach provides high-throughput capabilities without sacrificing the nuanced reasoning required for multi-step logical deduction. Integration via API allows you to leverage trillion-parameter intelligence for RAG pipelines or autonomous tool-use without the prohibitive hardware overhead of hosting a dense 2T model locally.

text generationAPI

qwen3.8-2.4t-a95b

qwen
1048576 ctx

For developers working with large-scale reasoning tasks, Qwen3.8 2.4T A95B represents a significant step in sparse Mixture-of-Experts (MoE) architecture. While the total parameter count sits at 2.4 trillion, the model only activates 95 billion parameters per token, offering a high-performance profile that balances massive knowledge density with much lower inference latency than a dense model of equivalent scale. This makes it particularly effective for complex multi-step reasoning, sophisticated code generation, and high-context information retrieval. With a massive 1M token context window, it is designed to ingest entire codebases or lengthy documentation sets without losing coherence. Compared to standard dense models, you get the intelligence of a massive frontier model with the throughput efficiency required for production-grade RAG pipelines and autonomous agent workflows. It is an ideal choice for integrating high-level cognitive capabilities into existing software stacks via API.

text generationAPI

seed-2-1-turbo

bytedance-seed
262144 ctx

Seed 2.1 Turbo is a specialized multimodal model from ByteDance designed to bridge the gap between high-level reasoning and practical execution. Unlike standard LLMs that struggle with complex, multi-step processes, this model is architected specifically for long-horizon agent workflows and end-to-end software delivery. For developers, this means moving beyond simple code completion toward autonomous task execution where the model can interpret visual UI elements alongside textual requirements. It excels in environments requiring high context retention, supporting a 262k token window which is critical for navigating large codebases or lengthy documentation. Whether you are building autonomous coding agents or complex automation pipelines that require visual grounding, Seed 2.1 Turbo offers a streamlined API integration for workflows that demand both logical precision and multimodal understanding.

text generationAPI

dots-3-note-preview:free

dots-studio
512000 ctx

Dots-3-Note-Preview is a specialized Mixture-of-Experts (MoE) model designed for high-efficiency text generation. While it sits within a massive 280B parameter architecture, it operates with only 16B active parameters per token, offering a pragmatic balance between reasoning depth and inference speed. For developers, this means you get the sophisticated pattern recognition of a large-scale model without the prohibitive latency typically associated with dense models of this magnitude. With a massive 512,000 context window, it is purpose-built for long-form document analysis, complex codebase summarization, and large-scale data extraction tasks where maintaining long-range dependencies is critical. Unlike standard dense models, its MoE structure allows for more granular task specialization, making it an excellent candidate for RAG pipelines and multi-step reasoning workflows. It serves as an accessible entry point into the Dots 3 ecosystem, optimized for developers who need high-throughput performance on complex, context-heavy workloads.

text generationAPI

qwen3.8-27b:free

qwen
262144 ctx

Qwen3.8-27B is a high-density, open-weight vision-language model engineered for developers who require a balance between reasoning depth and computational efficiency. Unlike standard LLMs, this model integrates multimodal capabilities directly into its architecture, making it highly effective for tasks involving visual reasoning, complex document parsing, and spatial understanding. For engineers building autonomous agents, the model's ability to manage long-running workflows and extended context windows is a significant advantage. It excels in specialized domains such as automated code generation, technical research, and professional-grade data extraction. While larger models offer more raw knowledge, the 27B parameter scale provides a sweet spot for production environments where low-latency inference and high throughput are critical. Whether you are integrating it via API for agentic tasks or deploying it for multimodal RAG pipelines, Qwen3.8 offers a robust, flexible backbone for sophisticated AI applications.

text generationAPI

qwen3.8-27b

qwen
1000000 ctx

Qwen3.8-27B is a dense, open-weight vision-language model designed for developers who need a balance between high-reasoning capabilities and efficient deployment. Unlike purely text-based LLMs, this model integrates multimodal perception, making it capable of processing visual data alongside complex text instructions. It is specifically architected to handle professional-grade workflows, including sophisticated coding tasks, scientific research, and long-context agentic reasoning. For engineers building autonomous agents, the model's ability to maintain coherence over extended task sequences is a significant advantage. While larger models offer higher raw intelligence, the 27B parameter count provides a sweet spot for low-latency integration and cost-effective scaling in production environments. Whether you are implementing visual document parsing or building multi-step reasoning pipelines, Qwen3.8 offers a robust, flexible foundation that competes closely with much larger proprietary systems.

text generationAPI

glm-5.3:batch

z-ai
1048576 ctx

GLM-5.3:batch is a specialized reasoning model engineered specifically for high-complexity software engineering workflows and long-horizon agentic orchestration. Unlike standard chat models optimized for quick interactions, this iteration focuses on maintaining logical consistency across massive datasets, supported by a robust 1-million-token context window. For developers, this means you can ingest entire codebases, extensive documentation, or multi-file logs into a single prompt without losing structural coherence. The 'batch' designation suggests an optimization for high-throughput processing, making it an ideal candidate for automated code reviews, large-scale refactoring tasks, or complex debugging agents that require deep reasoning over extended sequences. While many models struggle with 'lost in the middle' phenomena during long-context retrieval, GLM-5.3 is architected to handle the dependencies inherent in large-scale software architecture, providing a reliable backend for autonomous engineering agents.

text generationAPI

glm-5.3

z-ai
1310720 ctx

GLM-5.3 is a high-reasoning model specifically architected for developers working on complex software engineering workflows and autonomous agent orchestration. Unlike general-purpose chat models, this iteration focuses on long-horizon task execution, making it suitable for multi-step debugging, codebase analysis, and automated system design. The standout technical feature is the 1M-token context window, which allows you to ingest entire repositories or massive documentation sets without losing structural coherence. For integration, it functions via API, providing a scalable backend for building agentic tools that require deep logical consistency over extended sessions. If your roadmap includes building self-correcting coding assistants or sophisticated RAG pipelines that demand high precision in long-context retrieval, GLM-5.3 offers a robust alternative to existing industry standards by prioritizing reasoning depth over simple pattern matching.

text generationAPI

hy-mt2-7b

tencent
8192 ctx

Hy-MT2-7B is a specialized 7B-parameter model from Tencent designed specifically for high-precision machine translation. Unlike general-purpose LLMs that treat translation as a standard text-generation task, this model is architected to handle complex linguistic workflows. It supports 33 major language pairs alongside five specific Chinese dialects and minority languages, making it a robust choice for localized applications in the APAC region. For developers, the standout feature is its granular control over translation logic; it supports structured, delimiter-based, and glossary-based workflows, allowing you to enforce terminology consistency and maintain specific stylistic tones via prompt-driven constraints. With an 8k context window, it can manage longer segments while maintaining contextual coherence. It is ideal for integrating into localization pipelines, technical documentation workflows, or any application requiring high-fidelity cross-lingual mapping with strict adherence to predefined glossaries.

text generationAPI

hy-mt2-30b-a3b

tencent
8192 ctx

For developers building localization pipelines or multilingual interfaces, hy-mt2-30b-a3b represents a specialized shift from general-purpose LLMs toward high-precision translation. Unlike standard models that often struggle with linguistic nuance or structural consistency, this model is architected specifically for translation workflows. It supports 33 language pairs alongside niche Chinese dialects and minority languages, filling a critical gap for localized products in the APAC region. What sets it apart is the built-in support for structured data handling; you can integrate it into existing pipelines using delimiter-based or glossary-based workflows to ensure brand consistency and technical accuracy. Whether you are automating software localization or building real-time translation tools, the model's ability to respect context and specific terminology makes it a more reliable choice than generic chat models for production-grade NMT (Neural Machine Translation) tasks.

text generationAPI

hy-mt2-1.8b

tencent
8192 ctx

For developers building localized applications or real-time communication tools, hy-mt2-1.8b offers a highly efficient, specialized solution for multilingual translation. Unlike general-purpose LLMs that struggle with linguistic nuance at small scales, this 1.8B-parameter model is purpose-built for high-fidelity translation across 33 language pairs, including specific support for Chinese dialects and minority languages. What sets it apart from standard NMT models is its architectural support for complex workflows: you can pass structured delimiters, enforce specific glossaries to maintain brand terminology, and apply style-guided constraints to match a target persona. With an 8k context window, it handles more than just sentence-level snippets, allowing for better contextual awareness in longer passages. It is an ideal candidate for edge deployment or high-throughput microservices where low latency and strict adherence to technical terminology are more critical than broad generative reasoning.

text generationAPI

deepseek-v4-flash-vision-exp:batch

deepseek
1048576 ctx

DeepSeek-V4-Flash-Vision-Exp is an experimental multimodal iteration of the V4 Flash architecture, specifically designed to bridge high-speed text processing with robust visual reasoning. For developers building agentic workflows, this model provides a significant upgrade by allowing the same engine that handles complex logic and tool-calling to also interpret visual inputs. Unlike many vision models that sacrifice reasoning depth for speed, this version maintains the core performance characteristics of the V4 Flash series, making it ideal for real-time applications like automated UI testing, visual document parsing, and multi-modal agent orchestration. It is optimized for high-throughput batch processing, offering a cost-effective way to integrate vision into existing text-based pipelines without the latency overhead typically associated with larger vision-language models. If your stack requires low-latency multimodal feedback loops, this model offers a highly competitive integration path.

text generationAPI

deepseek-v4-flash-vision-exp

deepseek
1048576 ctx

DeepSeek-V4-Flash-Vision-Exp is an experimental multimodal iteration of the V4 Flash architecture, designed to bridge high-speed text processing with robust visual reasoning. For developers, this model offers a significant upgrade over the text-only base by enabling native image understanding within a low-latency framework. It is specifically optimized for agentic workflows where visual context—such as UI screenshots, diagrams, or document scans—must be parsed alongside complex text instructions. Unlike larger, heavier vision models that sacrifice throughput for depth, this 'Flash' variant prioritizes rapid inference and high context window utilization, making it ideal for real-time applications like automated visual QA, accessibility tools, or multimodal RAG pipelines. While it remains in an experimental phase, it maintains the core text capabilities and agentic instruction-following performance of the standard V4 Flash, providing a cost-effective way to integrate vision into existing text-based automation loops.

text generationAPI

muse-spark-1.2-contributor

meta
1048576 ctx

Muse Spark 1.2 Contributor is a specialized reasoning model from Meta designed to bridge the gap between high-performance logic and production-scale cost efficiency. While the flagship Muse Spark models excel at complex architectural reasoning, the Contributor tier is optimized for developers who need consistent logical throughput without the premium price tag. It maintains a massive 1M token context window, making it highly effective for deep codebase analysis, long-form documentation parsing, and multi-file repository reasoning. For developers building agentic workflows or automated debugging tools, this model offers a strategic middle ground: it provides the structured reasoning required for complex task decomposition while significantly lowering the inference overhead. If your application requires high-volume logical processing where latency and cost-per-token are critical constraints, this model serves as a highly scalable alternative to larger, more expensive reasoning engines.

text generationAPI

glm-5.3-flash:batch

z-ai
1048576 ctx

GLM-5.3-Flash:batch is a high-throughput, multimodal model engineered specifically for developers building autonomous agents and complex coding workflows. Unlike standard LLMs that struggle with context drift, this model utilizes a hybrid sparse and linear attention architecture. This design allows it to maintain high precision across its massive 1M+ token window, making it ideal for analyzing entire codebases or processing lengthy document sets without the typical quadratic computational cost. For developers, the 'batch' designation signifies an optimization for asynchronous, large-scale processing tasks where latency is secondary to cost-efficiency and volume. It bridges the gap between lightweight models and heavy-duty reasoning engines, offering a specialized middle ground for long-horizon reasoning and multimodal data extraction. If your roadmap involves agentic loops that require consistent state tracking over long sequences, this model provides a scalable architecture to support those requirements via API.

text generationAPI

glm-5.3-flash

z-ai
1310720 ctx

GLM-5.3-Flash is a specialized multimodal model designed for developers prioritizing high-throughput and low-latency execution. Unlike standard dense models, it utilizes a hybrid sparse and linear attention architecture, which allows it to maintain high retrieval accuracy across its extensive 1.3M token context window without the typical quadratic compute penalty. For engineers building autonomous agents or complex coding assistants, this architecture is critical for long-horizon reasoning and maintaining state over massive codebases or documentation sets. While many 'flash' models sacrifice reasoning depth for speed, GLM-5.3-Flash is optimized specifically for agentic workflows and multi-step task execution. It integrates easily via API, making it a viable alternative for production environments where cost-efficiency and long-context stability are more important than raw parameter count.

text generationAPI

qwen3.8-flash

qwen
1000000 ctx

Qwen3.8-Flash is a multimodal reasoning model designed for developers who need high-speed intelligence without sacrificing complex cognitive capabilities. Unlike standard text-only LLMs, this model excels at bridging the gap between visual inputs and logical execution. It is specifically optimized for agentic workflows, making it a strong candidate for autonomous desktop interaction and complex tool-use scenarios. For developers working on large-scale data, its ability to perform deep codebase analysis and long-video reasoning provides a significant edge in context retention. Whether you are building automated UI testing agents, sophisticated document parsing pipelines, or real-time visual assistants, Qwen3.8-Flash offers a low-latency solution that handles multimodal tokens—ranging from charts to video frames—with high precision. It positions itself as a high-throughput alternative for production environments where speed and multimodal reasoning are non-negotiable.

text generationAPI

ling-3.0-flash-fin:free

inclusionai
262144 ctx

Ling 3.0 Flash Fin is a specialized Mixture-of-Experts (MoE) model engineered specifically for the financial services sector. While it leverages the efficiency of the Ling 3.0 Flash architecture, it is fine-tuned to handle the nuances of investment analysis, market sentiment, and complex financial reasoning. With 5.1B active parameters out of a 124B total pool, it offers a strategic balance between high-speed inference and deep domain intelligence. For developers, this means significantly lower latency compared to dense large-scale models without sacrificing the precision required for financial data processing. The model supports a massive 262k context window, making it ideal for ingesting lengthy quarterly reports, regulatory filings, or extensive historical datasets. Whether you are building automated sentiment analysis pipelines, sophisticated fintech chatbots, or quantitative research assistants, this model provides a highly specialized alternative to general-purpose LLMs that often struggle with niche financial terminology and logic.

text generationAPI

ling-3.0-flash-fin

inclusionai
262144 ctx

Ling 3.0 Flash Fin is a finance-specialized mixture-of-experts model from InclusionAI, activating only 5.1B of its 124B total parameters per forward pass. This sparse architecture keeps inference latency and cost closer to a 5B dense model while retaining the knowledge capacity of a much larger system. The model is tuned for investment research, risk assessment, earnings-call summarization, regulatory compliance checks, and portfolio commentary generation. Its 256K token context window lets you feed full annual reports, prospectuses, or multi-year transcript histories in a single request. Access is provided via a REST API (OpenAI-compatible endpoints), so integration into existing Python, TypeScript, or LangChain workflows is straightforward. Compared to general-purpose LLMs, Flash Fin shows stronger numeric reasoning, domain-specific terminology handling, and citation discipline on financial benchmarks. Compared to larger finance models (e.g., BloombergGPT, FinGPT), it offers faster response times and lower per-token pricing due to the MoE design, making it practical for high-throughput production pipelines such as real-time alerting or batch document processing. Rate limits and data residency follow InclusionAI's standard API terms; no on-premise weights are distributed.

text generationAPI

hy4-preview

tencent
1048576 ctx

hy4-preview is Tencent's 49B-activation mixture-of-experts model built for coding agents and multi-step tool workflows. It handles 1M-token contexts, supports parallel tool calls, and excels at repository-scale code understanding. Compared to other 49B-class models, it balances strong coding performance with efficient inference via MoE routing. The API integrates easily into existing agent frameworks, supporting standard JSON-based function calling and streaming. It's well-suited for automated refactoring, codebase Q&A, and long-running dev tasks.

text generationAPI

granite-4.2-8b

ibm-granite
131072 ctx

Granite 4.2 8B is IBM's latest dense reasoning model, specifically engineered for developers building agentic workflows and complex logic-driven applications. Unlike standard chat models, this 8B parameter iteration focuses heavily on structured reasoning, making it a strong candidate for multi-step problem solving in mathematics and software engineering. For developers, the primary value lies in its ability to handle code generation and multilingual dialogue while maintaining a high degree of reliability in logical chains. With a substantial 131k context window, it supports deep document analysis and long-form reasoning tasks without immediate memory loss. While smaller than massive frontier models, its efficiency makes it an ideal choice for low-latency integrations and specialized agent tasks where cost-effective, high-precision reasoning is required. It bridges the gap between lightweight edge models and heavy-duty reasoning engines, offering a scalable middle ground for production environments.

text generationAPI

muse-spark-1.3

meta
1048576 ctx

Muse Spark 1.3 is Meta's latest specialized multimodal reasoning model, engineered specifically for high-autonomy agentic workflows. Unlike standard LLMs that struggle with context drift during long-running tasks, this model is optimized for state management across extended execution loops. For developers building multi-agent systems or complex automated coding pipelines, Spark 1.3 provides the necessary stability to maintain logical consistency over long horizons. Its architecture excels at multimodal reasoning, allowing it to bridge the gap between visual inputs and structured code generation. While many models prioritize raw throughput, Muse Spark 1.3 prioritizes task persistence and reasoning depth, making it a superior choice for autonomous agents that require deep integration into existing DevOps or software engineering lifecycles. It effectively addresses the 'forgetting' problem common in long-context windows by prioritizing structural coherence during iterative processes.

text generationAPI

muse-spark-1.3-contributor

meta
1048576 ctx

Muse Spark 1.3 Contributor is Meta’s specialized, cost-optimized tier designed specifically for developers building high-frequency, iterative workflows. Unlike heavy-duty reasoning models intended for single-shot complex tasks, this model is tuned for the 'connective tissue' of AI development: agentic loops, multi-agent coordination, and real-time coding assistance. It excels at maintaining context across long-running processes, supported by a massive 1M token context window that allows for deep ingestion of entire codebases or massive documentation sets. For teams in the experimentation phase, it offers a high-throughput alternative to larger models, making it ideal for testing agentic reasoning patterns or fine-tuning multi-step task execution without the prohibitive latency or cost of flagship models. It serves as a reliable backbone for developers who need consistent, low-latency intelligence to drive autonomous workflows.

text generationAPI
Email