Global AI chat room · 17 online now Join now
DIRECTORY / 02

AI Models | Open-Source LLM Directory

Discover and compare open-source LLMs, language models and multimodal models by capability, scale, license, downloads and provenance.

Compare modelsFind the right building block for your next workflow
Directory overview
591
curated entries
24 topic groupsLive
02 / MODEL INDEX

Find the right model for the job

Context first, better decisions. Every entry keeps the signal that matters.

CURATED DIRECTORY591 results

qwen3.8-max-0902

qwen
1000000 ctx

For developers working with high-density multimodal workloads, qwen3.8-max-0902 represents a significant leap in parameter scale and architectural efficiency. Built on a 2.4-trillion-parameter Mixture-of-Experts (MoE) framework, this model is designed to handle complex reasoning tasks that require both deep linguistic nuance and sophisticated visual understanding. Unlike standard text-only LLMs, this snapshot integrates native support for image and video inputs, making it a versatile backbone for applications involving video captioning, visual reasoning, or automated content analysis. From an integration standpoint, the model is optimized for high-throughput API environments, offering a massive 1-million-token context window that solves the common bottleneck of long-document processing and multi-frame video analysis. While many models struggle with coherence in long-form context, the MoE architecture here allows for high-performance inference without the typical latency penalties of dense models of this magnitude. It is a robust choice for engineers building sophisticated agentic workflows or complex multimodal RAG pipelines.

text generationAPI

ling-3.0-flash-sante:free

inclusionai
262144 ctx

For developers building in the MedTech and digital health sectors, Ling 3.0 Flash Sante offers a specialized Mixture-of-Experts (MoE) architecture designed to balance domain-specific accuracy with low-latency performance. While many general-purpose models struggle with the nuance of clinical terminology, this model leverages 5.1B active parameters to maintain high throughput while focusing its intelligence on medical reasoning and health-related text generation. With a massive 262,144 context window, it is particularly well-suited for analyzing lengthy electronic health records (EHRs), synthesizing clinical research papers, or powering patient-facing triage interfaces. Unlike monolithic models that incur heavy compute costs for every query, the MoE design ensures that you get specialized medical insights without the typical latency overhead, making it an ideal candidate for real-time integration into healthcare APIs and diagnostic support tools.

text generationAPI

nex-n2.5-pro:free

nex-agi
262144 ctx

Nex-N2.5-Pro is an agentic-focused model engineered specifically for autonomous task execution and complex software engineering workflows. Unlike standard chat models that focus on single-turn completions, this model is optimized for high-context reasoning within a visual feedback loop. This makes it particularly effective for developers working on large-scale refactoring, multi-file feature implementation, and autonomous debugging where the model must interpret code changes and iterate based on compiler or runtime feedback. With a massive 262k context window, it can ingest entire repositories to maintain architectural awareness. For integration, it functions as a reasoning engine that can be plugged into agentic frameworks to bridge the gap between high-level goal setting and verified code deployment. It positions itself as a tool for building autonomous coding agents rather than just a simple autocomplete assistant.

text generationAPI

nex-n2.5-mini:free

nex-agi
262144 ctx

Nex-N2.5-mini is a specialized agentic model designed for developers who need more than just code completion. Unlike standard LLMs that provide static snippets, this model is architected for autonomous task execution through a visual feedback loop. It excels at navigating complex, multi-file codebases, implementing logic across various modules, and verifying its own output against execution results. For engineers building automated DevOps pipelines or sophisticated AI coding assistants, the model’s ability to bridge the gap between intent and verified implementation is its primary differentiator. It operates with a substantial 262k context window, making it highly effective for large-scale repository analysis and long-form debugging sessions. While smaller in parameter scale, its focus on agentic workflows makes it a practical choice for low-latency, high-autonomy integration into existing IDEs and CI/CD environments.

text generationAPI

mercury-2.5

inception
260000 ctx

Mercury 2.5 represents a paradigm shift in inference architecture by moving away from traditional sequential token generation. Developed by Inception, this model utilizes a diffusion-based approach (dLLM) to produce and refine multiple tokens in parallel. For developers, this translates to a significant reduction in time-to-first-token and overall latency, making it particularly effective for high-throughput reasoning tasks. Unlike standard autoregressive models that struggle with long-form coherence during rapid generation, Mercury 2.5's parallel refinement process allows it to maintain structural integrity across its 260k context window. It is best suited for real-time agentic workflows, complex logical reasoning, and applications where low-latency response is critical. Integration is handled via API, allowing you to plug this non-sequential reasoning engine into existing pipelines without rearchitecting your entire inference stack.

text generationAPI

deepseek-v4.1-flash

deepseek
1048576 ctx

DeepSeek-V4.1-Flash represents a significant architectural shift for developers seeking high-throughput reasoning without the typical latency overhead of dense models. Moving away from standard decoder-only setups, this model utilizes a Causal Encoder-Decoder (CED) framework optimized via a sparse Mixture-of-Experts (MoE) design. For engineers, this means you get the intelligence of a much larger model while only activating a fraction of the parameters—8B on input and 16B during processing—making it exceptionally efficient for real-time applications. It is particularly well-suited for high-volume tasks like automated code reviews, complex data extraction, and real-time agentic workflows where low time-to-first-token is critical. Unlike many general-purpose models that struggle with instruction following under heavy load, the CED architecture provides a more structured approach to long-context reasoning. If your stack requires a balance of massive context windows (up to 1M tokens) and cost-effective inference, this model offers a highly competitive alternative to the current industry giants.

text generationAPI

ling-3.0-flash-vl:free

inclusionai
262144 ctx

Ling 3.0 Flash VL is a high-efficiency multimodal model designed for developers needing a balance between rapid inference and sophisticated visual reasoning. Built on a Mixture-of-Experts (MoE) architecture with 5.5B active parameters, it optimizes compute costs without sacrificing the depth required for complex tasks. Unlike standard text-only LLMs, this version integrates native visual perception, allowing it to process and interpret image data alongside text instructions in a single context window. For developers, this means seamless integration into workflows involving automated visual inspection, document parsing, or multimodal chat interfaces. With a massive 262,144 token context window, it excels at analyzing long-form visual documents and multi-image sequences. If your use case demands low-latency responses and high-throughput visual processing—similar to the Gemini Flash or GPT-4o-mini tier—this model provides a highly competitive, cost-effective alternative for production-grade applications.

text generationAPI

ling-3.0-flash-vl

inclusionai
262144 ctx

Ling 3.0 Flash VL is a high-efficiency multimodal model designed for developers who need a balance between rapid inference speeds and sophisticated visual reasoning. Built on a Mixture-of-Experts (MoE) architecture with 5.5B active parameters, it optimizes computational overhead without sacrificing the nuance required for complex language tasks. Unlike standard text-only models, this version features native visual perception, allowing it to interpret images and documents directly within the same context window. For developers, this means seamless integration for use cases like automated visual inspection, document parsing, and multimodal RAG pipelines. While it maintains a lightweight footprint suitable for high-throughput applications, its 131k context window provides the headroom necessary for processing extensive datasets. It serves as a pragmatic alternative to heavier proprietary vision models, offering lower latency for real-time interactive applications.

text generationAPI

fugu-max

sakana
1000000 ctx

Fugu-max represents a shift away from the standard monolithic LLM architecture, moving instead toward a learned multi-agent orchestration framework. Developed by Sakana AI, this model functions as an intelligent router that dynamically directs sub-tasks to specialized agents within the Fugu ecosystem. For developers, this means you aren't just hitting a single weight set; you are interacting with a system optimized for task decomposition and efficient resource allocation. It is specifically engineered for high cost-performance, making it an ideal candidate for complex pipelines where latency and token expenditure must be balanced against reasoning depth. Whether you are building autonomous agents or complex RAG workflows, Fugu-max provides a scalable way to manage multi-step logic without the overhead of manually coding every routing decision. It bridges the gap between simple chat completion and full-scale agentic workflows through its native orchestration capabilities.

text generationAPI

fugu-ultra-v2

sakana
1000000 ctx

Fugu-ultra-v2 represents a shift from monolithic LLM architectures toward a learned multi-agent orchestration framework. Instead of relying on a single massive parameter set to solve every task, this model functions as a central controller trained to intelligently delegate sub-tasks to specialized agents. For developers, this means higher reasoning accuracy and better handling of complex, multi-step workflows that typically cause standard models to hallucinate or lose context. While traditional models struggle with long-chain logic, Fugu-ultra-v2 leverages its orchestration layer to manage task decomposition and execution dynamically. It is particularly suited for autonomous agentic workflows, complex software engineering tasks, and multi-modal reasoning pipelines. Integration is handled via a standard API, making it a drop-in replacement for developers looking to move beyond simple chat interfaces toward sophisticated, self-correcting AI systems.

text generationAPI

schematron-v2-small

inference-net
128000 ctx

Schematron-v2-small is a specialized 3B-parameter model engineered specifically for high-fidelity HTML-to-JSON structural extraction. Unlike general-purpose LLMs that often struggle with nested DOM structures or lose context in long documents, this model is optimized for strict schema adherence. It utilizes a JSON schema as the primary instruction mechanism, ensuring that the output is not just textually accurate but programmatically valid for downstream pipelines. With a massive 128k context window, it can ingest entire complex web pages without truncation, making it ideal for web scraping, automated data ingestion, and transforming unstructured web content into structured databases. For developers, this means significantly lower post-processing overhead and higher reliability when dealing with volatile HTML layouts compared to larger, more expensive frontier models.

text generationAPI

schematron-v2-turbo

inference-net
128000 ctx

Schematron-v2-turbo is a specialized 3B-parameter model engineered specifically for high-throughput HTML-to-JSON extraction. Unlike general-purpose LLMs that struggle with structural consistency, this model is optimized for deterministic data parsing. It operates via a schema-first approach, requiring developers to provide a JSON schema within the response_format parameter to govern the output. This design significantly reduces parsing errors and eliminates the need for post-extraction cleanup. With a massive 128k context window, it can ingest entire web pages or complex document fragments in a single pass. For developers building web scrapers, automated data pipelines, or competitive intelligence tools, this model offers a much higher tokens-per-second ratio than larger frontier models while maintaining the structural integrity required for production-grade ETL workflows.

text generationAPI

pareto

unbiased
262144 ctx

Pareto is a multimodal composite model engineered specifically for developers building complex, autonomous systems. Unlike standard LLMs optimized solely for chat, Pareto is architected to handle the high-reasoning demands of agentic workflows and deep technical research. It features a massive 262k context window, making it a viable candidate for processing entire codebases or extensive documentation sets without losing coherence. For engineers, the primary value lies in its multimodal capabilities and its stability during multi-step tool use, which is often a bottleneck in agentic loops. Whether you are integrating it via API for automated coding assistants or deploying it as the brain of a research agent, Pareto aims to bridge the gap between general-purpose reasoning and specialized technical execution. It positions itself as a high-context, high-reliability backbone for production-grade AI agents.

text generationAPI

glm-5.3-flashx

z-ai
1048576 ctx

For developers building latency-sensitive applications, GLM-5.3-FlashX represents a significant shift toward high-throughput multimodal inference. Unlike standard LLMs that struggle with high-frequency streaming, this model is optimized for raw speed, hitting up to 200 tokens per second. It utilizes a hybrid sparse and linear attention architecture, which effectively balances long-context processing with computational efficiency. This makes it an ideal candidate for real-time agentic workflows, live multimodal captioning, and high-volume data extraction pipelines where response time is a critical KPI. While many models sacrifice reasoning depth for speed, the FlashX variant maintains the core multimodal capabilities of the GLM-5.3 family, providing a predictable API for integrating vision and text into low-latency production environments. If your stack requires rapid-fire reasoning or high-concurrency text generation without the typical overhead of dense transformer architectures, this model offers a highly competitive performance-to-cost ratio.

text generationAPI

ternary-bonsai-2-27b

prism-ml
262144 ctx

For developers building agentic workflows or complex reasoning pipelines, ternary-bonsai-2-27b offers a high-density alternative to larger, more expensive models. Built on the Qwen3.8 architecture, this 27B parameter model leverages ternary compression to maintain high reasoning performance while optimizing inference efficiency. It is specifically tuned for heavy-lift tasks including advanced mathematics, multi-step coding logic, and precise tool calling. Unlike standard lightweight models, it features a massive 262K-token context window, making it suitable for analyzing entire codebases or long-form technical documentation in a single pass. It also integrates multimodal capabilities for image understanding, allowing for unified vision-language processing. If your stack requires a balance between low-latency response times and the ability to handle deep logical reasoning without the overhead of a 70B+ parameter model, this is a highly competitive candidate for your production environment.

text generationAPI
Email