Global AI chat room · 18 online now Join now
DIRECTORY / 02

AI Models | Open-Source LLM Directory

Discover and compare open-source LLMs, language models and multimodal models by capability, scale, license, downloads and provenance.

Compare modelsFind the right building block for your next workflow
Directory overview
591
curated entries
24 topic groupsLive
02 / MODEL INDEX

Find the right model for the job

Context first, better decisions. Every entry keeps the signal that matters.

CURATED DIRECTORY591 results

Spark-X2.5-4B

XHToken
Model

Spark-X2.5-4B is a compact, high-efficiency text generation model designed for developers prioritizing low latency and minimal hardware footprints. At 4B parameters, it strikes a strategic balance between computational overhead and reasoning capability, making it an ideal candidate for edge deployment or local integration within resource-constrained environments. Unlike massive frontier models that require significant GPU clusters, Spark-X2.5 is optimized for rapid inference cycles, making it suitable for real-time applications like conversational agents, automated content drafting, and structured data extraction. For teams building microservices or mobile-integrated AI features, this model offers a lightweight alternative to larger architectures without sacrificing the core linguistic nuances required for production-grade text tasks. Released under the Apache-2.0 license, it provides the legal flexibility necessary for commercial integration and fine-tuning workflows.

text generationapache-2.0
1.3K starsView details

Qwen3.8-27B-OBLITERATED

OBLITERATUS
Model

Qwen3.8-27B-OBLITERATED is a high-performance text generation model optimized for developers requiring a balance between reasoning depth and deployment efficiency. Built on the Qwen architecture, this 27B parameter variant is specifically fine-tuned to minimize latency while maintaining high instruction-following accuracy. For engineers working with constrained hardware, the 27B scale offers a sweet spot: it provides significantly more nuanced context handling than 7B models without the massive VRAM overhead of 70B+ architectures. It is particularly effective for complex RAG (Retrieval-Augmented Generation) pipelines, structured data extraction, and multi-turn conversational agents. Since it is released under the Apache-2.0 license, it is highly suitable for commercial integration and local fine-tuning. If you are transitioning from smaller models and finding them lacking in logical consistency, this model serves as a robust middle-ground upgrade for production-ready AI applications.

text generationapache-2.0
1.3K starsView details

gemma-3-1b-it

google
Not specified

gemma-3-1b-it is a text generation model published on Hugging Face. It is primarily used with transformers and should be evaluated against the model card, license and deployment requirements before production use.

text generationgemma
1.1K starsView details

Qwen3-Coder-30B-A3B-Instruct-GGUF

unsloth
Not specified

Qwen3 Coder 30B A3B Instruct is a specialized Mixture-of-Experts (MoE) model optimized for high-performance programming tasks. By utilizing an active parameter count of 3B within a 30B total parameter architecture, it delivers the reasoning capabilities of a large model with the inference speed and memory efficiency of a much smaller one. For developers, this means a significant reduction in VRAM requirements without sacrificing complex logic handling or multi-language syntax accuracy. It is particularly effective for autonomous code generation, refactoring legacy systems, and acting as a local copilot. This GGUF quantization makes it highly accessible for local deployment via llama.cpp or Ollama, allowing seamless integration into IDEs without relying on cloud APIs.

text generationapache-2.0
1.0K starsView details

NeoHorse-1-9B

TokenRhythm
Model

NeoHorse-1-9B is a compact, high-efficiency text generation model designed for developers who need a balance between low latency and reasoning performance. At the 9B parameter scale, it sits in the 'sweet spot' for deployment on consumer-grade hardware or edge devices without requiring massive data center clusters. Unlike larger, cumbersome models, NeoHorse is optimized for streamlined integration into existing RAG pipelines, automated content workflows, and agentic frameworks where quick inference cycles are critical. While it doesn't aim to compete with trillion-parameter behemoths on broad world knowledge, its architecture is tuned for coherent instruction following and structured output generation. For teams working under strict memory constraints or those building specialized microservices, this model offers a highly portable alternative to larger closed-source APIs, providing more control over the deployment environment under the Apache-2.0 license.

text generationapache-2.0
1.0K starsView details

Nex-N2.5-mini

nex-agi
Model

Nex-N2.5-mini is a lightweight text-generation model designed for developers who need high-speed inference without the overhead of massive parameter counts. Released under the Apache-2.0 license, it is built for easy integration into existing production pipelines where latency and cost-efficiency are critical. Unlike larger foundation models that require significant GPU resources, this 'mini' variant is optimized for edge deployment and high-throughput tasks such as real-time chat completion, automated summarization, and structured data extraction. For developers working within the Hugging Face ecosystem, it offers a seamless transition from prototyping to deployment, providing a predictable performance profile for instruction-following tasks. While it may not match the deep reasoning capabilities of trillion-parameter models, its strength lies in its efficiency-to-performance ratio, making it an ideal candidate for microservices and local LLM implementations where resource constraints are a primary concern.

text generationapache-2.0
839 starsView details

Qwen2.5-1.5B-Instruct

Qwen
Not specified

Qwen2.5 1.5B Instruct is a lightweight, instruction-tuned LLM designed for high-efficiency deployment. Despite its small parameter count, it punches above its weight in coding and mathematics, making it an ideal candidate for edge computing, local IDE integration, or as a fast routing layer in a multi-model pipeline. It is optimized for low-latency inference and fits comfortably within limited VRAM envelopes without sacrificing significant reasoning capabilities. For developers, this model offers a practical balance between performance and resource consumption, supporting a wide range of downstream tasks like text summarization and structured data extraction via a permissive Apache-2.0 license.

text generationapache-2.0
837 starsView details

Hemmingway-1

Altworld
Model

Hemmingway-1 is a specialized text-generation model developed by Altworld, designed for developers looking to implement concise, high-impact linguistic patterns in their applications. While many modern LLMs struggle with verbosity and 'filler' language, this model appears architected to prioritize structural clarity and stylistic precision. For developers building content automation tools, creative writing assistants, or sophisticated chatbots, Hemmingway-1 offers a distinct alternative to the standard chat-heavy models. It is particularly suited for workflows requiring high stylistic consistency and minimal token waste. Integration is straightforward via the Hugging Face ecosystem, making it a viable candidate for local deployment or fine-tuning on niche datasets. Note that the CC-BY-NC-4.0 license restricts commercial use, so it is best utilized for research, prototyping, or non-commercial creative projects where stylistic nuance is a primary requirement.

text generationcc-by-nc-4.0
770 starsView details

Qwen3-32B

Qwen
Not specified

Qwen3 32B is a mid-sized LLM designed to balance high-reasoning performance with efficient deployment. For developers, the 32B parameter count is a strategic sweet spot, offering capabilities that approach larger frontier models while remaining runnable on consumer-grade hardware or single-node enterprise GPUs. It excels in complex coding tasks, mathematical reasoning, and multilingual processing, making it a versatile backbone for RAG pipelines or autonomous agents. With an Apache-2.0 license, it provides the flexibility needed for commercial integration without restrictive proprietary overhead. Compared to smaller models, it demonstrates significantly lower hallucination rates in structured data extraction and a stronger grasp of nuanced instruction following.

text generationapache-2.0
747 starsView details

Qwen3-4B

Qwen
Not specified

Qwen3-4B is a compact text generation model from the Qwen family, designed for efficient deployment in resource-constrained environments while maintaining competitive performance on common NLP tasks. It integrates smoothly with the Hugging Face Transformers library, making it straightforward to plug into existing pipelines for tasks like summarization, code assistance, and conversational AI. With an Apache 2.0 license, it offers flexibility for both research and commercial use, though developers should validate outputs and check the model card for limitations. Compared to larger models, Qwen3-4B trades some accuracy for faster inference and lower memory usage, making it a practical choice for edge devices, local development, or high-throughput serving scenarios where latency and cost matter.

text generationapache-2.0
703 starsView details

Ornith-1.0-9B-GGUF

ornith-ai
Not specified

Ornith-1.0-9B-GGUF is a text generation model released by ornith-ai under an MIT license. It’s distributed in GGUF format, which means it’s optimized for CPU inference and works well with llama.cpp and similar runtimes. For developers, this translates to a lightweight option you can run locally without a GPU, making it suitable for prototyping, personal projects, or edge deployments where resources are limited. The model is compatible with the Hugging Face ecosystem, so integration with existing transformer-based pipelines is straightforward. While it’s not positioned as a high-end model for complex reasoning or large-scale enterprise tasks, it’s a practical choice for developers who need a no-fuss, permissively licensed model that’s easy to set up and quick to iterate with. As always, review the model card and test it against your specific use case before deploying in production.

text generationmit
659 starsView details

Qwen2.5-0.5B-Instruct

Qwen
Not specified

Qwen2.5-0.5B-Instruct is a lightweight, instruction-tuned LLM designed for environments where memory overhead and latency are critical. Despite its small footprint, it punches above its weight in basic reasoning and text transformation tasks, making it an ideal candidate for edge deployment, on-device processing, or as a specialized agent in a larger router-based architecture. For developers, this model offers a highly efficient alternative to larger models for simple classification, summarization, and structured data extraction. It integrates seamlessly with standard transformers pipelines and is licensed under Apache-2.0, ensuring flexibility for commercial production. While it lacks the deep world knowledge of its larger siblings, its speed and low VRAM requirements make it a practical tool for high-throughput applications.

text generationapache-2.0
628 starsView details

Qwen3.6-35B-A3B-NVFP4

nvidia
Not specified

The Qwen3.6 35B A3B NVFP4 is a specialized iteration of the Qwen series, optimized specifically for NVIDIA hardware using the NVFP4 quantization format. For developers, the primary draw here is the efficiency gain; by leveraging 4-bit floating point precision, this model significantly reduces VRAM overhead without the drastic perplexity loss typically seen in integer quantization. It is designed for high-throughput text generation and complex reasoning tasks where latency is critical. Integration is streamlined for NVIDIA TensorRT-LLM environments, making it an ideal candidate for production-grade RAG pipelines or agentic workflows where you need the intelligence of a mid-sized model but the speed of a much smaller one.

text generationapache-2.0
624 starsView details

MiMo-V2.6-Pro-RL

XiaomiMiMo
Model

MiMo-V2.6-Pro-RL is a text-generation model from XiaomiMiMo that targets developer workflows rather than general chat. It’s built for coding assistance, documentation, and structured text output, with a focus on integration into existing toolchains via standard APIs. The model supports common tokenization schemes and runs on widely available frameworks, making it approachable for teams already using Hugging Face pipelines or self-hosted inference servers. Compared to heavier proprietary alternatives, it trades some raw scale for faster iteration and clearer licensing — the MIT license means fewer deployment restrictions. Performance-wise, expect solid reasoning on code-related prompts and decent multilingual support, though it may lag behind frontier models on creative or highly open-ended tasks. If you’re building internal dev tools, automating documentation, or need a lightweight assistant for code generation, this is worth testing alongside your current stack. Always check the model card for specific limitations and intended use cases before deploying in production.

text generationmit
598 starsView details

Qwen-2.5-1B-RLCD

harshatheg
Model

Qwen-2.5-1B-RLCD is a specialized, lightweight text generation model optimized through Reinforcement Learning from Contrastive Differentiation (RLCD). While many small-scale models struggle with instruction following and nuanced reasoning, this 1B-parameter variant is specifically fine-tuned to refine its output alignment, making it an ideal candidate for edge computing and resource-constrained environments. For developers, this means you can deploy a highly responsive agent on local hardware or mobile devices without the latency overhead of much larger LLMs. It excels in tasks requiring high-speed inference, such as real-time autocomplete, basic intent classification, and structured data extraction. Compared to standard base models of similar size, the RLCD training process provides a more disciplined response pattern, reducing the likelihood of repetitive or nonsensical outputs. It integrates seamlessly into existing Hugging Face workflows and is licensed under Apache-2.0, ensuring flexibility for both commercial and research applications.

text generationapache-2.0
573 starsView details

Qwen2.5-3B-Instruct

Qwen
Not specified

Qwen2.5-3B-Instruct is a compact instruction-tuned model from the Qwen family, designed for efficient text generation on resource-constrained setups. At roughly 3B parameters, it balances performance and speed, making it suitable for local development, edge deployment, or lightweight APIs. It works seamlessly with Hugging Face Transformers and supports standard input formats, so integration into existing NLP pipelines is straightforward. Compared to larger models, it trades some reasoning depth for faster inference and lower memory usage, which is ideal for prototyping, multilingual tasks, or embedding-based applications. Developers should review the model card and license for production constraints, as usage terms may vary. Overall, it's a practical choice for teams needing a responsive, small-footprint generative model without heavy infrastructure demands.

text generationother
572 starsView details

Qwen3-1.7B

Qwen
Not specified

Qwen3 1.7B is a compact, high-efficiency language model designed for developers who need strong performance without the overhead of massive parameter counts. Unlike larger LLMs, this model is optimized for low-latency inference and edge deployment, making it an ideal candidate for on-device applications, real-time autocomplete systems, or as a specialized agent in a multi-model pipeline. It balances a small memory footprint with surprising reasoning capabilities, allowing for easy integration into existing workflows via standard APIs or local hosting. For developers, it offers a cost-effective alternative for high-throughput tasks where full-scale frontier models would be overkill or too slow.

text generationapache-2.0
524 starsView details

MiMo-V2.6-Flash-RL

XiaomiMiMo
Model

MiMo-V2.6-Flash-RL is a specialized text-generation model optimized for low-latency environments where speed and reasoning efficiency are paramount. Built on the MiMo architecture, this 'Flash' iteration is specifically tuned using Reinforcement Learning (RL) to refine its decision-making processes and output coherence. For developers, this means a model that excels in high-throughput applications like real-time conversational agents, automated content summarization, and rapid instruction following. Unlike larger, heavier LLMs that prioritize exhaustive knowledge at the cost of inference time, MiMo-V2.6-Flash-RL targets the sweet spot between computational overhead and logical accuracy. It is designed for seamless integration into existing pipelines via Hugging Face, making it a viable candidate for edge deployment or scalable cloud microservices where minimizing time-to-first-token is a critical KPI. While the parameter count remains undisclosed, the RL-tuned architecture suggests a significant leap in following complex, multi-step prompts compared to standard base models.

text generationmit
519 starsView details

GLM-5.3-CYBERSECURITY-FP8

dealignai
Model

GLM-5.3-CYBERSECURITY-FP8 is a specialized text-generation model optimized for security-centric workflows. Unlike general-purpose LLMs, this version is fine-tuned to handle the nuances of cybersecurity tasks, making it a practical tool for automated vulnerability research, threat intelligence synthesis, and security log analysis. By utilizing FP8 quantization, the model offers a significantly reduced memory footprint, allowing developers to deploy it on consumer-grade hardware or edge devices without the massive VRAM requirements of standard high-parameter models. For integration, it follows the standard Hugging Face ecosystem, ensuring compatibility with existing inference pipelines and orchestration frameworks. While general models often struggle with the specific syntax of exploit code or technical security documentation, this model is architected to maintain higher precision in these domains. It is best suited for developers building automated SOC assistants, security auditing tools, or real-time anomaly detection systems where latency and resource efficiency are critical.

text generationmit
502 starsView details

Ornith-1.5-9B-GGUF

ornith-ai
Not specified

Ornith-1.5-9B-GGUF is a quantized text generation model optimized for local deployment and edge computing. Built on a 9-billion parameter architecture, it strikes a balance between reasoning depth and low-latency performance, making it ideal for developers working within hardware-constrained environments. By utilizing the GGUF format, this model is specifically designed for seamless integration with llama.cpp and other high-efficiency inference engines, allowing for efficient CPU and GPU offloading. While smaller than flagship frontier models, its architecture is tuned for high-throughput tasks such as automated content generation, structured data extraction, and conversational agents. For developers, the primary value lies in its portability and the MIT license, which simplifies integration into commercial pipelines without the heavy overhead of larger parameter models. It serves as a pragmatic middle ground for those needing reliable text generation that can run locally on consumer-grade hardware.

text generationmit
407 starsView details

Ternary-Bonsai-2-27B-mlx-2bit

prism-ml
Model

Ternary-Bonsai-2-27B-mlx-2bit is a highly compressed text generation model specifically optimized for the MLX framework. By utilizing a 2-bit ternary quantization scheme, this model drastically reduces its memory footprint, making it an ideal candidate for local execution on Apple Silicon hardware. While standard 27B parameter models typically require significant VRAM, this MLX-specific build allows developers to run sophisticated reasoning and generation tasks on consumer-grade MacBooks without sacrificing much throughput. It is best suited for developers building privacy-focused local agents, edge-based text processing pipelines, or prototyping complex workflows where high-speed inference on macOS is the priority. Compared to standard FP16 or 4-bit deployments, you will see a massive reduction in memory overhead, though you should benchmark the quantization loss against your specific downstream tasks to ensure semantic integrity remains within your required thresholds.

text generationapache-2.0
404 starsView details

AliceAI-Foundation-80B-A3B-Base

yandex
Model

AliceAI-Foundation-80B-A3B-Base is a high-parameter foundation model designed for robust text generation tasks. Built on an 80B architecture, it leverages a Mixture-of-Experts (MoE) approach, specifically utilizing 3B active parameters per token, which offers a strategic balance between computational efficiency and deep reasoning capabilities. For developers, this means you can achieve high-quality outputs without the massive inference overhead typically associated with dense 80B models. It is particularly suited for complex instruction following, creative writing, and structured data generation. Released under the Apache-2.0 license, it provides the legal flexibility required for commercial integration and fine-tuning. Whether you are building RAG pipelines or specialized agents, this model serves as a versatile backbone that competes well in the mid-to-large scale parameter class by optimizing the throughput-to-intelligence ratio.

text generationapache-2.0
358 starsView details

K2-Horizon-MoVA-36B-A4B

IFM
36b

K2-Horizon-MoVA-36B-A4B is a high-efficiency Mixture-of-Experts (MoE) model designed to bridge the gap between mid-sized parameter counts and high-performance reasoning. By utilizing a 36B architecture with an active parameter count of approximately 4B, it offers a highly optimized compute-to-performance ratio. For developers, this means you can achieve sophisticated text generation and logical reasoning capabilities without the massive VRAM overhead typically required by dense 30B+ models. It is particularly well-suited for deployment in resource-constrained environments or edge-cloud hybrid setups where low latency and high throughput are critical. Built under the Apache 2.0 license, it provides a flexible foundation for fine-tuning on domain-specific datasets or integrating into RAG (Retrieval-Augmented Generation) pipelines. Compared to standard dense models, the MoE architecture allows for faster inference speeds, making it a strong candidate for real-time conversational agents and automated content workflows.

text generationapache-2.0
340 starsView details

MiniCPM5-2B-GGUF

openbmb
Model

MiniCPM5-2B-GGUF is a quantized version of the MiniCPM5 series, specifically optimized for efficient deployment on edge devices and consumer-grade hardware. While many small language models struggle with coherence, this 2B parameter model aims to bridge the gap between lightweight footprints and high-quality reasoning. For developers, the GGUF format is the primary draw, allowing for seamless integration with llama.cpp and other inference engines that leverage CPU and GPU offloading. This makes it an ideal candidate for local RAG (Retrieval-Augmented Generation) pipelines, on-device chatbots, and low-latency automation tasks where cloud API costs or privacy concerns are limiting factors. Compared to standard FP16 models, this GGUF implementation significantly reduces VRAM requirements without a proportional loss in logic, making it a practical choice for mobile or IoT-based AI applications.

text generationapache-2.0
326 starsView details
Email