Global AI chat room · 18 online now Join now
DIRECTORY / 02

AI Models | Open-Source LLM Directory

Discover and compare open-source LLMs, language models and multimodal models by capability, scale, license, downloads and provenance.

Compare modelsFind the right building block for your next workflow
Directory overview
591
curated entries
24 topic groupsLive
02 / MODEL INDEX

Find the right model for the job

Context first, better decisions. Every entry keeps the signal that matters.

CURATED DIRECTORY591 results

Mistral-7B-v0.1

mistralai
Model

Mistral-7B-v0.1 represents a significant shift in the efficiency-to-performance ratio for small language models. For developers working with constrained compute environments or looking to deploy locally, this model offers a high-density reasoning capability that punches well above its 7B parameter weight class. Unlike many models in this size category that struggle with long-range dependencies or complex instruction following, Mistral utilizes a sliding window attention mechanism to optimize throughput and context handling. This makes it an ideal backbone for RAG (Retrieval-Augmented Generation) pipelines, local chatbots, and automated code completion tools. Because it is released under the Apache 2.0 license, it provides the legal flexibility required for commercial integration without the heavy overhead of proprietary APIs. Whether you are fine-tuning for a specific domain or deploying via vLLM for high-concurrency inference, Mistral-7B provides a robust, scalable foundation for production-grade NLP applications.

text generationapache-2.0
4.2K starsView details

gpt2

openai-community
Not specified

GPT-2 is a foundational transformer-based language model that marked a shift toward zero-shot learning in NLP. For developers, its primary value today lies in its lightweight architecture and permissive MIT license, making it an ideal candidate for local deployment, fine-tuning on niche datasets, or serving as a baseline for comparative benchmarks. Unlike modern massive LLMs, GPT-2 is computationally efficient, allowing for rapid iteration and hosting on modest hardware without relying on expensive API calls. It excels at basic text completion and structured pattern replication, though it lacks the complex reasoning of its successors. Integration is straightforward via the Hugging Face Transformers library, providing a stable environment for those building specialized text-generation pipelines where latency and privacy outweigh the need for state-of-the-art general intelligence.

text generationmit
4.1K starsView details

DeepSeek-V4-Flash-0731

deepseek-ai
Model

DeepSeek-V4-Flash-0731 is a high-throughput, low-latency text generation model designed for developers who need to balance reasoning depth with rapid inference speeds. While many large-scale models struggle with latency in real-time applications, the 'Flash' architecture optimizes the token generation pipeline, making it a strong candidate for agentic workflows, real-time chat interfaces, and high-volume data extraction tasks. For developers working within the Hugging Face ecosystem, integration is straightforward via the transformers library, allowing for seamless deployment in local or cloud-based environments. Compared to standard heavy-parameter models, this version prioritizes efficiency and cost-effectiveness without sacrificing the core linguistic capabilities required for complex instruction following. It is particularly well-suited for developers building scalable microservices where response time is a critical KPI.

text generationmit
4.0K starsView details

GLM-4 9B Chat

Tsinghua
9B

GLM-4 9B Chat is a lightweight, high-performance bilingual model optimized for seamless Chinese-English reasoning and instruction following. Built by the Zhipu AI team from Tsinghua University, this 9B parameter model strikes a strategic balance between computational efficiency and cognitive depth, making it an ideal candidate for edge deployment or high-throughput microservices where latency is critical. Unlike larger monolithic models, the 9B architecture is designed for developers who need robust multilingual capabilities—specifically handling nuanced semantic shifts between English and Chinese—without the massive infrastructure overhead. It excels in structured data extraction, code assistance, and conversational agent workflows. For integration, it follows standard API patterns, allowing for easy replacement of larger LLMs in specialized pipelines where a smaller, faster footprint is required to maintain low cost-per-token while preserving logical coherence.

text generationGLM-4
3.8K starsView details

Yi-1.5 34B

01.AI
34B

Yi-1.5 34B is a high-performance bilingual model from 01.AI designed to bridge the gap between English and Chinese language processing. For developers, the 34B parameter scale offers a strategic sweet spot: it provides significantly more reasoning depth and nuance than smaller 7B models while maintaining much lower latency and inference costs than massive 70B+ architectures. Unlike many models that treat non-English languages as an afterthought, Yi-1.5 is optimized for high-fidelity cross-lingual tasks, making it ideal for localization pipelines, multilingual chatbots, and complex document summarization involving both languages. Released under the Apache 2.0 license, it is highly accessible for commercial integration. Whether you are fine-tuning for specific domain knowledge or deploying via standard inference engines, this model serves as a robust mid-sized backbone for applications requiring sophisticated linguistic intelligence without the overhead of flagship-scale deployments.

text generationApache 2.0
3.6K starsView details

Edge0-35B-A3B-preview

Edge0
Model

Edge0-35B-A3B-preview is a specialized text-generation model designed for developers seeking a balance between high-performance reasoning and efficient deployment. Built on a 35B parameter architecture, it utilizes an activation-efficient design (A3B) that optimizes throughput without sacrificing the nuance required for complex instruction following. For developers working on edge computing or resource-constrained environments, this model offers a compelling middle ground between lightweight small language models and massive, high-latency frontier models. It is particularly well-suited for RAG pipelines, automated code documentation, and structured data extraction tasks. Since it is released under the Apache-2.0 license, it provides the legal flexibility necessary for commercial integration and fine-tuning. Compared to standard dense models of similar size, the preview architecture aims to deliver lower inference latency, making it a strong candidate for real-time application backends where response speed is as critical as linguistic accuracy.

text generationapache-2.0
3.6K starsView details

phi-2

microsoft
Model

Phi-2 is Microsoft's lightweight, high-performance small language model (SLM) designed to punch significantly above its weight class. While massive LLMs dominate general benchmarks, Phi-2 focuses on high-quality reasoning and logic within a compact parameter footprint. For developers, this means you can deploy sophisticated text generation, code completion, and logical reasoning tasks on edge devices or local environments without the massive VRAM overhead required by larger models. It excels in structured tasks and follows instructions with surprising precision for its size. Unlike general-purpose giants, Phi-2 is optimized for efficiency, making it an ideal backbone for local RAG pipelines, mobile integration, or specialized microservices where latency and compute costs are critical constraints. If your workflow requires high-density intelligence rather than raw scale, Phi-2 offers a highly competitive alternative to much larger proprietary APIs.

text generationmit
3.5K starsView details

gemma-7b

google
Model

Gemma-7b is a lightweight, open-weight transformer model from Google, engineered to deliver high-performance text generation within a manageable parameter footprint. For developers, the primary value lies in its efficiency; it provides a sophisticated balance between reasoning capabilities and low-latency inference, making it ideal for deployment on edge devices or consumer-grade GPUs. Unlike massive frontier models that require extensive cloud infrastructure, Gemma-7b is optimized for local integration and fine-tuning on domain-specific datasets. It excels in tasks such as code completion, summarization, and structured data extraction. While it lacks the raw breadth of much larger models, its architecture is highly responsive to instruction tuning, allowing developers to bake specific logic or stylistic constraints directly into their applications. If you are looking to build private, cost-effective AI agents or specialized NLP pipelines without the overhead of massive API costs, Gemma-7b serves as a robust foundation.

text generationgemma
3.4K starsView details

Mistral-7B-Instruct-v0.2

mistralai
Model

Mistral-7B-Instruct-v0.2 represents a significant milestone in high-performance, small-scale language modeling. For developers, the primary draw is its ability to punch far above its weight class, delivering reasoning and instruction-following capabilities that rival much larger proprietary models. This version improves upon its predecessor with enhanced stability and a refined vocabulary, making it ideal for low-latency edge deployment or fine-tuning on domain-specific datasets. Because it is released under the Apache 2.0 license, it offers the flexibility required for commercial integration without the restrictive overhead of closed-source APIs. Whether you are building sophisticated RAG pipelines, automating complex chat interfaces, or optimizing local inference on consumer-grade hardware, this model provides a highly efficient foundation that balances computational cost with sophisticated linguistic intelligence.

text generationapache-2.0
3.2K starsView details

DeepSeek-V3-0324

deepseek-ai
Model

DeepSeek-V3-0324 is a high-performance text generation model engineered for complex reasoning and large-scale language tasks. For developers integrating LLMs into production pipelines, this model offers a competitive alternative to proprietary closed-source APIs by providing robust instruction-following capabilities and efficient inference potential. It excels in coding assistance, mathematical reasoning, and structured data extraction, making it a versatile choice for building autonomous agents or sophisticated RAG (Retrieval-Augmented Generation) systems. Unlike many models that require heavy fine-tuning for specialized logic, V3 demonstrates strong zero-shot performance across technical domains. Its MIT license simplifies deployment in commercial environments, allowing for deep integration into local infrastructure or cloud-based microservices without the restrictive overhead of proprietary ecosystems. Whether you are optimizing for latency in a chat application or accuracy in a complex reasoning engine, this model provides a scalable foundation for high-throughput text processing.

text generationmit
3.2K starsView details

Llama-3.3-70B-Instruct

meta-llama
Model

Llama-3.3-70B-Instruct represents a significant efficiency milestone for developers working with mid-sized parameter models. While it maintains the 70B footprint, it delivers performance parity with much larger frontier models, making it a sweet spot for high-throughput applications. For engineers, this means you get near-GPT-4 level reasoning and instruction following without the massive latency or infrastructure costs associated with 400B+ parameter architectures. It excels in complex reasoning, coding assistance, and structured data extraction. Because it adheres to the Llama-3 ecosystem, integration is seamless via standard libraries like Transformers, vLLM, or Ollama. Whether you are deploying on-premises to maintain data sovereignty or scaling via managed APIs, this model offers a highly optimized balance of intelligence-per-watt, making it ideal for production-grade RAG pipelines and autonomous agent workflows.

text generationllama3.3
3.0K starsView details

Kimi-K2-Instruct

moonshotai
Model

Kimi-K2-Instruct is the latest instruction-tuned iteration from Moonshot AI, specifically engineered to handle complex reasoning and long-context instruction following. For developers working in multilingual environments, particularly those requiring high proficiency in Chinese and English, this model offers a robust alternative to mainstream Western LLMs. Unlike general-purpose chat models, K2-Instruct is optimized for structured output and logical consistency, making it a strong candidate for agentic workflows, automated coding assistance, and sophisticated data extraction tasks. While specific parameter counts remain proprietary, the model's performance profile suggests a focus on high-density reasoning rather than mere pattern matching. Integration is straightforward via Hugging Face, allowing for seamless deployment within existing inference pipelines. If your roadmap involves building RAG systems or autonomous agents that require nuanced command adherence, Kimi-K2-Instruct provides a competitive edge in logic-heavy applications.

text generationother
3.0K starsView details

starcoder

bigcode
Model

StarCoder is a specialized large language model engineered specifically for code intelligence and programming tasks. Unlike general-purpose LLMs that struggle with long-range syntax dependencies, StarCoder is trained on a massive, diverse corpus of source code, making it highly proficient in autocompletion, code explanation, and multi-language translation. For developers, the primary value lies in its ability to integrate directly into IDE workflows via LSP (Language Server Protocol) or custom plugins. Whether you are building a local co-pilot, automating unit test generation, or implementing complex refactoring tools, StarCoder provides a robust foundation. It is designed to be lightweight enough for efficient deployment while maintaining high accuracy across dozens of programming languages. Compared to monolithic proprietary models, StarCoder offers a more transparent, open-weights alternative that allows for fine-tuning on private repositories, ensuring your codebase's specific patterns and internal APIs are respected during inference.

text generationbigcode-openrail-m
3.0K starsView details

QwQ-32B

Qwen
Model

QwQ-32B is a specialized reasoning model from the Qwen ecosystem designed to bridge the gap between standard LLMs and high-compute reasoning agents. Unlike general-purpose chat models, QwQ is optimized for complex logical workflows, including multi-step mathematical problem solving, advanced code generation, and intricate symbolic reasoning. For developers, this means a significant step up in accuracy for tasks that typically trigger 'hallucinations' in smaller models. At 32B parameters, it offers a sweet spot for deployment: it provides sophisticated chain-of-thought capabilities that rival much larger models while remaining efficient enough to run on consumer-grade or mid-range enterprise hardware. It is particularly useful for building autonomous agents, automated debugging tools, or complex data extraction pipelines where logical consistency is more critical than creative prose. The Apache-2.0 license makes it highly accessible for commercial integration and fine-tuning within existing RAG or agentic frameworks.

text generationapache-2.0
3.0K starsView details

Baichuan2 13B

Baichuan
13B

Baichuan2-13B is a specialized open-weights model designed to bridge the gap between parameter efficiency and high-level cognitive reasoning. While many 13B models struggle with logical consistency, Baichuan2 stands out through its optimized training on diverse datasets, resulting in superior commonsense reasoning and linguistic nuance. For developers working in multilingual environments, it offers a robust alternative to Western-centric models, particularly when handling Chinese-language context or cross-cultural semantic logic. Because it is released under the Apache 2.0 license, it is highly suitable for commercial integration and fine-tuning into specialized pipelines such as automated customer support, content summarization, or lightweight RAG (Retrieval-Augmented Generation) systems. If you are looking for a model that balances a manageable memory footprint with sophisticated reasoning capabilities, Baichuan2-13B is a strong candidate for edge deployment or local inference.

text generationApache 2.0
2.9K starsView details

Ternary-Bonsai-2-27B-gguf

prism-ml
Model

Ternary-Bonsai-2-27B is a specialized large language model optimized for efficient deployment via the GGUF format. Designed for developers working within resource-constrained environments, this 27B parameter model strikes a balance between high-level reasoning and local hardware compatibility. Unlike massive frontier models that require multi-GPU clusters, this model is tailored for edge computing and consumer-grade workstations using llama.cpp or similar inference engines. Its primary strength lies in its architectural efficiency, making it an ideal candidate for RAG (Retrieval-Augmented Generation) pipelines, local coding assistants, and private document processing where data sovereignty is critical. For developers, the GGUF quantization allows for fine-tuned control over the memory-to-performance tradeoff, enabling high-throughput text generation without the overhead of massive VRAM requirements. Whether you are building autonomous agents or integrating LLMs into desktop applications, this model provides a scalable foundation for production-ready, localized AI workflows.

text generationapache-2.0
2.3K starsView details

Qwen2.5-7B-Instruct

Qwen
Model

Qwen2.5 7B Instruct is a dense decoder-only model designed for high-efficiency deployment without sacrificing reasoning depth. For developers, the primary draw is its balanced performance-to-size ratio, making it an ideal candidate for edge computing or local hosting where VRAM is constrained. It demonstrates significant improvements in structured data generation, coding proficiency, and mathematical reasoning compared to its predecessors. Unlike larger frontier models, it offers a low-latency response cycle and is compatible with standard LLM frameworks, simplifying integration into existing RAG pipelines or agentic workflows. It competes directly with other 7B-class models by providing stronger multilingual support and more reliable instruction following, reducing the need for extensive prompt engineering.

text generationapache-2.0
2.2K starsView details

Qwen3-8B

Qwen
Model

Qwen3 8B is a compact yet powerful language model designed for high-efficiency deployment without sacrificing reasoning depth. For developers, the 8B parameter scale represents a sweet spot for local hosting and low-latency API integration, making it ideal for RAG pipelines, autonomous agents, and edge computing. It demonstrates significant improvements in coding proficiency and multilingual nuance compared to its predecessors, offering a competitive alternative to Llama-class models in the same weight category. With an Apache-2.0 license, it provides the flexibility needed for commercial scaling. Whether you are building a specialized domain chatbot or a complex tool-calling workflow, Qwen3 8B balances throughput and intelligence to minimize infrastructure overhead while maintaining high accuracy.

text generationapache-2.0
2.0K starsView details

Xing4.0-29B-A4B

XingChen-AGI
Model

Xing4.0-29B-A4B is a specialized text-generation model designed for developers seeking a balance between high-performance reasoning and efficient deployment. Built on a 29B parameter architecture, this model is optimized to provide nuanced linguistic understanding while maintaining a manageable footprint for modern GPU clusters. Unlike massive frontier models that require extreme compute, Xing4.0 targets the 'sweet spot' of parameter scaling, making it an ideal candidate for fine-tuning on domain-specific datasets or integrating into RAG (Retrieval-Augmented Generation) pipelines. For engineers working with Apache-2.0 licensed software, it offers a permissive environment for both commercial and research applications. While it may not match the raw scale of trillion-parameter models, its efficiency in instruction following and structured output generation makes it a highly competitive choice for developers building autonomous agents or sophisticated conversational interfaces.

text generationapache-2.0
1.8K starsView details

MiniCPM5-2B

openbmb
Model

MiniCPM5-2B is a highly efficient, small-scale language model designed for developers prioritizing low-latency performance and edge-device deployment. While many models focus on massive parameter counts, this 2B-class model optimizes the power-to-performance ratio, making it an ideal candidate for local integration where GPU memory is constrained. It excels in text generation tasks and instruction following, offering a streamlined alternative to larger LLMs for specialized workflows like real-time chatbots, local summarization, or embedded agentic tasks. For developers working within the Apache-2.0 ecosystem, it provides a permissive framework for commercial integration. Compared to standard lightweight models, MiniCPM5-2B focuses on maintaining high reasoning density despite its compact footprint, ensuring that developers don't have to sacrifice much linguistic nuance for the sake of speed and reduced infrastructure costs.

text generationapache-2.0
1.7K starsView details

Llama-3.2-1B-Instruct

meta-llama
Not specified

Llama 3.2 1B Instruct is a lightweight, instruction-tuned model designed for high-efficiency deployment on edge devices and mobile hardware. Unlike its larger siblings, this model prioritizes low latency and a small memory footprint without sacrificing basic reasoning capabilities. It is particularly effective for narrow, task-specific applications such as text summarization, simple entity extraction, and basic conversational interfaces where local execution is required to ensure privacy or reduce API costs. For developers, it offers a viable path to integrate LLM functionality into client-side applications, serving as an ideal candidate for quantization and deployment via frameworks like llama.cpp or MLC LLM. While it lacks the deep world knowledge of larger parameter models, its performance-to-size ratio makes it a strong tool for orchestration and preprocessing pipelines.

text generationllama3.2
1.7K starsView details

Qwen3-0.6B

Qwen
Not specified

Qwen3 0.6B is a highly compact language model designed for efficiency and low-latency deployment. At under one billion parameters, it is optimized for edge computing and resource-constrained environments where VRAM is limited. Unlike larger LLMs, this model is built for specific, high-throughput tasks such as basic text classification, entity extraction, and simple dialogue management. It integrates seamlessly into existing pipelines via Apache-2.0 licensing, making it an ideal candidate for on-device integration or as a fast drafting layer in a larger agentic workflow. Developers can expect a lightweight footprint that allows for rapid iteration and deployment without the need for heavy GPU clusters.

text generationapache-2.0
1.6K starsView details

Ternary-Bonsai-27B-gguf

prism-ml
Model

Ternary-Bonsai-27B-gguf is a quantized text-generation model from the prism-ml team, designed for efficient inference without relying on dense weight matrices. By using ternary representations, it achieves a smaller memory footprint while keeping competitive generation quality compared to traditional dense models. Developers working on edge devices, local tooling, or cost-sensitive API deployments will find it useful for tasks like chatbots, code assistance, and lightweight content generation. The model ships in GGUF format, so it integrates smoothly with llama.cpp-based runtimes and popular local inference stacks. It runs well on consumer GPUs and CPUs, making it a practical option when you need a balance between performance and resource usage. That said, always check the model card for licensing (Apache 2.0) and intended use guidelines before deploying in production.

text generationapache-2.0
1.4K starsView details

Qwen3.8-27B-Uncensored-GGUF

JonathanColetti
Model

For developers working with local LLM deployments, this model offers a specialized branch of the Qwen architecture optimized for high-compliance and unrestricted instruction following. By utilizing the GGUF format, it is specifically engineered for efficient inference on consumer-grade hardware via llama.cpp or similar backends. Unlike standard enterprise models that may trigger false positives on complex creative writing or technical edge cases, this version removes restrictive alignment layers to provide more direct, unfiltered responses. At 27B parameters, it hits a 'sweet spot' for developers: it provides significantly higher reasoning capabilities and nuance than 7B models while remaining small enough to run on high-end enthusiast GPUs or large-memory Mac silicon. It is particularly useful for roleplay engines, uncensored creative writing tools, and complex agentic workflows where strict safety guardrails often break logic flow or prevent deep technical exploration.

text generationapache-2.0
1.3K starsView details
Email