Global AI chat room · 11 online now Join now
DIRECTORY / 02

AI Models | Open-Source LLM Directory

Discover and compare open-source LLMs, language models and multimodal models by capability, scale, license, downloads and provenance.

Compare modelsFind the right building block for your next workflow
Directory overview
839
curated entries
24 topic groupsLive
02 / MODEL INDEX

Find the right model for the job

Context first, better decisions. Every entry keeps the signal that matters.

CURATED DIRECTORY839 results

Bonsai-2-27B-Ternary-CRACK-GGUF

dealignai
Model

Bonsai-2-27B-Ternary-CRACK-GGUF is a specialized quantization of the Bonsai-2 architecture, optimized for deployment via the GGUF format. For developers working with resource-constrained environments, this model offers a unique middle ground between massive parameter counts and edge-device limitations. By utilizing ternary weight logic, it achieves a significantly reduced memory footprint without the typical performance collapse seen in standard 4-bit integer quantizations. This makes it particularly effective for local RAG (Retrieval-Augmented Generation) pipelines and private LLM deployments where VRAM is at a premium. Unlike standard dense models, the ternary approach focuses on high-efficiency inference, allowing you to run a 27B-class capability on consumer-grade hardware. If your workflow involves integrating LLMs into local desktop applications or specialized edge servers, this model provides a high-throughput, low-latency alternative to larger, more cumbersome weights.

text generationapache-2.0
140 starsView details

Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF

ISTA-DASLab
Model

Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF is a specialized multimodal model designed for developers seeking efficient image-to-text generation. Built on the Qwen architecture, this variant optimizes speed and precision for coding and technical documentation tasks. It processes visual inputs—such as code snippets, diagrams, or UI screenshots—and translates them into structured text output. The GGUF format ensures seamless integration with local inference engines like llama.cpp, making it ideal for edge deployment or private environments without heavy GPU dependencies. Unlike general-purpose vision models, this iteration focuses on 'Flash' performance, reducing latency for real-time applications. The GSQ and RCO tags suggest advanced quantization and optimization techniques, balancing accuracy with reduced memory footprint. Developers can leverage it for automated code documentation, visual debugging assistance, or converting technical diagrams into markdown. Licensed under Apache-2.0, it offers flexibility for commercial and open-source projects. With over 5,000 downloads and growing community interest, it represents a practical choice for teams needing reliable, lightweight vision-language capabilities. Its design prioritizes usability in constrained environments, allowing rapid prototyping of multimodal features. This model bridges the gap between raw visual data and actionable text, streamlining workflows where traditional OCR falls short.

image text to textapache-2.0
135 starsView details

Qwen3 Embedding 0.6B

Qwen
Model

Qwen3 Embedding 0.6B is a compact, high-efficiency feature extraction model designed for developers building RAG pipelines and semantic search systems. At 0.6B parameters, it strikes a balance between low latency and high representational accuracy, making it suitable for edge deployment or high-throughput production environments where larger embedding models introduce too much overhead. It maps text into dense vectors, enabling efficient similarity searches and clustering. For developers, this means faster indexing and lower memory costs without sacrificing the semantic nuance required for complex query retrieval. It integrates seamlessly into standard vector databases and is released under the Apache-2.0 license, ensuring flexibility for commercial applications.

feature-extractionapache-2.0
133 starsView details

gemma-4-E2B-it-qat-q4_0-gguf

google
Model

For developers looking to integrate multimodal capabilities into edge or local environments, the gemma-4-E2B-it-qat-q4_0-gguf represents a highly optimized deployment of Google's latest Gemma 4 architecture. This specific build utilizes Quantization-Aware Training (QAT) and the GGUF format, making it purpose-built for efficient inference on consumer-grade hardware via llama.cpp or similar runtimes. Unlike standard text-only LLMs, this 'any-to-any' model handles diverse input modalities, allowing for more complex reasoning tasks involving cross-modal data. While larger models offer higher reasoning ceilings, this quantized version prioritizes a high performance-to-latency ratio, making it ideal for real-time applications like local voice assistants, vision-integrated chatbots, or automated content analysis where memory constraints are a primary concern. It bridges the gap between heavy cloud-based multimodal APIs and the need for private, low-latency local execution.

any to anyapache-2.0
132 starsView details

Qwen3.8-27B-GGUF

byteshape
Model

Qwen3.8-27B-GGUF is a multimodal model that takes both images and text as input and generates text output, making it suitable for tasks like visual question answering, image captioning, and document understanding. Developed by byteshape and hosted on Hugging Face, it supports the GGUF format which enables efficient CPU-based inference without requiring high-end GPUs. This makes it accessible for developers working in resource-constrained environments or those who want to deploy locally. The model follows an Apache-2.0 license, offering flexibility for both open-source and commercial applications. Compared to larger cloud-only models, Qwen3.8-27B-GGUF trades some scale for portability and ease of integration, especially when paired with GGUF-compatible backends like llama.cpp. Developers can leverage existing Hugging Face pipelines or convert the model for use in custom inference servers. While it may not match the performance of enterprise-grade multimodal systems, its balance of capability and deployability makes it a practical choice for prototyping and edge deployments. Always consult the model card for specific limitations and recommended usage guidelines.

image text to textapache-2.0
123 starsView details

DeepSeek-V4.1-Flash-Abliterated-Cybersecurity-Unleashed

drowzeys
Model

DeepSeek-V4.1-Flash-Abliterated-Cybersecurity-Unleashed is a multimodal model designed for high-speed image-to-text and text-to-text processing, specifically optimized for cybersecurity research workflows. Unlike standard general-purpose models that often trigger heavy safety refusals during technical analysis, this version uses ablation techniques to minimize friction when analyzing code, network logs, or visual security documentation. For developers building automated threat intelligence tools or security auditing pipelines, this model offers a lower-latency alternative to larger parameter models while maintaining high reasoning capabilities in specialized domains. It is particularly useful for parsing complex visual data—such as architectural diagrams or dashboard screenshots—and converting them into actionable structured text. While it provides greater flexibility for security-centric prompting, developers should integrate it via standard Hugging Face workflows and ensure it is deployed within controlled environments suitable for sensitive technical analysis.

image text to textmit
121 starsView details

GOT OCR2 0

stepfun-ai
Model

GOT OCR2.0 is a specialized vision-language model designed to bridge the gap between raw image data and structured text extraction. Unlike traditional OCR engines that rely on rigid pipeline architectures, this model treats document understanding as a generative task, allowing it to handle complex layouts, handwritten notes, and multi-lingual documents with higher spatial awareness. For developers, this means a streamlined integration process where visual parsing and semantic understanding happen in a single pass. It is particularly effective for automating data entry from non-standardized forms and digitizing legacy archives where traditional rule-based systems typically fail. Released under the Apache-2.0 license, it offers the flexibility needed for commercial deployment and custom fine-tuning on domain-specific datasets.

ocrapache-2.0
120 starsView details

Qwen3 VL 4B Instruct

Qwen
Model

Qwen3 VL 4B Instruct is a compact yet powerful vision-language model designed for efficient multimodal processing. Unlike larger VLMs that demand massive VRAM, this 4B parameter version balances latency and reasoning, making it ideal for edge deployment or as a specialized agent in a larger pipeline. It excels at high-resolution image understanding, document parsing, and visual grounding, allowing developers to build applications that can accurately interpret complex layouts or extract data from images. With an Apache-2.0 license, it offers significant flexibility for commercial integration. Compared to previous iterations, it demonstrates improved spatial awareness and a more refined instruction-following capability in image-text tasks, providing a reliable alternative for developers who need a lightweight model without sacrificing significant accuracy.

image-text-to-textapache-2.0
119 starsView details

ThinkingCap-Qwen3.8-27B

bottlecapai
Model

ThinkingCap-Qwen3.8-27B is a specialized multimodal model designed for high-fidelity image-to-text reasoning. Built on the Qwen architecture, this 27B parameter model bridges the gap between visual perception and complex linguistic reasoning, making it particularly effective for tasks that require more than simple captioning. For developers, this means moving beyond basic OCR toward deep visual understanding, such as interpreting complex diagrams, analyzing spatial relationships in UI screenshots, or performing structured data extraction from visual documents. While many lightweight vision models struggle with nuance, the 27B scale provides the necessary cognitive depth to handle multi-step reasoning based on visual input. It is an ideal candidate for integration into RAG pipelines involving visual assets or as a reasoning engine for automated visual inspection workflows. Integration via Hugging Face makes it accessible for local deployment or fine-tuning on domain-specific visual datasets.

image text to textother
119 starsView details

deberta-v3-base-prompt-injection-v2

protectai
Not specified

For developers building LLM-based applications, securing the prompt layer is becoming as critical as the model logic itself. DeBERTa-v3-base-prompt-injection-v2 is a specialized text classification model designed to detect adversarial prompt injection attempts before they reach your core reasoning engine. Built on the DeBERTa-v3 architecture, it offers a highly efficient balance between classification accuracy and inference latency, making it suitable for real-time middleware integration. Unlike general-purpose safety filters, this model is fine-tuned specifically to identify the linguistic patterns characteristic of injection attacks. It integrates seamlessly into existing Hugging Face transformer pipelines, allowing you to implement a robust defensive layer within your existing API workflows. Whether you are deploying an autonomous agent or a customer-facing chatbot, this model serves as a lightweight, high-performance gatekeeper to mitigate the risk of unauthorized instruction overrides.

text classificationapache-2.0
117 starsView details

Swift-1.5-Qwen3.8-27B-GGUF

ukisai
Model

Swift-1.5-Qwen3.8-27B-GGUF is a multimodal model that processes both image and text inputs to generate coherent text responses. Built on the Qwen3 architecture and quantized in GGUF format, it enables efficient deployment on consumer-grade hardware while maintaining strong performance in vision-language tasks. Developers can use it for image captioning, visual question answering, and multimodal chat applications without requiring high-end GPUs. Its GGUF packaging supports seamless integration with llama.cpp and related inference backends, offering flexibility across CPU, GPU, and metal accelerators. While not the largest model in its class, it strikes a practical balance between capability and accessibility, making it suitable for prototyping and production use in resource-constrained environments.

image text to textother
114 starsView details

Qwen3.8-27B-Splash

incoai
Model

Qwen3.8-27B-Splash is a mid-sized text generation model designed to strike a balance between high-reasoning capabilities and deployment efficiency. For developers working with limited VRAM or edge computing environments, the 27B parameter count offers a sweet spot: it provides significantly more nuance and instruction-following stability than 7B models without the massive infrastructure overhead of 70B+ architectures. Built on the robust Qwen lineage, this iteration is optimized for complex text synthesis and logical workflows. It is particularly useful for RAG (Retrieval-Augmented Generation) pipelines where precise context adherence is required, or as a specialized agentic core for tool-calling tasks. Since it is released under the Apache-2.0 license, it is highly suitable for commercial integration and fine-tuning. While it lacks the massive scale of frontier models, its performance-to-compute ratio makes it a strong candidate for production-grade applications requiring low-latency responses.

text generationapache-2.0
113 starsView details

Ternary-Bonsai-2-27B-Abliterated-GGUF

Hikari07jp
Model

Ternary-Bonsai-2-27B-Abliterated is a specialized GGUF quantization of the 27B parameter Bonsai-2 architecture, fine-tuned specifically to minimize refusal triggers and alignment constraints. For developers working on uncensored roleplay, creative writing, or complex instruction-following tasks where standard safety filters often cause false positives, this model offers a significantly more permissive reasoning path. By utilizing the GGUF format, it is optimized for efficient deployment on consumer-grade hardware via llama.cpp, making it highly accessible for local inference. While it maintains the core logic and linguistic capabilities of the base model, the 'abliterated' technique fundamentally alters its response patterns to ensure higher compliance with user prompts. It sits in a sweet spot for developers who need a medium-sized model that balances high-quality prose with a lack of restrictive guardrails, making it a robust choice for private, local-first AI applications.

text generationapache-2.0
111 starsView details

Qwen3.5 4B

Qwen
Model

Qwen3.5 4B is a compact multimodal model designed for efficient image-and-text processing. Unlike larger LLMs that require massive compute, this 4B parameter version focuses on high-density reasoning and visual understanding, making it ideal for edge deployment or low-latency applications. Developers can leverage it for automated image captioning, visual Q&A, and document parsing where fast inference is critical. It follows the Apache-2.0 license, ensuring flexibility for commercial integration. Compared to previous iterations, it balances a smaller memory footprint with improved instruction following, providing a viable alternative for those needing multimodal capabilities without the overhead of a 70B+ parameter model.

image-text-to-textapache-2.0
110 starsView details

IQuest-Q1

IQuestLab
Model

IQuest-Q1 is a text-generation model on Hugging Face by IQuestLab. Downloads 130, likes 109. Read the model card for license and intended use before deploying.

text generationother
109 starsView details

Naive-N0.5-Flash

NaiveAI
Model

Naive-N0.5-Flash is a text-generation model on Hugging Face by NaiveAI. Downloads 605, likes 92. Read the model card for license and intended use before deploying.

text generationmit
108 starsView details

MiMo-V2.6-Distill-Qwen-9B-GGUF

bartowski
Model

MiMo-V2.6-Distill-Qwen-9B-GGUF is a distilled multimodal model optimized for efficient image-to-text reasoning. Built on the Qwen architecture, this 9B parameter model bridges the gap between heavy vision-language models and lightweight edge deployments. By utilizing the GGUF format, it is specifically tailored for quantized inference, making it highly compatible with llama.cpp and various local execution environments. For developers, this means you can run sophisticated visual question answering (VQA), detailed image captioning, and document parsing tasks on consumer-grade hardware without the massive VRAM overhead of larger vision transformers. While it lacks the raw scale of massive frontier models, its distilled nature offers a high performance-to-latency ratio, making it an ideal candidate for real-time vision agents or integrated multimodal chatbots where speed and local privacy are non-negotiable.

image text to textSee model card
107 starsView details

robertuito-sentiment-analysis

pysentimiento
Not specified

Robertuito is a specialized text-classification model designed specifically for sentiment analysis of Spanish-language content. Unlike general-purpose LLMs, it is fine-tuned to handle the nuances of social media discourse, including slang, irony, and the informal linguistic patterns common in Spanish tweets. For developers building social listening tools or customer feedback pipelines, it provides a lightweight, high-accuracy alternative to larger models. It integrates easily into Python workflows via the pysentimiento library, offering a streamlined API for classifying text as positive, negative, or neutral without the latency overhead of massive transformer architectures.

text classificationSee model card
104 starsView details

mms-300m-1130-forced-aligner

MahmoudAshraf
Not specified

mms-300m-1130-forced-aligner is a automatic speech recognition model published on Hugging Face. It is primarily used with transformers and should be evaluated against the model card, license and deployment requirements before production use.

automatic speech recognitioncc-by-nc-4.0
103 starsView details

LFM2.5-2.6B-Qwen3.8-Turbo-Brilliance-Power-X12-NEO-MAX-GGUF

DavidAU
Model

LFM2.5-2.6B-Qwen3.8-Turbo-Brilliance-Power-X12-NEO-MAX-GGUF is a text-generation model on Hugging Face by DavidAU. Downloads 6,489, likes 95. Read the model card for license and intended use before deploying.

text generationapache-2.0
101 starsView details

Realtime-Venus

inclusionAI
Model

Realtime-Venus is an emerging any-to-any multimodal model designed for low-latency, cross-modal interaction. Unlike standard text-to-text or text-to-speech models that rely on cascading discrete modules, this architecture aims to handle diverse input-output streams within a unified framework. For developers, this means a significant reduction in pipeline complexity when building real-time agents, voice assistants, or interactive media tools. While the specific parameter count remains undisclosed, its Apache-2.0 license makes it highly accessible for commercial integration and fine-tuning. If you are working on applications requiring seamless transitions between audio, text, and potentially visual data, Venus offers a streamlined alternative to traditional multi-step inference chains. It is particularly relevant for those looking to minimize the 'turn-taking' latency typical in current conversational AI implementations.

any to anyapache-2.0
101 starsView details

Qwen3.6 27B FP8

Qwen
Model

Qwen3.6 27B FP8 is a high-efficiency multimodal model designed for developers who need a balance between reasoning depth and deployment agility. By utilizing FP8 quantization, this version significantly reduces VRAM overhead and increases throughput without the drastic perplexity loss often seen in 4-bit alternatives. It excels in image-to-text tasks, including complex document parsing, visual reasoning, and structured data extraction from images. For developers, this means easier integration into existing pipelines on consumer-grade hardware or optimized cloud instances. Compared to larger dense models, the 27B parameter count offers a sweet spot for low-latency applications that still require sophisticated understanding of interleaved visual and textual contexts.

image-text-to-textapache-2.0
91 starsView details

Sharp-Spark-X2.5-4B-GGUF

peculiar-ragdoll
Model

Sharp-Spark-X2.5-4B-GGUF is a compact, high-efficiency text generation model optimized for edge deployment and local workflows. Built on a 4-billion parameter architecture, it strikes a pragmatic balance between reasoning capabilities and low computational overhead. For developers working with constrained hardware or latency-sensitive applications, this GGUF-quantized version is specifically designed for seamless integration with llama.cpp and similar inference engines. Unlike larger models that require massive VRAM, this model allows for rapid prototyping of RAG pipelines, local chat interfaces, and automated content generation on consumer-grade hardware. While it may not match the deep nuance of 70B+ parameter models, its speed and small memory footprint make it a highly competitive choice for specialized, task-oriented deployments where throughput and local privacy are the primary engineering constraints.

text generationapache-2.0
91 starsView details

gemma-4-E2B-it-GGUF

ggml-org
Model

The gemma-4-E2B-it-GGUF is a quantized implementation of the Gemma 4 architecture, optimized specifically for local deployment via the llama.cpp ecosystem. Unlike standard text-only models, this iteration leverages any-to-any capabilities, allowing developers to build pipelines that process and reason across multiple modalities including text, images, and audio. For engineers working under hardware constraints, the GGUF format provides a critical advantage by enabling efficient memory management through 4-bit or 8-bit quantization without significant logic degradation. This makes it an ideal candidate for edge computing, privacy-focused local assistants, or RAG workflows where low latency and offline availability are non-negotiable. While larger proprietary APIs offer massive scale, this model provides a highly portable, open-weight alternative for developers who need granular control over their inference stack and deployment environment.

any to anyapache-2.0
89 starsView details
Email