Global AI chat room · 18 online now Join now
DIRECTORY / 02

AI Models | Open-Source LLM Directory

Discover and compare open-source LLMs, language models and multimodal models by capability, scale, license, downloads and provenance.

Compare modelsFind the right building block for your next workflow
Directory overview
839
curated entries
24 topic groupsLive
02 / MODEL INDEX

Find the right model for the job

Context first, better decisions. Every entry keeps the signal that matters.

CURATED DIRECTORY839 results

Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16

nvidia
Model

Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16 represents NVIDIA's push into efficient, multi-modal reasoning architectures. Unlike standard text-only LLMs, this 'any-to-any' model is designed to handle diverse input and output modalities, making it a versatile candidate for complex, multi-step reasoning tasks that require more than just pattern matching. For developers, the primary value proposition lies in its specialized reasoning capabilities paired with a compact 30B parameter footprint, optimized for BF16 precision. This balance suggests it can be deployed in environments where low latency and high intelligence are required without the massive overhead of trillion-parameter models. Whether you are building autonomous agents, complex multimodal assistants, or advanced data processing pipelines, this model offers a streamlined integration path for developers looking to move beyond simple chat interfaces into true cognitive task automation.

any to anyother
434 starsView details

Swift-Qwen3.8-27B-GGUF

ukisai
Model

Swift-Qwen3.8-27B-GGUF is a quantized multimodal model designed for efficient vision-language tasks. Built on the Qwen architecture, this 27B parameter model bridges the gap between high-reasoning capabilities and local deployment feasibility. By utilizing the GGUF format, it is specifically optimized for llama.cpp and other CPU/GPU hybrid inference engines, making it an ideal choice for developers building local RAG pipelines or edge-based visual assistants. Unlike massive proprietary vision models, this version offers a streamlined balance of spatial understanding and text generation, allowing for complex image captioning, document parsing, and visual reasoning without the latency of heavy cloud APIs. For developers working with constrained hardware, the quantization provides a significant reduction in VRAM requirements while maintaining the structural intelligence necessary for nuanced multimodal interaction.

image text to textother
432 starsView details

multilingual-e5-small

intfloat
Not specified

Multilingual E5 Small is a lightweight, high-efficiency embedding model designed for cross-lingual sentence similarity and semantic search. Unlike larger LLMs, this model focuses specifically on mapping text from multiple languages into a shared vector space, making it an ideal choice for developers building RAG (Retrieval-Augmented Generation) pipelines or clustering systems where low latency and minimal memory overhead are critical. It balances performance with a small footprint, allowing for deployment on edge devices or CPU-only environments without sacrificing significant retrieval accuracy. Integration is straightforward via standard sentence-transformer libraries, providing a scalable alternative to proprietary embedding APIs for international applications.

sentence similaritymit
419 starsView details

Ornith-1.5-9B-GGUF

ornith-ai
Not specified

Ornith-1.5-9B-GGUF is a quantized text generation model optimized for local deployment and edge computing. Built on a 9-billion parameter architecture, it strikes a balance between reasoning depth and low-latency performance, making it ideal for developers working within hardware-constrained environments. By utilizing the GGUF format, this model is specifically designed for seamless integration with llama.cpp and other high-efficiency inference engines, allowing for efficient CPU and GPU offloading. While smaller than flagship frontier models, its architecture is tuned for high-throughput tasks such as automated content generation, structured data extraction, and conversational agents. For developers, the primary value lies in its portability and the MIT license, which simplifies integration into commercial pipelines without the heavy overhead of larger parameter models. It serves as a pragmatic middle ground for those needing reliable text generation that can run locally on consumer-grade hardware.

text generationmit
407 starsView details

Ternary-Bonsai-2-27B-mlx-2bit

prism-ml
Model

Ternary-Bonsai-2-27B-mlx-2bit is a highly compressed text generation model specifically optimized for the MLX framework. By utilizing a 2-bit ternary quantization scheme, this model drastically reduces its memory footprint, making it an ideal candidate for local execution on Apple Silicon hardware. While standard 27B parameter models typically require significant VRAM, this MLX-specific build allows developers to run sophisticated reasoning and generation tasks on consumer-grade MacBooks without sacrificing much throughput. It is best suited for developers building privacy-focused local agents, edge-based text processing pipelines, or prototyping complex workflows where high-speed inference on macOS is the priority. Compared to standard FP16 or 4-bit deployments, you will see a massive reduction in memory overhead, though you should benchmark the quantization loss against your specific downstream tasks to ensure semantic integrity remains within your required thresholds.

text generationapache-2.0
404 starsView details

Prompt-Guard-86M

meta-llama
Not specified

Prompt Guard 86M is a lightweight, specialized classifier designed to secure LLM pipelines by detecting prompt injections and jailbreak attempts. Unlike general-purpose models, this 86M-parameter model is optimized for low-latency inference, making it an ideal first-pass filter before requests hit your primary generative model. It categorizes inputs into 'safe' or 'unsafe' based on adversarial patterns, allowing developers to implement programmatic guards without sacrificing system performance. It integrates easily into existing middleware or API gateways, providing a critical layer of defense against malicious user inputs that seek to bypass system instructions.

text classificationllama3.1
403 starsView details

Chroma-4B

FlashLabs
Model

Chroma-4B, developed by FlashLabs, is an 'any-to-any' multimodal model designed to bridge the gap between diverse data modalities within a compact 4B parameter footprint. For developers working on resource-constrained environments or edge computing, this model offers a streamlined alternative to massive, multi-stage pipelines. Unlike traditional models that require separate encoders for different tasks, Chroma-4B aims to handle cross-modal transformations natively, making it highly suitable for unified sensory processing tasks. While the exact architecture details are evolving, its Apache-2.0 license provides the flexibility needed for commercial integration and fine-tuning. If your workflow involves complex interactions between text, vision, or audio, Chroma-4B serves as a versatile foundation for building integrated, multi-sensory applications without the overhead of much larger foundational models.

any to anyapache-2.0
388 starsView details

multilingual-e5-base

intfloat
Not specified

Multilingual-e5-base is a high-performance text embedding model designed for cross-lingual semantic search and retrieval. Unlike generative LLMs, this model maps text into a dense vector space, making it ideal for developers building RAG (Retrieval-Augmented Generation) pipelines or clustering systems across multiple languages. It excels at sentence-similarity tasks, allowing you to match queries to documents even when they are in different languages. Integration is straightforward via Hugging Face Transformers or Sentence-Transformers, offering a lightweight footprint that balances latency with retrieval accuracy. Compared to larger proprietary embeddings, it provides a transparent, MIT-licensed alternative that can be self-hosted to ensure data privacy and reduce API costs.

sentence similaritymit
387 starsView details

Qwen3.8-Flash-Next-GSQ-RCO-GGUF

ISTA-DASLab
Model

Qwen3.8-Flash-Next-GSQ-RCO-GGUF is a specialized multimodal model optimized for high-speed image-to-text reasoning and visual understanding. Built on the Qwen architecture and quantized via GGUF, this iteration is specifically designed for developers requiring low-latency performance on consumer-grade hardware or edge devices. Unlike standard large-scale vision models that demand massive VRAM, this 'Flash' variant prioritizes throughput and efficient inference without sacrificing significant spatial reasoning capabilities. It is particularly effective for real-time visual captioning, document parsing, and visual QA workflows where response time is a critical KPI. For teams integrating vision capabilities into local applications, the GGUF format ensures seamless compatibility with llama.cpp and other lightweight inference engines, making it a highly practical choice for local-first AI deployments and privacy-sensitive environments.

image text to textapache-2.0
380 starsView details

Qwen3 VL 8B Instruct

Qwen
Model

Qwen3 VL 8B Instruct is a versatile vision-language model designed for developers needing efficient, high-performance multimodal capabilities without the overhead of massive parameter counts. It excels at bridging the gap between visual perception and textual reasoning, making it ideal for complex OCR tasks, document parsing, and real-time image analysis. Unlike general-purpose LLMs, this model is optimized for precise spatial understanding and detailed visual grounding. With an Apache-2.0 license, it offers significant flexibility for commercial deployment. It integrates seamlessly into existing AI pipelines via standard inference frameworks, providing a competitive alternative to larger proprietary models by balancing latency with high-accuracy visual interpretation.

image-text-to-textapache-2.0
378 starsView details

gte-multilingual-base

Alibaba-NLP
Not specified

For developers building cross-lingual search or semantic retrieval systems, gte-multilingual-base offers a robust foundation for mapping diverse languages into a shared vector space. Unlike monolingual models that require translation layers, this architecture is designed to handle sentence similarity tasks directly across multiple languages, making it ideal for globalized RAG (Retrieval-Augmented Generation) pipelines and multilingual FAQ bots. It integrates seamlessly with the sentence-transformers library, allowing for straightforward implementation of cosine similarity workflows. While it serves as a high-performance 'base' model, developers should benchmark its embedding density against specific domain datasets to ensure retrieval precision. Compared to larger, general-purpose LLMs, this model provides a more computationally efficient path for high-throughput semantic search tasks where latency and memory footprint are critical constraints.

sentence similarityapache-2.0
377 starsView details

xlm-roberta-base-language-detection

papluca
Not specified

For developers building multilingual applications, accurately identifying input language is a foundational step for routing tasks to specialized downstream models. The xlm-roberta-base-language-detection model leverages the robust XLM-RoBERTa architecture to provide high-precision language identification across a wide array of global scripts. Unlike simple N-gram or dictionary-based detectors, this transformer-based approach captures semantic and structural nuances, making it more resilient to code-switching and noisy text. It is designed for seamless integration via the Hugging Face Transformers library, making it easy to plug into existing NLP pipelines. While it is optimized for classification speed, developers should note its parameter footprint relative to the base XLM-R model. It is best suited for pre-processing stages in content moderation, multilingual search indexing, or automated translation workflows where reliable language tagging is a prerequisite for performance.

text classificationmit
376 starsView details

Qwen2.5-Omni-3B

Qwen
Model

Qwen2.5-Omni-3B represents a significant step toward efficient, native multimodal intelligence for edge and local deployments. Unlike traditional pipelines that chain separate vision and audio encoders to a language model, this 'any-to-any' architecture is designed to process and generate across multiple modalities within a unified framework. For developers, the 3B parameter count is the sweet spot: it offers enough reasoning capacity for complex instruction following while remaining small enough to run on consumer-grade hardware or mobile environments with low latency. You can leverage this model for real-time voice assistants, visual reasoning tasks, or interactive multimodal agents where context switching between text, vision, and audio must be seamless. Compared to larger, monolithic models, Qwen2.5-Omni-3B prioritizes high-speed inference and architectural fluidity, making it an ideal backbone for integrated applications that require more than just text-based interaction.

any to anyother
358 starsView details

AliceAI-Foundation-80B-A3B-Base

yandex
Model

AliceAI-Foundation-80B-A3B-Base is a high-parameter foundation model designed for robust text generation tasks. Built on an 80B architecture, it leverages a Mixture-of-Experts (MoE) approach, specifically utilizing 3B active parameters per token, which offers a strategic balance between computational efficiency and deep reasoning capabilities. For developers, this means you can achieve high-quality outputs without the massive inference overhead typically associated with dense 80B models. It is particularly suited for complex instruction following, creative writing, and structured data generation. Released under the Apache-2.0 license, it provides the legal flexibility required for commercial integration and fine-tuning. Whether you are building RAG pipelines or specialized agents, this model serves as a versatile backbone that competes well in the mid-to-large scale parameter class by optimizing the throughput-to-intelligence ratio.

text generationapache-2.0
358 starsView details

NuExtract3

numind
Not specified

NuExtract3 is a specialized image-to-text model designed for structured information extraction. Unlike general-purpose VLMs that often struggle with precision or hallucinate during data parsing, NuExtract3 focuses on transforming unstructured visual data into machine-readable formats. It is particularly effective for developers building automated pipelines for invoice processing, form digitization, and document analysis where schema adherence is critical. With an Apache-2.0 license, it offers the flexibility for commercial integration without restrictive overhead. Developers can integrate it into existing OCR workflows to replace brittle rule-based parsing with a more robust, neural extraction layer that maintains high fidelity to the source document.

image to textapache-2.0
349 starsView details

DeepSeek-V4.1-Flash-UNCENSORED-FP8

dealignai
Model

DeepSeek-V4.1-Flash-UNCENSORED-FP8 is a high-throughput multimodal model optimized for low-latency vision-language tasks. Built on the Flash architecture and quantized to FP8, it strikes a balance between rapid inference speeds and significant memory savings, making it ideal for edge deployment or cost-sensitive scaling. Unlike standard vision models that struggle with restrictive alignment, this iteration is tuned for high instruction-following fidelity across diverse visual contexts without heavy-handed filtering. For developers, this means more reliable performance in complex OCR, visual reasoning, and document analysis workflows where precision is non-negotiable. It integrates seamlessly into standard Hugging Face pipelines, offering a streamlined path for those needing to process image-text pairs in real-time applications such as automated visual inspection or interactive multimodal agents.

image text to textmit
348 starsView details

Qwen3-ASR-0.6B

Qwen
Not specified

Qwen3 ASR 0.6B is a compact, high-efficiency automatic speech recognition model designed for low-latency transcription tasks. At 0.6 billion parameters, it is optimized for edge deployment and resource-constrained environments where full-scale models are impractical. Developers can integrate this model into real-time voice pipelines, accessibility tools, or lightweight virtual assistants without sacrificing significant accuracy. Unlike larger ASR frameworks, it offers a lean memory footprint, making it an ideal candidate for on-device processing or high-throughput server-side scaling. Licensed under Apache-2.0, it provides the flexibility needed for both commercial integration and custom fine-tuning on domain-specific audio datasets.

automatic speech recognitionapache-2.0
342 starsView details

K2-Horizon-MoVA-36B-A4B

IFM
36b

K2-Horizon-MoVA-36B-A4B is a high-efficiency Mixture-of-Experts (MoE) model designed to bridge the gap between mid-sized parameter counts and high-performance reasoning. By utilizing a 36B architecture with an active parameter count of approximately 4B, it offers a highly optimized compute-to-performance ratio. For developers, this means you can achieve sophisticated text generation and logical reasoning capabilities without the massive VRAM overhead typically required by dense 30B+ models. It is particularly well-suited for deployment in resource-constrained environments or edge-cloud hybrid setups where low latency and high throughput are critical. Built under the Apache 2.0 license, it provides a flexible foundation for fine-tuning on domain-specific datasets or integrating into RAG (Retrieval-Augmented Generation) pipelines. Compared to standard dense models, the MoE architecture allows for faster inference speeds, making it a strong candidate for real-time conversational agents and automated content workflows.

text generationapache-2.0
340 starsView details

all-MiniLM-L12-v2

sentence-transformers
Not specified

all-MiniLM-L12-v2 is a lightweight, high-performance sentence-transformer designed for generating dense vector embeddings. Unlike larger LLMs, it is optimized specifically for semantic search and sentence similarity tasks, mapping text into a 384-dimensional space. For developers, this means significantly lower latency and memory overhead during inference, making it ideal for edge deployment or as the retrieval engine in RAG (Retrieval-Augmented Generation) pipelines. It balances speed and accuracy effectively, providing a robust alternative to heavier models when building vector databases or clustering large datasets where millisecond response times are critical.

sentence similarityapache-2.0
330 starsView details

MiniCPM5-2B-GGUF

openbmb
Model

MiniCPM5-2B-GGUF is a quantized version of the MiniCPM5 series, specifically optimized for efficient deployment on edge devices and consumer-grade hardware. While many small language models struggle with coherence, this 2B parameter model aims to bridge the gap between lightweight footprints and high-quality reasoning. For developers, the GGUF format is the primary draw, allowing for seamless integration with llama.cpp and other inference engines that leverage CPU and GPU offloading. This makes it an ideal candidate for local RAG (Retrieval-Augmented Generation) pipelines, on-device chatbots, and low-latency automation tasks where cloud API costs or privacy concerns are limiting factors. Compared to standard FP16 models, this GGUF implementation significantly reduces VRAM requirements without a proportional loss in logic, making it a practical choice for mobile or IoT-based AI applications.

text generationapache-2.0
326 starsView details

gemma-4-31B-it-assistant

google
Model

The Gemma 4 31B-it-assistant is a high-parameter, instruction-tuned model designed for complex, multimodal workflows. Unlike standard text-only LLMs, this 'any-to-any' architecture allows developers to build applications that seamlessly process and reason across diverse data modalities. At 31B parameters, it strikes a strategic balance between high-level reasoning capabilities and deployment efficiency, making it suitable for edge-cloud hybrid architectures or high-throughput local inference. For developers, the primary value lies in its versatility: you can leverage it for sophisticated cross-modal retrieval, complex instruction following, and integrated multimodal reasoning tasks. Released under the Apache 2.0 license, it offers the flexibility needed for commercial integration without the constraints of restrictive proprietary licenses. Whether you are building advanced agents or multimodal RAG pipelines, this model provides a robust foundation for non-linear data processing.

any to anyapache-2.0
325 starsView details

opt-125m

facebook
Not specified

opt-125m is Meta's compact text generation model designed for developers who need a lightweight solution that runs efficiently on modest hardware. With 125M parameters, it trades raw scale for speed and accessibility, making it suitable for prototyping, edge deployment, or scenarios where larger models are impractical. It integrates smoothly with Hugging Face Transformers and can be fine-tuned for tasks like summarization, classification, or chatbots. While it won't match the quality of billion-parameter models, its low resource footprint and permissive license make it attractive for experimentation and small-scale production use. Developers should check the model card and license carefully before deployment, as usage restrictions may apply.

text generationother
323 starsView details

Qwen3-Omni-30B-A3B-Thinking

Qwen
Model

Qwen3-Omni-30B-A3B-Thinking represents a significant shift toward true multimodal reasoning. Unlike standard LLMs that rely on separate vision or audio encoders, this 'any-to-any' architecture is designed to process and generate across multiple modalities natively. For developers, the standout feature is the integrated 'thinking' process, which allows the model to perform complex, multi-step chain-of-thought reasoning before outputting a response. This makes it particularly effective for sophisticated tasks like interleaved multimodal dialogue, complex visual reasoning, and real-time audio interaction. While the 30B parameter scale offers a sweet spot between high-level intelligence and deployment efficiency, the true value lies in its ability to handle non-textual inputs without the latency typical of modular pipelines. Whether you are building autonomous agents or advanced multimodal interfaces, this model provides a unified backbone that reduces the need for complex orchestration of multiple specialized models.

any to anyother
323 starsView details

distilbert-base-multilingual-cased-sentiments-student

lxyuan
Not specified

The distilbert-base-multilingual-cased-sentiments-student is a lightweight, distilled version of BERT designed for efficient sentiment analysis across multiple languages. By leveraging knowledge distillation, it maintains a high degree of accuracy while significantly reducing latency and memory overhead compared to full-scale transformer models. For developers, this makes it an ideal candidate for real-time production environments, edge deployments, or applications requiring rapid inference without sacrificing multilingual support. It integrates seamlessly into standard Hugging Face pipelines, allowing for quick deployment in customer feedback loops, social media monitoring, and global sentiment tracking across diverse linguistic datasets.

text classificationapache-2.0
317 starsView details
Email