Global AI chat room · 11 online now Join now
DIRECTORY / 02

AI Models | Open-Source LLM Directory

Discover and compare open-source LLMs, language models and multimodal models by capability, scale, license, downloads and provenance.

Compare modelsFind the right building block for your next workflow
Directory overview
839
curated entries
24 topic groupsLive
02 / MODEL INDEX

Find the right model for the job

Context first, better decisions. Every entry keeps the signal that matters.

CURATED DIRECTORY839 results

gemma-4-12B-it-qat-q4_0-gguf

google
Model

The Gemma 4 12B Instruct model represents a significant step forward in the lightweight, open-weights ecosystem, specifically optimized for high-performance local deployment. This quantized GGUF version is tailored for developers who need to balance reasoning depth with hardware constraints, making it ideal for edge computing or consumer-grade GPU setups. Unlike standard text-only LLMs, this model architecture supports 'any-to-any' modalities, allowing you to build sophisticated pipelines that process diverse input types within a single inference pass. For developers working with llama.cpp or similar local inference engines, this 12B parameter model offers a sweet spot: it provides much higher instruction-following accuracy than 7B models while maintaining a significantly lower VRAM footprint than 30B+ architectures. Whether you are integrating it into a local RAG system, an automated coding assistant, or a multimodal agent, the model's ability to handle complex context makes it a versatile tool for production-grade local AI applications.

any to anyapache-2.0
312 starsView details

SenseNova-U1-8B-MoT

sensenova
Model

SenseNova-U1-8B-MoT enters the ecosystem as a compact, highly versatile any-to-any multimodal model. For developers working within resource-constrained environments or edge computing scenarios, this 8B-parameter architecture offers a significant leap in cross-modal reasoning without the massive overhead of larger frontier models. Unlike standard LLMs that rely on separate encoders for vision or audio, the MoT (Mixture-of-Tokens) approach suggests a more unified processing pipeline, enabling smoother transitions between different data modalities. This makes it particularly effective for building interactive agents, real-time multimodal assistants, or complex sensory-input applications. It is released under the Apache-2.0 license, providing the legal flexibility required for commercial integration and fine-tuning. If you are looking to move beyond text-only pipelines and need a model that can natively handle diverse input streams while remaining easy to deploy via Hugging Face, SenseNova-U1-8B-MoT is a strong candidate for your stack.

any to anyapache-2.0
290 starsView details

Qwen3 Embedding 8B

Qwen
Model

Qwen3 Embedding 8B is a high-capacity feature extraction model designed for developers building sophisticated RAG pipelines and semantic search systems. Unlike smaller embedding models, the 8B parameter scale allows for deeper nuance in vector representation, significantly improving retrieval accuracy for complex, long-form queries. It is optimized for dense vectorization across diverse datasets, making it a strong candidate for cross-lingual applications and high-dimensional similarity searches. Integration is straightforward via standard embedding APIs, fitting seamlessly into existing vector databases like Milvus or Pinecone. Compared to previous iterations, this model prioritizes a better balance between representation density and inference latency, providing a robust backbone for enterprise-grade knowledge retrieval without the overhead of a full LLM.

feature-extractionapache-2.0
278 starsView details

twitter-xlm-roberta-base-sentiment

cardiffnlp
Not specified

The twitter-xlm-roberta-base-sentiment model is a multilingual transformer designed specifically for sentiment analysis on short-form social media text. Built on the XLM-RoBERTa architecture, it excels at detecting polarity (positive, negative, neutral) across multiple languages, making it an ideal choice for global brand monitoring or real-time community feedback loops. Unlike standard BERT models, this version is fine-tuned on noisy Twitter data, meaning it handles emojis, slang, and irregular syntax more effectively. For developers, it integrates seamlessly via the Hugging Face ecosystem, offering a lightweight footprint that balances inference speed with cross-lingual accuracy without requiring language-specific preprocessing pipelines.

text classificationSee model card
277 starsView details

LensVLM-9B

apple
Model

LensVLM-9B is a specialized vision-language model from Apple designed to bridge the gap between visual perception and linguistic reasoning. Built on a 9-billion parameter architecture, it functions as an image-text-to-text engine, making it highly effective for tasks requiring nuanced visual understanding, such as detailed image captioning, visual question answering (VQA), and document parsing. For developers, the primary value lies in its balance between model footprint and reasoning depth; it is lightweight enough for efficient deployment in edge-adjacent environments while maintaining the sophisticated semantic grasp typically seen in much larger multimodal models. Unlike general-purpose LLMs that may struggle with spatial grounding, LensVLM is optimized for high-fidelity visual grounding. It integrates seamlessly into existing Hugging Face workflows, offering a robust foundation for building intelligent agents, accessibility tools, or automated visual inspection pipelines. If your stack requires a model that can 'see' and 'reason' without the massive overhead of a 70B+ parameter model, this is a highly competitive candidate for your production pipeline.

image text to textapple-amlr
265 starsView details

Realistic_Vision_V5.1_noVAE

SG161222
Not specified

Realistic Vision V5.1 (noVAE) is a fine-tuned Stable Diffusion checkpoint optimized specifically for photorealism. Unlike base models that often struggle with 'plastic' skin textures or anatomical inconsistencies, this version excels at rendering high-fidelity human portraits, architectural photography, and cinematic lighting. The 'noVAE' designation means the Variational Autoencoder is not baked into the model file, giving developers more flexibility to swap VAEs to manage color saturation and image clarity based on their specific pipeline needs. It is an ideal choice for integrating realistic asset generation into apps, game development, or synthetic dataset creation where visual authenticity is prioritized over stylized art.

text to imagecreativeml-openrail-m
263 starsView details

Z-Image-Turbo-GGUF

unsloth
Not specified

Z-Image-Turbo-GGUF is a specialized text-to-image model optimized for the GGUF format, making it a highly efficient choice for developers looking to run diffusion tasks on consumer-grade hardware. Unlike standard high-parameter models that require massive VRAM, this version leverages quantization to balance inference speed with visual fidelity. It is particularly useful for edge computing applications, local workstations, or integrated environments where memory constraints are a primary concern. For developers working within the llama.cpp or ggml ecosystems, this model provides a streamlined path to integrating generative image capabilities into existing local pipelines. While it serves as a high-speed alternative to larger monolithic architectures, you should benchmark its prompt adherence against your specific creative requirements before scaling in a production environment.

text to imageapache-2.0
263 starsView details

needle3

Cactus-Compute
Model

needle3 is a specialized text-generation model released by Cactus-Compute under the Apache-2.0 license, making it highly accessible for commercial and open-source integration. While many general-purpose models suffer from 'lost in the middle' phenomena, needle3 is architected to prioritize context retention and precise information retrieval within long-form sequences. For developers, this means more reliable performance in RAG (Retrieval-Augmented Generation) pipelines and complex document analysis tasks where factual accuracy is non-negotiable. Unlike monolithic proprietary APIs, needle3 offers the flexibility of local deployment via Hugging Face, allowing you to fine-tune the weights for specific domain expertise or latency requirements. It is an ideal candidate for building agentic workflows that require high-fidelity reasoning over large datasets without the overhead of massive parameter counts found in larger frontier models.

text generationapache-2.0
263 starsView details

SenseNova-U1.5-8B-MoT

sensenova
Model

SenseNova-U1.5-8B-MoT is a specialized any-to-any multimodal model designed for versatile cross-modal processing. Unlike standard LLMs that rely on separate vision or audio encoders, this architecture is built to handle diverse input-output modalities within a unified framework. For developers, the 8B parameter scale strikes a critical balance between high-performance reasoning and deployment efficiency, making it suitable for edge integration or low-latency microservices. The 'MoT' (Mixture-of-Tokens) approach suggests an optimized way of handling heterogeneous data streams, allowing for more granular attention across different modalities. This makes it a strong candidate for complex automation tasks, such as real-time multimedia analysis, interactive voice assistants, or vision-language reasoning pipelines. Released under the Apache-2.0 license, it offers the flexibility required for commercial integration without the typical proprietary constraints found in larger multimodal ecosystems.

any to anyapache-2.0
261 starsView details

LightOnOCR-1B-1025

lightonai
Not specified

LightOnOCR-1B-1025 is an open image-to-text model from LightOn AI, built for optical character recognition and document understanding. It takes images as input and returns extracted text, making it useful for digitizing scanned documents, forms, receipts, or any visual content containing text. The model is published under the permissive Apache-2.0 license, so it can be freely used in commercial and research applications without licensing concerns. It integrates smoothly with Hugging Face transformers, allowing developers to load and run it with standard pipelines for token classification or sequence-to-sequence tasks. While the parameter count isn't specified, its focus on OCR suggests it's optimized for accuracy in reading text rather than general image understanding. Compared to larger multimodal models, LightOnOCR-1B-1025 is lightweight and specialized, offering faster inference and lower resource usage for text extraction workflows. Developers should evaluate the model card for supported languages, input resolution limits, and recommended preprocessing steps before deployment.

image to textapache-2.0
256 starsView details

Qwen3-Omni-30B-A3B-Captioner

Qwen
Model

Qwen3-Omni-30B-A3B-Captioner is a specialized any-to-any multimodal model designed to bridge the gap between diverse sensory inputs and high-fidelity descriptive outputs. Unlike standard text-only LLMs, this model is architected to process complex, multi-modal data streams and generate precise, context-aware captions. For developers building automated content moderation, accessibility tools, or advanced visual search engines, this model offers a significant upgrade in descriptive granularity. Its 'any-to-any' capability suggests a flexible integration path for pipelines that require translating non-textual signals into structured linguistic data. While many captioning models struggle with nuanced spatial reasoning or temporal changes in video, the Qwen3 architecture is optimized for high-density information extraction. It serves as a robust backbone for developers looking to implement sophisticated vision-language tasks without the overhead of massive, general-purpose multimodal giants, providing a more efficient parameter-to-performance ratio for specialized captioning workflows.

any to anyother
248 starsView details

bge-reranker-base

BAAI
Not specified

The bge-reranker-base is a cross-encoder model designed to refine the results of initial vector searches. Unlike bi-encoders used for retrieval, this model evaluates the specific relevance between a query and a document pair, significantly reducing false positives in RAG pipelines. It is particularly effective for developers building high-precision knowledge bases where the top-k results from a vector database need re-scoring to ensure the most contextually accurate information is passed to the LLM. Integration is straightforward via the Sentence-Transformers library or Hugging Face, fitting seamlessly into existing retrieval-augmented generation workflows to boost hit rates without requiring massive index rebuilds.

text classificationmit
246 starsView details

Qwen Image Edit 2509

Qwen
Model

Qwen Image Edit 2509 is a specialized image-to-image model designed for precise visual manipulation. Unlike general generative models, this iteration focuses on maintaining structural consistency while executing specific modifications based on user prompts. For developers, this means a more reliable workflow for tasks like object replacement, style transfer, and localized editing without the common issue of 'hallucinating' the entire scene. It integrates easily into existing AI pipelines via standard API calls and is released under the Apache-2.0 license, offering significant flexibility for commercial deployment and custom fine-tuning. Compared to previous versions, it demonstrates improved adherence to spatial constraints and better preservation of original image details during the editing process.

image-to-imageapache-2.0
246 starsView details

voice-activity-detection

pyannote
Not specified

Pyannote's Voice Activity Detection (VAD) is a specialized tool designed to distinguish human speech from silence or background noise in audio streams. For developers building speech-to-text pipelines or voice assistants, this model serves as a critical preprocessing layer to reduce computational overhead by filtering out non-speech segments before they hit heavier ASR engines. Unlike simple energy-based thresholds, this model handles complex acoustic environments more robustly, making it ideal for long-form audio transcription and speaker diarization workflows. It integrates easily into Python-based stacks and is released under the permissive MIT license, allowing for flexible commercial deployment and modification.

automatic speech recognitionmit
241 starsView details

Agnes-3.0-Flash

Agnes-AI
Model

Agnes-3.0-Flash is a multimodal vision-language model designed for high-throughput image-to-text workflows. Unlike heavy-parameter vision models that struggle with latency, this 'Flash' iteration focuses on optimizing the inference-to-accuracy ratio, making it suitable for real-time applications. For developers, this means you can deploy it in pipelines requiring rapid visual reasoning, such as automated image captioning, visual document parsing, or UI element detection. It follows an Apache-2.0 license, ensuring it is production-ready for commercial integration without restrictive legal overhead. While it may not match the deep reasoning of massive proprietary models, its strength lies in its speed and ease of integration via Hugging Face, providing a lightweight alternative for developers building responsive, vision-aware agents or automated content moderation tools.

image text to textapache-2.0
239 starsView details

RealVisXL_V5.0

SG161222
Not specified

RealVisXL V5.0 is a specialized fine-tune of SDXL designed for developers and creators prioritizing photorealism over stylized art. Unlike base models that often struggle with skin textures and lighting physics, V5.0 optimizes for high-fidelity human anatomy and environmental accuracy. It is particularly effective for generating synthetic datasets, architectural visualizations, and e-commerce assets where visual authenticity is critical. Integration is straightforward via standard Diffusers pipelines or ComfyUI workflows, maintaining compatibility with existing SDXL LoRAs and ControlNets. Compared to previous iterations, V5.0 demonstrates improved prompt adherence and a significant reduction in common anatomical artifacts, making it a reliable engine for production-grade image generation.

text to imageopenrail++
230 starsView details

whisperkit-coreml

argmaxinc
Not specified

WhisperKit CoreML brings OpenAI's Whisper speech-to-text capabilities directly to Apple silicon, optimizing inference for macOS, iOS, and iPadOS. Unlike generic wrappers, this implementation leverages CoreML to utilize the Neural Engine, significantly reducing CPU overhead and battery drain during transcription. For developers, this means the ability to implement high-accuracy, offline ASR (Automatic Speech Recognition) without relying on cloud APIs, ensuring user privacy and low latency. It is particularly effective for building real-time transcription tools, accessibility features, or voice-controlled interfaces where local execution is critical. Integration is streamlined for Swift environments, offering a performant alternative to PyTorch-based deployments on Apple hardware.

automatic speech recognitionmit
228 starsView details

finbert-tone

yiyanghkust
Not specified

FinBERT-Tone is a specialized text-classification model fine-tuned specifically for sentiment analysis within the financial domain. Unlike general-purpose NLP models that often struggle with the nuanced language of markets—where words like 'volatile' or 'bearish' carry specific weights—this model is optimized to categorize financial text into positive, negative, or neutral tones. For developers building algorithmic trading bots, portfolio monitors, or market sentiment dashboards, it provides a reliable way to quantify qualitative data from earnings reports, news feeds, and analyst notes. It integrates easily into standard PyTorch or Hugging Face pipelines, offering a lightweight alternative to LLMs for high-throughput sentiment labeling tasks.

text classificationSee model card
223 starsView details

K2-Horizon-7B

IFM
7b

K2-Horizon-7B is a streamlined 7-billion parameter text generation model optimized for efficiency and deployment flexibility. Developed by IFM under the Apache-2.0 license, it offers a permissive framework for both commercial and research integration. For developers, the primary value lies in its balance between computational overhead and reasoning capabilities, making it an ideal candidate for edge computing or local fine-tuning pipelines where VRAM is a constraint. While larger models may dominate complex reasoning benchmarks, K2-Horizon-7B is engineered for high-throughput tasks such as automated content generation, dialogue management, and structured data extraction. Its architecture allows for seamless integration into existing LLM stacks via Hugging Face, providing a lightweight alternative to much heavier models without sacrificing the core logic required for most production-grade NLP workflows.

text generationapache-2.0
222 starsView details

trocr-base-printed

microsoft
Not specified

TrOCR-base-printed is a transformer-based optical character recognition (OCR) model designed specifically for printed text. Unlike traditional OCR pipelines that rely on separate text detection and recognition stages, TrOCR leverages a vision transformer (ViT) encoder and a language model decoder to map image patches directly to text sequences. This end-to-end architecture makes it particularly effective for high-accuracy transcription of documents, labels, and digitized archives. For developers, it offers a streamlined integration path via the Hugging Face ecosystem, providing a robust alternative to Tesseract or cloud-based OCR APIs when local deployment and Apache-2.0 licensing are priorities.

image to textSee model card
220 starsView details

VoxCPM2

openbmb
Model

VoxCPM2 is an open-source text-to-speech (TTS) model from OpenBMB, designed for developers needing high-fidelity voice synthesis with an Apache-2.0 license. Unlike closed-API solutions, VoxCPM2 allows for local deployment, giving you full control over data privacy and latency. It is optimized for natural prosody and clear articulation, making it suitable for integrating into accessibility tools, automated content creation, or interactive AI agents. For developers, the primary draw is the balance between computational efficiency and output quality, allowing it to scale across various hardware configurations without requiring massive GPU clusters for inference.

text-to-speechapache-2.0
214 starsView details

OrcaSAQ-2-27B

orcarouter
Model

OrcaSAQ-2-27B is a mid-sized text generation model designed to balance computational efficiency with high-reasoning capabilities. For developers working within constrained hardware environments, the 27B parameter count offers a strategic sweet spot—providing significantly more nuanced instruction following and logical depth than standard 7B models without the massive VRAM requirements of 70B+ architectures. While it operates primarily in the text generation domain, its architecture is optimized for tasks requiring structured output and complex context handling. Compared to larger frontier models, OrcaSAQ-2-27B is built for high-throughput applications where latency and deployment costs are critical factors. It is an ideal candidate for fine-tuning on domain-specific datasets or integrating into RAG (Retrieval-Augmented Generation) pipelines where precise information extraction is required. Released under the Apache-2.0 license, it provides the legal flexibility necessary for both research and commercial production environments.

text generationapache-2.0
210 starsView details

gemma-4-E4B-it-qat-GGUF

unsloth
Model

The gemma-4-E4B-it-qat-GGUF model is a quantized, instruction-tuned iteration of the Gemma 4 architecture, optimized specifically for local deployment via the GGUF format. Developed through Unsloth's optimization pipeline, this version focuses on high-efficiency inference, making it ideal for developers working within constrained hardware environments or edge computing scenarios. Unlike standard full-precision models, this Quantization-Aware Training (QAT) variant aims to minimize the perplexity loss typically associated with 4-bit or 8-bit compression, ensuring that reasoning capabilities remain intact while significantly reducing VRAM requirements. For developers building RAG pipelines, local chat interfaces, or automated agents, this model offers a streamlined integration path using llama.cpp or similar runtimes. It bridges the gap between high-performance multimodal reasoning and the practical necessity of low-latency, on-device execution.

any to anyapache-2.0
207 starsView details

gemma-4-E4B-it-ultra-uncensored-heretic-GGUF

llmfan46
Model

For developers working with edge computing or local LLM deployments, this GGUF-quantized iteration of the Gemma 4 architecture offers a specialized approach to multimodal processing. Unlike standard text-only models, this version is designed for 'any-to-any' capabilities, meaning it can handle diverse input modalities within a single streamlined workflow. By utilizing the GGUF format, it is optimized for high-performance inference on consumer-grade hardware via llama.cpp or similar backends, significantly reducing the VRAM overhead typically associated with multimodal models. While the 'heretic' designation suggests a fine-tuned weights profile aimed at reducing alignment constraints, the core value for engineers lies in its versatile integration potential across text and non-text data streams. It serves as a robust baseline for developers building autonomous agents or multimodal RAG systems that require low-latency responses without constant cloud API dependency.

any to anyapache-2.0
204 starsView details
Email