Global AI chat room · 17 online now Join now
DIRECTORY / 02

AI Models | Open-Source LLM Directory

Discover and compare open-source LLMs, language models and multimodal models by capability, scale, license, downloads and provenance.

Compare modelsFind the right building block for your next workflow
Directory overview
67
curated entries
24 topic groupsLive
02 / MODEL INDEX

Find the right model for the job

Context first, better decisions. Every entry keeps the signal that matters.

CURATED DIRECTORY67 results

Qwen3.8-27B

Qwen
Model

Qwen3.8-27B is a multimodal model designed to bridge the gap between visual perception and complex linguistic reasoning. Unlike standard text-only LLMs, this architecture processes image-text inputs to generate high-fidelity textual outputs, making it a versatile tool for developers building vision-centric applications. At 27 billion parameters, it strikes a strategic balance between computational efficiency and deep reasoning capabilities, offering a middle ground for those who find 7B models too shallow but 70B+ models too resource-intensive for real-time inference. For engineers, this means lower latency and reduced VRAM requirements while maintaining strong performance in tasks like visual document understanding, automated image captioning, and complex scene reasoning. Released under the Apache-2.0 license, it is highly accessible for commercial integration and fine-tuning. Whether you are building visual QA systems or automated content moderation pipelines, Qwen3.8-27B provides a robust, production-ready foundation that integrates seamlessly into existing Hugging Face workflows.

image text to textapache-2.0
16.6K starsView details

Kimi-K3

moonshotai
Model

Kimi-K3, developed by moonshotai, is a multimodal model designed to bridge the gap between visual perception and complex textual reasoning. Unlike standard vision-language models that focus solely on captioning, K3 is architected for high-fidelity image-text-to-text tasks, making it a strong candidate for developers building advanced document parsing, visual question answering (VQA), or automated UI inspection tools. For engineers integrating this into existing pipelines, the model offers a sophisticated understanding of spatial relationships and text embedded within images. While many models struggle with dense visual information, Kimi-K3 shows significant promise in maintaining context across multimodal inputs. It is particularly relevant for developers working in OCR-heavy industries or those building intelligent agents that require a 'visual eye' to interpret complex charts, diagrams, and structured layouts. As an open-weight resource on Hugging Face, it provides a flexible foundation for fine-tuning specific domain expertise without the overhead of proprietary API constraints.

image text to textother
11.5K starsView details

Qwen3.8-Flash-Next

Qwen
Model

Qwen3.8-Flash-Next is a high-speed multimodal model designed for developers needing low-latency reasoning across both visual and textual inputs. Unlike heavy-duty vision-language models that struggle with real-time constraints, this 'Flash' iteration optimizes the trade-off between inference speed and spatial understanding. It is particularly effective for applications involving OCR, visual document parsing, and real-time scene description where sub-second response times are critical. For developers integrating via Hugging Face, it offers a streamlined path for building agentic workflows that require 'seeing' and 'reasoning' simultaneously. While it may not match the deep zero-shot reasoning of much larger parameter models, its efficiency makes it a superior choice for edge-case deployment, high-throughput pipelines, and cost-sensitive production environments where latency is the primary bottleneck.

image text to textother
5.8K starsView details

Unlimited-OCR

baidu
Model

Unlimited OCR is a specialized image-to-text model designed for high-accuracy character recognition across diverse visual layouts. Unlike general-purpose LLMs that may struggle with precise spatial positioning or rare glyphs, this model focuses on converting visual text into structured digital strings with minimal hallucination. For developers, it serves as a reliable preprocessing layer for RAG pipelines, automated document digitization, and accessibility tools. It is particularly effective for extracting data from scanned PDFs, receipts, and complex signage where maintaining text integrity is critical. Integration is straightforward via standard API calls, offering a lightweight alternative to massive multimodal models when the primary goal is raw text extraction rather than visual reasoning.

image text to textmit
4.3K starsView details

LLaVA 1.5 7B

LLaVA
7B

LLaVA 1.5 7B is a streamlined multimodal model designed to bridge the gap between visual perception and linguistic reasoning. Unlike traditional vision-language models that often struggle with spatial reasoning or complex instructions, LLaVA 1.5 leverages a projection layer to align visual features from a CLIP encoder with the semantic space of a Llama 2 backbone. For developers, this means a highly efficient 7B parameter footprint that delivers surprisingly high performance in visual question answering (VQA), image captioning, and document understanding. It is particularly useful for edge deployment or as a modular component in larger agentic workflows where low latency is critical. While it may not match the sheer scale of proprietary closed-source giants, its open-weight nature and Llama 2-based architecture make it easy to fine-tune on domain-specific datasets, offering a level of control and cost-efficiency that is ideal for specialized computer vision tasks.

image text to textLlama 2
4.1K starsView details

gemma-4-31B-it

google
Model

Gemma 4 31B IT is a mid-sized, instruction-tuned multimodal model designed for developers who need a balance between high-reasoning capabilities and deployment efficiency. Unlike smaller edge models, the 31B parameter count provides the depth necessary for complex logical tasks and nuanced text generation, while its vision-language integration allows it to process image inputs directly for multimodal RAG or visual analysis. Operating under the Apache-2.0 license, it offers significant flexibility for commercial integration. For developers, this model serves as a powerful alternative to massive frontier models when latency and hosting costs are concerns, yet the task requires more than what a 7B or 9B model can reliably handle.

image text to textapache-2.0
4.0K starsView details

DeepSeek-V4.1-Flash

deepseek-ai
Model

DeepSeek-V4.1-Flash is a high-efficiency multimodal model designed for low-latency image-to-text and text-to-text workflows. Unlike massive monolithic models that sacrifice speed for reasoning depth, this 'Flash' iteration prioritizes throughput and rapid inference, making it an ideal candidate for real-time applications like visual question answering (VQA), automated image captioning, and document parsing. For developers building production-grade pipelines, the model offers a streamlined integration path via Hugging Face, supporting standard vision-language architectures. While it may not match the extreme reasoning capabilities of its larger siblings, its performance-to-cost ratio is optimized for high-volume tasks where latency is a critical bottleneck. It is particularly useful for developers needing to process visual data streams or automate metadata extraction without the overhead of heavy compute resources.

image text to textmit
3.9K starsView details

Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive

HauhauCS
Model

For developers building multimodal applications, Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive offers a specialized approach to vision-language tasks. Based on the Qwen architecture, this 35B parameter model is fine-tuned for high-fidelity image-to-text reasoning and complex instruction following. Unlike standard restricted models, this version is optimized for unfiltered responses, making it a practical choice for researchers and developers working on edge cases, creative writing, or datasets where safety alignment might otherwise suppress nuanced or raw information. It excels in visual reasoning, OCR, and descriptive captioning. Integration is straightforward via Hugging Face, supporting standard multimodal pipelines. While the 'Aggressive' tuning increases response volatility, it provides a significant advantage for developers needing high-entropy outputs that aren't constrained by heavy-handed RLHF filters, allowing for more direct and precise alignment with complex user prompts.

image text to textapache-2.0
3.8K starsView details

DeepSeek-OCR

deepseek-ai
Model

DeepSeek OCR is a specialized vision-language model designed to bridge the gap between raw image pixels and structured text. Unlike traditional OCR engines that rely on rigid layout analysis, this model leverages deep learning to handle complex documents, handwritten notes, and non-standard formatting with high fidelity. For developers, this means fewer pre-processing steps and better accuracy on noisy data. It is particularly effective for automating data extraction from invoices, digitizing legacy archives, and building accessible interfaces for visual content. Integration is streamlined via a standard API, allowing it to fit easily into existing RAG pipelines or document processing workflows where precise text recovery is critical.

image text to textmit
3.4K starsView details

LocateAnything-3B

nvidia
Model

LocateAnything-3B is a specialized vision-language model from NVIDIA designed specifically for high-precision spatial grounding. Unlike general-purpose multimodal models that provide broad descriptions, this 3B-parameter architecture focuses on the 'where' as much as the 'what.' It excels at mapping natural language queries to specific bounding boxes within an image, making it a critical tool for developers building object detection pipelines, visual search engines, or automated robotic vision systems. For engineers looking to integrate grounding capabilities without the massive computational overhead of larger models, its 3B footprint offers a highly efficient middle ground between lightweight detectors and heavy LLMs. It is particularly useful for zero-shot object localization tasks where pre-defined labels are insufficient and you need to detect arbitrary objects via text prompts. Integration is straightforward via Hugging Face, making it a plug-and-play option for RAG-based vision workflows or complex scene understanding applications.

image text to textother
3.1K starsView details

Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled

Jackrong
Model

Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled is a specialized multimodal model designed to bridge the gap between high-level reasoning and efficient parameter counts. By distilling advanced reasoning capabilities into a 27B architecture, this model targets developers who need complex visual-to-text reasoning without the massive latency or compute overhead of flagship-scale models. It excels in multi-step visual logic, where the model must interpret spatial relationships or textual data within images to provide coherent, structured outputs. For engineers building agentic workflows or sophisticated OCR-based analysis tools, this model offers a middle ground: the intuitive instruction-following typical of top-tier reasoning models, paired with the deployment flexibility of a medium-sized open-weights model. It is particularly useful for integration into local pipelines where throughput and reasoning depth must be balanced.

image text to textapache-2.0
2.9K starsView details

Qwen3.6-35B-A3B

Qwen
Model

Qwen3.6 35B A3B is a multimodal model designed for efficient image-text processing, balancing high-parameter intelligence with an optimized architecture. For developers, this model serves as a versatile engine for visual question answering, document parsing, and complex image reasoning tasks. It is particularly useful for building pipelines that require deep semantic understanding of visual inputs without the overhead of massive frontier models. With an Apache-2.0 license, it offers significant flexibility for commercial deployment and fine-tuning. Integration is streamlined for those already using the Qwen ecosystem, providing a competitive alternative to other open-weight multimodal models in terms of accuracy-to-latency ratios.

image text to textapache-2.0
2.9K starsView details

Kimi-K2.5

moonshotai
Model

Kimi-K2.5 is a multimodal model from moonshotai designed to bridge the gap between visual perception and complex linguistic reasoning. Unlike standard text-only LLMs, this model processes image-text inputs to generate high-fidelity textual outputs, making it a strong candidate for vision-language tasks. For developers, the primary value lies in its ability to handle document intelligence, visual question answering (VQA), and scene understanding within a single inference pipeline. While specific parameter counts aren't disclosed, its popularity on Hugging Face suggests robust performance in real-world multimodal benchmarks. Integration is straightforward via the Transformers ecosystem, allowing you to plug it into existing RAG pipelines that require visual context. If your roadmap involves extracting structured data from complex diagrams or building sophisticated visual assistants, Kimi-K2.5 offers a specialized alternative to more generalized multimodal giants.

image text to textother
2.9K starsView details

SigLIP SO400M

Google
400M

...

image text to textApache 2.0
2.8K starsView details

Qwythos-9B-Claude-Mythos-5-1M-GGUF

empero-ai
Model

Qwythos-9B is a specialized multimodal model designed for developers working at the intersection of vision and language. Built on a 9B parameter architecture, this model is optimized for image-to-text and text-to-text tasks, offering a compact footprint that makes it ideal for local deployment or edge computing environments. Unlike massive proprietary models, this GGUF-quantized version is tailored for efficient inference, allowing developers to integrate sophisticated visual reasoning into applications without massive VRAM overhead. It excels in scenarios requiring context-aware image description, visual question answering, and complex reasoning based on visual inputs. For those building RAG pipelines or automated content moderation tools, the model provides a highly accessible entry point for multimodal workflows. While it lacks the raw scale of trillion-parameter models, its performance-to-size ratio makes it a pragmatic choice for developers prioritizing low latency and cost-effective scaling in production-ready local environments.

image text to textapache-2.0
2.8K starsView details

GLM-5.3-Flash

zai-org
Model

GLM-5.3-Flash is a high-speed multimodal model designed for developers requiring low-latency vision-language processing. Unlike heavy-weight vision transformers, this model optimizes the bridge between visual input and textual reasoning, making it an ideal candidate for real-time applications like automated image captioning, visual document parsing, and UI element detection. For international teams, the model's efficiency is its primary selling point; it provides a streamlined inference path that reduces compute overhead without sacrificing the contextual accuracy needed for complex scene understanding. It integrates seamlessly into existing Hugging Face workflows and is released under the MIT license, offering significant flexibility for commercial deployment. If your roadmap involves building responsive agents that need to 'see' and respond instantly, this model offers a highly competitive performance-to-cost ratio compared to larger, more cumbersome multimodal architectures.

image text to textmit
2.6K starsView details

Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF

DavidAU
Model

Qwen3.6 27B Fable Fusion is a specialized large language model designed for developers who need high-parameter reasoning without the restrictive guardrails of standard corporate models. By utilizing a GGUF quantization, it is optimized for local deployment on consumer-grade hardware, offering a strong balance between memory efficiency and cognitive depth. This version is particularly suited for creative writing, complex roleplay, and unfiltered data synthesis where strict adherence to a specific persona or 'heretic' logic is required. It integrates seamlessly into existing llama.cpp or Ollama pipelines, providing a robust alternative for those building applications that require raw, unconstrained text generation and nuanced context handling.

image text to textapache-2.0
2.4K starsView details

Qwen3.6-27B

Qwen
Model

Qwen3.6 27B is a multimodal model designed to bridge the gap between high-performance reasoning and efficient deployment. By integrating image-to-text and text-to-text capabilities, it allows developers to build applications that process visual data and complex queries within a single pipeline. Unlike larger frontier models, the 27B parameter scale offers a strategic balance, providing enough capacity for sophisticated nuance while remaining accessible for self-hosting on consumer-grade or mid-tier enterprise GPUs. It is particularly effective for automated visual inspection, document parsing, and multimodal RAG systems. With an Apache-2.0 license, it provides significant flexibility for commercial integration without the restrictive overhead of proprietary APIs.

image text to textapache-2.0
2.3K starsView details

Qwen3.5-9B

Qwen
Model

Qwen3.5 9B is a versatile multimodal model designed to bridge the gap between lightweight efficiency and high-reasoning capabilities. Unlike standard LLMs, this model natively handles image-to-text and text-to-text tasks, making it an ideal candidate for developers building visual QA systems, automated document parsing, or accessible UI assistants. With a 9B parameter footprint, it offers a competitive performance-to-latency ratio, allowing for deployment on consumer-grade hardware or scaled cloud environments without the overhead of massive frontier models. It integrates seamlessly into existing pipelines via the Apache-2.0 license, providing the flexibility needed for commercial modification and deployment. Compared to previous iterations, it emphasizes improved spatial understanding and more precise grounding in visual contexts.

image text to textapache-2.0
2.1K starsView details

gemma-3-27b-it

google
Model

Gemma-3-27b-it represents a significant step forward in Google's open-weights ecosystem, specifically targeting the intersection of vision and language. Unlike text-only predecessors, this model is natively multimodal, allowing you to process complex visual inputs alongside textual instructions. At 27 billion parameters, it hits a 'sweet spot' for developers: it is large enough to handle sophisticated reasoning and nuanced visual understanding, yet lightweight enough to be deployed on accessible high-end consumer hardware or optimized cloud instances. For developers building RAG pipelines with visual data, automated image captioning systems, or intelligent UI agents, this model offers a high performance-to-compute ratio. It integrates seamlessly into existing Hugging Face workflows and is designed to be fine-tuned for domain-specific vision tasks. Compared to larger proprietary models, it provides a more controlled, cost-effective path for local deployment without sacrificing the reasoning depth required for complex multimodal instruction following.

image text to textgemma
2.0K starsView details

Muse-Glimmer-30B

meta-models
Model

Muse-Glimmer-30B is a mid-sized multimodal model designed for high-fidelity image-to-text reasoning and complex visual instruction following. For developers building vision-language applications, this 30B parameter architecture offers a strategic middle ground between lightweight edge models and massive, computationally expensive frontier models. It excels at tasks requiring deep semantic understanding of visual inputs, such as detailed image captioning, visual question answering (VQA), and extracting structured data from complex diagrams or UI screenshots. Because it is released under the Apache-2.0 license, it is particularly attractive for commercial integration and fine-tuning within private infrastructure. Compared to smaller vision encoders, Muse-Glimmer provides significantly better nuance in descriptive accuracy, making it a strong candidate for automated content moderation, accessibility tools, and sophisticated visual search engines where precision is critical.

image text to textapache-2.0
2.0K starsView details

ZDTaichu5.0-9B

TaichuAI
Model

ZDTaichu5.0-9B is a compact, multimodal model designed for efficient image-to-text and visual reasoning tasks. Built on a 9B parameter architecture, it strikes a pragmatic balance between computational overhead and high-fidelity visual understanding. For developers, this means you can deploy sophisticated vision-language capabilities on consumer-grade hardware or edge devices without the latency typical of much larger foundational models. The model excels at tasks ranging from detailed image captioning and visual question answering (VQA) to complex document parsing where spatial context is critical. Unlike massive general-purpose models that require heavy cloud infrastructure, ZDTaichu5.0 offers a streamlined integration path for developers building real-time visual assistants, automated tagging systems, or accessibility tools. If your workflow requires a model that is fast, relatively lightweight, and capable of grounding text in visual data, this is a highly viable candidate for your local inference stack.

image text to textSee model card
1.9K starsView details

Florence-2-large

microsoft
Model

...

image text to textmit
1.9K starsView details

Qwen3.8-27B-GSQ-RCO-GGUF

ISTA-DASLab
Model

For developers building multimodal applications, Qwen3.8-27B-GSQ-RCO-GGUF represents a highly optimized entry point into high-performance vision-language tasks. This model bridges the gap between complex image understanding and text generation, making it suitable for automated visual inspection, document parsing, and sophisticated captioning pipelines. Unlike standard LLMs, this version is specifically fine-tuned for integrated image-text reasoning, allowing for nuanced context extraction from visual inputs. The GGUF quantization is a key differentiator here; it allows you to run this 27B parameter model on consumer-grade hardware or edge devices with significantly reduced VRAM requirements without a catastrophic loss in perplexity. If you are looking to integrate vision capabilities into local workflows or private cloud environments via llama.cpp or similar runtimes, this model offers a balanced trade-off between inference speed and reasoning depth that outperforms many larger, unquantized alternatives.

image text to textapache-2.0
1.8K starsView details
Email