Global AI chat room · 12 online now Join now
DIRECTORY / 02

AI Models | Open-Source LLM Directory

Discover and compare open-source LLMs, language models and multimodal models by capability, scale, license, downloads and provenance.

Compare modelsFind the right building block for your next workflow
Directory overview
839
curated entries
24 topic groupsLive
02 / MODEL INDEX

Find the right model for the job

Context first, better decisions. Every entry keeps the signal that matters.

CURATED DIRECTORY839 results

pix2text-mfr

breezedeus
Not specified

pix2text mfr is a specialized image-to-text model designed for high-accuracy mathematical formula recognition. Unlike general OCR, this model focuses on the structural complexity of LaTeX-style notation, making it ideal for digitizing academic papers, technical documentation, and educational content. For developers, it serves as a robust bridge between raw image data and editable markup, streamlining the pipeline for converting screenshots or PDFs into structured mathematical text. It integrates efficiently into document processing workflows where precision in symbolic representation is critical, offering a lightweight alternative to heavy multimodal LLMs for specific formula extraction tasks.

image to textmit
58 starsView details

wav2vec2-large-xlsr-53-arabic

jonatasgrosman
Not specified

For developers building voice-driven applications in the MENA region, wav2vec2-large-xlsr-53-arabic offers a robust foundation for Automatic Speech Recognition (ASR). Built on the XLS-R architecture, this model leverages cross-lingual pre-training to handle the nuances of Arabic phonetics more effectively than standard monolingual models. It is specifically optimized for transcribing spoken Arabic into text, making it a primary candidate for voice assistants, automated captioning, and transcription services. Integration is straightforward via the Hugging Face Transformers library, allowing for seamless deployment in Python-based workflows. While it excels at capturing acoustic patterns, developers should note that performance may vary across different regional dialects. For production environments, we recommend fine-tuning the model on your specific domain-specific datasets to maximize word error rate (WER) improvements and ensure high accuracy in specialized technical or conversational contexts.

automatic speech recognitionapache-2.0
55 starsView details

faster-whisper-small

Systran
Not specified

For developers building real-time transcription services or batch processing pipelines, faster-whisper-small offers a high-efficiency alternative to the standard OpenAI Whisper implementation. By leveraging the CTranslate2 inference engine, this model achieves significantly lower latency and reduced memory footprints without sacrificing the core acoustic modeling capabilities of the original architecture. It is specifically optimized for CPU and GPU deployment where throughput is a critical KPI. While the 'small' parameter size is a strategic trade-off favoring speed and low-resource environments, it remains highly effective for clear audio in standard languages. Integrating this into your stack is straightforward via Hugging Face, making it an ideal candidate for edge computing, voice-command interfaces, or scalable microservices where cost-per-inference must be minimized.

automatic speech recognitionmit
52 starsView details

MiniCPM V

openbmb
Model

MiniCPM V is a compact yet powerful vision-language model designed for efficient multimodal processing. Unlike monolithic VLM architectures, it focuses on high-performance visual question answering (VQA) while maintaining a small enough footprint for deployment in resource-constrained environments. For developers, this means a lower barrier to entry for integrating image-to-text capabilities into applications without requiring massive GPU clusters. It excels at interpreting visual context and translating it into structured text, making it ideal for automating image tagging, accessibility tools, and visual data extraction. Built under the Apache-2.0 license, it offers the flexibility needed for commercial integration and customization.

visual-question-answeringApache-2.0
52 starsView details

paraphrase-mpnet-base-v2

sentence-transformers
Not specified

The paraphrase-mpnet-base-v2 is a high-performance sentence-transformer model optimized for mapping sentences and paragraphs to a dense vector space. Unlike general-purpose LLMs, this model is purpose-built for semantic similarity and clustering tasks, leveraging an MPNet architecture to balance the strengths of Masked Language Modeling (MLM) and Permuted Language Modeling (PLM). For developers, this means highly accurate embeddings that capture nuanced meaning rather than just keyword overlap. It is an ideal drop-in replacement for BERT-based encoders in RAG pipelines, semantic search engines, and duplicate detection systems. Integration is straightforward via the sentence-transformers library, offering a computationally efficient alternative to larger models while maintaining state-of-the-art retrieval performance on benchmark datasets.

sentence similarityapache-2.0
51 starsView details

xor

juspay
Model

xor is a text-generation model on Hugging Face by juspay. Downloads 2,346, likes 49. Read the model card for license and intended use before deploying.

text generationapache-2.0
49 starsView details

Qwen3.5 2B

Qwen
Model

Qwen3.5 2B is a compact, multimodal model designed for high-efficiency deployment in edge computing and resource-constrained environments. Unlike larger LLMs, this 2B parameter model balances a small memory footprint with strong image-text understanding, making it ideal for real-time visual analysis, OCR tasks, and interactive AI agents. It follows the Apache-2.0 license, offering developers significant flexibility for commercial integration. For engineers, this means the ability to run sophisticated vision-language tasks locally on consumer hardware or mobile devices without sacrificing the reasoning capabilities typically found in larger models. It serves as a versatile drop-in for pipelines requiring fast inference and low latency across diverse visual inputs.

image-text-to-textapache-2.0
43 starsView details

Qwen3.6 27B AWQ INT4

cyankiwi
Model

Qwen3.6 27B AWQ INT4 is a quantized multimodal model designed for developers needing a high-performance balance between reasoning capabilities and VRAM efficiency. By utilizing 4-bit AWQ quantization, this version significantly lowers the hardware barrier for deploying a 27B parameter model without substantial loss in perplexity or accuracy. It excels in vision-language tasks, allowing for seamless integration into pipelines that require complex image analysis, document parsing, and interleaved text-image reasoning. For developers, this means faster inference speeds and the ability to run the model on consumer-grade GPUs or tighter cloud instances compared to the full-precision weights, making it an ideal candidate for production-ready RAG applications and multimodal agents.

image-text-to-textapache-2.0
40 starsView details

gemma-4-E4B-it-ultra-uncensored-heretic

llmfan46
Model

The gemma-4-E4B-it-ultra-uncensored-heretic is a specialized any-to-any multimodal model designed for developers who require high-flexibility input processing. Built on the Gemma 4 architecture, this iteration is optimized for cross-modal tasks, allowing for seamless interaction across text, image, and audio inputs. Unlike standard text-only LLMs, this model is tailored for complex workflows where sensory data integration is critical, such as automated media analysis or multimodal reasoning engines. While the 'uncensored' designation suggests a reduction in restrictive safety filtering—making it a candidate for research into edge-case reasoning and unconstrained creative generation—developers should approach deployment with an understanding of its specific alignment profile. It is an ideal choice for integration into local pipelines where strict API-based guardrails might interfere with specialized domain tasks or raw data interpretation. For those working within the Apache-2.0 ecosystem, it offers a permissive foundation for both commercial and research-oriented fine-tuning.

any to anyapache-2.0
38 starsView details

gemma 4 26B A4B it

google
Model

Gemma 4 26B A4B is a multimodal model from Google, designed for developers needing high-performance image-to-text and text-to-text capabilities within an open-weights framework. Unlike smaller edge models, the 26B parameter scale allows for more nuanced reasoning and complex visual analysis while remaining deployable on consumer-grade hardware or private clouds. It is particularly effective for automated document parsing, visual QA, and augmenting RAG pipelines with image-based context. With an Apache-2.0 license, it offers significant flexibility for commercial integration, providing a competitive alternative to proprietary multimodal APIs by reducing latency and eliminating per-token costs for high-volume inference.

image-text-to-textapache-2.0
34 starsView details

Edge-4B-TELL

ginigen-ai
Model

Edge-4B-TELL is a compact, any-to-any multimodal model designed for high-efficiency edge deployment. Unlike standard text-only LLMs, this architecture handles diverse input-output modalities, making it a versatile candidate for local processing on resource-constrained hardware. With a 4-billion parameter footprint, it strikes a pragmatic balance between reasoning capabilities and low-latency execution. For developers building IoT solutions, mobile applications, or offline assistants, Edge-4B-TELL offers a way to implement complex multimodal workflows without relying on heavy cloud APIs. While it is currently in its early stages of community adoption, its Apache-2.0 license provides the legal flexibility required for commercial integration. If your roadmap involves on-device sensory processing or cross-modal interaction, this model serves as a lightweight foundation for testing low-overhead multimodal pipelines.

any to anyapache-2.0
34 starsView details

PhysBrain1.5-8B

DeepCybo
Model

PhysBrain1.5-8B is a compact, any-to-any multimodal model designed for cross-modal reasoning and processing. While many models specialize in text-to-text or vision-to-text, this 8B parameter architecture is built to handle diverse input-output modalities, making it a versatile candidate for complex multi-sensory tasks. For developers, the primary value lies in its ability to bridge different data types within a single inference pass, which is critical for robotics, physical simulation analysis, or advanced sensor fusion applications. Compared to much larger proprietary models, PhysBrain offers a more efficient footprint for edge deployment or specialized fine-tuning, allowing for lower latency in real-time environments. It is particularly suited for developers building agents that require a unified understanding of heterogeneous data streams rather than chaining separate specialized models together.

any to anySee model card
27 starsView details

Qwen3.6 27B NVFP4

unsloth
Model

Qwen3.6 27B NVFP4 is a high-efficiency multimodal model optimized for developers who need a balance between reasoning power and deployment speed. By utilizing NVFP4 quantization, this version significantly reduces memory overhead without sacrificing the core capabilities of the 27B parameter architecture, making it viable for consumer-grade GPUs or constrained cloud environments. It excels at image-to-text tasks, including complex visual reasoning, document parsing, and interleaved multimodal understanding. For developers, this means faster inference cycles and lower latency when integrating vision-language capabilities into RAG pipelines or automated content analysis tools. Compared to full-precision alternatives, it offers a streamlined path to production for high-throughput applications while maintaining the robust performance expected from the Qwen series.

image-text-to-textapache-2.0
24 starsView details

OTel-2.0-LLM-31B-IT

farbodtavakkoli
Not specified

OTel 2.0 LLM 31B IT is an instruction-tuned model designed for developers needing a balance between high-parameter reasoning and deployment efficiency. With 31 billion parameters, it sits in a sweet spot for complex text generation and logical synthesis tasks that typically overwhelm smaller 7B or 13B models, yet it remains manageable for mid-tier GPU clusters. It is particularly effective for automating documentation, synthesizing technical logs, and building RAG-based pipelines where precision and context adherence are critical. Integrated via standard Apache-2.0 licensing, it offers an open-weight alternative for teams avoiding proprietary lock-in while requiring a model capable of nuanced instruction following and structured output generation.

text generationapache-2.0
20 starsView details

CLIP ViT B 32 laion2B s34B b79K

laion
Model

CLIP ViT-B/32 (trained on LAION-2B) is a robust vision-language model designed for high-performance image-text alignment. By leveraging a Vision Transformer (ViT) backbone, it maps images and text into a shared embedding space, allowing developers to calculate cosine similarity for efficient retrieval and zero-shot classification. This specific variant is optimized for scale, making it ideal for building semantic search engines, automated image tagging systems, or as a visual encoder for larger multimodal architectures. Compared to smaller CLIP models, it offers a strong balance between inference latency and retrieval accuracy, integrating seamlessly into PyTorch and Hugging Face pipelines for rapid deployment in production environments.

image-text-retrievalmit

fashion clip

patrickjohncyh
Model

Fashion CLIP is a domain-specific adaptation of the CLIP architecture, fine-tuned specifically for the fashion industry to bridge the gap between visual imagery and textual descriptions. Unlike general-purpose vision-language models, this model is optimized for the nuances of apparel, recognizing specific textures, garment cuts, and style attributes that generic models often overlook. For developers, this makes it an ideal engine for building high-accuracy visual search tools, automated product tagging, or recommendation systems where precise image-text alignment is critical. It integrates seamlessly into existing PyTorch or Hugging Face pipelines, offering a drop-in replacement for standard CLIP embeddings when working with clothing datasets to significantly reduce retrieval noise and improve Mean Average Precision (mAP).

image-text-retrievalmit

CLIP convnext base w laion2B s13B b82K augreg

laion
Model

This model is a high-performance vision-language encoder based on the ConvNeXt architecture, trained on the massive LAION-2B dataset. Unlike traditional ViT-based CLIP models, it leverages a pure convolutional backbone to extract spatial features, often providing better inductive biases for image recognition tasks. It is specifically optimized for image-text retrieval, zero-shot classification, and semantic search. For developers, this means a robust tool for building visual search engines or content moderation systems where precise alignment between natural language queries and image embeddings is critical. Integration is straightforward via standard CLIP interfaces, offering a competitive alternative to Transformer-based encoders when deployment efficiency or specific spatial feature extraction is required.

image-text-retrievalmit

flan t5 base

google
Model

Flan-T5 Base is an instruction-tuned version of the original T5 encoder-decoder framework, designed for developers who need a lightweight yet versatile text-to-text model. Unlike standard T5, Flan-T5 is trained on a vast collection of tasks phrased as instructions, significantly improving its zero-shot performance across NLU and NLG benchmarks. It is particularly effective for constrained environments where latency and memory overhead are concerns, serving as a reliable baseline for text summarization, classification, and question answering. Because it follows a standard Seq2Seq architecture, it integrates seamlessly with the Hugging Face Transformers library, making it easy to fine-tune on domain-specific datasets without requiring massive compute clusters.

text2text-generationapache-2.0

chronos t5 base

amazon
Model

Chronos-T5 Base is a specialized time-series forecasting model that treats numerical sequences as language. By leveraging a T5-based encoder-decoder architecture, it reframes forecasting as a text-to-text problem, allowing it to perform zero-shot predictions on unseen datasets without requiring traditional retraining. For developers, this means a significant reduction in the cold-start problem for time-series analysis. It is particularly effective for forecasting trends across diverse domains where historical data is sparse or inconsistent. Integration is straightforward for those familiar with the Hugging Face ecosystem, offering a scalable alternative to traditional statistical models like ARIMA or Prophet by applying transformer-based attention to temporal patterns.

text2text-generationapache-2.0

chronos t5 tiny

amazon
Model

Chronos-T5 Tiny is a lightweight, text-to-text model based on the T5 architecture, optimized for efficiency and rapid deployment. Unlike general-purpose LLMs, this model is designed for specific sequence-to-sequence tasks where low latency and minimal compute overhead are critical. Developers can integrate it into edge environments or as a specialized microservice for tasks like text normalization, simple translation, or structured data extraction. By leveraging the T5 framework, it maintains a predictable tokenization process and stable performance, making it a viable alternative to larger models when the task complexity doesn't justify the memory footprint of a multi-billion parameter system.

text2text-generationapache-2.0

parrot paraphraser on T5

prithivida
Model

The Parrot Paraphraser is a specialized text-to-text model built on the T5 architecture, designed specifically for high-quality sentence rewriting. Unlike general-purpose LLMs that may drift from the original meaning, this model focuses on maintaining semantic equivalence while altering syntactic structure. For developers, it serves as a lightweight utility for data augmentation, avoiding repetitive phrasing in automated content pipelines, or preprocessing text for NLP training sets. It integrates easily into existing Hugging Face pipelines and provides a predictable, deterministic alternative to larger generative models when the goal is strictly paraphrasing rather than creative expansion.

text2text-generationApache-2.0

prot t5 xl uniref50

Rostlab
Model

ProtT5-XL-UniRef50 is a specialized encoder-decoder transformer trained on the UniRef50 protein database, designed specifically for protein sequence representation. Unlike general-purpose LLMs, this model treats amino acid sequences as a language, allowing developers to leverage its pre-trained weights for downstream biological tasks such as secondary structure prediction, protein-protein interaction analysis, and mutation effect estimation. It integrates easily into PyTorch or Hugging Face pipelines as a text2text-generation model, providing a robust foundation for those building bioinformatics tools who need a model that understands the evolutionary and structural context of proteins without training from scratch.

text2text-generationApache-2.0

chronos t5 small

amazon
Model

Chronos-T5 Small is a specialized time-series forecasting model based on the T5 architecture, treating numerical sequences as language tokens. Unlike traditional statistical models, it leverages a pretrained transformer backbone to perform zero-shot forecasting across diverse datasets without requiring extensive retraining. For developers, this means a streamlined pipeline for predicting trends and anomalies where historical data is available but labeled training sets are scarce. It integrates easily into Python-based ML workflows, offering a lightweight alternative for edge deployment or rapid prototyping of forecasting services compared to larger, computationally expensive models.

text2text-generationapache-2.0

unifiedqa t5 small

allenai
Model

UnifiedQA T5-Small is a lightweight text-to-text model fine-tuned for a broad spectrum of question-answering tasks. Unlike specialized QA models, it treats various formats—such as multiple-choice, extractive, and open-domain QA—as a single unified problem. For developers, this means a consistent interface for diverse retrieval tasks without needing task-specific architectures. Given its small parameter footprint, it is ideal for edge deployment, low-latency inference, or as a baseline for distillation. It integrates seamlessly with the Hugging Face Transformers library, making it easy to drop into existing Python pipelines for rapid prototyping or lightweight production services where compute resources are constrained.

text2text-generationApache-2.0
Email