Global AI chat room · 17 online now Join now
DIRECTORY / 02

AI Models | Open-Source LLM Directory

Discover and compare open-source LLMs, language models and multimodal models by capability, scale, license, downloads and provenance.

Compare modelsFind the right building block for your next workflow
Directory overview
839
curated entries
24 topic groupsLive
02 / MODEL INDEX

Find the right model for the job

Context first, better decisions. Every entry keeps the signal that matters.

CURATED DIRECTORY839 results

Whisper Small

OpenAI
244M

Whisper Small is a streamlined version of OpenAI’s robust speech-to-text architecture, optimized specifically for developers targeting edge computing and low-latency environments. While larger iterations of Whisper prioritize absolute accuracy at the cost of massive VRAM requirements, the Small model strikes a pragmatic balance by utilizing 244M parameters. This makes it viable for deployment on consumer-grade hardware or mobile-adjacent devices without sacrificing significant word error rate (WER) performance. For engineers building real-time transcription services, voice assistants, or automated captioning tools, this model offers a high throughput-to-accuracy ratio. It integrates seamlessly into existing Python-based ML pipelines and supports a wide array of multilingual tasks. If your use case requires local inference where cloud API latency or data privacy is a concern, Whisper Small serves as an efficient middle ground between the lightweight 'Base' model and the heavy 'Large' variants.

automatic speech recognitionMIT
3.8K starsView details

Flan-T5 Large

Google
780M

...

text2text generationApache 2.0
3.8K starsView details

GLM-4 9B Chat

Tsinghua
9B

GLM-4 9B Chat is a lightweight, high-performance bilingual model optimized for seamless Chinese-English reasoning and instruction following. Built by the Zhipu AI team from Tsinghua University, this 9B parameter model strikes a strategic balance between computational efficiency and cognitive depth, making it an ideal candidate for edge deployment or high-throughput microservices where latency is critical. Unlike larger monolithic models, the 9B architecture is designed for developers who need robust multilingual capabilities—specifically handling nuanced semantic shifts between English and Chinese—without the massive infrastructure overhead. It excels in structured data extraction, code assistance, and conversational agent workflows. For integration, it follows standard API patterns, allowing for easy replacement of larger LLMs in specialized pipelines where a smaller, faster footprint is required to maintain low cost-per-token while preserving logical coherence.

text generationGLM-4
3.8K starsView details

Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive

HauhauCS
Model

For developers building multimodal applications, Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive offers a specialized approach to vision-language tasks. Based on the Qwen architecture, this 35B parameter model is fine-tuned for high-fidelity image-to-text reasoning and complex instruction following. Unlike standard restricted models, this version is optimized for unfiltered responses, making it a practical choice for researchers and developers working on edge cases, creative writing, or datasets where safety alignment might otherwise suppress nuanced or raw information. It excels in visual reasoning, OCR, and descriptive captioning. Integration is straightforward via Hugging Face, supporting standard multimodal pipelines. While the 'Aggressive' tuning increases response volatility, it provides a significant advantage for developers needing high-entropy outputs that aren't constrained by heavy-handed RLHF filters, allowing for more direct and precise alignment with complex user prompts.

image text to textapache-2.0
3.8K starsView details

speaker-diarization-3.1

pyannote
Not specified

Speaker Diarization 3.1, powered by pyannote, is a specialized framework designed to solve the 'who spoke when' problem in audio processing. Unlike standard ASR which only provides text, this model identifies distinct speaker identities and maps their timestamps across a recording. For developers, this is critical for building automated meeting minutes, multi-party interview transcripts, or voice-activated analytics. It integrates efficiently into speech pipelines, acting as a pre-processing or parallel layer to transcription engines. Compared to basic clustering methods, version 3.1 offers improved precision in speaker change detection and overlap handling, making it a robust choice for noisy, real-world audio environments where speaker turns are rapid.

automatic speech recognitionmit
3.7K starsView details

Janus-Pro-7B

deepseek-ai
Model

Janus-Pro-7B is a versatile 'any-to-any' multimodal model from the DeepSeek team, designed to bridge the gap between text and visual reasoning. Unlike traditional models that treat vision as a secondary input, Janus-Pro is built for seamless cross-modal generation and understanding. For developers, this means you can move beyond simple image captioning into complex tasks like high-fidelity image synthesis, visual document parsing, and sophisticated spatial reasoning within a single 7B parameter framework. Its architecture is optimized for efficiency, making it a strong candidate for edge deployment or integration into agentic workflows where both visual perception and creative output are required. While larger models offer brute-force reasoning, Janus-Pro provides a highly competitive performance-to-latency ratio, making it ideal for real-time applications like interactive UI assistants or automated visual content pipelines. It is released under the MIT license, ensuring high flexibility for commercial integration.

any to anymit
3.7K starsView details

e5-large-v2

intfloat
Not specified

The e5-large-v2 is a high-performance text embedding model designed for dense vector representation. Unlike generative LLMs, this model focuses on mapping text into a continuous vector space where semantic similarity is measured via cosine similarity. It is particularly effective for Retrieval-Augmented Generation (RAG) pipelines, semantic search, and clustering tasks. Built on a transformer architecture and optimized for sentence-level embeddings, it offers a strong balance between dimensionality and retrieval accuracy. For developers, it serves as a lightweight, MIT-licensed alternative to proprietary embedding APIs, allowing for local deployment and full control over data privacy without sacrificing significant mAP (mean Average Precision) in information retrieval benchmarks.

sentence similaritymit
3.6K starsView details

Yi-1.5 34B

01.AI
34B

Yi-1.5 34B is a high-performance bilingual model from 01.AI designed to bridge the gap between English and Chinese language processing. For developers, the 34B parameter scale offers a strategic sweet spot: it provides significantly more reasoning depth and nuance than smaller 7B models while maintaining much lower latency and inference costs than massive 70B+ architectures. Unlike many models that treat non-English languages as an afterthought, Yi-1.5 is optimized for high-fidelity cross-lingual tasks, making it ideal for localization pipelines, multilingual chatbots, and complex document summarization involving both languages. Released under the Apache 2.0 license, it is highly accessible for commercial integration. Whether you are fine-tuning for specific domain knowledge or deploying via standard inference engines, this model serves as a robust mid-sized backbone for applications requiring sophisticated linguistic intelligence without the overhead of flagship-scale deployments.

text generationApache 2.0
3.6K starsView details

Edge0-35B-A3B-preview

Edge0
Model

Edge0-35B-A3B-preview is a specialized text-generation model designed for developers seeking a balance between high-performance reasoning and efficient deployment. Built on a 35B parameter architecture, it utilizes an activation-efficient design (A3B) that optimizes throughput without sacrificing the nuance required for complex instruction following. For developers working on edge computing or resource-constrained environments, this model offers a compelling middle ground between lightweight small language models and massive, high-latency frontier models. It is particularly well-suited for RAG pipelines, automated code documentation, and structured data extraction tasks. Since it is released under the Apache-2.0 license, it provides the legal flexibility necessary for commercial integration and fine-tuning. Compared to standard dense models of similar size, the preview architecture aims to deliver lower inference latency, making it a strong candidate for real-time application backends where response speed is as critical as linguistic accuracy.

text generationapache-2.0
3.6K starsView details

phi-2

microsoft
Model

Phi-2 is Microsoft's lightweight, high-performance small language model (SLM) designed to punch significantly above its weight class. While massive LLMs dominate general benchmarks, Phi-2 focuses on high-quality reasoning and logic within a compact parameter footprint. For developers, this means you can deploy sophisticated text generation, code completion, and logical reasoning tasks on edge devices or local environments without the massive VRAM overhead required by larger models. It excels in structured tasks and follows instructions with surprising precision for its size. Unlike general-purpose giants, Phi-2 is optimized for efficiency, making it an ideal backbone for local RAG pipelines, mobile integration, or specialized microservices where latency and compute costs are critical constraints. If your workflow requires high-density intelligence rather than raw scale, Phi-2 offers a highly competitive alternative to much larger proprietary APIs.

text generationmit
3.5K starsView details

OpenJourney v4

PromptHero
860M

OpenJourney v4 is a specialized fine-tuned Stable Diffusion model engineered to replicate the high-fidelity aesthetic and stylistic nuance typically associated with Midjourney. For developers building generative art pipelines or creative tooling, this model bridges the gap between open-source flexibility and premium commercial output. With 860M parameters, it is optimized for high-quality text-to-image synthesis, focusing on complex lighting, texture rendering, and compositional depth. Unlike base SD models that often require extensive prompt engineering to achieve professional results, OpenJourney v4 is tuned to interpret descriptive natural language more effectively. It integrates seamlessly into existing Diffusers workflows and ComfyUI environments, making it a practical choice for developers looking to implement sophisticated image generation without the restrictive API constraints or costs of proprietary closed-source models.

text to imageCreativeML Open RAIL++-M
3.5K starsView details

gemma-7b

google
Model

Gemma-7b is a lightweight, open-weight transformer model from Google, engineered to deliver high-performance text generation within a manageable parameter footprint. For developers, the primary value lies in its efficiency; it provides a sophisticated balance between reasoning capabilities and low-latency inference, making it ideal for deployment on edge devices or consumer-grade GPUs. Unlike massive frontier models that require extensive cloud infrastructure, Gemma-7b is optimized for local integration and fine-tuning on domain-specific datasets. It excels in tasks such as code completion, summarization, and structured data extraction. While it lacks the raw breadth of much larger models, its architecture is highly responsive to instruction tuning, allowing developers to bake specific logic or stylistic constraints directly into their applications. If you are looking to build private, cost-effective AI agents or specialized NLP pipelines without the overhead of massive API costs, Gemma-7b serves as a robust foundation.

text generationgemma
3.4K starsView details

bge large en v1.5

BAAI
Model

The bge-large-en-v1.5 is a high-performance embedding model designed for transforming text into dense vectors for retrieval-augmented generation (RAG) and semantic search. Unlike generative LLMs, this model specializes in feature extraction, mapping queries and documents into a shared latent space where cosine similarity correlates strongly with semantic relevance. It is specifically optimized to resolve common retrieval pitfalls found in earlier versions, offering improved stability and ranking accuracy. For developers, it serves as a lightweight, MIT-licensed alternative to proprietary embedding APIs, making it ideal for local deployment in vector databases like Milvus, Pinecone, or FAISS to power efficient knowledge bases and document clustering pipelines.

feature-extractionmit
3.4K starsView details

Bark

Suno
1.2B

Bark is a transformer-based text-to-audio model designed to move beyond simple speech synthesis by generating highly realistic acoustic environments. Unlike standard TTS engines that focus solely on phoneme accuracy, Bark utilizes a language modeling approach to produce non-verbal cues such as laughter, sighs, and hesitation, alongside background noise and music. For developers, the 1.2B parameter architecture offers a significant leap in expressive prosody, making it ideal for immersive storytelling, game NPC dialogue, and automated content creation. While it requires more compute than lightweight TTS models, its ability to handle multi-lingual inputs and complex audio textures makes it a versatile tool for high-fidelity audio pipelines. Integrating Bark into your stack allows for the generation of nuanced, human-like audio from raw text prompts, bridging the gap between robotic speech and organic soundscapes.

text to audioMIT
3.4K starsView details

DeepSeek-OCR

deepseek-ai
Model

DeepSeek OCR is a specialized vision-language model designed to bridge the gap between raw image pixels and structured text. Unlike traditional OCR engines that rely on rigid layout analysis, this model leverages deep learning to handle complex documents, handwritten notes, and non-standard formatting with high fidelity. For developers, this means fewer pre-processing steps and better accuracy on noisy data. It is particularly effective for automating data extraction from invoices, digitizing legacy archives, and building accessible interfaces for visual content. Integration is streamlined via a standard API, allowing it to fit easily into existing RAG pipelines or document processing workflows where precise text recovery is critical.

image text to textmit
3.4K starsView details

Mistral-7B-Instruct-v0.2

mistralai
Model

Mistral-7B-Instruct-v0.2 represents a significant milestone in high-performance, small-scale language modeling. For developers, the primary draw is its ability to punch far above its weight class, delivering reasoning and instruction-following capabilities that rival much larger proprietary models. This version improves upon its predecessor with enhanced stability and a refined vocabulary, making it ideal for low-latency edge deployment or fine-tuning on domain-specific datasets. Because it is released under the Apache 2.0 license, it offers the flexibility required for commercial integration without the restrictive overhead of closed-source APIs. Whether you are building sophisticated RAG pipelines, automating complex chat interfaces, or optimizing local inference on consumer-grade hardware, this model provides a highly efficient foundation that balances computational cost with sophisticated linguistic intelligence.

text generationapache-2.0
3.2K starsView details

BART Large CNN

Facebook
406M

BART Large CNN is a sequence-to-sequence transformer model specifically fine-tuned on the CNN/Daily Mail dataset for abstractive summarization. Unlike general-purpose LLMs, it is optimized to condense long-form documents into concise, coherent summaries while maintaining factual consistency. With 406M parameters, it offers a lightweight footprint compared to modern frontier models, making it highly efficient for production environments where latency and cost are critical. Developers can easily integrate it via the Hugging Face Transformers library for automated news aggregation, document synthesis, or internal knowledge base pruning. It excels in tasks requiring structured summaries rather than open-ended creative generation.

summarizationApache 2.0
3.2K starsView details

DeepSeek-V3-0324

deepseek-ai
Model

DeepSeek-V3-0324 is a high-performance text generation model engineered for complex reasoning and large-scale language tasks. For developers integrating LLMs into production pipelines, this model offers a competitive alternative to proprietary closed-source APIs by providing robust instruction-following capabilities and efficient inference potential. It excels in coding assistance, mathematical reasoning, and structured data extraction, making it a versatile choice for building autonomous agents or sophisticated RAG (Retrieval-Augmented Generation) systems. Unlike many models that require heavy fine-tuning for specialized logic, V3 demonstrates strong zero-shot performance across technical domains. Its MIT license simplifies deployment in commercial environments, allowing for deep integration into local infrastructure or cloud-based microservices without the restrictive overhead of proprietary ecosystems. Whether you are optimizing for latency in a chat application or accuracy in a complex reasoning engine, this model provides a scalable foundation for high-throughput text processing.

text generationmit
3.2K starsView details

CodeQwen 7B

Alibaba
7B

CodeQwen 7B is a specialized transformer model optimized specifically for code intelligence, bridging the gap between high-performance reasoning and efficient deployment. Developed by Alibaba, this 7B parameter model excels in bilingual contexts, making it particularly effective for developers working across Chinese and English documentation or codebases. Unlike general-purpose LLMs that often struggle with syntax precision in smaller footprints, CodeQwen is fine-tuned for code generation, completion, and logic debugging. For engineering teams, the primary value lies in its balance of performance and resource requirements; it is lightweight enough to run on local workstations or edge devices while maintaining competitive benchmarks in common programming languages. Its Apache 2.0 license ensures seamless integration into commercial workflows without the friction of restrictive proprietary terms. Whether you are building an IDE extension, an automated code review agent, or a local copilot, CodeQwen provides a robust, low-latency foundation for production-grade developer tools.

code generationApache 2.0
3.1K starsView details

LocateAnything-3B

nvidia
Model

LocateAnything-3B is a specialized vision-language model from NVIDIA designed specifically for high-precision spatial grounding. Unlike general-purpose multimodal models that provide broad descriptions, this 3B-parameter architecture focuses on the 'where' as much as the 'what.' It excels at mapping natural language queries to specific bounding boxes within an image, making it a critical tool for developers building object detection pipelines, visual search engines, or automated robotic vision systems. For engineers looking to integrate grounding capabilities without the massive computational overhead of larger models, its 3B footprint offers a highly efficient middle ground between lightweight detectors and heavy LLMs. It is particularly useful for zero-shot object localization tasks where pre-defined labels are insufficient and you need to detect arbitrary objects via text prompts. Integration is straightforward via Hugging Face, making it a plug-and-play option for RAG-based vision workflows or complex scene understanding applications.

image text to textother
3.1K starsView details

Llama-3.3-70B-Instruct

meta-llama
Model

Llama-3.3-70B-Instruct represents a significant efficiency milestone for developers working with mid-sized parameter models. While it maintains the 70B footprint, it delivers performance parity with much larger frontier models, making it a sweet spot for high-throughput applications. For engineers, this means you get near-GPT-4 level reasoning and instruction following without the massive latency or infrastructure costs associated with 400B+ parameter architectures. It excels in complex reasoning, coding assistance, and structured data extraction. Because it adheres to the Llama-3 ecosystem, integration is seamless via standard libraries like Transformers, vLLM, or Ollama. Whether you are deploying on-premises to maintain data sovereignty or scaling via managed APIs, this model offers a highly optimized balance of intelligence-per-watt, making it ideal for production-grade RAG pipelines and autonomous agent workflows.

text generationllama3.3
3.0K starsView details

Kimi-K2-Instruct

moonshotai
Model

Kimi-K2-Instruct is the latest instruction-tuned iteration from Moonshot AI, specifically engineered to handle complex reasoning and long-context instruction following. For developers working in multilingual environments, particularly those requiring high proficiency in Chinese and English, this model offers a robust alternative to mainstream Western LLMs. Unlike general-purpose chat models, K2-Instruct is optimized for structured output and logical consistency, making it a strong candidate for agentic workflows, automated coding assistance, and sophisticated data extraction tasks. While specific parameter counts remain proprietary, the model's performance profile suggests a focus on high-density reasoning rather than mere pattern matching. Integration is straightforward via Hugging Face, allowing for seamless deployment within existing inference pipelines. If your roadmap involves building RAG systems or autonomous agents that require nuanced command adherence, Kimi-K2-Instruct provides a competitive edge in logic-heavy applications.

text generationother
3.0K starsView details

starcoder

bigcode
Model

StarCoder is a specialized large language model engineered specifically for code intelligence and programming tasks. Unlike general-purpose LLMs that struggle with long-range syntax dependencies, StarCoder is trained on a massive, diverse corpus of source code, making it highly proficient in autocompletion, code explanation, and multi-language translation. For developers, the primary value lies in its ability to integrate directly into IDE workflows via LSP (Language Server Protocol) or custom plugins. Whether you are building a local co-pilot, automating unit test generation, or implementing complex refactoring tools, StarCoder provides a robust foundation. It is designed to be lightweight enough for efficient deployment while maintaining high accuracy across dozens of programming languages. Compared to monolithic proprietary models, StarCoder offers a more transparent, open-weights alternative that allows for fine-tuning on private repositories, ensuring your codebase's specific patterns and internal APIs are respected during inference.

text generationbigcode-openrail-m
3.0K starsView details

QwQ-32B

Qwen
Model

QwQ-32B is a specialized reasoning model from the Qwen ecosystem designed to bridge the gap between standard LLMs and high-compute reasoning agents. Unlike general-purpose chat models, QwQ is optimized for complex logical workflows, including multi-step mathematical problem solving, advanced code generation, and intricate symbolic reasoning. For developers, this means a significant step up in accuracy for tasks that typically trigger 'hallucinations' in smaller models. At 32B parameters, it offers a sweet spot for deployment: it provides sophisticated chain-of-thought capabilities that rival much larger models while remaining efficient enough to run on consumer-grade or mid-range enterprise hardware. It is particularly useful for building autonomous agents, automated debugging tools, or complex data extraction pipelines where logical consistency is more critical than creative prose. The Apache-2.0 license makes it highly accessible for commercial integration and fine-tuning within existing RAG or agentic frameworks.

text generationapache-2.0
3.0K starsView details
Email