Global AI chat room · 14 online now Join now
DIRECTORY / 02

AI Models | Open-Source LLM Directory

Discover and compare open-source LLMs, language models and multimodal models by capability, scale, license, downloads and provenance.

Compare modelsFind the right building block for your next workflow
Directory overview
839
curated entries
24 topic groupsLive
02 / MODEL INDEX

Find the right model for the job

Context first, better decisions. Every entry keeps the signal that matters.

CURATED DIRECTORY839 results

sqlcoder

Ollama
Model

SQLCoder is a specialized large language model engineered specifically for text-to-SQL tasks. Unlike general-purpose LLMs that often struggle with complex schema mapping or hallucinate non-existent table relationships, SQLCoder is fine-tuned to translate natural language queries into accurate, executable SQL code. For developers building data analytics interfaces or automated reporting tools, this model provides a high-degree of precision in structured query generation. It is designed for local inference via Ollama, making it an ideal choice for privacy-conscious environments where sending sensitive database schemas to external APIs is a security risk. While it lacks the broad conversational versatility of models like Llama 3, its performance on complex JOIN operations and dialect-specific syntax makes it a superior tool for database-centric workflows. Integration is straightforward for anyone already using the Ollama ecosystem, allowing for rapid prototyping of natural language interfaces over existing relational databases.

text generationSee Ollama library

llama4

Ollama
Model

Llama 4 represents the next evolution in Meta's open-weights ecosystem, optimized specifically for high-performance local inference via tools like Ollama. For developers, this model moves beyond simple chat interfaces, offering enhanced reasoning capabilities and improved instruction-following that make it viable for complex agentic workflows and autonomous coding assistants. Unlike massive cloud-hosted APIs, Llama 4 is designed to balance parameter efficiency with deep semantic understanding, allowing you to deploy sophisticated NLP pipelines on edge hardware or private infrastructure without data egress concerns. Whether you are fine-tuning for domain-specific tasks or building RAG (Retrieval-Augmented Generation) systems, Llama 4 provides a highly predictable latency profile and a robust architecture that integrates seamlessly into existing Python-based AI stacks. It positions itself as a direct, locally-controllable competitor to proprietary frontier models, prioritizing developer autonomy and architectural transparency.

text generationSee Ollama library

moondream

Ollama
Model

Moondream is a compact, high-efficiency vision-language model designed for developers who need to integrate image understanding into edge devices or local workflows without the heavy overhead of multi-billion parameter models. Unlike massive multimodal LLMs that require significant VRAM, moondream is optimized for speed and low latency, making it ideal for real-time applications like automated visual tagging, accessibility tools, or IoT sensor analysis. It excels at descriptive captioning and answering specific questions about visual inputs. For developers using Ollama, it offers a streamlined path to local inference, allowing you to build privacy-focused vision pipelines that run entirely on-device. While it may lack the deep reasoning depth of larger models, its performance-to-size ratio makes it a highly practical choice for specialized computer vision tasks where resource constraints are a primary concern.

text generationSee Ollama library

dolphin-mixtral

Ollama
Model

Dolphin-Mixtral is a high-performance fine-tuned variant of the Mixtral architecture, specifically optimized for instruction following and complex reasoning through the Dolphin dataset. For developers working on local-first applications, this model offers a significant leap in conversational fluidity and task adherence compared to base MoE (Mixture-of-Experts) models. Unlike standard commercial APIs, Dolphin-Mixtral is designed for uncensored, versatile deployment, making it ideal for specialized agentic workflows, creative coding assistants, and complex data extraction where strict guardrails might otherwise impede logical output. It integrates seamlessly into local stacks via Ollama, allowing for private, low-latency inference on consumer-grade hardware. While it requires careful hardware allocation due to its sparse MoE structure, the trade-off is a model that excels at nuanced multi-turn dialogues and sophisticated logical deduction without the overhead of massive dense models.

text generationSee Ollama library

qwen3-coder-next

Ollama
Model

qwen3-coder-next is a specialized iteration in the Qwen series designed specifically for high-performance programming tasks. For developers looking to integrate advanced reasoning into their local workflows, this model offers a robust alternative to cloud-dependent APIs. It excels in code completion, complex debugging, and translating logic across multiple programming languages. Unlike general-purpose LLMs, its architecture is optimized for structural syntax accuracy and algorithmic reasoning, making it a reliable engine for IDE extensions or automated CI/CD linting tools. Since it is available via Ollama, you can deploy it locally to ensure data privacy and minimize latency during development cycles. While specific parameter counts vary by version, the model's primary strength lies in its ability to maintain context within large codebases and follow strict architectural patterns. It is a pragmatic choice for teams building local-first developer tools or requiring high-speed, offline code intelligence.

text generationSee Ollama library

cogito

Ollama
Model

Cogito is a text-generation model optimized for local inference via the Ollama framework. Unlike massive cloud-hosted LLMs, Cogito is designed for developers who prioritize data privacy, low latency, and offline availability. While specific parameter counts vary depending on the version pulled, the model's architecture is tuned for efficient resource utilization on consumer-grade hardware. It serves as a versatile backbone for building local RAG (Retrieval-Augmented Generation) pipelines, automated coding assistants, or private chat interfaces where sending sensitive data to external APIs is not an option. For integration, it leverages the standard Ollama API, making it trivial to swap into existing workflows using RESTful calls or local SDKs. If you are looking to prototype or deploy edge-based AI without the overhead of massive infrastructure, Cogito provides a reliable, containerized starting point.

text generationSee Ollama library

dolphin-llama3

Ollama
Model

Dolphin-Llama3 is an uncensored fine-tune of Meta's Llama 3 architecture, specifically optimized through the Dolphin dataset to prioritize instruction-following without the typical refusal behaviors found in base models. For developers, this means a significant reduction in 'as an AI language model' refusals, making it highly effective for complex reasoning, creative writing, and roleplay scenarios where strict safety alignment often interferes with the intended logic. While it inherits the high-performance reasoning and linguistic capabilities of the Llama 3 backbone, the Dolphin tuning shifts the model's persona toward a more compliant and versatile assistant. It is designed for local deployment via Ollama, making it an ideal candidate for privacy-sensitive applications, local RAG pipelines, or edge computing environments where data sovereignty is a priority. Unlike proprietary APIs, you have full control over the system prompt and temperature settings to shape its behavior without external moderation layers.

text generationSee Ollama library

embeddinggemma

Ollama
Model

embeddinggemma is a specialized embedding model based on the Gemma architecture, optimized for converting unstructured text into high-dimensional vector representations. Unlike generative models designed for chat, this model is purpose-built for downstream semantic tasks. For developers building RAG (Retrieval-Augmented Generation) pipelines, semantic search engines, or clustering workflows, embeddinggemma provides a locally hostable solution via Ollama, ensuring data privacy and low latency. It excels at capturing nuanced contextual relationships within text, making it a strong candidate for improving document retrieval accuracy in vector databases. Because it runs locally, you can integrate it into your existing microservices architecture without relying on external API calls, significantly reducing operational costs and infrastructure complexity. It serves as a robust middle-layer component for any application requiring sophisticated natural language understanding through vector embeddings.

text generationSee Ollama library

smollm

Ollama
Model

SmolLM is a family of lightweight, high-performance language models designed specifically for local execution and edge computing. Unlike massive frontier models that require high-end data center GPUs, SmolLM is optimized to run efficiently on consumer-grade hardware, including laptops and mobile devices, via tools like Ollama. For developers, this means significantly lower latency and reduced infrastructure costs when implementing text generation tasks. While it lacks the broad world knowledge of multi-billion parameter models, it excels in specific, constrained environments such as code completion, structured data extraction, and local chat interfaces where privacy and offline availability are non-negotiable. It is an ideal candidate for developers building agentic workflows that require high-frequency, low-cost reasoning cycles without the overhead of cloud API calls.

text generationSee Ollama library

qwen3.8

Ollama
Model

qwen3.8 is a versatile text-generation model optimized for local inference via the Ollama ecosystem. Designed for developers who prioritize data privacy and low-latency execution, this model allows you to run sophisticated LLM workflows entirely on your own hardware without relying on external APIs. While specific parameter counts vary by quantization level, the architecture is engineered to balance reasoning capabilities with computational efficiency. It is particularly effective for building local RAG (Retrieval-Augmented Generation) pipelines, automating code documentation, and powering edge-based chat interfaces. For integration, its availability in the Ollama library means you can deploy it with a single command, making it an ideal candidate for testing prompt engineering or prototyping agentic workflows in a sandboxed environment. Unlike massive cloud-hosted models, qwen3.8 offers a predictable cost structure and total control over your inference stack.

text generationSee Ollama library

gemma3n

Ollama
Model

Gemma3n is a specialized iteration of the Gemma family designed for high-performance local inference via the Ollama ecosystem. For developers building privacy-first applications or edge computing solutions, this model provides a streamlined path to deploying sophisticated text generation capabilities without relying on external APIs. While specific parameter counts vary depending on the quantized version pulled, the architecture is optimized for low-latency reasoning and instruction following. Unlike massive cloud-hosted models, Gemma3n is built to balance computational efficiency with linguistic nuance, making it ideal for local RAG (Retrieval-Augmented Generation) pipelines, automated coding assistants, and private data summarization. Integration is straightforward for anyone already utilizing the Ollama runtime, allowing for rapid prototyping of agentic workflows. It serves as a robust middle-ground for those needing more intelligence than tiny SLMs but requiring significantly less VRAM than flagship frontier models.

text generationSee Ollama library

qwq

Ollama
Model

qwq is a specialized text generation model optimized for local deployment via the Ollama ecosystem. Designed for developers who prioritize data privacy and low-latency inference, it allows for high-performance language tasks without relying on external cloud APIs. While specific parameter counts vary depending on the quantized version pulled, the model is built to handle complex reasoning, structured data extraction, and conversational logic. For integration, qwq follows the standard Ollama API patterns, making it a drop-in replacement for existing LLM workflows in local development environments. Compared to massive proprietary models, qwq offers a more efficient footprint, making it ideal for edge computing, automated testing pipelines, and private RAG (Retrieval-Augmented Generation) implementations where keeping data on-premise is a non-negotiable requirement.

text generationSee Ollama library

llava-llama3

Ollama
Model

llava-llama3 represents a significant step forward in local multimodal reasoning by marrying the LLaVA vision-language architecture with the Llama 3 backbone. For developers building privacy-first applications, this model enables seamless processing of both text and visual inputs directly on local hardware via Ollama. Unlike standard text-only LLMs, llava-llama3 can perform complex visual reasoning, such as describing images, extracting text from documents, or identifying objects within a scene. It is particularly useful for edge computing scenarios where cloud latency or data privacy concerns prohibit the use of proprietary APIs. While performance scales with your VRAM, the integration via Ollama makes it trivial to deploy into existing Python or JavaScript workflows. Compared to earlier LLaVA iterations, the Llama 3 integration offers improved instruction following and more coherent linguistic output, making it a robust choice for developers prototyping multimodal agents or automated visual inspection tools.

text generationSee Ollama library

glm-5.1

Ollama
Model

GLM-5.1 is a versatile text generation model optimized for local inference via the Ollama ecosystem. For developers building privacy-first applications or edge-based AI solutions, this model offers a streamlined path to deploying high-performance language capabilities without relying on external APIs. While specific parameter counts vary by quantized version, the architecture is designed to balance reasoning depth with computational efficiency. It excels in standard NLP tasks such as code completion, structured data extraction, and conversational logic. Unlike massive cloud-hosted models, GLM-5.1 is engineered for developers who need low-latency responses and full control over their data pipeline. Integration is straightforward through the Ollama CLI and API, making it an ideal candidate for local RAG (Retrieval-Augmented Generation) workflows and automated content pipelines where data sovereignty is a primary requirement.

text generationSee Ollama library

minimax-m2.7

Ollama
Model

Minimax-m2.7 is a specialized text generation model optimized for local deployment via the Ollama framework. For developers building privacy-centric applications or latency-sensitive workflows, this model provides a streamlined path to running high-quality inference on local hardware without relying on external APIs. While the exact parameter count is abstracted in the current library listing, its architecture is tuned for efficient reasoning and structured text output. It is particularly well-suited for developers working on RAG (Retrieval-Augmented Generation) pipelines, local coding assistants, or automated content processing where data sovereignty is a priority. Unlike massive cloud-based LLMs, m2.7 focuses on a balanced performance-to-compute ratio, making it a versatile choice for edge computing environments or local development testing before scaling to larger production clusters.

text generationSee Ollama library

translategemma

Ollama
Model

Translategemma is a specialized fine-tuned model designed specifically for high-fidelity machine translation tasks. Unlike general-purpose LLMs that may struggle with linguistic nuances or suffer from 'translationese,' this model is optimized to preserve semantic intent and stylistic consistency across language pairs. For developers building localization pipelines, real-time chat interfaces, or multilingual content engines, translategemma offers a streamlined alternative to massive, multi-modal models. It is specifically architected for local inference via Ollama, making it an ideal candidate for privacy-sensitive applications where data cannot leave the local environment. While general models excel at reasoning, translategemma prioritizes lexical accuracy and grammatical fluency, providing a more predictable output for structured translation workflows. Integration is straightforward for anyone already using the Ollama ecosystem, allowing for rapid prototyping of multilingual features without the latency or cost overhead of external APIs.

text generationSee Ollama library

mistral-small3.2

Ollama
Model

Mistral-Small-3.2 is a precision-engineered model designed for developers who need a balance between high-level reasoning and low-latency execution. Unlike massive frontier models that demand heavy compute, this iteration focuses on optimizing the parameter-to-performance ratio, making it an ideal candidate for local deployment via Ollama. It excels in structured data extraction, complex instruction following, and code reasoning tasks where speed is just as critical as accuracy. For teams building agentic workflows or RAG pipelines, Mistral-Small offers a predictable, efficient middle ground that minimizes inference costs without sacrificing the nuance required for sophisticated text generation. It is particularly well-suited for edge computing or local development environments where hardware constraints are a factor, providing a robust alternative to larger, more cumbersome models.

text generationSee Ollama library

falcon3

Ollama
Model

Falcon3 represents the next evolution in the Falcon series, optimized specifically for high-performance local inference via Ollama. For developers building privacy-first or edge-based applications, this model offers a robust alternative to closed-source APIs by providing low-latency text generation directly on your hardware. Unlike general-purpose monolithic models, Falcon3 is engineered to balance parameter efficiency with reasoning capabilities, making it ideal for RAG (Retrieval-Augmented Generation) pipelines, local code assistance, and structured data extraction. Integration is seamless for those already using the Ollama ecosystem, allowing you to transition from prototyping to local deployment with minimal configuration changes. While specific parameter counts vary by quantized version, the architecture focuses on high throughput and reduced memory overhead, ensuring it remains accessible for developers working on consumer-grade GPUs or high-end workstations.

text generationSee Ollama library

llama2-uncensored

Ollama
Model

For developers working with local LLMs, llama2-uncensored represents a specialized fine-tune of the Llama 2 architecture designed to strip away the standard safety refusals and alignment constraints. While the base Llama 2 model is highly capable, its strict instruction-following can sometimes trigger false positives in safety filters, hindering creative writing, complex roleplay, or unfiltered data analysis. This version is optimized for high-fidelity instruction following without the 'as an AI language model' caveats. It is ideal for building autonomous agents that require uninhibited reasoning or for developers experimenting with edge-case scenarios where standard censorship would break the logic flow. Since it is available via Ollama, integration into local workflows is seamless via a simple API, making it a practical choice for privacy-conscious developers who need raw, unmoderated text generation on their own hardware.

text generationSee Ollama library

mixtral

Ollama
Model

Mixtral is a high-performance Mixture-of-Experts (MoE) model designed to bridge the gap between massive dense models and lightweight local deployment. For developers, the primary value lies in its efficiency: it utilizes a sparse architecture that activates only a fraction of its total parameters during inference, delivering GPT-3.5-level reasoning speeds with significantly lower computational overhead. This makes it an ideal candidate for local integration via Ollama, where latency and hardware constraints are critical. You can leverage Mixtral for complex tasks like multi-turn dialogue, sophisticated code generation, and high-context retrieval-augmented generation (RAG) pipelines. Unlike standard dense models, Mixtral offers a superior performance-to-compute ratio, allowing you to run a highly capable reasoning engine on consumer-grade hardware without the massive VRAM requirements typically associated with high-parameter models. It is an excellent choice for building privacy-focused, offline-capable AI agents.

text generationSee Ollama library

nemotron-3-super

Ollama
Model

Nemotron-3-super is a specialized text generation model optimized for local deployment via Ollama. While many large-scale models require heavy cloud infrastructure, this iteration is designed to balance high-reasoning capabilities with the efficiency required for edge computing and local development environments. For developers, the primary value lies in its ability to handle complex instruction-following and structured data generation tasks without the latency or privacy concerns of external APIs. It serves as a robust backbone for building local RAG (Retrieval-Augmented Generation) pipelines, autonomous agents, and automated coding assistants. When integrating, you should prioritize testing its specific parameter scaling against your hardware constraints, as performance varies significantly based on your local quantization settings. Compared to standard general-purpose models, Nemotron-3-super focuses on high-fidelity output suitable for production-grade logic tasks.

text generationSee Ollama library

starcoder2

Ollama
Model

StarCoder2 is a high-performance code intelligence model family designed specifically for the software development lifecycle. Unlike general-purpose LLMs that treat code as just another language, StarCoder2 is optimized for deep syntax understanding, long-context repository comprehension, and low-latency completions. For developers building IDE extensions, automated code review tools, or local copilot clones, this model offers a significant leap in accuracy across dozens of programming languages. It is particularly effective for tasks ranging from simple autocomplete to complex refactoring and unit test generation. Because it is available for local inference via Ollama, you can integrate it directly into your private workflows without the latency or privacy concerns of cloud-based APIs. Compared to earlier iterations, StarCoder2 demonstrates superior reasoning in polyglot environments, making it a robust choice for modern, multi-language microservices architectures.

text generationSee Ollama library

granite3.1-moe

Ollama
Model

Granite3.1-MoE is a Mixture-of-Experts model designed for efficient, high-performance text generation. For developers working in resource-constrained environments or seeking low-latency local inference, the MoE architecture provides a strategic advantage by activating only a subset of parameters per token. This allows for sophisticated reasoning and complex instruction following without the massive compute overhead of dense models of similar capacity. It is particularly well-suited for RAG (Retrieval-Augmented Generation) pipelines, code assistance, and automated data extraction where throughput is critical. Available via Ollama, it integrates seamlessly into existing local workflows, making it a practical choice for privacy-conscious applications or edge computing scenarios. Compared to standard dense models, Granite3.1-MoE offers a superior balance of intelligence-per-watt, enabling more complex logic on consumer-grade hardware.

text generationSee Ollama library

orca-mini

Ollama
Model

Orca-mini is a lightweight, distilled language model designed specifically for efficient local inference. For developers working with resource-constrained environments—such as edge devices, mobile hardware, or local development machines without high-end GPUs—this model offers a pragmatic alternative to massive parameter models. While it lacks the broad reasoning depth of its larger counterparts, orca-mini excels at structured text generation and basic instruction following tasks. It is optimized for low-latency responses, making it an excellent candidate for prototyping conversational interfaces, local RAG (Retrieval-Augmented Generation) pipelines, or basic text processing workflows. Since it is available via the Ollama library, integration into your existing local stack is seamless, allowing you to test agentic workflows or specialized fine-tuning ideas without the overhead of cloud-based API costs or privacy concerns.

text generationSee Ollama library
Email