Global AI chat room · 12 online now Join now
DIRECTORY / 02

AI Models | Open-Source LLM Directory

Discover and compare open-source LLMs, language models and multimodal models by capability, scale, license, downloads and provenance.

Compare modelsFind the right building block for your next workflow
Directory overview
591
curated entries
24 topic groupsLive
02 / MODEL INDEX

Find the right model for the job

Context first, better decisions. Every entry keeps the signal that matters.

CURATED DIRECTORY591 results

smollm

Ollama
Model

SmolLM is a family of lightweight, high-performance language models designed specifically for local execution and edge computing. Unlike massive frontier models that require high-end data center GPUs, SmolLM is optimized to run efficiently on consumer-grade hardware, including laptops and mobile devices, via tools like Ollama. For developers, this means significantly lower latency and reduced infrastructure costs when implementing text generation tasks. While it lacks the broad world knowledge of multi-billion parameter models, it excels in specific, constrained environments such as code completion, structured data extraction, and local chat interfaces where privacy and offline availability are non-negotiable. It is an ideal candidate for developers building agentic workflows that require high-frequency, low-cost reasoning cycles without the overhead of cloud API calls.

text generationSee Ollama library

qwen3.8

Ollama
Model

qwen3.8 is a versatile text-generation model optimized for local inference via the Ollama ecosystem. Designed for developers who prioritize data privacy and low-latency execution, this model allows you to run sophisticated LLM workflows entirely on your own hardware without relying on external APIs. While specific parameter counts vary by quantization level, the architecture is engineered to balance reasoning capabilities with computational efficiency. It is particularly effective for building local RAG (Retrieval-Augmented Generation) pipelines, automating code documentation, and powering edge-based chat interfaces. For integration, its availability in the Ollama library means you can deploy it with a single command, making it an ideal candidate for testing prompt engineering or prototyping agentic workflows in a sandboxed environment. Unlike massive cloud-hosted models, qwen3.8 offers a predictable cost structure and total control over your inference stack.

text generationSee Ollama library

gemma3n

Ollama
Model

Gemma3n is a specialized iteration of the Gemma family designed for high-performance local inference via the Ollama ecosystem. For developers building privacy-first applications or edge computing solutions, this model provides a streamlined path to deploying sophisticated text generation capabilities without relying on external APIs. While specific parameter counts vary depending on the quantized version pulled, the architecture is optimized for low-latency reasoning and instruction following. Unlike massive cloud-hosted models, Gemma3n is built to balance computational efficiency with linguistic nuance, making it ideal for local RAG (Retrieval-Augmented Generation) pipelines, automated coding assistants, and private data summarization. Integration is straightforward for anyone already utilizing the Ollama runtime, allowing for rapid prototyping of agentic workflows. It serves as a robust middle-ground for those needing more intelligence than tiny SLMs but requiring significantly less VRAM than flagship frontier models.

text generationSee Ollama library

qwq

Ollama
Model

qwq is a specialized text generation model optimized for local deployment via the Ollama ecosystem. Designed for developers who prioritize data privacy and low-latency inference, it allows for high-performance language tasks without relying on external cloud APIs. While specific parameter counts vary depending on the quantized version pulled, the model is built to handle complex reasoning, structured data extraction, and conversational logic. For integration, qwq follows the standard Ollama API patterns, making it a drop-in replacement for existing LLM workflows in local development environments. Compared to massive proprietary models, qwq offers a more efficient footprint, making it ideal for edge computing, automated testing pipelines, and private RAG (Retrieval-Augmented Generation) implementations where keeping data on-premise is a non-negotiable requirement.

text generationSee Ollama library

llava-llama3

Ollama
Model

llava-llama3 represents a significant step forward in local multimodal reasoning by marrying the LLaVA vision-language architecture with the Llama 3 backbone. For developers building privacy-first applications, this model enables seamless processing of both text and visual inputs directly on local hardware via Ollama. Unlike standard text-only LLMs, llava-llama3 can perform complex visual reasoning, such as describing images, extracting text from documents, or identifying objects within a scene. It is particularly useful for edge computing scenarios where cloud latency or data privacy concerns prohibit the use of proprietary APIs. While performance scales with your VRAM, the integration via Ollama makes it trivial to deploy into existing Python or JavaScript workflows. Compared to earlier LLaVA iterations, the Llama 3 integration offers improved instruction following and more coherent linguistic output, making it a robust choice for developers prototyping multimodal agents or automated visual inspection tools.

text generationSee Ollama library

glm-5.1

Ollama
Model

GLM-5.1 is a versatile text generation model optimized for local inference via the Ollama ecosystem. For developers building privacy-first applications or edge-based AI solutions, this model offers a streamlined path to deploying high-performance language capabilities without relying on external APIs. While specific parameter counts vary by quantized version, the architecture is designed to balance reasoning depth with computational efficiency. It excels in standard NLP tasks such as code completion, structured data extraction, and conversational logic. Unlike massive cloud-hosted models, GLM-5.1 is engineered for developers who need low-latency responses and full control over their data pipeline. Integration is straightforward through the Ollama CLI and API, making it an ideal candidate for local RAG (Retrieval-Augmented Generation) workflows and automated content pipelines where data sovereignty is a primary requirement.

text generationSee Ollama library

minimax-m2.7

Ollama
Model

Minimax-m2.7 is a specialized text generation model optimized for local deployment via the Ollama framework. For developers building privacy-centric applications or latency-sensitive workflows, this model provides a streamlined path to running high-quality inference on local hardware without relying on external APIs. While the exact parameter count is abstracted in the current library listing, its architecture is tuned for efficient reasoning and structured text output. It is particularly well-suited for developers working on RAG (Retrieval-Augmented Generation) pipelines, local coding assistants, or automated content processing where data sovereignty is a priority. Unlike massive cloud-based LLMs, m2.7 focuses on a balanced performance-to-compute ratio, making it a versatile choice for edge computing environments or local development testing before scaling to larger production clusters.

text generationSee Ollama library

translategemma

Ollama
Model

Translategemma is a specialized fine-tuned model designed specifically for high-fidelity machine translation tasks. Unlike general-purpose LLMs that may struggle with linguistic nuances or suffer from 'translationese,' this model is optimized to preserve semantic intent and stylistic consistency across language pairs. For developers building localization pipelines, real-time chat interfaces, or multilingual content engines, translategemma offers a streamlined alternative to massive, multi-modal models. It is specifically architected for local inference via Ollama, making it an ideal candidate for privacy-sensitive applications where data cannot leave the local environment. While general models excel at reasoning, translategemma prioritizes lexical accuracy and grammatical fluency, providing a more predictable output for structured translation workflows. Integration is straightforward for anyone already using the Ollama ecosystem, allowing for rapid prototyping of multilingual features without the latency or cost overhead of external APIs.

text generationSee Ollama library

mistral-small3.2

Ollama
Model

Mistral-Small-3.2 is a precision-engineered model designed for developers who need a balance between high-level reasoning and low-latency execution. Unlike massive frontier models that demand heavy compute, this iteration focuses on optimizing the parameter-to-performance ratio, making it an ideal candidate for local deployment via Ollama. It excels in structured data extraction, complex instruction following, and code reasoning tasks where speed is just as critical as accuracy. For teams building agentic workflows or RAG pipelines, Mistral-Small offers a predictable, efficient middle ground that minimizes inference costs without sacrificing the nuance required for sophisticated text generation. It is particularly well-suited for edge computing or local development environments where hardware constraints are a factor, providing a robust alternative to larger, more cumbersome models.

text generationSee Ollama library

falcon3

Ollama
Model

Falcon3 represents the next evolution in the Falcon series, optimized specifically for high-performance local inference via Ollama. For developers building privacy-first or edge-based applications, this model offers a robust alternative to closed-source APIs by providing low-latency text generation directly on your hardware. Unlike general-purpose monolithic models, Falcon3 is engineered to balance parameter efficiency with reasoning capabilities, making it ideal for RAG (Retrieval-Augmented Generation) pipelines, local code assistance, and structured data extraction. Integration is seamless for those already using the Ollama ecosystem, allowing you to transition from prototyping to local deployment with minimal configuration changes. While specific parameter counts vary by quantized version, the architecture focuses on high throughput and reduced memory overhead, ensuring it remains accessible for developers working on consumer-grade GPUs or high-end workstations.

text generationSee Ollama library

llama2-uncensored

Ollama
Model

For developers working with local LLMs, llama2-uncensored represents a specialized fine-tune of the Llama 2 architecture designed to strip away the standard safety refusals and alignment constraints. While the base Llama 2 model is highly capable, its strict instruction-following can sometimes trigger false positives in safety filters, hindering creative writing, complex roleplay, or unfiltered data analysis. This version is optimized for high-fidelity instruction following without the 'as an AI language model' caveats. It is ideal for building autonomous agents that require uninhibited reasoning or for developers experimenting with edge-case scenarios where standard censorship would break the logic flow. Since it is available via Ollama, integration into local workflows is seamless via a simple API, making it a practical choice for privacy-conscious developers who need raw, unmoderated text generation on their own hardware.

text generationSee Ollama library

mixtral

Ollama
Model

Mixtral is a high-performance Mixture-of-Experts (MoE) model designed to bridge the gap between massive dense models and lightweight local deployment. For developers, the primary value lies in its efficiency: it utilizes a sparse architecture that activates only a fraction of its total parameters during inference, delivering GPT-3.5-level reasoning speeds with significantly lower computational overhead. This makes it an ideal candidate for local integration via Ollama, where latency and hardware constraints are critical. You can leverage Mixtral for complex tasks like multi-turn dialogue, sophisticated code generation, and high-context retrieval-augmented generation (RAG) pipelines. Unlike standard dense models, Mixtral offers a superior performance-to-compute ratio, allowing you to run a highly capable reasoning engine on consumer-grade hardware without the massive VRAM requirements typically associated with high-parameter models. It is an excellent choice for building privacy-focused, offline-capable AI agents.

text generationSee Ollama library

nemotron-3-super

Ollama
Model

Nemotron-3-super is a specialized text generation model optimized for local deployment via Ollama. While many large-scale models require heavy cloud infrastructure, this iteration is designed to balance high-reasoning capabilities with the efficiency required for edge computing and local development environments. For developers, the primary value lies in its ability to handle complex instruction-following and structured data generation tasks without the latency or privacy concerns of external APIs. It serves as a robust backbone for building local RAG (Retrieval-Augmented Generation) pipelines, autonomous agents, and automated coding assistants. When integrating, you should prioritize testing its specific parameter scaling against your hardware constraints, as performance varies significantly based on your local quantization settings. Compared to standard general-purpose models, Nemotron-3-super focuses on high-fidelity output suitable for production-grade logic tasks.

text generationSee Ollama library

starcoder2

Ollama
Model

StarCoder2 is a high-performance code intelligence model family designed specifically for the software development lifecycle. Unlike general-purpose LLMs that treat code as just another language, StarCoder2 is optimized for deep syntax understanding, long-context repository comprehension, and low-latency completions. For developers building IDE extensions, automated code review tools, or local copilot clones, this model offers a significant leap in accuracy across dozens of programming languages. It is particularly effective for tasks ranging from simple autocomplete to complex refactoring and unit test generation. Because it is available for local inference via Ollama, you can integrate it directly into your private workflows without the latency or privacy concerns of cloud-based APIs. Compared to earlier iterations, StarCoder2 demonstrates superior reasoning in polyglot environments, making it a robust choice for modern, multi-language microservices architectures.

text generationSee Ollama library

granite3.1-moe

Ollama
Model

Granite3.1-MoE is a Mixture-of-Experts model designed for efficient, high-performance text generation. For developers working in resource-constrained environments or seeking low-latency local inference, the MoE architecture provides a strategic advantage by activating only a subset of parameters per token. This allows for sophisticated reasoning and complex instruction following without the massive compute overhead of dense models of similar capacity. It is particularly well-suited for RAG (Retrieval-Augmented Generation) pipelines, code assistance, and automated data extraction where throughput is critical. Available via Ollama, it integrates seamlessly into existing local workflows, making it a practical choice for privacy-conscious applications or edge computing scenarios. Compared to standard dense models, Granite3.1-MoE offers a superior balance of intelligence-per-watt, enabling more complex logic on consumer-grade hardware.

text generationSee Ollama library

orca-mini

Ollama
Model

Orca-mini is a lightweight, distilled language model designed specifically for efficient local inference. For developers working with resource-constrained environments—such as edge devices, mobile hardware, or local development machines without high-end GPUs—this model offers a pragmatic alternative to massive parameter models. While it lacks the broad reasoning depth of its larger counterparts, orca-mini excels at structured text generation and basic instruction following tasks. It is optimized for low-latency responses, making it an excellent candidate for prototyping conversational interfaces, local RAG (Retrieval-Augmented Generation) pipelines, or basic text processing workflows. Since it is available via the Ollama library, integration into your existing local stack is seamless, allowing you to test agentic workflows or specialized fine-tuning ideas without the overhead of cloud-based API costs or privacy concerns.

text generationSee Ollama library

deepseek-coder-v2

Ollama
Model

DeepSeek-Coder-V2 represents a significant shift in open-weights coding models, utilizing a Mixture-of-Experts (MoE) architecture to balance high-tier reasoning with computational efficiency. For developers, this means access to a model that rivals proprietary benchmarks in code completion, debugging, and complex architectural reasoning, while remaining viable for local deployment via Ollama. Unlike dense models that struggle with long-context dependency, V2 is optimized for massive codebases, supporting extensive context windows that allow for better repository-wide understanding. It excels in over 300 programming languages, making it a versatile tool for polyglot environments. While it requires careful hardware consideration due to its MoE structure, its ability to integrate into existing IDE workflows and CI/CD pipelines via local APIs makes it a powerful alternative to closed-source assistants. If you are looking to build private, low-latency coding tools without sending proprietary logic to external servers, this is a primary candidate for your stack.

text generationSee Ollama library

mistral-small

Ollama
Model

Mistral-Small is a high-efficiency model designed for developers who need a balance between reasoning capabilities and low-latency performance. Unlike larger parameter models that demand significant VRAM, Mistral-Small is optimized for production environments where throughput and cost-effectiveness are critical. It excels at structured tasks such as JSON extraction, code generation, and complex instruction following, making it an ideal candidate for agentic workflows and RAG pipelines. For developers working locally via Ollama, this model offers a streamlined deployment path, allowing for rapid prototyping without the overhead of massive hardware requirements. While it may not match the deep creative nuance of its larger siblings, its predictable logic and fast inference speeds make it a superior choice for scalable, task-oriented applications where reliability is the primary metric.

text generationSee Ollama library

snowflake-arctic-embed

Ollama
Model

Snowflake Arctic Embed is a high-performance embedding model designed specifically for enterprise-grade retrieval tasks. Unlike general-purpose LLMs, this model focuses on mapping text into high-dimensional vector spaces with extreme precision, making it a specialized tool for RAG (Retrieval-Augmented Generation) pipelines. For developers building semantic search engines or long-context knowledge bases, Arctic Embed offers a significant upgrade in retrieval accuracy and latency compared to older transformer-based encoders. It is optimized for integration via Ollama, allowing for seamless local inference without the overhead of cloud-based API costs or data privacy concerns. Whether you are fine-tuning a vector database or implementing hybrid search, this model provides the dense vector representations necessary to bridge the gap between natural language queries and unstructured data repositories.

text generationSee Ollama library

codegemma

Ollama
Model

CodeGemma is a specialized model family fine-tuned specifically for code completion, generation, and logical reasoning. Unlike general-purpose LLMs that might struggle with strict syntax, CodeGemma is optimized to understand complex programming patterns and provide contextually relevant snippets across multiple languages. For developers, the primary value lies in its ability to run locally via Ollama, ensuring that your proprietary codebase never leaves your local environment—a critical requirement for enterprise security and privacy. It serves as an excellent lightweight alternative to massive cloud-based models, making it ideal for IDE integrations, automated unit test generation, and real-time code explanation. While it may lack the vast breadth of a trillion-parameter model, its efficiency in low-latency environments makes it a practical choice for local development workflows and CI/CD pipeline automation.

text generationSee Ollama library

all-minilm

Ollama
Model

For developers building high-performance RAG (Retrieval-Augmented Generation) pipelines or semantic search engines, all-minilm offers a lightweight, specialized solution for high-speed text embedding. Unlike massive generative models, this model is optimized for dimensionality reduction and vector representation, making it ideal for local deployment where latency and memory footprint are critical constraints. It excels at transforming raw text into dense vectors that capture semantic meaning, allowing you to perform efficient similarity searches across large datasets. Because it is available via Ollama, integration into existing local workflows is seamless, providing a standardized API for embedding tasks without the overhead of cloud-based providers. While it lacks the conversational reasoning of larger LLMs, its utility in the pre-processing and retrieval stages of an AI pipeline is significant for developers prioritizing edge computing and privacy.

text generationSee Ollama library

olmo2

Ollama
Model

OLMo2 is a recent release from the Allen Institute for AI (AI2) designed to advance the transparency and accessibility of open-source language modeling. Unlike many proprietary models, OLMo2 is built with a focus on scientific rigor, providing researchers and developers with a more predictable foundation for fine-tuning and evaluation. For developers, the primary value lies in its architecture's efficiency for local inference via Ollama, making it a strong candidate for privacy-sensitive applications or edge computing environments. While it may not match the raw scale of massive closed-source models, its performance-to-parameter ratio is optimized for text generation tasks like summarization, code assistance, and structured data extraction. If your workflow requires a model that is easy to inspect, deploy locally, and free from the 'black box' constraints of commercial APIs, OLMo2 offers a highly reliable alternative for building specialized downstream agents.

text generationSee Ollama library

deepseek-v3

Ollama
Model

DeepSeek-V3 represents a significant leap in open-weights architecture, designed to challenge the performance ceilings of much larger proprietary models. For developers, the primary value proposition lies in its Mixture-of-Experts (MoE) design, which optimizes computational efficiency by activating only a fraction of its parameters during inference. This makes it a highly viable candidate for complex reasoning tasks, sophisticated code generation, and nuanced multilingual processing without the massive overhead typically associated with frontier-class models. Unlike standard dense models, V3 offers a better performance-to-latency ratio, making it suitable for high-throughput production environments. It integrates seamlessly into existing workflows via Ollama, allowing for local testing and private deployment. Whether you are building autonomous agents or fine-tuning for domain-specific logic, DeepSeek-V3 provides a robust, scalable backbone that competes directly with top-tier closed models in both logic and instruction following.

text generationSee Ollama library

smollm2

Ollama
Model

SmolLM2 is a high-performance small language model series designed specifically for efficient local execution and edge computing. Unlike massive frontier models that require heavy GPU clusters, SmolLM2 is optimized for developers building low-latency applications, mobile integrations, or privacy-centric local tools. It excels at text generation, summarization, and basic reasoning tasks while maintaining a footprint small enough to run on consumer-grade hardware or even mobile devices via Ollama. For developers, the primary value proposition lies in its high throughput-to-parameter ratio, making it an ideal candidate for RAG (Retrieval-Augmented Generation) pipelines where quick retrieval and processing are prioritized over complex multi-step logic. While it may not match the deep reasoning capabilities of a 70B parameter model, its ability to provide coherent, instruction-following outputs within a constrained memory budget makes it a versatile tool for microservices and local prototyping.

text generationSee Ollama library
Email