Global AI chat room · 14 online now Join now
DIRECTORY / 02

AI Models | Open-Source LLM Directory

Discover and compare open-source LLMs, language models and multimodal models by capability, scale, license, downloads and provenance.

Compare modelsFind the right building block for your next workflow
Directory overview
839
curated entries
24 topic groupsLive
02 / MODEL INDEX

Find the right model for the job

Context first, better decisions. Every entry keeps the signal that matters.

CURATED DIRECTORY839 results

gpt-5.6-luna-pro:batch

openai
1050000 ctx

For developers working on high-stakes reasoning tasks, gpt-5.6-luna-pro:batch offers a specialized optimization of the Luna architecture. Unlike standard inference endpoints, this model utilizes the 'pro' reasoning mode, specifically tuned for deep logical deduction, complex mathematical problem-solving, and intricate code architecture planning. It is designed for workflows where accuracy and depth of thought outweigh the need for raw millisecond latency. The 'batch' designation indicates this is optimized for asynchronous processing, making it a cost-effective solution for large-scale data transformation, synthetic data generation, or automated code auditing where results can be processed in bulk. If your pipeline requires more than just pattern matching—specifically, if you need a model to 'think' through multi-step constraints before outputting—this configuration provides a significant step up from standard text generation models.

text generationAPI

gpt-5.6-luna-pro

openai
1050000 ctx

GPT-5.6 Luna Pro is a specialized iteration of the Luna architecture, specifically optimized for deep reasoning tasks via the 'pro' reasoning mode. For developers building agentic workflows or complex logic engines, this model represents a shift from standard next-token prediction toward structured cognitive processing. Unlike the base Luna model, the Pro tier is designed to minimize logical hallucinations during multi-step problem solving, making it ideal for code synthesis, mathematical verification, and intricate architectural planning. It supports a massive 1.05M context window, allowing you to ingest entire repositories or extensive technical documentation without losing coherence. Integration is straightforward via the standard OpenAI API, requiring only a parameter adjustment to toggle the enhanced reasoning engine. While it carries higher latency and cost than standard models, the trade-off is a significant leap in accuracy for high-stakes, autonomous reasoning tasks.

text generationAPI

gemini-3.5-flash-lite:batch

google
1048576 ctx

Gemini 3.5 Flash Lite (Batch) is engineered for developers building high-throughput, agentic workflows where latency and cost-efficiency are the primary constraints. Unlike larger flagship models designed for broad reasoning, this model is optimized as a specialized subagent. It excels at executing granular, discrete tasks—such as data extraction, classification, or tool-calling—within a larger multi-agent orchestration. By utilizing the batch processing mode, developers can significantly reduce costs for non-real-time workloads while maintaining a massive 1M token context window. This makes it an ideal choice for processing large datasets or managing background asynchronous tasks in complex pipelines. If your architecture requires a 'worker' model to handle high-volume, repetitive logic without the overhead of a heavy-duty LLM, this version provides the necessary speed and scale for production-grade automation.

text generationAPI

gemini-3.5-flash-lite

google
1048576 ctx

Gemini 3.5 Flash Lite is engineered for developers building high-density, multi-agent architectures where latency and cost-per-token are the primary constraints. Unlike larger flagship models designed for broad reasoning, this model is optimized for 'agentic subtasks'—the granular, repetitive operations that occur within a larger workflow, such as data extraction, tool calling, or intent classification. It features a massive 1M token context window, allowing it to ingest large documentation sets or long conversation histories without losing coherence. For teams scaling agentic loops, Flash Lite offers a sweet spot: it provides the specialized instruction-following required for autonomous tool execution while maintaining the rapid inference speeds necessary to prevent bottlenecking complex pipelines. It is a strategic choice for developers moving from prototyping to production-scale agentic deployments.

text generationAPI

gemini-3.6-flash:batch

google
1048576 ctx

Gemini 3.6 Flash (Batch) is optimized for developers prioritizing high-throughput processing and cost-efficiency in large-scale asynchronous workflows. Unlike standard real-time endpoints, this batch-optimized variant is engineered for massive datasets where latency is secondary to volume and unit cost. It excels in agentic reasoning and complex code generation tasks, providing a significant leap in instruction following compared to previous Flash iterations. For developers building automated testing suites, large-scale content moderation pipelines, or batch-processing data enrichment tools, this model offers a massive 1M token context window, allowing you to ingest entire repositories or extensive documentation in a single request. While it lacks the sub-second responsiveness of real-time models, its ability to produce production-ready, 'low-edit' code makes it a powerful backbone for background automation and complex developer tooling integrations.

text generationAPI

gemini-3.6-flash

google
1048576 ctx

Gemini 3.6 Flash is engineered specifically for developers prioritizing low-latency execution and high-throughput agentic workflows. Unlike larger, heavier models that trade speed for depth, this iteration optimizes the balance between reasoning density and response time, making it an ideal backbone for real-time application logic and automated coding assistants. With a massive 1M token context window, it handles massive codebase ingestion and complex multi-turn reasoning without the typical performance degradation seen in smaller models. For teams building autonomous agents or integrating LLMs into web/app lifecycles, the model offers a significant reduction in 'edit overhead'—meaning its first-pass outputs are structurally sound and ready for production. It integrates seamlessly via Google's existing API ecosystem, providing a predictable, scalable solution for developers who need high intelligence at a fraction of the latency cost of flagship-class models.

text generationAPI

gemini-3.7-flash:batch

google
1048576 ctx

Gemini 3.7 Flash: Batch is engineered for developers building high-throughput, agentic workflows where latency-to-cost efficiency is the primary constraint. Unlike standard real-time endpoints, this batch-optimized variant is designed for asynchronous processing of massive datasets, making it ideal for large-scale data extraction, batch code refactoring, or long-form document analysis. It retains the core strengths of the 3.7 architecture—specifically its advanced multi-step reasoning and multimodal capabilities—but shifts the focus toward massive context windows and reliable complex instruction following. For teams integrating LLMs into automated pipelines, this model offers a way to scale complex reasoning tasks without the overhead of synchronous API calls. It sits in a sweet spot for developers who need 'smart' reasoning for bulk processing rather than just simple pattern matching, providing a robust alternative to smaller, less capable models when dealing with high-volume, multi-turn logic.

text generationAPI

gemini-3.7-flash

google
1048576 ctx

Gemini 3.7 Flash is engineered for developers building high-throughput, agentic applications where latency is a critical bottleneck. Moving beyond simple chat interfaces, this model is optimized for multi-step reasoning and complex workflow orchestration. Its standout feature is the massive 1M token context window, allowing you to ingest entire codebases or massive documentation sets for RAG-based tasks without losing coherence. For those working in automated coding environments or building autonomous agents, the model provides a high density of intelligence per millisecond. Compared to larger frontier models, Flash prioritizes speed and cost-efficiency while maintaining the multimodal capabilities necessary for processing vision and text inputs simultaneously. It is an ideal choice for real-time tool use, automated debugging, and structured data extraction where rapid response cycles are non-negotiable.

text generationAPI

gemini-3.8-flash:batch

google
1048576 ctx

Gemini 3.8 Flash (Batch) is engineered for developers needing high-throughput reasoning without the latency penalties of larger frontier models. This iteration focuses heavily on the 'agentic' layer, showing measurable improvements in multi-step planning and complex software engineering workflows compared to the 3.7 series. For teams building autonomous agents or automated code review pipelines, the model provides a more reliable logic engine for long-context tasks. The batch processing optimization makes it particularly cost-effective for non-real-time, high-volume workloads like large-scale data extraction, document summarization, or asynchronous batch testing. While it maintains the characteristic speed of the Flash lineage, the core upgrade lies in its ability to maintain coherence through intricate, multi-turn reasoning chains, making it a competitive choice for developers moving beyond simple chat interfaces into structured, agent-driven automation.

text generationAPI

gemini-3.8-flash

google
1048576 ctx

Gemini 3.8 Flash is engineered for developers who need high-velocity inference without sacrificing complex reasoning capabilities. While previous Flash iterations focused primarily on low-latency throughput, this version introduces significant architectural improvements specifically targeting agentic workflows and software engineering tasks. It excels in multi-step logic and autonomous tool use, making it a strong candidate for building autonomous agents or automated code review pipelines. With a massive 1M token context window, it handles large-scale codebase ingestion and long-form document analysis with ease. Compared to its predecessors, you will notice a marked reduction in logic errors during complex instruction following. For teams integrating via API, this model offers a sweet spot between the raw intelligence of the Pro series and the cost-efficiency required for high-volume, real-time applications.

text generationAPI

gpt-6-astra-pro:batch

openai
1050000 ctx

GPT-6 Astra Pro: Batch is a specialized high-throughput deployment designed for developers tackling massive-scale reasoning tasks. Unlike the standard Astra model, this version utilizes the 'pro' reasoning mode, specifically optimized for deep logical inference, multi-step problem solving, and complex code synthesis. While it shares the same foundational architecture as the standard Astra, the 'pro' setting prioritizes accuracy and depth over raw latency, making it ideal for asynchronous batch processing where quality is non-negotiable. For teams building agentic workflows, automated data extraction pipelines, or large-scale synthetic data generators, this model offers a significant upgrade in cognitive reliability. It integrates seamlessly via the OpenAI-compatible API, supporting a massive 1.05M token context window. If your use case involves analyzing entire codebases or processing long-form technical documentation in bulk, this batch endpoint provides the most cost-effective way to leverage top-tier reasoning without the overhead of real-time interaction requirements.

text generationAPI

gpt-6-astra-pro

openai
1050000 ctx

GPT-6 Astra Pro is a specialized iteration of the Astra architecture, specifically optimized for high-stakes reasoning tasks. While the standard Astra model excels at rapid inference and general-purpose dialogue, the 'Pro' designation refers to the `reasoning.mode` parameter being locked to its highest tier. For developers, this means a significant reduction in logical hallucinations and a superior ability to handle multi-step mathematical, coding, or architectural problems. It is designed for workflows where accuracy outweighs raw latency, such as automated code auditing, complex data synthesis, or deep logical verification. Integration is seamless for those already using the OpenAI API ecosystem, requiring minimal changes to existing prompts while providing a much more rigorous cognitive backbone for agentic workflows. If your application requires a model that 'thinks' before it speaks, this is the tier to deploy.

text generationAPI

gpt-6-astra:batch

openai
1050000 ctx

GPT-6 Astra:Batch is a high-throughput version of OpenAI’s flagship reasoning engine, specifically optimized for large-scale, asynchronous processing. Unlike standard chat endpoints, this model is architected for long-horizon tasks where latency is secondary to depth and accuracy. For developers, this means you can offload massive workloads—such as automated codebase refactoring, large-scale scientific data synthesis, or exhaustive document auditing—without hitting the typical concurrency bottlenecks of real-time APIs. With a massive 1.05M token context window, it excels at maintaining coherence across entire repositories or multi-hundred-page technical manuals. While it lacks the instant responsiveness of smaller models, its value proposition lies in its ability to execute complex, multi-step reasoning chains autonomously. It is best integrated into batch processing pipelines where you need high-fidelity analytical outputs for datasets that would overwhelm traditional LLMs.

text generationAPI

gpt-6-astra

openai
1050000 ctx

GPT-6 Astra represents a significant shift toward agentic, long-horizon reasoning for developers building complex autonomous systems. Unlike previous iterations optimized for chat or quick retrieval, Astra is architected for deep, multi-step workflows such as end-to-end software engineering, complex scientific modeling, and exhaustive document synthesis. With a massive 1,050,000 token context window, it effectively eliminates the 'forgetting' problem in large-scale codebase analysis or multi-document research projects. For integration, the model is designed to act as a reasoning engine rather than just a text generator, making it ideal for developers implementing RAG-heavy architectures or automated DevOps pipelines. While previous models excelled at pattern matching, Astra focuses on logical consistency over extended execution paths, providing a more reliable foundation for applications requiring high-fidelity planning and execution.

text generationAPI

magistral

Ollama
Model

Magistral is a specialized text generation model optimized for local deployment via the Ollama framework. Designed for developers who prioritize data sovereignty and low-latency inference, it serves as a robust alternative to cloud-dependent APIs. While specific parameter counts vary by version, the model is engineered for efficient resource utilization on consumer-grade hardware. Developers can integrate Magistral into local RAG (Retrieval-Augmented Generation) pipelines, private coding assistants, or automated content workflows without exposing sensitive data to external servers. Compared to massive frontier models, Magistral trades broad general knowledge for high performance in specific generative tasks, making it an ideal choice for edge computing environments and privacy-centric application architectures. Its primary advantage lies in the seamless integration with the Ollama ecosystem, allowing for rapid prototyping and deployment through standard CLI tools.

text generationSee Ollama library

granite4

Ollama
Model

Granite4 is a versatile text-generation model optimized for local inference via the Ollama ecosystem. Designed with a focus on efficiency, it allows developers to deploy high-performance language capabilities directly on edge hardware or local workstations without relying on external APIs. While specific parameter counts vary across the model family, the architecture is tuned for low-latency reasoning and structured text generation, making it an ideal candidate for RAG (Retrieval-Augmented Generation) pipelines and local coding assistants. Unlike massive cloud-hosted models, Granite4 prioritizes a manageable footprint, enabling seamless integration into privacy-sensitive workflows or offline development environments. For developers building autonomous agents or local automation tools, it offers a predictable, cost-effective alternative to proprietary LLMs, provided you verify the specific quantization and license terms within the Ollama library before deployment.

text generationSee Ollama library

command-r

Ollama
Model

Command R is a high-performance model specifically engineered for enterprise-grade RAG (Retrieval-Augmented Generation) and complex tool-use workflows. Unlike general-purpose models that prioritize creative prose, Command R is optimized for long-context reasoning and precise citation, making it a top-tier choice for developers building agentic workflows or knowledge-base interfaces. It excels at processing large amounts of retrieved data and translating that information into grounded, actionable outputs. For developers using Ollama, it offers a streamlined path to local inference, allowing you to test sophisticated multi-step reasoning and API-calling capabilities without the latency or privacy concerns of cloud-based providers. If your stack requires a model that can reliably navigate external documentation and interact with structured tools, Command R provides the necessary stability and instruction-following accuracy.

text generationSee Ollama library

phi

Ollama
Model

Phi is a family of lightweight, high-performance small language models (SLMs) designed for efficient local inference. Unlike massive frontier models that require high-end data center GPUs, Phi is optimized to run on consumer-grade hardware and edge devices via frameworks like Ollama. For developers, this means significantly lower latency and reduced operational costs when deploying text generation tasks. While its parameter count is smaller than industry giants, Phi punches above its weight class in reasoning, logic, and coding tasks by leveraging high-quality synthetic training data. It is an ideal choice for developers building privacy-first applications, local RAG (Retrieval-Augmented Generation) pipelines, or embedded AI features where bandwidth and compute resources are constrained. Integration is straightforward through standard API patterns, making it a practical alternative to cloud-dependent LLMs for prototyping and production-scale edge deployment.

text generationSee Ollama library

granite-code

Ollama
Model

Granite-Code is a specialized model family designed specifically for the software development lifecycle. Unlike general-purpose LLMs that attempt to master everything, this model is optimized for code generation, refactoring, and technical reasoning. For developers working in privacy-sensitive environments, its availability via Ollama makes it an ideal candidate for local-first workflows, allowing you to run powerful coding assistants without sending proprietary source code to external APIs. It excels at understanding complex syntax across multiple programming languages and can be integrated directly into local IDE extensions or CLI tools. While general models often struggle with strict logic in niche languages, Granite-Code focuses on high-fidelity code completion and debugging, making it a practical tool for building automated CI/CD pipelines or enhancing local development environments with low latency.

text generationSee Ollama library

dolphin-phi

Ollama
Model

Dolphin-phi is a fine-tuned derivative of the Microsoft Phi series, optimized through the Dolphin dataset to enhance instruction-following capabilities and reasoning. For developers working in resource-constrained environments, this model offers a high performance-to-size ratio, making it an ideal candidate for local edge deployment via Ollama. Unlike base models that may struggle with complex prompts, the Dolphin tuning focuses on unfiltered, direct responses, which is particularly useful for building agentic workflows and specialized coding assistants. While it lacks the massive parameter count of frontier models, its efficiency allows for low-latency inference on consumer-grade hardware. Integration is straightforward through the Ollama API, making it a practical choice for developers prototyping local LLM applications, privacy-focused chat interfaces, or offline RAG pipelines where data sovereignty is a priority.

text generationSee Ollama library

hermes3

Ollama
Model

Hermes3 is a high-performance fine-tuned model designed for developers who need advanced reasoning and instruction-following capabilities within a local inference environment. Unlike general-purpose base models, Hermes3 is optimized to handle complex multi-turn dialogues and nuanced task execution, making it a strong candidate for building autonomous agents or sophisticated RAG pipelines. For developers working with Ollama, it offers a streamlined path to deploying a model that punches above its weight class in logic and creative synthesis. While specific parameter counts vary across versions, the model is engineered for high efficiency, allowing for rapid prototyping without the latency or privacy concerns of proprietary APIs. It serves as an excellent open-source alternative for those needing a controllable, local engine for structured data extraction, code assistance, or complex logical reasoning tasks.

text generationSee Ollama library

phi4-reasoning

Ollama
Model

Phi-4-reasoning marks a significant shift in the small language model (SLM) landscape by prioritizing deep logical reasoning over mere pattern matching. Designed for developers who need high-level cognitive capabilities without the massive footprint of a frontier model, this iteration excels in complex chain-of-thought tasks, mathematical problem-solving, and structured code generation. Unlike standard SLMs that often struggle with multi-step logic, Phi-4-reasoning utilizes specialized training to navigate intricate instruction sets. For local deployment via Ollama, it offers a high performance-to-parameter ratio, making it an ideal candidate for edge computing, private RAG pipelines, and local agentic workflows where latency and data privacy are critical. If your stack requires a model that can 'think' through a problem rather than just predicting the next token, this is a highly efficient alternative to larger, resource-heavy architectures.

text generationSee Ollama library

dolphin-mistral

Ollama
Model

Dolphin-Mistral is a fine-tuned derivative of the Mistral architecture, specifically optimized through the Dolphin dataset to enhance instruction-following capabilities and conversational fluidity. For developers working on local-first applications, this model offers a significant step up from base Mistral by reducing refusal rates and improving complex reasoning tasks. Unlike standard models that may be overly constrained by safety alignment, Dolphin is designed to be more compliant with diverse user intents, making it ideal for creative writing, complex coding assistance, and nuanced roleplay scenarios. Because it is hosted via Ollama, integration into local workflows is seamless, allowing for high-performance inference on consumer-grade hardware without the latency or privacy concerns of cloud APIs. It serves as a robust middle-ground for those needing a model that is lightweight enough for edge deployment but intelligent enough to handle multi-turn logic and structured data extraction.

text generationSee Ollama library

glm-4.7-flash

Ollama
Model

For developers building latency-sensitive applications, glm-4.7-flash offers a strategic middle ground between high-speed throughput and reasoning depth. Designed for efficient text generation, this model is optimized for high-frequency tasks where response time is as critical as accuracy. Unlike massive parameter models that struggle with real-time constraints, this 'flash' iteration prioritizes rapid token generation, making it an ideal candidate for chat interfaces, real-time summarization, and automated content pipelines. Because it is available via Ollama, you can integrate it directly into local development workflows or edge computing environments without managing complex cloud dependencies. While it may not match the heavy-duty reasoning of its larger counterparts in complex multi-step logic, its performance-to-latency ratio makes it a highly competitive choice for scalable, production-ready agentic workflows and high-volume API integrations.

text generationSee Ollama library
Email