Global AI chat room · 11 online now Join now
DIRECTORY / 02

AI Models | Open-Source LLM Directory

Discover and compare open-source LLMs, language models and multimodal models by capability, scale, license, downloads and provenance.

Compare modelsFind the right building block for your next workflow
Directory overview
591
curated entries
24 topic groupsLive
02 / MODEL INDEX

Find the right model for the job

Context first, better decisions. Every entry keeps the signal that matters.

CURATED DIRECTORY591 results

gemini-3.1-flash-lite

google
1048576 ctx

For developers building high-throughput applications, Gemini 3.1 Flash Lite represents a strategic shift toward cost-efficient, low-latency multimodal processing. Unlike larger flagship models that prioritize deep reasoning at the expense of speed, this model is engineered specifically for high-volume agentic workflows and real-time interaction. It handles a massive 1M token context window, making it uniquely capable of processing long-form PDFs, extensive video files, or complex audio streams without the typical latency penalties. While it may lack the heavy-duty reasoning of the Pro tier, its strength lies in its integration readiness for RAG pipelines, automated data extraction, and multi-modal sensory tasks. If your stack requires a lightweight model to act as a fast-acting agent or a scalable parser for unstructured data, this model offers a highly competitive price-to-performance ratio for production environments.

text generationAPI

gemini-3.5-flash:batch

google
1048576 ctx

Gemini 3.5 Flash: Batch is engineered for developers who need to balance high-throughput processing with sophisticated reasoning. While the standard Flash model excels at low-latency interactions, this batch-optimized version is specifically designed for large-scale asynchronous workloads where cost-efficiency and massive context handling are the primary drivers. It maintains high proficiency in code generation and complex logic, making it an ideal backbone for agentic workflows that require parallel execution across massive datasets. With a 1M token context window, you can ingest entire repositories or extensive documentation sets without losing coherence. Unlike standard real-time endpoints, this model is optimized for jobs where you can trade immediate response for significantly reduced per-token costs, making it a strategic choice for data labeling, large-scale summarization, and automated code refactoring pipelines.

text generationAPI

gemini-3.5-flash

google
1048576 ctx

Gemini 3.5 Flash is engineered for developers who need to balance high-reasoning capabilities with aggressive latency requirements. Unlike previous lightweight models that often struggled with complex logic, this iteration delivers near-Pro performance specifically optimized for code generation, debugging, and structured data extraction. Its standout feature is the massive 1M token context window, making it an ideal candidate for large-scale codebase analysis or processing massive document sets in a single pass. For engineers building agentic workflows, the model's efficiency in parallel execution allows for rapid multi-step reasoning without the typical cost overhead of larger frontier models. It is best utilized in high-throughput production environments where speed-to-first-token and cost-per-request are critical KPIs, providing a competitive middle ground between ultra-fast small models and heavy-duty reasoning engines.

text generationAPI

gemini-3-pro-image

google
131072 ctx

Gemini-3-Pro-Image represents a significant shift from traditional diffusion-based generators by leveraging the underlying multimodal reasoning of the Gemini 3 Pro architecture. For developers, this means moving beyond simple prompt-to-image workflows toward complex, instruction-based image manipulation and semantic understanding. Unlike previous iterations that struggled with spatial reasoning or text rendering, this model excels at grounding visual elements in real-world logic, making it ideal for sophisticated design automation, asset generation for gaming, and high-fidelity marketing content. Integration is handled via a robust API, supporting a large 131k context window that allows for multi-turn conversational editing and complex scene descriptions. While standard models often require heavy prompt engineering to achieve specific compositions, Gemini-3-Pro-Image interprets intent more naturally, reducing the iterative loop required to achieve production-ready visual outputs.

text generationAPI

gemini-3.1-flash-image

google
131072 ctx

Gemini 3.1 Flash Image (codenamed Nano Banana 2) is engineered for developers who need high-fidelity visual synthesis without the typical latency overhead of larger diffusion models. While previous iterations often forced a trade-off between prompt adherence and generation speed, this model bridges that gap by delivering Pro-tier aesthetic quality at Flash-class inference speeds. It is particularly optimized for real-time applications, such as dynamic UI asset generation, interactive gaming environments, and rapid prototyping workflows. For integration, the model supports complex image-to-image editing and sophisticated text-to-image pipelines via a streamlined API. Compared to standard heavyweight models, it offers a significantly lower time-to-first-token for visual data, making it a superior choice for production environments where user experience depends on near-instantaneous visual feedback.

text generationAPI

gemini-3.1-flash-lite-image

google
65536 ctx

For developers building high-throughput applications, Gemini 3.1 Flash Lite Image is engineered to solve the latency-cost bottleneck inherent in visual generative workflows. Unlike larger, more compute-heavy diffusion models, this 'Lite' iteration focuses on rapid-fire inference cycles, making it ideal for real-time UI prototyping, automated asset generation, and dynamic content scaling within production pipelines. It excels in scenarios where millisecond response times are more critical than extreme photorealism, such as generating quick visual placeholders or iterative design explorations. Integration is straightforward via Google's existing API ecosystem, allowing for seamless inclusion in existing LLM-driven agentic workflows. While it may not match the granular detail of heavyweight models for high-end artistic production, its efficiency makes it the pragmatic choice for developers needing to scale image generation across millions of requests without prohibitive overhead.

text generationAPI

gpt-5.6-sol:batch

openai
1050000 ctx

For developers building autonomous agents or complex software pipelines, GPT-5.6-sol:batch represents a significant shift toward high-reliability reasoning. Unlike standard chat models optimized for conversational flow, this iteration is architected specifically for deep logic and long-context execution. It excels in multi-step reasoning tasks where state management and instruction adherence are critical, making it a primary candidate for automated code refactoring, complex CLI automation, and agentic workflows that require navigating large codebases. With a 1.05M token context window, it handles massive technical documentation or entire repository structures without losing coherence. While previous models often struggled with the 'drift' seen in long-running agentic loops, the Sol series demonstrates improved stability in executing sequential commands. It is best utilized as a backend reasoning engine for developer tools rather than a simple chatbot interface, offering the precision required for production-grade automation.

text generationAPI

gpt-5.6-sol

openai
1050000 ctx

GPT-5.6 Sol represents a significant architectural shift toward agentic reasoning and deep-logic execution. For developers, this isn't just another chat model; it is a specialized engine designed for high-autonomy workflows. While previous iterations excelled at pattern matching, Sol is optimized for multi-step reasoning and complex codebase manipulation. It demonstrates a marked improvement in command-line proficiency and long-context coherence, making it a primary candidate for building autonomous coding agents or complex DevOps automation tools. Integrating Sol into your stack allows for more reliable tool-calling and structured output in environments where precision is non-negotiable. Compared to standard LLMs, Sol minimizes the 'drift' often seen in long-chain reasoning, providing a more stable foundation for developers building sophisticated, self-correcting software agents.

text generationAPI

gpt-5.6-sol-pro:batch

openai
1050000 ctx

For developers tackling high-stakes logic or deep reasoning, gpt-5.6-sol-pro:batch offers a specialized tier of the Sol architecture. Unlike standard inference modes, this model utilizes a 'pro' reasoning setting designed to prioritize accuracy and multi-step deduction over raw speed. This makes it particularly effective for complex code refactoring, mathematical proofs, and intricate architectural planning where the cost of a hallucination outweighs the latency of a longer response. By utilizing the batch processing endpoint, you can significantly reduce costs for non-real-time workloads like dataset labeling, large-scale document analysis, or automated unit test generation. While standard models excel at conversational fluency, this version is optimized for the 'thinking' phase of the pipeline, providing a more robust backbone for autonomous agents and sophisticated RAG workflows that require deep semantic understanding.

text generationAPI

gpt-5.6-sol-pro

openai
1050000 ctx

GPT-5.6 Sol Pro is a specialized iteration of the Sol architecture, specifically optimized for high-stakes reasoning tasks via the `reasoning.mode: pro` configuration. For developers, the primary distinction lies in its compute-intensive approach to problem-solving; while the standard Sol model offers speed, the Pro mode allocates more internal processing to verify logic and reduce hallucination rates in complex workflows. This makes it particularly effective for autonomous agent orchestration, multi-step code synthesis, and advanced mathematical verification where accuracy outweighs raw latency. It integrates seamlessly into existing OpenAI-compatible pipelines, requiring only a parameter adjustment to unlock its full reasoning depth. If your application demands deep logical consistency rather than just rapid text generation, this model serves as a high-fidelity backbone for your logic layer.

text generationAPI

gpt-5.6-terra:batch

openai
1050000 ctx

GPT-5.6 Terra:Batch is a specialized mid-tier model designed for high-throughput asynchronous processing. Positioned strategically between the heavyweight Sol flagship and the lightweight Luna tier, Terra offers a pragmatic sweet spot for developers needing reliable reasoning without the latency or cost overhead of top-tier models. It excels in agentic workflows, complex code generation, and large-scale data reasoning tasks. With a substantial 1.05M token context window, it is particularly effective for analyzing entire repositories or processing massive document sets in batch mode. For teams building autonomous agents or automated CI/CD pipelines, this model provides the necessary intelligence to handle multi-step logic while maintaining a scalable cost structure. Unlike the flagship models that prioritize absolute peak performance, Terra is optimized for consistency and efficiency in repetitive, high-volume production environments.

text generationAPI

gpt-5.6-terra

openai
1050000 ctx

GPT-5.6 Terra is designed as a mid-range workhorse for developers who need a sweet spot between high-end reasoning and operational cost-efficiency. While the Sol tier handles massive complexity and Luna focuses on throughput, Terra is optimized for high-frequency agentic workflows and iterative coding tasks. It features a massive 1.05 million token context window, making it particularly effective for analyzing entire codebases or long-form documentation in a single pass. For engineering teams, this means you can deploy it for autonomous agents, complex debugging, and RAG-heavy applications without the prohibitive latency or pricing of flagship-class models. It offers a significant step up in logical consistency over the Luna tier, making it a reliable choice for production environments where reliability and context retention are non-negotiable.

text generationAPI

gpt-5.6-terra-pro:batch

openai
1050000 ctx

For developers tackling high-stakes logic or deep architectural planning, GPT-5.6 Terra Pro:Batch offers a specialized reasoning tier. Unlike standard inference modes, this model utilizes the 'pro' reasoning setting, specifically optimized for multi-step problem solving and complex code synthesis. While the base Terra model is highly capable, the 'pro' mode forces a more exhaustive internal chain-of-thought, making it ideal for debugging intricate distributed systems or generating mathematically rigorous proofs. This batch-oriented deployment is designed for high-throughput workflows where latency is secondary to absolute accuracy. If your pipeline requires heavy-duty reasoning—such as automated unit test generation or complex data transformation logic—this model provides a significant step up in reliability compared to general-purpose chat models. Integration is seamless via the standard OpenAI-compatible API, allowing you to toggle the reasoning depth through the reasoning.mode parameter.

text generationAPI

gpt-5.6-terra-pro

openai
1050000 ctx

GPT-5.6 Terra Pro is a specialized iteration of the Terra architecture, specifically optimized for high-stakes reasoning tasks via a dedicated 'pro' mode. Unlike standard inference paths, this model is engineered to prioritize depth and logical consistency over raw generation speed. For developers, this means a significant reduction in logical fallacies and hallucinations when handling multi-step mathematical proofs, complex code refactoring, or intricate architectural planning. It operates within a massive 1.05M token context window, making it ideal for analyzing entire codebases or massive documentation sets in a single pass. While the latency is higher than the base Terra model, the trade-off is a measurable increase in accuracy for non-trivial problem-solving. Integration is straightforward via the standard OpenAI API, requiring only the adjustment of the reasoning mode parameter to unlock the enhanced cognitive capabilities.

text generationAPI

gpt-5.6-luna:batch

openai
1050000 ctx

GPT-5.6 Luna:Batch is a specialized iteration within the GPT-5.6 series, architected specifically for high-throughput production environments. While larger flagship models focus on deep reasoning, Luna is optimized for the 'middle tier' of developer workflows: tasks that require reliable logic but demand low latency and reduced token costs. For engineers building scalable applications, this model serves as an ideal engine for high-volume classification, real-time chat interfaces, and lightweight agentic loops where rapid response times are critical to user experience. It manages a significant 1.05M context window, allowing for extensive document processing without the overhead of heavier models. If your stack requires processing massive datasets or managing thousands of concurrent low-complexity sessions, Luna provides a pragmatic balance between intelligence and operational efficiency, making it a superior choice for cost-sensitive scaling compared to standard frontier models.

text generationAPI

gpt-5.6-luna

openai
1050000 ctx

GPT-5.6 Luna is a high-throughput, low-latency model designed specifically for developers building high-volume applications where speed and cost-efficiency are non-negotiable. While larger frontier models focus on deep reasoning, Luna is optimized for the 'execution layer' of your stack. It excels in latency-sensitive environments such as real-time chat interfaces, large-scale text classification, and the iterative loops required for lightweight agentic workflows. With a 1.05M context window, it handles massive datasets or long conversation histories without the typical performance degradation seen in smaller models. For developers, this means you can offload repetitive, high-frequency tasks to Luna to optimize your token spend while reserving heavier models for complex logic. Integration remains seamless via OpenAI's standard API, making it a plug-and-play upgrade for existing pipelines that require a balance of reasoning capability and rapid response times.

text generationAPI

gpt-5.6-luna-pro:batch

openai
1050000 ctx

For developers working on high-stakes reasoning tasks, gpt-5.6-luna-pro:batch offers a specialized optimization of the Luna architecture. Unlike standard inference endpoints, this model utilizes the 'pro' reasoning mode, specifically tuned for deep logical deduction, complex mathematical problem-solving, and intricate code architecture planning. It is designed for workflows where accuracy and depth of thought outweigh the need for raw millisecond latency. The 'batch' designation indicates this is optimized for asynchronous processing, making it a cost-effective solution for large-scale data transformation, synthetic data generation, or automated code auditing where results can be processed in bulk. If your pipeline requires more than just pattern matching—specifically, if you need a model to 'think' through multi-step constraints before outputting—this configuration provides a significant step up from standard text generation models.

text generationAPI

gpt-5.6-luna-pro

openai
1050000 ctx

GPT-5.6 Luna Pro is a specialized iteration of the Luna architecture, specifically optimized for deep reasoning tasks via the 'pro' reasoning mode. For developers building agentic workflows or complex logic engines, this model represents a shift from standard next-token prediction toward structured cognitive processing. Unlike the base Luna model, the Pro tier is designed to minimize logical hallucinations during multi-step problem solving, making it ideal for code synthesis, mathematical verification, and intricate architectural planning. It supports a massive 1.05M context window, allowing you to ingest entire repositories or extensive technical documentation without losing coherence. Integration is straightforward via the standard OpenAI API, requiring only a parameter adjustment to toggle the enhanced reasoning engine. While it carries higher latency and cost than standard models, the trade-off is a significant leap in accuracy for high-stakes, autonomous reasoning tasks.

text generationAPI

gemini-3.5-flash-lite:batch

google
1048576 ctx

Gemini 3.5 Flash Lite (Batch) is engineered for developers building high-throughput, agentic workflows where latency and cost-efficiency are the primary constraints. Unlike larger flagship models designed for broad reasoning, this model is optimized as a specialized subagent. It excels at executing granular, discrete tasks—such as data extraction, classification, or tool-calling—within a larger multi-agent orchestration. By utilizing the batch processing mode, developers can significantly reduce costs for non-real-time workloads while maintaining a massive 1M token context window. This makes it an ideal choice for processing large datasets or managing background asynchronous tasks in complex pipelines. If your architecture requires a 'worker' model to handle high-volume, repetitive logic without the overhead of a heavy-duty LLM, this version provides the necessary speed and scale for production-grade automation.

text generationAPI

gemini-3.5-flash-lite

google
1048576 ctx

Gemini 3.5 Flash Lite is engineered for developers building high-density, multi-agent architectures where latency and cost-per-token are the primary constraints. Unlike larger flagship models designed for broad reasoning, this model is optimized for 'agentic subtasks'—the granular, repetitive operations that occur within a larger workflow, such as data extraction, tool calling, or intent classification. It features a massive 1M token context window, allowing it to ingest large documentation sets or long conversation histories without losing coherence. For teams scaling agentic loops, Flash Lite offers a sweet spot: it provides the specialized instruction-following required for autonomous tool execution while maintaining the rapid inference speeds necessary to prevent bottlenecking complex pipelines. It is a strategic choice for developers moving from prototyping to production-scale agentic deployments.

text generationAPI

gemini-3.6-flash:batch

google
1048576 ctx

Gemini 3.6 Flash (Batch) is optimized for developers prioritizing high-throughput processing and cost-efficiency in large-scale asynchronous workflows. Unlike standard real-time endpoints, this batch-optimized variant is engineered for massive datasets where latency is secondary to volume and unit cost. It excels in agentic reasoning and complex code generation tasks, providing a significant leap in instruction following compared to previous Flash iterations. For developers building automated testing suites, large-scale content moderation pipelines, or batch-processing data enrichment tools, this model offers a massive 1M token context window, allowing you to ingest entire repositories or extensive documentation in a single request. While it lacks the sub-second responsiveness of real-time models, its ability to produce production-ready, 'low-edit' code makes it a powerful backbone for background automation and complex developer tooling integrations.

text generationAPI

gemini-3.6-flash

google
1048576 ctx

Gemini 3.6 Flash is engineered specifically for developers prioritizing low-latency execution and high-throughput agentic workflows. Unlike larger, heavier models that trade speed for depth, this iteration optimizes the balance between reasoning density and response time, making it an ideal backbone for real-time application logic and automated coding assistants. With a massive 1M token context window, it handles massive codebase ingestion and complex multi-turn reasoning without the typical performance degradation seen in smaller models. For teams building autonomous agents or integrating LLMs into web/app lifecycles, the model offers a significant reduction in 'edit overhead'—meaning its first-pass outputs are structurally sound and ready for production. It integrates seamlessly via Google's existing API ecosystem, providing a predictable, scalable solution for developers who need high intelligence at a fraction of the latency cost of flagship-class models.

text generationAPI

gemini-3.7-flash:batch

google
1048576 ctx

Gemini 3.7 Flash: Batch is engineered for developers building high-throughput, agentic workflows where latency-to-cost efficiency is the primary constraint. Unlike standard real-time endpoints, this batch-optimized variant is designed for asynchronous processing of massive datasets, making it ideal for large-scale data extraction, batch code refactoring, or long-form document analysis. It retains the core strengths of the 3.7 architecture—specifically its advanced multi-step reasoning and multimodal capabilities—but shifts the focus toward massive context windows and reliable complex instruction following. For teams integrating LLMs into automated pipelines, this model offers a way to scale complex reasoning tasks without the overhead of synchronous API calls. It sits in a sweet spot for developers who need 'smart' reasoning for bulk processing rather than just simple pattern matching, providing a robust alternative to smaller, less capable models when dealing with high-volume, multi-turn logic.

text generationAPI

gemini-3.7-flash

google
1048576 ctx

Gemini 3.7 Flash is engineered for developers building high-throughput, agentic applications where latency is a critical bottleneck. Moving beyond simple chat interfaces, this model is optimized for multi-step reasoning and complex workflow orchestration. Its standout feature is the massive 1M token context window, allowing you to ingest entire codebases or massive documentation sets for RAG-based tasks without losing coherence. For those working in automated coding environments or building autonomous agents, the model provides a high density of intelligence per millisecond. Compared to larger frontier models, Flash prioritizes speed and cost-efficiency while maintaining the multimodal capabilities necessary for processing vision and text inputs simultaneously. It is an ideal choice for real-time tool use, automated debugging, and structured data extraction where rapid response cycles are non-negotiable.

text generationAPI
Email