Global AI chat room · 18 online now Join now
DIRECTORY / 02

AI Models | Open-Source LLM Directory

Discover and compare open-source LLMs, language models and multimodal models by capability, scale, license, downloads and provenance.

Compare modelsFind the right building block for your next workflow
Directory overview
591
curated entries
24 topic groupsLive
02 / MODEL INDEX

Find the right model for the job

Context first, better decisions. Every entry keeps the signal that matters.

CURATED DIRECTORY591 results

gpt-6-luna

openai
1050000 ctx

GPT-6 Luna is OpenAI's latest mid-tier text generation model, designed for developers who need speed and cost efficiency without sacrificing too much capability. It slots below GPT-6 Sol in the GPT-6 lineup, making it ideal for high-volume applications like chatbots, content classification, and lightweight agentic workflows. With a 1050000-token context window, Luna handles long inputs effectively while keeping latency low, which is crucial for responsive user experiences. The model is optimized for seamless API integration, supporting common frameworks and tools, so teams can deploy it quickly in production environments. Compared to its bigger sibling, Luna trades some advanced reasoning for faster inference and lower costs, making it a smart choice for scalable, budget-conscious projects where throughput matters more than peak performance.

text generationAPI

gpt-6-luna-pro:batch

openai
1050000 ctx

GPT-6 Luna Pro: Batch is a specialized high-throughput endpoint designed for developers who need deep reasoning capabilities without the latency overhead of real-time interaction. Built on the GPT-6 Luna architecture, this version leverages the 'pro' reasoning mode, which optimizes the model's internal chain-of-thought processes for complex logic, advanced mathematical derivation, and intricate code synthesis. Unlike standard chat completions, the batch implementation is specifically engineered for asynchronous workloads where cost-efficiency and high-quality reasoning are prioritized over immediate response times. It is an ideal choice for processing large-scale datasets, automated code auditing, or long-form document analysis where the depth of thought is critical. Integrating this model is straightforward for those already utilizing the OpenAI API ecosystem, offering a seamless transition for scaling reasoning-heavy pipelines via batch processing workflows.

text generationAPI

gpt-6-luna-pro

openai
1050000 ctx

GPT-6 Luna Pro is a specialized iteration of the Luna architecture, optimized specifically for high-stakes reasoning workflows. While the standard Luna model handles general-purpose tasks, this version utilizes a dedicated 'pro' reasoning mode to navigate multi-step logic, complex mathematical proofs, and deep architectural planning. For developers, this means a significant reduction in logical hallucinations when the model is tasked with debugging intricate codebases or synthesizing vast amounts of technical documentation. It supports a massive 1.05 million token context window, making it an ideal engine for long-form codebase analysis and RAG-heavy applications. Integration is seamless via standard OpenAI-compatible API endpoints, requiring only a parameter adjustment to toggle the enhanced reasoning capabilities. If your application demands more than just pattern matching—specifically deep cognitive processing—this model serves as a high-precision alternative to standard frontier models.

text generationAPI

grok-4.7

x-ai
500000 ctx

Grok-4.7 is xAI's latest language model tuned for developers who need reliable code generation and agentic automation. It builds on Grok 4.6 with better context handling (up to 500K tokens), improved self-verification, and stronger performance on long-running engineering tasks like test-driven development, debugging, and multi-step refactoring. The model excels at understanding large codebases, generating accurate unit tests, and maintaining consistency across extended conversations—useful for both interactive pair-programming and autonomous agent workflows. It supports standard APIs (OpenAI-compatible endpoints), making integration straightforward for most modern toolchains. While not open-source, its permissive API access and competitive pricing make it accessible for commercial and hobbyist projects alike. In benchmarks, Grok-4.7 performs well against other leading models like GPT-4o and Claude 3.5 in coding-specific tasks, especially when iterative validation and context persistence matter. If you're building AI-powered dev tools or automating complex software pipelines, it's worth evaluating for its depth of context and self-checking capabilities.

text generationAPI

mimo-v2.6-pro

xiaomi
1050000 ctx

MiMo-V2.6-Pro is Xiaomi's latest large-scale text generation model, built with over 1 trillion parameters and a context window of 1,048,576 tokens. It's designed for developers working on complex tasks like long-document understanding, multilingual content creation, and code generation across multiple programming languages. The model supports standard APIs, making it easy to integrate into existing applications or agent-based workflows. While it competes with other top-tier models in terms of scale and performance, MiMo-V2.6-Pro stands out with its extremely long context handling, which is useful for processing entire books, technical documents, or lengthy codebases in a single pass. It also offers solid reasoning capabilities and supports both Chinese and English, making it a strong choice for bilingual or Asia-focused development teams. The model is currently available via API under a commercial-friendly license, so developers can experiment or deploy without legal hurdles.

text generationAPI

mimo-v2.6-flash

xiaomi
1048576 ctx

MiMo-V2.6-Flash is Xiaomi's latest open-source text generation model, built for developer flexibility. With 309B total parameters and 15B activated per token via a Mixture-of-Experts setup, it balances strong performance with efficient inference. The hybrid attention mechanism supports long contexts—up to 1M tokens—making it suitable for complex document analysis, code generation, and multilingual tasks. As an open-weight model under an API license, it integrates well with existing toolchains and allows customization for specific use cases. While not the largest model available, its architecture prioritizes practical efficiency, offering competitive results without excessive computational overhead. Ideal for developers looking to build or extend applications with a locally runnable, high-capacity language model.

text generationAPI

mimo-v2.6-pro-ultraspeed

xiaomi
1048576 ctx

MiMo-V2.6-Pro-UltraSpeed is Xiaomi's optimized inference variant of the 1T-parameter MiMo-V2.6-Pro foundation model. Unlike typical fast-distilled variants that sacrifice quality, UltraSpeed retains the full model's generation quality while targeting significantly faster throughput for latency-sensitive applications. It's positioned for developers building real-time AI services where response speed directly impacts user experience - chatbots, interactive coding assistants, streaming content generation, and edge inference scenarios. The model supports a 1M token context window and uses Xiaomi's API-based licensing, making it accessible via standard HTTP endpoints with pay-per-use pricing. Integration follows familiar OpenAI-compatible patterns, so existing toolchains mostly work with minimal adapter changes. Compared to other speed-focused models in the Chinese ecosystem (like ByteDance's faster variants or Alibaba's turbo editions), UltraSpeed differentiates through its quality-speed balance - it doesn't aggressively compress the model, instead optimizing the inference pipeline and token sampling. This makes it a pragmatic choice when you need both fast responses and coherent long-form outputs, particularly for Chinese-language applications where it shows strong performance. The tradeoff is vendor lock-in through Xiaomi's API rather than open-weights distribution.

text generationAPI

gpt-oss

Ollama
Model

gpt-oss is a text generation model available through Ollama's local inference library. It's designed for developers who want to run models on their own hardware without relying on external APIs. The model fits into Ollama's ecosystem, meaning you can pull it directly with a simple command and integrate it into existing workflows that already use Ollama's tooling. Since it's local, you get control over data privacy and can avoid network latency for certain use cases. Check the official Ollama page for the latest tags, model size, and license details before pulling, as these can change. For developers building prototypes, experimenting with prompts, or embedding generation tasks, gpt-oss offers a straightforward way to test locally without cloud dependencies. It's best suited for lighter text generation tasks where you value having the model close to your code and data.

text generationSee Ollama library

gpt-3.5-turbo:batch

openai
16385 ctx

For developers managing high-volume asynchronous workloads, the GPT-3.5-turbo:batch endpoint offers a cost-effective way to process large datasets without blocking real-time application threads. While the standard turbo model is optimized for low-latency chat interactions, the batch variant is specifically architected for non-urgent tasks where throughput matters more than immediate response. It excels at bulk text transformations, large-scale data labeling, and synthetic dataset generation. By decoupling the request from the immediate response cycle, you can significantly reduce API costs while maintaining high reliability for background jobs. Compared to real-time inference, this is your go-to tool for offline processing pipelines where you need to scale up text generation or summarization tasks across millions of tokens without hitting strict rate limits or paying premium latency prices.

text generationAPI

gpt-3.5-turbo

openai
16385 ctx

GPT-3.5 Turbo is OpenAI's optimized model for conversational AI and general text generation tasks. It excels at understanding and generating natural language, making it ideal for chatbots, content creation, and code assistance. With a context window of 16,385 tokens, it handles moderately complex prompts efficiently. Training data up to September 2021 means it's suitable for tasks not requiring the very latest information. The model is accessible via API, allowing easy integration into applications. Compared to larger models like GPT-4, it's faster and more cost-effective for many use cases, though it may lack some advanced reasoning capabilities. It supports multiple programming languages and can assist with debugging, documentation, and code generation. Its strength lies in balancing performance, speed, and cost, making it a practical choice for developers building interactive applications or automating text-based workflows.

text generationAPI

gpt-3.5-turbo-16k

openai
16385 ctx

GPT-3.5-turbo-16k is OpenAI's expanded-context variant of the popular GPT-3.5-turbo model, offering a 16,385-token context window—roughly four times that of its predecessor. This makes it well-suited for handling longer documents, multi-turn conversations, or complex prompts that require retaining more information in a single request. While it comes at a higher cost per token, it eliminates the need to chunk or summarize input, which can improve performance in tasks like code analysis, document summarization, and extended dialogue systems. It supports the same instruction-following and text-generation capabilities as GPT-3.5-turbo and integrates seamlessly via the OpenAI API, making it easy to swap in for existing applications. Compared to GPT-4, it's faster and more cost-effective while still delivering solid performance, though with lower accuracy and reasoning depth. It's a practical choice for developers who need more context without jumping to larger, pricier models.

text generationAPI

gpt-3.5-turbo-instruct

openai
4095 ctx

GPT-3.5-turbo-instruct is a variant of the GPT-3.5 Turbo model, specifically tuned for instructional prompts and text generation tasks. Unlike its chat-optimized counterpart, this model focuses on processing and generating coherent, context-aware text based on detailed instructions. It's well-suited for developers building applications that require robust natural language understanding and generation, such as content creation tools, code generation assistants, or automated documentation systems. With a context window of 4095 tokens and training data up to September 2021, it offers reliable performance for tasks within its knowledge scope. Integration is straightforward via the OpenAI API, making it accessible for developers looking to leverage powerful language capabilities without the overhead of chat-specific formatting. Compared to other models in its class, it strikes a balance between performance and efficiency, ideal for instruction-following use cases where conversational overhead isn't needed. Whether you're generating structured text, summarizing documents, or building domain-specific tools, this model provides a solid foundation with predictable behavior and strong developer support.

text generationAPI

gpt-3.5-turbo-0613

openai
4095 ctx

GPT-3.5 Turbo (0613) is OpenAI's optimized chat model, built for fast, reliable conversational AI. It handles natural language and code generation with solid accuracy, making it a practical choice for chat assistants, content drafting, and lightweight coding tasks. With a 4095-token context and training data up to September 2021, it balances performance and efficiency for most general-purpose applications. It's well-suited for integration via the OpenAI API, supporting developers who need a responsive, cost-effective model without heavy computational overhead. While newer models offer improved reasoning and larger contexts, GPT-3.5 Turbo remains a strong option for prototyping and scaling chat-based experiences quickly.

text generationAPI

gemini-2.5-pro-preview

google
1048576 ctx

Gemini 2.5 Pro is Google's latest reasoning-focused model built for complex problem solving. It handles large contexts (up to 1M tokens), making it strong for deep document analysis, long conversations, and codebases. The model supports multimodal input, so you can feed it text and images together for tasks like diagram understanding or visual QA. It's accessible via the Gemini API with client libraries for common languages, fitting into existing developer workflows. In benchmarks, it performs well on coding (SWE-bench-style tasks), math, and STEM reasoning, often trailing only top-tier closed models. It's a solid choice when you need strong accuracy on technical workloads without managing your own infrastructure.

text generationAPI

gemini-2.5-pro:batch

google
1048576 ctx

Gemini 2.5 Pro: Batch is a high-throughput iteration of Google’s flagship reasoning model, specifically optimized for large-scale asynchronous processing. Unlike standard real-time endpoints, this version is engineered for developers running massive workloads where latency is secondary to cost-efficiency and massive context utilization. The model features an expansive 1M token context window, making it a powerhouse for analyzing entire codebases, long-form technical documentation, or massive datasets in a single pass. Its core strength lies in its 'thinking' architecture, which provides a significant edge in complex logical reasoning, multi-step mathematical proofs, and advanced debugging tasks. For developers building automated data pipelines, large-scale content synthesis tools, or deep architectural analysis agents, this batch model offers a scalable way to leverage state-of-the-art intelligence without the premium overhead of synchronous API calls. It integrates seamlessly into existing Google Cloud workflows, making it a logical choice for heavy-duty backend processing.

text generationAPI

gemini-2.5-pro

google
1048576 ctx

Gemini 2.5 Pro represents a significant shift toward agentic reasoning in large language models. Unlike standard LLMs that predict the next token linearly, this model incorporates an explicit 'thinking' phase, allowing it to process complex logic, multi-step mathematical proofs, and intricate debugging tasks with much higher reliability. For developers, the standout feature is the massive 1M+ token context window, which makes it a powerhouse for codebase analysis, long-document retrieval, and processing entire video files in a single prompt. While previous iterations excelled at general chat, the 2.5 Pro architecture is specifically tuned for high-precision technical workflows. It integrates seamlessly via Google's existing API ecosystem, making it a direct competitor to specialized reasoning models. If your use case involves deep architectural reasoning or massive data ingestion rather than simple text completion, this is the model to prioritize in your stack.

text generationAPI

gemini-2.5-flash:batch

google
1048576 ctx

Gemini 2.5 Flash:Batch is a specialized deployment of Google’s high-throughput model, engineered specifically for large-scale asynchronous processing. Unlike standard real-time endpoints, this batch version is optimized for high-volume workloads where latency is secondary to cost-efficiency and massive throughput. It integrates advanced reasoning and 'thinking' steps directly into its architecture, making it particularly effective for complex logic, code refactoring, and mathematical verification at scale. For developers, this means you can offload heavy computational tasks—such as processing massive datasets, large-scale document analysis, or bulk code audits—without the overhead of per-request latency constraints. It maintains the same 1M token context window as its real-time counterparts, allowing for deep reasoning across vast amounts of data. If your workflow involves periodic, high-volume data transformations or deep analysis of large repositories, this model provides a superior balance of intelligence and operational economy.

text generationAPI

gemini-2.5-flash

google
1048576 ctx

Gemini 2.5 Flash is engineered as a high-throughput, low-latency workhorse for developers requiring a balance of speed and deep reasoning. Unlike standard lightweight models that often sacrifice logic for velocity, this iteration introduces integrated 'thinking' processes, making it significantly more reliable for complex code generation, mathematical derivation, and multi-step scientific reasoning. With a massive 1M token context window, it excels at processing massive codebases, long-form documentation, or extensive datasets in a single pass. For integration, it fits seamlessly into existing Google Cloud and Vertex AI workflows, offering a scalable solution for real-time agentic workflows where reasoning depth is non-negotiable but latency must remain minimal. If your use case involves autonomous debugging, complex data extraction, or building sophisticated RAG pipelines, this model provides the computational density needed without the overhead of larger frontier models.

text generationAPI

gemini-2.5-flash-lite:batch

google
1048576 ctx

For developers building high-volume, latency-sensitive applications, Gemini 2.5 Flash-Lite represents a strategic shift toward extreme efficiency without sacrificing core reasoning capabilities. Unlike larger flagship models that prioritize deep nuance, this 'lite' iteration is specifically engineered for high-throughput workloads where cost-per-token and response speed are the primary constraints. It excels in scenarios like real-time data extraction, high-frequency classification, and large-scale summarization tasks that would be prohibitively expensive or slow on heavier architectures. The 'batch' optimization suggests it is particularly well-suited for asynchronous processing pipelines where you need to ingest massive datasets and receive structured outputs at scale. While you might trade off some complex multi-step logical depth found in the Pro series, the trade-off is a massive gain in operational velocity and significantly lower overhead for production-grade agentic workflows.

text generationAPI

gemini-2.5-flash-lite

google
1048576 ctx

Gemini 2.5 Flash-Lite is engineered for developers prioritizing high-frequency, low-latency applications where cost-per-token is a critical constraint. While larger models in the Gemini family handle complex, multi-step reasoning, Flash-Lite is purpose-built for speed and high throughput. It excels in real-time scenarios such as conversational agents, real-time data extraction, and high-volume classification tasks. For integration, it maintains the standard Gemini API ecosystem, allowing for seamless transitions from prototyping on Pro models to production deployment on Lite. Compared to previous lightweight iterations, this model offers a superior balance of reasoning capabilities and token generation speed, making it an ideal choice for edge-case logic within massive-scale pipelines without the overhead of a heavy-duty LLM.

text generationAPI

gpt-oss-20b

openai
131072 ctx

For developers seeking a balance between high-performance reasoning and deployment flexibility, gpt-oss-20b offers a compelling middle ground. Built on a Mixture-of-Experts (MoE) architecture, this 21B parameter model utilizes only 3.6B active parameters per token, significantly reducing inference latency and compute overhead without sacrificing the depth of a larger dense model. The Apache 2.0 license makes it an ideal candidate for commercial applications where data sovereignty and local hosting are priorities. With a massive 131k context window, it excels at long-form document analysis, complex codebase reasoning, and multi-turn conversational agents. Compared to standard dense models of similar size, you'll notice much higher throughput, making it particularly effective for scaling RAG pipelines or real-time agentic workflows where cost-per-token and speed are critical constraints.

text generationAPI

gpt-oss-120b:batch

openai
131072 ctx

For developers building complex autonomous systems, gpt-oss-120b:batch introduces a high-efficiency Mixture-of-Experts (MoE) architecture that balances massive scale with low-latency execution. With 117B total parameters but only 5.1B activated per token, it offers the reasoning depth of a large-scale model while maintaining the throughput necessary for production-grade agentic workflows. This model is specifically tuned for high-context reasoning and multi-step task decomposition, making it ideal for RAG pipelines, automated code generation, and complex decision-making agents. Unlike monolithic dense models that incur heavy compute costs, this MoE approach allows for cost-effective scaling in batch processing environments. It is designed to integrate seamlessly into existing API-driven infrastructures, providing a robust backbone for developers who need high-intelligence outputs without the traditional latency penalties of massive dense architectures.

text generationAPI

gpt-oss-120b

openai
131072 ctx

gpt-oss-120b is a high-density Mixture-of-Experts (MoE) model designed to bridge the gap between massive parameter counts and production-grade inference efficiency. By activating only 5.1B parameters per token, it offers a unique value proposition: the reasoning capabilities of a large-scale model with the low latency typically associated with much smaller architectures. For developers, this means improved throughput for agentic workflows and complex multi-step reasoning tasks without the massive compute overhead of dense 100B+ models. With a 131k context window, it is well-suited for long-form document analysis, RAG pipelines, and maintaining state in complex autonomous agents. Unlike standard dense models, its MoE structure allows for specialized knowledge retrieval during the forward pass, making it particularly effective for coding, mathematical reasoning, and structured data extraction. It is built for integration into existing API-driven stacks where high reliability and reasoning depth are non-negotiable.

text generationAPI

gpt-5-nano:batch

openai
400000 ctx

For developers building latency-sensitive applications, gpt-5-nano:batch represents a strategic shift toward high-throughput, low-latency inference. While it lacks the deep multi-step reasoning capabilities of the flagship GPT-5 models, it is purpose-built for high-volume tasks where speed and cost-efficiency are the primary constraints. This model excels in real-time text processing, autocomplete features, and rapid classification tasks within high-concurrency environments. With a 400,000 token context window, it maintains a surprisingly large receptive field for its size, making it viable for processing long documentation snippets or large batches of structured data. Integration is straightforward via standard API endpoints, making it an ideal candidate for edge-case logic, data preprocessing pipelines, or as a 'routing' layer to determine if a query requires a more computationally expensive model. If your workflow prioritizes millisecond response times over complex logical deduction, this is your primary workhorse.

text generationAPI
Email