Global AI chat room · 12 online now Join now
DIRECTORY / 02

AI Models | Open-Source LLM Directory

Discover and compare open-source LLMs, language models and multimodal models by capability, scale, license, downloads and provenance.

Compare modelsFind the right building block for your next workflow
Directory overview
839
curated entries
24 topic groupsLive
02 / MODEL INDEX

Find the right model for the job

Context first, better decisions. Every entry keeps the signal that matters.

CURATED DIRECTORY839 results

flan t5 small

google
Model

Flan-T5 Small is a lightweight, encoder-decoder model designed for efficient text-to-text generation. Unlike the base T5, the Flan version is instruction-tuned, meaning it performs significantly better on zero-shot tasks without requiring extensive fine-tuning. For developers, this model is an ideal choice for low-latency applications or edge deployment where memory is constrained. It excels at focused NLP tasks such as classification, basic summarization, and question answering. While it lacks the reasoning depth of larger LLMs, its small footprint makes it an excellent candidate for distillation targets or as a specialized component within a larger modular pipeline via the Hugging Face Transformers library.

text2text-generationapache-2.0

CLIP ViT L 14 laion2B s32B b82K

laion
Model

The CLIP ViT-L/14 (laion2B s32B b82K) is a high-performance vision-language model optimized for cross-modal retrieval and zero-shot classification. By leveraging a Vision Transformer (ViT) backbone and training on a massive, filtered subset of the LAION-2B dataset, it creates a shared embedding space where images and text are mathematically aligned. For developers, this means you can perform semantic image searches or categorize visual data using natural language queries without needing labeled training sets. It serves as a robust foundation for building image search engines, automated tagging systems, or as a visual encoder for generative AI pipelines. Compared to smaller CLIP variants, the L/14 architecture offers a superior balance of granularity and inference speed, making it suitable for production-grade retrieval tasks.

image-text-retrievalmit

paraphrase MiniLM L12 v2

sentence-transformers
Model

The paraphrase-MiniLM-L12-v2 is a lightweight, high-performance sentence transformer optimized for mapping sentences to a dense vector space. Unlike general-purpose LLMs, this model is specifically tuned for semantic textual similarity (STS), making it an ideal choice for developers building RAG pipelines, semantic search engines, or clustering systems where latency and resource overhead are critical constraints. It strikes a strong balance between embedding quality and inference speed, delivering performance comparable to much larger models while remaining small enough to deploy on edge devices or CPU-only environments. Integration is straightforward via the sentence-transformers library, allowing for rapid vectorization of large datasets without requiring massive GPU clusters.

text2text-generationapache-2.0

paraphrase multilingual MiniLM L12 v2 onnx Q

Qdrant
Model

The paraphrase-multilingual-MiniLM-L12-v2 (ONNX quantized) is a lightweight, high-efficiency sentence transformer designed for cross-lingual semantic similarity tasks. Unlike large generative models, this model focuses on mapping text from over 100 languages into a shared vector space, making it ideal for clustering, semantic search, and duplicate detection across different languages. The ONNX quantization significantly reduces the memory footprint and latency, allowing for high-throughput deployment on CPUs without requiring heavy GPU resources. For developers, this means an easy integration path for RAG pipelines or multilingual chatbots where low-latency embedding generation is critical.

text2text-generationapache-2.0

st polish paraphrase from mpnet

sdadas
Model

The 'st polish paraphrase from mpnet' is a specialized text-to-text model designed to refine and rewrite Polish text while maintaining original semantic meaning. Built upon the MPNet architecture, it leverages sentence-transformer embeddings to ensure high-quality paraphrasing that avoids the common pitfalls of generic translation models. For developers, this is an ideal utility for augmenting NLP datasets, removing redundancy in user-generated content, or implementing 'rewrite' features in localization pipelines. It integrates easily into existing Python-based LLM workflows, offering a lightweight alternative to massive generative models when the goal is precise stylistic polishing rather than creative generation.

text2text-generationlgpl

ImageTextRetrieval

ohgnues
Model

ImageTextRetrieval is a specialized model designed for cross-modal alignment, allowing developers to perform efficient semantic searches across image and text datasets. Unlike standard classification models, this architecture maps both visual and textual inputs into a shared embedding space. This makes it ideal for building reverse image search engines, automated tagging systems, or content discovery tools where natural language queries must retrieve relevant visual assets. It integrates easily into RAG (Retrieval-Augmented Generation) pipelines by serving as the encoder for vector databases, offering a lightweight alternative to massive multimodal LLMs when the primary goal is retrieval speed and precision rather than generative output.

image-text-retrievalApache-2.0

gpt-6.1-sol:batch

openai
1050000 ctx

Introducing GPT-6.1 Sol, the latest iteration in the GPT-6 series from OpenAI, designed to excel in agentic coding and document-heavy professional tasks. Positioned below the flagship GPT-6 Astra, this model offers a robust set of capabilities for developers seeking advanced text generation. With 1050000 context tokens and API licensing, GPT-6.1 Sol is tailored for seamless integration into various development environments, making it a versatile tool for international developers.

text generationAPI

gpt-6.1-sol

openai
1050000 ctx

GPT-6.1 Sol is an upgrade to GPT-6 Sol from OpenAI, positioned below the flagship GPT-6 Astra in the GPT-6 series. It is suited for agentic coding, computer use, document-heavy professional...

text generationAPI

gpt-6.1-sol-pro:batch

openai
1050000 ctx

GPT-6.1 Sol Pro is the same underlying model as [GPT-6.1 Sol](https://openrouter.ai/openai/gpt-6.1-sol), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. **Cost note:** pro mode spends far more...

text generationAPI

gpt-6.1-sol-pro

openai
1050000 ctx

GPT-6.1 Sol Pro is the same underlying model as [GPT-6.1 Sol](https://openrouter.ai/openai/gpt-6.1-sol), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. **Cost note:** pro mode spends far more...

text generationAPI

jev-router

typesafe
1000000 ctx

For developers managing large-scale LLM deployments, the primary challenge isn't just model intelligence, but the trade-off between latency, reasoning depth, and API costs. Jev-router addresses this by acting as an intelligent orchestration layer sitting atop the TypeSafe System One model. Instead of routing every prompt to a heavy, expensive reasoning model, it dynamically analyzes incoming requests to determine the optimal balance of reasoning effort and model capacity required for the task. This makes it particularly effective for complex agentic workflows where some steps require deep logic while others only need rapid instruction following. Integration is straightforward via the Jev API, allowing you to offload the complexity of model selection to the router. Compared to static routing solutions, Jev-router provides a more fluid, cost-efficient way to maintain high output quality without overpaying for unnecessary compute on simple queries.

text generationAPI

phi4-mini

Ollama
Model

phi4-mini is a compact, high-efficiency language model designed for developers who need to balance reasoning capabilities with low-latency local execution. Unlike massive frontier models that require significant GPU clusters, this model is optimized for edge deployment and local inference via Ollama. It excels in structured text generation, logic-heavy tasks, and code assistance where a smaller footprint is a requirement rather than a limitation. For developers building privacy-first applications or working in resource-constrained environments, phi4-mini offers a pragmatic alternative to cloud-based APIs. It integrates seamlessly into existing local workflows, allowing for rapid prototyping and deployment of agentic loops without the overhead of massive parameter counts. While it may not match the broad world knowledge of its larger siblings, its strength lies in its high performance-per-parameter ratio, making it an ideal engine for specialized, task-oriented pipelines.

text generationSee Ollama library

mistral-large-2512

mistralai
262144 ctx

Mistral Large 2512 represents a significant architectural leap for developers needing high-reasoning capabilities without the latency overhead of dense monolithic models. Built on a sparse Mixture-of-Experts (MoE) framework, it utilizes 41B active parameters within a 675B total parameter structure, striking an efficient balance between raw intelligence and inference speed. For engineers, the most compelling aspect is its Apache 2.0 licensing, which provides much-needed flexibility for commercial deployment compared to closed-source competitors. The model excels in complex multilingual reasoning, advanced coding tasks, and structured data extraction. With a massive 262k context window, it is purpose-built for deep document analysis and long-form codebase comprehension. Whether you are integrating via API or optimizing for specific logic-heavy workflows, this model offers a high-performance alternative to GPT-4 class models while maintaining a more developer-friendly ecosystem.

text generationAPI

perceptron-mk1.5

perceptron
36864 ctx

Perceptron-mk1.5 is a multimodal reasoning engine specifically architected for embodied AI and physical agents. Unlike standard LLMs that treat vision as a secondary modality, this model is designed to bridge the gap between high-level semantic reasoning and spatial awareness. It processes interleaved text, image, video, and audio streams to drive decision-making in real-world environments. For developers, the standout feature is its ability to output not just natural language, but structured spatial annotations including bounding boxes, polygons, and temporal tracking data. This makes it a critical component for robotics, autonomous systems, and augmented reality applications where an agent must identify, locate, and track objects across time. It integrates via API, providing a scalable way to add complex spatial intelligence to existing hardware stacks without the overhead of training custom vision-language models from scratch.

text generationAPI

glm-5.3-prime

z-ai
1000000 ctx

For developers building latency-sensitive applications, GLM-5.3-Prime offers a strategic middle ground between raw intelligence and execution speed. While maintaining the core reasoning capabilities of the standard GLM-5.3 architecture, this 'Prime' variant is specifically optimized for high-throughput inference. We are seeing 1.5x to 2x improvements in tokens per second, making it a viable candidate for real-time agentic workflows, high-volume chat interfaces, and automated content pipelines where response lag is a dealbreaker. The model supports a massive 1M-token context window, allowing you to ingest entire codebases or extensive documentation without losing coherence. Unlike standard models that might throttle during peak demand, the Prime architecture is engineered to maintain consistent velocity. If your stack requires deep semantic understanding but demands rapid-fire output for seamless user experiences, this is the model to integrate into your production environment.

text generationAPI

ember-1

fireworks
1048576 ctx

Ember-1 is a specialized reasoning model optimized for high-efficiency inference. Built on the Kimi K3 architecture, it addresses a common pain point in Large Reasoning Models (LRMs): the high latency and cost associated with long-form 'Chain of Thought' traces. By engineering more concise reasoning paths, Ember-1 achieves a significant reduction in token overhead—using approximately 40% fewer tokens for internal logic compared to standard reasoning models—without sacrificing the depth of its final output. For developers, this translates to faster time-to-first-token (TTFT) and lower API costs, making it a pragmatic choice for complex agentic workflows, multi-step logical deduction, and automated code reasoning. It bridges the gap between heavy-duty reasoning capabilities and the operational requirements of production-grade applications where token economy is critical.

text generationAPI

command-a-plus

cohere
192000 ctx

Command A+ is Cohere's latest high-performance model specifically architected for agentic workflows and complex enterprise automation. Unlike general-purpose chat models, this model is optimized for reliability in tool-use scenarios, supporting strict schema enforcement to minimize parsing errors during function calling. With a massive 192K context window, it can ingest extensive documentation, long-form codebase structures, or multi-modal inputs involving both text and images without losing coherence. For developers building autonomous agents or RAG-based systems, the primary advantage lies in its precision with structured data outputs and its ability to maintain state across deep reasoning chains. It bridges the gap between simple text generation and robust, production-ready orchestration, making it a strong candidate for integration into existing CI/CD pipelines, automated customer support systems, or complex data extraction workflows where accuracy is non-negotiable.

text generationAPI

solar-mini4

upstage
524288 ctx

Solar-mini4 is a highly optimized Mixture-of-Experts (MoE) model designed for developers who need to balance high-performance reasoning with low-latency execution. While it carries a 35B parameter footprint, its architecture utilizes only 3B active parameters per token, making it significantly faster and more cost-effective than dense models of similar scale. The standout feature is the massive 524K context window, which allows for deep document analysis, large-scale codebase ingestion, and long-form conversation memory without the typical performance degradation seen in smaller models. For engineers building autonomous agents, Solar-mini4 offers the throughput necessary for rapid tool-calling and iterative reasoning loops. It serves as an ideal middle ground for those who find 7B models too limited for complex instruction following, but find 70B+ models too slow or expensive for real-time production environments.

text generationAPI

aion-3.5

aion-labs
262144 ctx

Aion-3.5 is a specialized multi-model architecture designed specifically for complex narrative generation and roleplaying workflows. Unlike monolithic LLMs that struggle with character consistency, Aion-3.5 leverages a collaborative generation process rooted in the GLM family. It utilizes multiple specialized sub-models to handle different facets of storytelling, which significantly reduces the 'drift' often seen in long-form creative writing. For developers, the primary value lies in its massive 262k context window, making it highly capable of maintaining deep world-building lore and long-term character memory. While general-purpose models excel at instruction following, Aion-3.5 is optimized for nuanced dialogue, emotional intelligence, and maintaining persona stability. Integration is handled via API, making it a plug-and-play solution for developers building interactive fiction engines, sophisticated NPCs for gaming, or automated creative writing assistants.

text generationAPI

aion-3.5-mini

aion-labs
262144 ctx

Aion-3.5-mini is a specialized lightweight model engineered for high-fidelity narrative generation and complex roleplaying scenarios. Built upon the GLM architecture, it prioritizes character consistency and stylistic nuance, making it an ideal choice for developers building interactive fiction, NPC engines, or immersive storytelling platforms. While it functions as a cost-effective alternative to the larger Aion 3.5 flagship, it retains a massive 262k context window, allowing for long-form continuity without the typical memory decay seen in smaller models. For developers, this means you can maintain deep world-building data and extensive dialogue histories within a single prompt. Integration is straightforward via API, offering a high throughput-to-latency ratio that is critical for real-time applications. If your use case requires sophisticated persona management rather than just raw logic or code completion, this model provides a highly optimized balance of performance and operational economy.

text generationAPI

space-bunny-alpha

stealth
1000000 ctx

Space-bunny-alpha is a high-throughput multimodal model designed for developers who require a balance of low-latency inference and deep reasoning. Unlike standard LLMs that struggle with large-scale context, this model natively supports a 1M-token window, making it ideal for codebase analysis, long-document processing, and complex RAG pipelines. A standout feature is the adjustable reasoning effort, allowing you to toggle between rapid-fire responses for simple tasks and compute-intensive logic for complex debugging or architectural planning. It handles text and visual inputs seamlessly, providing a unified interface for multimodal applications. Whether you are building autonomous agents or integrating sophisticated coding assistants, the model's architecture prioritizes speed without sacrificing the logical depth required for production-grade software engineering.

text generationAPI

qwen3.8-max-prime

qwen
1000000 ctx

Qwen3.8-Max-Prime is a high-throughput optimization of the flagship Qwen3.8 Max model, specifically engineered for production environments where latency and scale are critical. While the standard Max model focuses on raw reasoning depth, the Prime variant is architected to handle higher request volumes and larger concurrent workloads without the typical performance degradation seen in dense models. It is a natively multimodal engine, capable of processing text, image, and video inputs within a massive 1-million-token context window. For developers, this means you can build sophisticated agents that reason over long-form video content or massive codebases with much higher reliability in high-traffic applications. Compared to standard API offerings, the Prime SKU prioritizes consistent throughput, making it an ideal choice for real-time multimodal RAG pipelines and complex automated reasoning workflows where speed-to-inference is just as vital as intelligence.

text generationAPI

gpt-oss-20b:batch

openai
131072 ctx

gpt-oss-20b:batch is OpenAI's open-weight 21B parameter text generation model under Apache 2.0. It uses a Mixture-of-Experts architecture with 3.6B active parameters per forward pass, making it efficient for batch processing tasks. The model supports 131,072 token context length, suitable for long document analysis and generation. Developers can integrate it directly into existing pipelines via standard transformer libraries or run it locally without relying on external APIs. Compared to other open models, it balances performance and resource efficiency, though it requires more compute than smaller distilled variants. It's particularly useful for content generation, code synthesis, and multilingual applications where customization and data privacy matter. The batch optimization makes it cost-effective for large-scale inference workloads.

text generationAPI

nex-n2.5-pro

nex-agi
262144 ctx

Nex-N2.5-pro is a specialized agentic model engineered for autonomous software engineering tasks. Unlike standard LLMs that primarily predict text, this model is optimized for closed-loop execution, specifically targeting multi-file repository manipulation and codebase exploration. It operates through a visual feedback loop, allowing it to observe execution errors or UI changes and iteratively self-correct its implementation. For developers, this means moving beyond simple snippet generation toward delegating complex, goal-oriented refactoring or feature implementation. It handles high-context environments with a 262k token window, making it suitable for deep integration into CI/CD pipelines or as the core engine for autonomous coding agents. While general-purpose models excel at chat, Nex-N2.5-pro is built for the 'plan-act-verify' cycle required in professional production environments.

text generationAPI
Email