Global AI chat room · 17 online now Join now
DIRECTORY / 02

AI Models | Open-Source LLM Directory

Discover and compare open-source LLMs, language models and multimodal models by capability, scale, license, downloads and provenance.

Compare modelsFind the right building block for your next workflow
Directory overview
15
curated entries
24 topic groupsLive
02 / MODEL INDEX

Find the right model for the job

Context first, better decisions. Every entry keeps the signal that matters.

CURATED DIRECTORY15 results

Whisper Large V3

OpenAI
1.5B

Whisper Large V3 represents the latest evolution in OpenAI's open-source speech-to-text lineage, optimized for high-fidelity transcription across diverse linguistic landscapes. For developers, the primary value proposition lies in its massive scale—1.5B parameters—which provides superior robustness against background noise and varying accents compared to previous iterations. Unlike many proprietary APIs, the MIT license allows for deep local integration and fine-tuning within private infrastructure, making it ideal for privacy-sensitive applications. It excels in multi-lingual transcription and translation tasks, offering a reliable foundation for building automated captioning, meeting assistants, or voice-command interfaces. While it requires more compute overhead than the 'base' or 'small' variants, the trade-off is a significant reduction in Word Error Rate (WER) for complex audio environments. It is best utilized in pipelines requiring high-accuracy long-form transcription where latency is secondary to precision.

automatic speech recognitionMIT
5.5K starsView details

Whisper Small

OpenAI
244M

Whisper Small is a streamlined version of OpenAI’s robust speech-to-text architecture, optimized specifically for developers targeting edge computing and low-latency environments. While larger iterations of Whisper prioritize absolute accuracy at the cost of massive VRAM requirements, the Small model strikes a pragmatic balance by utilizing 244M parameters. This makes it viable for deployment on consumer-grade hardware or mobile-adjacent devices without sacrificing significant word error rate (WER) performance. For engineers building real-time transcription services, voice assistants, or automated captioning tools, this model offers a high throughput-to-accuracy ratio. It integrates seamlessly into existing Python-based ML pipelines and supports a wide array of multilingual tasks. If your use case requires local inference where cloud API latency or data privacy is a concern, Whisper Small serves as an efficient middle ground between the lightweight 'Base' model and the heavy 'Large' variants.

automatic speech recognitionMIT
3.8K starsView details

speaker-diarization-3.1

pyannote
Not specified

Speaker Diarization 3.1, powered by pyannote, is a specialized framework designed to solve the 'who spoke when' problem in audio processing. Unlike standard ASR which only provides text, this model identifies distinct speaker identities and maps their timestamps across a recording. For developers, this is critical for building automated meeting minutes, multi-party interview transcripts, or voice-activated analytics. It integrates efficiently into speech pipelines, acting as a pre-processing or parallel layer to transcription engines. Compared to basic clustering methods, version 3.1 offers improved precision in speaker change detection and overlap handling, making it a robust choice for noisy, real-world audio environments where speaker turns are rapid.

automatic speech recognitionmit
3.7K starsView details

speaker-diarization-community-1

pyannote
Not specified

Speaker Diarization Community 1, powered by pyannote, is a specialized tool for the 'who spoke when' problem in audio processing. Unlike standard ASR that only transcribes text, this model partitions audio streams into segments based on speaker identity. It is particularly effective for multi-speaker environments such as podcasts, interviews, and meeting recordings where precise speaker attribution is required. For developers, it integrates well into speech-to-text pipelines to provide structured metadata, allowing for the creation of speaker-labeled transcripts. Compared to generic clustering methods, it offers a more robust framework for handling overlapping speech and varying acoustic conditions, operating under a permissive CC-BY-4.0 license for flexible deployment.

automatic speech recognitioncc-by-4.0
1.9K starsView details

Qwen3-ASR-1.7B

Qwen
Not specified

Qwen3 ASR 1.7B is a compact, efficient automatic speech recognition model designed for low-latency transcription and deployment in resource-constrained environments. At 1.7 billion parameters, it strikes a balance between computational overhead and accuracy, making it suitable for edge computing or as a specialized component in a larger voice-AI pipeline. Developers can leverage this model for real-time captioning, voice-command processing, and automated transcription services. Given its Apache-2.0 license, it offers significant flexibility for commercial integration. Compared to larger ASR models, Qwen3 ASR 1.7B prioritizes fast inference speeds and a smaller memory footprint without sacrificing the core robustness required for production-grade speech-to-text tasks.

automatic speech recognitionapache-2.0
1.1K starsView details

Voxtral-Mini-4B-Realtime-2602

mistralai
Not specified

Voxtral Mini 4B Realtime 2602 is a compact, low-latency speech-to-text model designed for high-throughput environments. Unlike larger ASR models that struggle with inference costs, this 4B parameter version optimizes for real-time streaming and edge deployment without sacrificing significant accuracy. It is particularly suited for developers building live transcription services, voice-driven interfaces, or accessibility tools where minimal lag is critical. Integration is streamlined via standard ASR pipelines, offering a lightweight alternative for those who need a balance between performance and resource consumption. Compared to heavier models, it reduces memory overhead while maintaining the robustness required for diverse acoustic environments.

automatic speech recognitionapache-2.0
986 starsView details

Qwen3-ASR-0.6B

Qwen
Not specified

Qwen3 ASR 0.6B is a compact, high-efficiency automatic speech recognition model designed for low-latency transcription tasks. At 0.6 billion parameters, it is optimized for edge deployment and resource-constrained environments where full-scale models are impractical. Developers can integrate this model into real-time voice pipelines, accessibility tools, or lightweight virtual assistants without sacrificing significant accuracy. Unlike larger ASR frameworks, it offers a lean memory footprint, making it an ideal candidate for on-device processing or high-throughput server-side scaling. Licensed under Apache-2.0, it provides the flexibility needed for both commercial integration and custom fine-tuning on domain-specific audio datasets.

automatic speech recognitionapache-2.0
342 starsView details

voice-activity-detection

pyannote
Not specified

Pyannote's Voice Activity Detection (VAD) is a specialized tool designed to distinguish human speech from silence or background noise in audio streams. For developers building speech-to-text pipelines or voice assistants, this model serves as a critical preprocessing layer to reduce computational overhead by filtering out non-speech segments before they hit heavier ASR engines. Unlike simple energy-based thresholds, this model handles complex acoustic environments more robustly, making it ideal for long-form audio transcription and speaker diarization workflows. It integrates easily into Python-based stacks and is released under the permissive MIT license, allowing for flexible commercial deployment and modification.

automatic speech recognitionmit
241 starsView details

whisperkit-coreml

argmaxinc
Not specified

WhisperKit CoreML brings OpenAI's Whisper speech-to-text capabilities directly to Apple silicon, optimizing inference for macOS, iOS, and iPadOS. Unlike generic wrappers, this implementation leverages CoreML to utilize the Neural Engine, significantly reducing CPU overhead and battery drain during transcription. For developers, this means the ability to implement high-accuracy, offline ASR (Automatic Speech Recognition) without relying on cloud APIs, ensuring user privacy and low latency. It is particularly effective for building real-time transcription tools, accessibility features, or voice-controlled interfaces where local execution is critical. Integration is streamlined for Swift environments, offering a performant alternative to PyTorch-based deployments on Apple hardware.

automatic speech recognitionmit
228 starsView details

mms-300m-1130-forced-aligner

MahmoudAshraf
Not specified

mms-300m-1130-forced-aligner is a automatic speech recognition model published on Hugging Face. It is primarily used with transformers and should be evaluated against the model card, license and deployment requirements before production use.

automatic speech recognitioncc-by-nc-4.0
103 starsView details

wav2vec2-large-xlsr-53-japanese

jonatasgrosman
Not specified

The wav2vec2-large-xlsr-53-japanese model is a robust automatic speech recognition (ASR) tool fine-tuned for Japanese audio. Built on Meta's cross-lingual XLSR framework, it leverages self-supervised pre-training across 53 languages to achieve high phonetic accuracy even with limited labeled Japanese data. For developers, this means a reliable solution for transcribing Japanese speech into text without needing to build a model from scratch. It integrates seamlessly with the Hugging Face Transformers library, making it easy to deploy in Python-based pipelines for applications like automated subtitling, voice command interfaces, or accessibility tools. Compared to general-purpose models, its specialized tuning for Japanese provides better handling of the language's specific acoustic properties.

automatic speech recognitionapache-2.0
87 starsView details

wav2vec2-large-xlsr-53-russian

jonatasgrosman
Not specified

The wav2vec2-large-xlsr-53-russian is a specialized automatic speech recognition (ASR) model based on Meta's cross-lingual wav2vec 2.0 architecture. Unlike general-purpose models, this version is fine-tuned specifically for the Russian language, making it highly effective for transcribing speech-to-text tasks where high linguistic precision is required. For developers, this model offers a robust alternative to proprietary APIs, allowing for local deployment and full control over data privacy. It integrates seamlessly into PyTorch and Hugging Face pipelines, making it straightforward to implement in voice-controlled applications, automated transcription services, or accessibility tools. While it requires more computational resources than distilled models, it provides superior accuracy for complex Russian phonetic structures compared to smaller, multilingual baselines.

automatic speech recognitionapache-2.0
77 starsView details

wav2vec2-large-xlsr-53-portuguese

jonatasgrosman
Not specified

The wav2vec2-large-xlsr-53-portuguese model is a robust Automatic Speech Recognition (ASR) tool fine-tuned specifically for the Portuguese language. Built upon Meta's cross-lingual wav2vec 2.0 framework, it leverages massive self-supervised pre-training across 53 languages to achieve high phonetic accuracy even with limited labeled Portuguese data. For developers, this means a reliable pipeline for converting Portuguese audio to text without needing to build a model from scratch. It integrates seamlessly into Hugging Face transformers pipelines, making it straightforward to deploy in transcription services, voice-command interfaces, or accessibility tools. Compared to generic multilingual models, this specialized version offers better word error rates (WER) for Portuguese dialects, providing a more precise output for production-grade NLP workflows.

automatic speech recognitionapache-2.0
58 starsView details

wav2vec2-large-xlsr-53-arabic

jonatasgrosman
Not specified

For developers building voice-driven applications in the MENA region, wav2vec2-large-xlsr-53-arabic offers a robust foundation for Automatic Speech Recognition (ASR). Built on the XLS-R architecture, this model leverages cross-lingual pre-training to handle the nuances of Arabic phonetics more effectively than standard monolingual models. It is specifically optimized for transcribing spoken Arabic into text, making it a primary candidate for voice assistants, automated captioning, and transcription services. Integration is straightforward via the Hugging Face Transformers library, allowing for seamless deployment in Python-based workflows. While it excels at capturing acoustic patterns, developers should note that performance may vary across different regional dialects. For production environments, we recommend fine-tuning the model on your specific domain-specific datasets to maximize word error rate (WER) improvements and ensure high accuracy in specialized technical or conversational contexts.

automatic speech recognitionapache-2.0
55 starsView details

faster-whisper-small

Systran
Not specified

For developers building real-time transcription services or batch processing pipelines, faster-whisper-small offers a high-efficiency alternative to the standard OpenAI Whisper implementation. By leveraging the CTranslate2 inference engine, this model achieves significantly lower latency and reduced memory footprints without sacrificing the core acoustic modeling capabilities of the original architecture. It is specifically optimized for CPU and GPU deployment where throughput is a critical KPI. While the 'small' parameter size is a strategic trade-off favoring speed and low-resource environments, it remains highly effective for clear audio in standard languages. Integrating this into your stack is straightforward via Hugging Face, making it an ideal candidate for edge computing, voice-command interfaces, or scalable microservices where cost-per-inference must be minimized.

automatic speech recognitionmit
52 starsView details
Email