Categories

AI Models Marketplace

Explore the latest and most popular AI models — LLM, vision, audio and more — to pick and integrate the right one.

whisper large v3 turbo

mit
Whisper large-v3-turbo is a streamlined version of OpenAI's state-of-the-art speech recognition model, engineered specif...
openai automatic-speech-recognition
7.9M 0

whisper base

apache-2.0
Whisper Base is a lightweight, open-source automatic speech recognition (ASR) model designed for efficient transcription...
openai automatic-speech-recognition
3.1M 0

whisper small

apache-2.0
Whisper Small is a versatile automatic speech recognition (ASR) model designed for developers who need a balance between...
openai automatic-speech-recognition
2.6M 0

wav2vec2 indonesian javanese sundanese

apache-2.0
This wav2vec2 model is a specialized ASR solution fine-tuned for the Indonesian linguistic landscape, providing speech-t...
indonesian-nlp automatic-speech-recognition
2.2M 0

Qwen3 ASR 1.7B

apache-2.0
Qwen3 ASR 1.7B is a compact, efficient automatic speech recognition model designed for low-latency transcription and dep...
Qwen automatic-speech-recognition
675.6K 162

Qwen3 ASR 0.6B

apache-2.0
Qwen3 ASR 0.6B is a compact, high-efficiency automatic speech recognition model designed for low-latency transcription t...
Qwen automatic-speech-recognition
168.8K 46

speaker diarization 3.1

mit
Speaker Diarization 3.1, powered by pyannote, is a specialized framework designed to solve the 'who spoke when' problem ...
pyannote automatic-speech-recognition
23.2K 11

speaker diarization community 1

cc-by-4.0
Speaker Diarization Community 1, powered by pyannote, is a specialized tool for the 'who spoke when' problem in audio pr...
pyannote automatic-speech-recognition
5.8K 1

Voxtral Mini 4B Realtime 2602

apache-2.0
Voxtral Mini 4B Realtime 2602 is a compact, low-latency speech-to-text model designed for high-throughput environments. ...
mistralai automatic-speech-recognition
2.9K 4

wav2vec2 large xlsr 53 japanese

apache-2.0
The wav2vec2-large-xlsr-53-japanese model is a robust automatic speech recognition (ASR) tool fine-tuned for Japanese au...
jonatasgrosman automatic-speech-recognition
2.0K 0

wav2vec2 large xlsr 53 dutch

apache-2.0
The wav2vec2-large-xlsr-53-dutch model is a specialized Automatic Speech Recognition (ASR) tool fine-tuned for the Dutch...
jonatasgrosman automatic-speech-recognition
2.0K 0

wav2vec2 large xlsr 53 portuguese

apache-2.0
The wav2vec2-large-xlsr-53-portuguese model is a robust Automatic Speech Recognition (ASR) tool fine-tuned specifically ...
jonatasgrosman automatic-speech-recognition
2.0K 0

voice activity detection

mit
Pyannote's Voice Activity Detection (VAD) is a specialized tool designed to distinguish human speech from silence or bac...
pyannote automatic-speech-recognition
1.9K 0

wav2vec2 large xlsr 53 russian

apache-2.0
The wav2vec2-large-xlsr-53-russian is a specialized automatic speech recognition (ASR) model based on Meta's cross-lingu...
jonatasgrosman automatic-speech-recognition
1.8K 0

wav2vec2 large xlsr 53 polish

apache-2.0
The wav2vec2-large-xlsr-53-polish model is a specialized Automatic Speech Recognition (ASR) tool fine-tuned for the Poli...
jonatasgrosman automatic-speech-recognition
1.4K 0

whisperkit coreml

Apache-2.0
WhisperKit CoreML brings OpenAI's Whisper speech-to-text capabilities directly to Apple silicon, optimizing inference fo...
argmaxinc automatic-speech-recognition
274 0
Join our Telegram