saad speech recognition hausa audio to text

ProviderBaghdad99
Categoryaudio-generation
Licenseapache-2.0
Downloads8
Stars0

Overview

The SAAD speech recognition model is a specialized audio-to-text tool designed specifically for the Hausa language. For developers building localized applications in West African markets, this model provides a critical bridge for accessibility and data digitization. Unlike general-purpose multilingual models that often struggle with regional dialects or low-resource languages, SAAD is tuned for higher accuracy in Hausa phonetic transcription. It is released under the Apache-2.0 license, making it highly flexible for commercial integration. Devs can leverage this for automated transcription services, voice-command interfaces, or preprocessing audio datasets for downstream NLP tasks without the overhead of massive, generalist LLMs.

Highlights

  • Optimized for high-accuracy Hausa audio transcription
  • Apache-2.0 license allows flexible commercial deployment
  • Ideal for localized voice-to-text application development
  • Efficient alternative to oversized multilingual speech models

Usage

Install
# Install Hugging Face transformers
pip install transformers torch
SDK Usage
# Load model with transformers
from transformers import AutoModel, AutoTokenizer

model = AutoModel.from_pretrained("Baghdad99/saad-speech-recognition-hausa-audio-to-text")
tokenizer = AutoTokenizer.from_pretrained("Baghdad99/saad-speech-recognition-hausa-audio-to-text")

Hugging Face Download

We recommend downloading the model via the Hugging Face CLI or Hub SDK.

Guidance:Before downloading, install huggingface_hub with:

Guidance
pip install -U huggingface_hub

CLI Download

Download the full repository

Download the full repository
huggingface-cli download Baghdad99/saad-speech-recognition-hausa-audio-to-text

Download a single file to a local folder (e.g. config.json into ./dir)

Download a single file to a local folder (e.g. config.json into ./dir)
huggingface-cli download Baghdad99/saad-speech-recognition-hausa-audio-to-text config.json --local-dir ./dir

See the official docs for more CLI options

SDK Download

SDK Download
# 模型下载
from huggingface_hub import snapshot_download
model_dir = snapshot_download('Baghdad99/saad-speech-recognition-hausa-audio-to-text')

Git Download

Make sure git-lfs is installed first

Git Download
git lfs install
git clone https://huggingface.co/Baghdad99/saad-speech-recognition-hausa-audio-to-text

To skip LFS large-file downloads, use:

Skip LFS
GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/Baghdad99/saad-speech-recognition-hausa-audio-to-text

Model files are hosted on the Hugging Face Hub — download directly via HF CLI / SDK / Git, not through this site.

PyTorch / Transformers Usage

Install Transformers

Install Transformers
pip install -U transformers torch

Load the model and run inference

Load the model and run inference
from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained('Baghdad99/saad-speech-recognition-hausa-audio-to-text')
tokenizer = AutoTokenizer.from_pretrained('Baghdad99/saad-speech-recognition-hausa-audio-to-text')

Full Documentation

来源: HuggingFace

---
language:

  • ha

license: apache-2.0
base_model: openai/whisper-small
tags:
  • generated_from_trainer

datasets:
  • mozilla-foundation/common_voice_13_0

metrics:
  • wer

model-index:
  • name: Hausa Whisper Small - Saad

results:
- task:
name: Automatic Speech Recognition
type: automatic-speech-recognition
dataset:
name: Common Voice 13
type: mozilla-foundation/common_voice_13_0
config: ha
split: test
args: ha
metrics:
- name: Wer
type: wer
value: 44.41266209000763
---

<!-- This model card has been generated automatically according to the information the Trainer had access to. You
should probably proofread and complete it, then remove this comment. -->

Hausa Whisper Small - Saad

This model is a fine-tuned version of openai/whisper-small on the Common Voice 13 dataset.
It achieves the following results on the evaluation set:

  • Loss: 0.7524

  • Wer Ortho: 47.7050

  • Wer: 44.4127

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 1e-05

  • train_batch_size: 16

  • eval_batch_size: 16

  • seed: 42

  • optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08

  • lr_scheduler_type: constant_with_warmup

  • lr_scheduler_warmup_steps: 50

  • training_steps: 500

  • mixed_precision_training: Native AMP

Training results

| Training Loss | Epoch | Step | Validation Loss | Wer Ortho | Wer |
|:-------------:|:-----:|:----:|:---------------:|:---------:|:-------:|
| 0.0104 | 3.18 | 500 | 0.7524 | 47.7050 | 44.4127 |

Framework versions

  • Transformers 4.35.0
  • Pytorch 2.1.0+cu118
  • Datasets 2.14.6
  • Tokenizers 0.14.1
Join our Telegram