Common Voice Gender Detection

ProviderprithivMLmods
Categoryaudio-classification
Licenseapache-2.0
Downloads2.8K
Stars0

Overview

Common Voice Gender Detection is a specialized audio classification model designed to identify the gender of a speaker from voice samples. For developers building accessibility tools, voice-driven UX, or demographic analytics, this model provides a streamlined way to categorize audio metadata without requiring a full speech-to-text pipeline. It is lightweight and optimized for classification tasks, making it suitable for integration into preprocessing workflows where speaker profiling is necessary. Compared to general-purpose audio models, its narrow focus on gender detection reduces computational overhead and simplifies the inference logic for developers implementing targeted audio analysis.

Highlights

  • Specialized in binary gender classification for audio samples.
  • Lightweight architecture ensures fast inference and low latency.
  • Apache-2.0 license allows for flexible commercial integration.
  • Ideal for speaker profiling and demographic data analysis.

Usage

Install
# Install Hugging Face transformers
pip install transformers torch
SDK Usage
# Load model with transformers
from transformers import AutoModel, AutoTokenizer

model = AutoModel.from_pretrained("prithivMLmods/Common-Voice-Gender-Detection")
tokenizer = AutoTokenizer.from_pretrained("prithivMLmods/Common-Voice-Gender-Detection")

Hugging Face Download

We recommend downloading the model via the Hugging Face CLI or Hub SDK.

Guidance:Before downloading, install huggingface_hub with:

Guidance
pip install -U huggingface_hub

CLI Download

Download the full repository

Download the full repository
huggingface-cli download prithivMLmods/Common-Voice-Gender-Detection

Download a single file to a local folder (e.g. config.json into ./dir)

Download a single file to a local folder (e.g. config.json into ./dir)
huggingface-cli download prithivMLmods/Common-Voice-Gender-Detection config.json --local-dir ./dir

See the official docs for more CLI options

SDK Download

SDK Download
# 模型下载
from huggingface_hub import snapshot_download
model_dir = snapshot_download('prithivMLmods/Common-Voice-Gender-Detection')

Git Download

Make sure git-lfs is installed first

Git Download
git lfs install
git clone https://huggingface.co/prithivMLmods/Common-Voice-Gender-Detection

To skip LFS large-file downloads, use:

Skip LFS
GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/prithivMLmods/Common-Voice-Gender-Detection

Model files are hosted on the Hugging Face Hub — download directly via HF CLI / SDK / Git, not through this site.

PyTorch / Transformers Usage

Install Transformers

Install Transformers
pip install -U transformers torch

Load the model and run inference

Load the model and run inference
from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained('prithivMLmods/Common-Voice-Gender-Detection')
tokenizer = AutoTokenizer.from_pretrained('prithivMLmods/Common-Voice-Gender-Detection')

Model Download

We recommend downloading the model via the ModelScope CLI or SDK.

Guidance:Before downloading, install ModelScope with:

Guidance
pip install modelscope

CLI Download

Download the full repository

Download the full repository
modelscope download --model prithivMLmods/Common-Voice-Gender-Detection

Download a single file to a local folder (e.g. README.md into ./dir)

Download a single file to a local folder (e.g. README.md into ./dir)
modelscope download --model prithivMLmods/Common-Voice-Gender-Detection README.md --local_dir ./dir

See the docs for more CLI options

SDK Download

SDK Download
# 模型下载
from modelscope import snapshot_download
model_dir = snapshot_download('prithivMLmods/Common-Voice-Gender-Detection')

Git Download

Make sure git-lfs is installed first

Git Download
git lfs install
git clone https://www.modelscope.cn/prithivMLmods/Common-Voice-Gender-Detection.git

To skip LFS large-file downloads, use:

Skip LFS
GIT_LFS_SKIP_SMUDGE=1 git clone https://www.modelscope.cn/prithivMLmods/Common-Voice-Gender-Detection.git

ModelScope 模型页直接下载模型文件;无需将模型文件放在本站服务器。

Notebook Quickstart

Install the ModelScope library

Install the ModelScope library
pip install "modelscope[audio,cv,nlp,multi-modal,science]" -f https://modelscope.oss-cn-beijing.aliyuncs.com/releases/repo.html

Load the model and run inference

Load the model and run inference
from modelscope.pipelines import pipeline
from modelscope.utils.constant import Tasks

p = pipeline('text-generation', 'prithivMLmods/Common-Voice-Gender-Detection')

Full Documentation

来源: HuggingFace

---
license: apache-2.0
language:

  • en

base_model:
  • facebook/wav2vec2-base-960h

pipeline_tag: audio-classification
library_name: transformers
tags:
  • voice-gender-detection

  • male

  • female

  • biology

  • SFT

---

!1

Common-Voice-Gender-Detection

> Common-Voice-Gender-Detection is a fine-tuned version of facebook/wav2vec2-base-960h for binary audio classification, specifically trained to detect speaker gender as female or male. This model leverages the Wav2Vec2ForSequenceClassification architecture for efficient and accurate voice-based gender classification.

> [!note]
Wav2Vec2: Self-Supervised Learning for Speech Recognition : https://arxiv.org/pdf/2006.11477

py
Classification Report:

precision recall f1-score support

female 0.9705 0.9916 0.9809 2622
male 0.9943 0.9799 0.9870 3923

accuracy 0.9846 6545
macro avg 0.9824 0.9857 0.9840 6545
weighted avg 0.9848 0.9846 0.9846 6545

!download.png

!download (1).png

---

Label Space: 2 Classes

code
Class 0: female  
Class 1: male

---

Install Dependencies

bash
pip install gradio transformers torch librosa hf_xet

---

Inference Code

python
import gradio as gr
from transformers import Wav2Vec2ForSequenceClassification, Wav2Vec2FeatureExtractor
import torch
import librosa

Load model and processor

model_name = "prithivMLmods/Common-Voice-Geneder-Detection" model = Wav2Vec2ForSequenceClassification.from_pretrained(model_name) processor = Wav2Vec2FeatureExtractor.from_pretrained(model_name)

Label mapping

id2label = { "0": "female", "1": "male" }

def classify_audio(audio_path):
# Load and resample audio to 16kHz
speech, sample_rate = librosa.load(audio_path, sr=16000)

# Process audio
inputs = processor(
speech,
sampling_rate=sample_rate,
return_tensors="pt",
padding=True
)

with torch.no_grad():
outputs = model(**inputs)
logits = outputs.logits
probs = torch.nn.functional.softmax(logits, dim=1).squeeze().tolist()

prediction = {
id2label[str(i)]: round(probs[i], 3) for i in range(len(probs))
}

return prediction

Gradio Interface

iface = gr.Interface( fn=classify_audio, inputs=gr.Audio(type="filepath", label="Upload Audio (WAV, MP3, etc.)"), outputs=gr.Label(num_top_classes=2, label="Gender Classification"), title="Common Voice Gender Detection", description="Upload an audio clip to classify the speaker's gender as female or male." )

if __name__ == "__main__":
iface.launch()

---

Demo Inference

> [!note]
male

<audio controls src="https://cdn-uploads.huggingface.co/production/uploads/65bb837dbfb878f46c77de4c/7woMf3_bgX_D99-1Uy3jH.mpga"></audio>

!Screenshot 2025-05-31 at 20-19-39 Common Voice Gender Detection.png

> [!note]
female

<audio controls src="https://cdn-uploads.huggingface.co/production/uploads/65bb837dbfb878f46c77de4c/0d2rDf_DT-gjRWBwiPbm_.mpga"></audio>

!Screenshot 2025-05-31 at 20-21-57 Common Voice Gender Detection.png

---

Intended Use

Common-Voice-Gender-Detection is designed for:

  • Speech Analytics – Assist in analyzing speaker demographics in call centers or customer service recordings.
  • Conversational AI Personalization – Adjust tone or dialogue based on gender detection for more personalized voice assistants.
  • Voice Dataset Curation – Automatically tag or filter voice datasets by speaker gender for better dataset management.
  • Research Applications – Enable linguistic and acoustic research involving gender-specific speech patterns.
  • Multimedia Content Tagging – Automate metadata generation for gender identification in podcasts, interviews, or video content.
Join our Telegram