mpse smoe speech enhancement

提供商AdityaRaikar
分类speech-enhancement
许可证mit
下载量11
星标0

简介

mpse-smoe 是一款专注于语音增强的轻量化模型,采用了高效的混合专家(MoE)架构。它旨在解决录音环境中的背景噪声干扰,在保持语音自然度的同时,通过动态激活参数来提升降噪质量。对于需要处理低质量音频、播客后期清理或实时语音增强的开发者来说,该模型提供了极高的性价比。相比于传统的全量参数模型,它在推理速度和内存占用上更有优势,非常适合部署在资源受限的边缘设备或集成到音频处理流水线中。

核心亮点

  • 基于 MoE 架构,实现高性能与低功耗的平衡
  • 有效去除背景噪声,提升语音清晰度与自然度
  • MIT 协议开源,方便开发者快速集成与商用
  • 适用于播客清理、会议记录及实时语音增强

使用方法

安装依赖
# 安装 Hugging Face transformers
pip install transformers torch
SDK 使用
# 使用 transformers 加载模型
from transformers import AutoModel, AutoTokenizer

model = AutoModel.from_pretrained("AdityaRaikar/mpse-smoe-speech-enhancement")
tokenizer = AutoTokenizer.from_pretrained("AdityaRaikar/mpse-smoe-speech-enhancement")

Hugging Face 下载

我们推荐使用命令行或者 Hugging Face Hub SDK 来进行模型的下载。

操作指引:在下载前,请先通过如下命令安装 huggingface_hub:

操作指引
pip install -U huggingface_hub

命令行下载

下载完整模型库

下载完整模型库
huggingface-cli download AdityaRaikar/mpse-smoe-speech-enhancement

下载单个文件到指定本地文件夹(以下载 config.json 到当前路径下 ./dir 目录为例)

下载单个文件到指定本地文件夹(以下载 config.json 到当前路径下 ./dir 目录为例)
huggingface-cli download AdityaRaikar/mpse-smoe-speech-enhancement config.json --local-dir ./dir

更多命令行下载选项,可参见官方文档

SDK 下载

SDK 下载
# 模型下载
from huggingface_hub import snapshot_download
model_dir = snapshot_download('AdityaRaikar/mpse-smoe-speech-enhancement')

Git 下载

请确保 lfs 已经被正确安装

Git 下载
git lfs install
git clone https://huggingface.co/AdityaRaikar/mpse-smoe-speech-enhancement

如果您希望跳过 lfs 大文件下载,可以使用如下命令

跳过 LFS
GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/AdityaRaikar/mpse-smoe-speech-enhancement

模型文件托管在 Hugging Face Hub,使用 HF CLI / SDK / Git 直接下载,不经过本站。

PyTorch / Transformers 使用

安装 Transformers

安装 Transformers
pip install -U transformers torch

模型加载和推理

模型加载和推理
from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained('AdityaRaikar/mpse-smoe-speech-enhancement')
tokenizer = AutoTokenizer.from_pretrained('AdityaRaikar/mpse-smoe-speech-enhancement')

完整文档

来源: HuggingFace

---
tags:
- speech-enhancement
- audio-to-audio
- MP-SENet
- supervised-mixture-of-experts
- S-MoE
- dual-bandwidth
- narrowband
- wideband
license: mit
---

MP-SENet + Supervised Mixture of Experts (S-MoE) for Dual-Bandwidth Speech Enhancement

Speech enhancement model handling both Narrowband (NB, 8kHz) and Wideband (WB, 16kHz) audio using Supervised Mixture of Experts.

Based on:

  • MP-SENet — STFT-domain magnitude-phase speech enhancement backbone

  • S-MoE — Supervised MoE with deterministic hard gating by bandwidth metadata

Pre-trained Baseline

The official pre-trained MP-SENet baseline (from yxlu-0102/MP-SENet) is included at pretrained/g_best_vb.pt:

  • Trained on VoiceBank+DEMAND, 16kHz, 100 epochs

  • 2.263M params, 247 state dict keys

  • No need to train baseline from scratch — S-MoE training initializes directly from this checkpoint

Training (S-MoE only)

bash
pip install -r requirements-gpu.txt

S-MoE training (downloads pre-trained baseline automatically)

SMOE_EPOCHS=60 BATCH_SIZE=2 GRAD_ACCUM_STEPS=2 python train.py

The script automatically:
1. Downloads VoiceBank+DEMAND (~5GB)
2. Pre-extracts WB/NB data into separate folders (no on-the-fly resampling)
3. Downloads official pre-trained MP-SENet baseline from this repo
4. Initializes S-MoE experts from baseline (shared → both WB+NB experts)
5. Trains S-MoE on joint WB+NB data for 60 epochs
6. Pushes trained model to HF Hub

Data Pipeline

All NB/WB data is pre-extracted to separate folders before training:

| Folder | Type | Processing | bandwidth_id |
|---|---|---|---|
| original_wb/ | WB | Original 16kHz (symlinks) | 0 |
| simple_nb/ | NB | 16k→8k→16k resampled | 1 |
| codec_nb/ | NB | G.711 A-law/mu-law (round-robin) | 1 |

Architecture

| Component | Baseline | S-MoE |
|---|---|---|
| Total params | 2.263M | 3.587M |
| Active params (inference) | 2.263M | 2.263M (same!) |
| FFN type | Single GRU-FFN | 2× GRU-FFN experts, hard gated by bandwidth_id |
| Attention | Shared | Shared (unchanged) |

Environment Variables

| Variable | Default | Description |
|---|---|---|
| SMOE_EPOCHS | 60 | Epochs for S-MoE training |
| BATCH_SIZE | 2 | GPU micro-batch size |
| GRAD_ACCUM_STEPS | 2 | Gradient accumulation (effective BS = BATCH_SIZE × GRAD_ACCUM_STEPS) |
| PESQ_EVERY_N | 10 | Compute PESQ for discriminator every N steps (expensive) |
| LR | 5e-4 | Learning rate |
| EVAL_INTERVAL | 5 | Evaluate every N epochs |
| BASELINE_REPO | AdityaRaikar/mpse-smoe-speech-enhancement | HF repo with pre-trained baseline |
| BASELINE_FILE | pretrained/g_best_vb.pt | Checkpoint filename in repo |

Files

| File | Description |
|---|---|
| train.py | Combined data prep + S-MoE training (self-contained) |
| pretrained/g_best_vb.pt | Official pre-trained MP-SENet baseline (VoiceBank+DEMAND) |
| smoe_models.py | Model definitions with detailed docstrings |
| data_preparation.py | Standalone NB/WB codec pipeline |
| config.json | Default hyperparameters |

References

bibtex
@article{lu2023mpse,
  title={MP-SENet: A Speech Enhancement Model with Parallel Denoising of Magnitude and Phase Spectra},
  author={Lu, Ye-Xin and Ai, Yang and Ling, Zhen-Hua},
  journal={arXiv preprint arXiv:2305.13686},
  year={2023}
}