MioCodec 25Hz 24kHz

提供商Aratako
分类audio-to-audio
许可证mit
下载量45
星标0

简介

MioCodec 25Hz 24kHz 是一款由 Aratako 提供的轻量级音频编解码模型,专注于高效的音频信号处理。它采用了 25Hz 的低采样率量化和 24kHz 的高质量重构,在保证音频还原度的同时极大地压缩了数据量。对于开发者而言,该模型非常适合用于实时语音传输、低带宽音频存储或作为 AI 语音管线的预处理模块。由于采用 MIT 协议且架构精简,其部署门槛较低,能够与主流的音频处理库无缝集成,是构建端到端语音 AI 应用的理想底层组件。

核心亮点

  • 低带宽占用,支持 24kHz 高保真音频还原
  • MIT 开源协议,企业级部署无压力
  • 极低延迟,适配实时语音交互场景
  • 轻量化设计,显著降低计算资源开销

使用方法

安装依赖
# 安装 Hugging Face transformers
pip install transformers torch
SDK 使用
# 使用 transformers 加载模型
from transformers import AutoModel, AutoTokenizer

model = AutoModel.from_pretrained("Aratako/MioCodec-25Hz-24kHz")
tokenizer = AutoTokenizer.from_pretrained("Aratako/MioCodec-25Hz-24kHz")

Hugging Face 下载

我们推荐使用命令行或者 Hugging Face Hub SDK 来进行模型的下载。

操作指引:在下载前,请先通过如下命令安装 huggingface_hub:

操作指引
pip install -U huggingface_hub

命令行下载

下载完整模型库

下载完整模型库
huggingface-cli download Aratako/MioCodec-25Hz-24kHz

下载单个文件到指定本地文件夹(以下载 config.json 到当前路径下 ./dir 目录为例)

下载单个文件到指定本地文件夹(以下载 config.json 到当前路径下 ./dir 目录为例)
huggingface-cli download Aratako/MioCodec-25Hz-24kHz config.json --local-dir ./dir

更多命令行下载选项,可参见官方文档

SDK 下载

SDK 下载
# 模型下载
from huggingface_hub import snapshot_download
model_dir = snapshot_download('Aratako/MioCodec-25Hz-24kHz')

Git 下载

请确保 lfs 已经被正确安装

Git 下载
git lfs install
git clone https://huggingface.co/Aratako/MioCodec-25Hz-24kHz

如果您希望跳过 lfs 大文件下载,可以使用如下命令

跳过 LFS
GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/Aratako/MioCodec-25Hz-24kHz

模型文件托管在 Hugging Face Hub,使用 HF CLI / SDK / Git 直接下载,不经过本站。

PyTorch / Transformers 使用

安装 Transformers

安装 Transformers
pip install -U transformers torch

模型加载和推理

模型加载和推理
from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained('Aratako/MioCodec-25Hz-24kHz')
tokenizer = AutoTokenizer.from_pretrained('Aratako/MioCodec-25Hz-24kHz')

模型下载

我们推荐使用命令行或者 ModelScope SDK 来进行模型的下载。

操作指引:在下载前,请先通过如下命令安装 ModelScope:

操作指引
pip install modelscope

命令行下载

下载完整模型库

下载完整模型库
modelscope download --model Aratako/MioCodec-25Hz-24kHz

下载单个文件到指定本地文件夹(以下载 README.md 到当前路径下 dir 目录为例)

下载单个文件到指定本地文件夹(以下载 README.md 到当前路径下 dir 目录为例)
modelscope download --model Aratako/MioCodec-25Hz-24kHz README.md --local_dir ./dir

更多更丰富的命令行下载选项,可参见具体文档

SDK 下载

SDK 下载
# 模型下载
from modelscope import snapshot_download
model_dir = snapshot_download('Aratako/MioCodec-25Hz-24kHz')

Git 下载

请确保 lfs 已经被正确安装

Git 下载
git lfs install
git clone https://www.modelscope.cn/Aratako/MioCodec-25Hz-24kHz.git

如果您希望跳过 lfs 大文件下载,可以使用如下命令

跳过 LFS
GIT_LFS_SKIP_SMUDGE=1 git clone https://www.modelscope.cn/Aratako/MioCodec-25Hz-24kHz.git

ModelScope 模型页直接下载模型文件;无需将模型文件放在本站服务器。

Notebook 快速开发

下载并安装 ModelScope library

下载并安装 ModelScope library
pip install "modelscope[audio,cv,nlp,multi-modal,science]" -f https://modelscope.oss-cn-beijing.aliyuncs.com/releases/repo.html

模型加载和推理

模型加载和推理
from modelscope.pipelines import pipeline
from modelscope.utils.constant import Tasks

p = pipeline('text-generation', 'Aratako/MioCodec-25Hz-24kHz')

完整文档

来源: HuggingFace

---
license: mit
language:

  • en

  • ja

  • nl

  • fr

  • de

  • it

  • pl

  • pt

  • es

  • ko

  • zh

tags:
  • speech

  • audio

  • tokenizer

datasets:
  • sarulab-speech/mls_sidon

  • mythicinfinity/Libriheavy-HQ

  • nvidia/hifitts-2

  • amphion/Emilia-Dataset

pipeline_tag: audio-to-audio
---

MioCodec-25Hz-24kHz: Lightweight Neural Audio Codec for Efficient Spoken Language Modeling

![GitHub](https://github.com/Aratako/MioCodec)

MioCodec-25Hz-24kHz is a lightweight and fast neural audio codec designed for efficient spoken language modeling. Based on the Kanade-Tokenizer implementation, this model features an integrated wave decoder (iSTFTHead) that directly synthesizes waveforms without requiring an external vocoder.

For higher audio fidelity at 44.1 kHz, see MioCodec-25Hz-44.1kHz.

🌟 Overview

MioCodec decomposes speech into two distinct components:

1. Content Tokens: Discrete representations that primarily capture linguistic information and phonetic content ("what" is being said) at a low frame rate (25 Hz).
2. Global Embeddings: A continuous vector representing broad acoustic characteristics ("how")—including speaker identity, recording environment, and microphone traits.

By disentangling these elements, MioCodec is ideal for Spoken Language Modeling.

Key features

  • Lightweight & Fast: Integrated wave decoder (iSTFTHead) enables direct waveform synthesis without an external vocoder.
  • Ultra-Low Bitrate: Achieves high-fidelity reconstruction at only 341 bps.
  • End-to-End Design: Single model architecture from audio input to waveform output.

📊 Model Comparison

| Model | Token Rate | Vocab Size | Bit Rate | Sample Rate | SSL Encoder | Vocoder | Parameters | Highlights |
| :--- | :---: | :---: | :---: | :---: | :---: | :---: | :---: | :--- |
| MioCodec-25Hz-24kHz | 25 Hz | 12,800 | 341 bps | 24 kHz | WavLM-base+ | - (iSTFTHead) | 132M | Lightweight, fast inference |
| MioCodec-25Hz-44.1kHz | 25 Hz | 12,800 | 341 bps | 44.1 kHz | WavLM-base+ | MioVocoder (Jointly Tuned) | 118M (w/o vocoder) | High-quality, high sample rate |
| kanade-25hz | 25 Hz | 12,800 | 341 bps | 24 kHz | WavLM-base+ | Vocos 24kHz | 118M (w/o vocoder) | Original 25Hz model |
| kanade-12.5hz | 12.5 Hz | 12,800 | 171 bps | 24 kHz | WavLM-base+ | Vocos 24kHz | 120M (w/o vocoder) | Original 12.5Hz model |

🚀 Quick Start

Installation

bash
# Install via pip
pip install git+https://github.com/Aratako/MioCodec

Or using uv

uv add git+https://github.com/Aratako/MioCodec

Basic Inference

Basic usage for encoding and decoding audio:

python
from miocodec import MioCodecModel, load_audio
import soundfile as sf

1. Load model

model = MioCodecModel.from_pretrained("Aratako/MioCodec-25Hz-24kHz").eval().cuda()

2. Load audio

waveform = load_audio("input.wav", sample_rate=model.config.sample_rate).cuda()

3. Encode Audio

features = model.encode(waveform)

4. Decode to Waveform (directly, no vocoder needed)

resynth = model.decode( content_token_indices=features.content_token_indices, global_embedding=features.global_embedding, )

5. Save

sf.write("output.wav", resynth.cpu().numpy(), model.config.sample_rate)

Voice Conversion (Zero-shot)

MioCodec allows you to swap speaker identities by combining the content tokens of a source with the global embedding of a reference.

python
source = load_audio("source_content.wav", sample_rate=model.config.sample_rate).cuda()
reference = load_audio("target_speaker.wav", sample_rate=model.config.sample_rate).cuda()

Perform conversion

vc_wave = model.voice_conversion(source, reference) sf.write("converted.wav", vc_wave.cpu().numpy(), model.config.sample_rate)

🏗️ Training Methodology

MioCodec-25Hz-24kHz was trained in two phases with an integrated wave decoder that directly synthesizes waveforms via iSTFT.

Phase 1: Feature Alignment

The model is trained to minimize both Multi-Resolution Mel-spectrogram loss and SSL feature reconstruction loss (using WavLM-base+). The wave decoder directly generates waveforms, and losses are computed on the reconstructed audio.

  • Multi-Resolution Mel Spectrogram Loss: Using window lengths of [32, 64, 128, 256, 512, 1024, 2048].
  • SSL Feature Reconstruction Loss: Using WavLM-base+ features.

Phase 2: Adversarial Refinement

Building upon Phase 1, adversarial training is introduced to improve perceptual quality. The training objectives include:

  • Multi-Resolution Mel Spectrogram Loss: Using window lengths of [32, 64, 128, 256, 512, 1024, 2048].
  • SSL Feature Reconstruction Loss: Using WavLM-base+ features.
  • Multi-Period Discriminator (MPD): Using periods of [2, 3, 5, 7, 11, 17, 23].
  • Multi-Scale STFT Discriminator (MS-STFTD): Using FFT sizes of [118, 190, 310, 502, 814, 1314, 2128, 3444].
  • RMS Loss: To stabilize energy and volume.

📚 Training Data

The training datasets are listed below:

| Language | Approx. Hours | Dataset |
| :--- | :--- | :--- |
| Japanese | ~22,500h | Various public HF datasets |
| English | ~500h | Libriheavy-HQ |
| English | ~4,000h | MLS-Sidon |
| English | ~9,000h | HiFiTTS-2 |
| English | ~27,000h | Emilia-YODAS |
| German | ~1,950h | MLS-Sidon |
| German | ~5,600h | Emilia-YODAS |
| Dutch | ~1,550h | MLS-Sidon |
| French | ~1,050h | MLS-Sidon |
| French | ~7,400h | Emilia-YODAS |
| Spanish | ~900h | MLS-Sidon |
| Italian | ~240h | MLS-Sidon |
| Portuguese | ~160h | MLS-Sidon |
| Polish | ~100h | MLS-Sidon |
| Korean | ~7,300h | Emilia-YODAS |
| Chinese | ~300h | Emilia-YODAS |

📜 Acknowledgements

  • Decoder Design: Inspired by XCodec2.

🖊️ Citation

bibtex
@misc{miocodec-25hz-24khz,
  author = {Chihiro Arata},
  title = {MioCodec: High-Fidelity Neural Audio Codec for Efficient Spoken Language Modeling},
  year = {2026},
  publisher = {Hugging Face},
  journal = {Hugging Face repository},
  howpublished = {\url{https://huggingface.co/Aratako/MioCodec-25Hz-24kHz}}
}