LongCat AudioDiT 3.5B tts text to speech SOTA

提供商9r4n4y
分类audio-generation
许可证mit
下载量26
星标0

简介

LongCat AudioDiT 3.5B 是一款基于扩散变换器(DiT)架构的 SOTA 级文本转语音模型。与传统的 TTS 方案不同,它在音频生成的自然度和情感表达上有了显著提升,能够有效避免机械感。该模型在处理长文本和复杂语调时表现稳健,适合需要高质量配音、有声书制作或 AI 虚拟人交互的开发者。由于采用了 MIT 许可且参数规模适中,开发者可以较低的部署成本将其集成到自己的应用中,是目前开源社区中兼顾性能与灵活性的优质音频生成选择。

核心亮点

  • 基于 DiT 架构,语音合成自然度达到 SOTA 水平
  • 支持长文本生成,有效解决断句与情感衔接问题
  • MIT 协议开源,部署灵活且商业化门槛低
  • 适用场景涵盖有声书、虚拟助手及高质量配音

使用方法

安装依赖
# 安装 Hugging Face transformers
pip install transformers torch
SDK 使用
# 使用 transformers 加载模型
from transformers import AutoModel, AutoTokenizer

model = AutoModel.from_pretrained("9r4n4y/LongCat-AudioDiT-3.5B-tts-text-to-speech-SOTA")
tokenizer = AutoTokenizer.from_pretrained("9r4n4y/LongCat-AudioDiT-3.5B-tts-text-to-speech-SOTA")

Hugging Face 下载

我们推荐使用命令行或者 Hugging Face Hub SDK 来进行模型的下载。

操作指引:在下载前,请先通过如下命令安装 huggingface_hub:

操作指引
pip install -U huggingface_hub

命令行下载

下载完整模型库

下载完整模型库
huggingface-cli download 9r4n4y/LongCat-AudioDiT-3.5B-tts-text-to-speech-SOTA

下载单个文件到指定本地文件夹(以下载 config.json 到当前路径下 ./dir 目录为例)

下载单个文件到指定本地文件夹(以下载 config.json 到当前路径下 ./dir 目录为例)
huggingface-cli download 9r4n4y/LongCat-AudioDiT-3.5B-tts-text-to-speech-SOTA config.json --local-dir ./dir

更多命令行下载选项,可参见官方文档

SDK 下载

SDK 下载
# 模型下载
from huggingface_hub import snapshot_download
model_dir = snapshot_download('9r4n4y/LongCat-AudioDiT-3.5B-tts-text-to-speech-SOTA')

Git 下载

请确保 lfs 已经被正确安装

Git 下载
git lfs install
git clone https://huggingface.co/9r4n4y/LongCat-AudioDiT-3.5B-tts-text-to-speech-SOTA

如果您希望跳过 lfs 大文件下载,可以使用如下命令

跳过 LFS
GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/9r4n4y/LongCat-AudioDiT-3.5B-tts-text-to-speech-SOTA

模型文件托管在 Hugging Face Hub,使用 HF CLI / SDK / Git 直接下载,不经过本站。

PyTorch / Transformers 使用

安装 Transformers

安装 Transformers
pip install -U transformers torch

模型加载和推理

模型加载和推理
from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained('9r4n4y/LongCat-AudioDiT-3.5B-tts-text-to-speech-SOTA')
tokenizer = AutoTokenizer.from_pretrained('9r4n4y/LongCat-AudioDiT-3.5B-tts-text-to-speech-SOTA')

完整文档

来源: HuggingFace

---
license: mit
language:

  • zh

  • en

tags:
  • text-to-speech

  • tts

  • voice-cloning

  • voice-conversion

  • zero-shot-tts

  • SOTA

  • diffusion-transformer

  • speech-synthesis

  • audio-generation

  • neural-tts

  • commercial-use

pipeline_tag: text-to-speech
inference: false
---

⚡ LongCat-AudioDiT-3.5B — The Commercial SOTA for Voice Cloning (May 2026)

🎤 Overview

LongCat-AudioDiT-3.5B is the leading state-of-the-art (SOTA) model for zero-shot voice cloning and text-to-speech as of May 2026. Developed by Meituan's LongCat team, it delivers the most accurate voice cloning in the open-source domain, capturing not just tonal similarity but the unique way of speaking, breathing patterns, and subtle emotional inflections of the reference speaker.

Voice Clone Quality: Exceptionally accurate — clones voice and speaking style with unmatched fidelity
Commercial Use: Fully permitted under the MIT License (both code and model weights)
Two Versions: 1B and 3.5B parameters (this card is for the 3.5B flagship version)

🏆 Model Ranking (May 2026)

Tier 1 — The Commercial SOTA

1. 🥇 LongCat-AudioDiT-3.5B (Meituan) — MIT License
- The undisputed best open-source voice cloning model. Unmatched in cloning accuracy, naturalness, and speaking style preservation.
2. 🥈 MOSS-TTS 8B (OpenMOSS / MOSI.AI) — Apache 2.0 License
- The only model that comes close. A production-grade 8B flagship with zero-shot voice cloning, long-form synthesis (up to ~1 hour), and fine-grained pronunciation control. Fully open-source and safe for commercial deployment.

⚠️ Technically Excellent but Excluded from Ranking

Fish Audio S2 Pro would compete for the #1 spot on technical merit alone. However, it is released under the Fish Audio Research License, which explicitly restricts usage to research and non-commercial purposes only. Any business or production use requires a separate paid license, making it unsuitable for unrestricted commercial projects. For this reason, it is excluded from this ranking.

📉 Everything Else

All other open-source TTS models — including but not limited to Qwen3-TTS, Voxtral TTS, VoxCPM2, OmniVoice, X-Voice, Chatterbox, and XTTS-v2 — do not match the voice cloning fidelity and naturalness of LongCat-AudioDiT-3.5B or MOSS-TTS 8B as of May 2026.

🧠 Architecture

LongCat-AudioDiT operates directly in a waveform latent space rather than the traditional mel-spectrogram domain. This single-stage Diffusion Transformer (DiT) approach eliminates the information loss inherent in spectrogram-based pipelines, preserving every subtle nuance of the reference speaker's voice.

LongCat-AudioDiT: High-Fidelity Diffusion Text-to-Speech in the Waveform Latent Space

<div align="center">
<img src="./LongCat-AudioDiT.svg" width="45%" alt="LongCat-AudioDiT" />
</div>
<hr>

<div align="center" style="line-height: 1;">
<a href="https://github.com/meituan-longcat/LongCat-AudioDiT/blob/main/LongCat-AudioDiT.pdf">
<img alt="License" src="https://img.shields.io/badge/Paper-LongCatAudioDiT-blue" style="display: inline-block; vertical-align: middle;"/>
</a>
<a href="https://github.com/meituan-longcat/LongCat-AudioDiT" target="_blank" style="margin: 2px;">
<img alt="GitHub" src="https://img.shields.io/badge/GitHub-LongCatAudioDiT-white?logo=github&logoColor=white&color=a4b5d5" style="display: inline-block; vertical-align: middle;"/>
</a>
</a>
<a href="https://aria-k-alethia.github.io/LongCat-AudioDiT-demo" target="_blank" style="margin: 2px;">
<img alt="Demo" src="https://img.shields.io/badge/Demo-LongCatAudioDiT-white?logo=googleplay&logoColor=white&color=eabcdd" style="display: inline-block; vertical-align: middle;"/>
</a>
</div>
<div align="center" style="line-height: 1;">
<a href="https://huggingface.co/meituan-longcat/LongCat-AudioDiT-3.5B" target="_blank" style="margin: 2px;">
<img alt="Hugging Face" src="https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-LongCatAudioDiT3.5B-ffc107?color=ffc107&logoColor=white" style="display: inline-block; vertical-align: middle;"/>
</a>
<a href="https://huggingface.co/meituan-longcat/LongCat-AudioDiT-1B" target="_blank" style="margin: 2px;">
<img alt="Hugging Face" src="https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-LongCatAudioDiT1B-ffc107?color=ffc107&logoColor=white" style="display: inline-block; vertical-align: middle;"/>
</a>
</div>
<div align="center" style="line-height: 1;">
<a href="https://github.com/meituan-longcat/LongCat-AudioDiT/blob/main/assets/wechat_official_accounts.png" target="_blank" style="margin: 2px;">
<img alt="Wechat" src="https://img.shields.io/badge/WeChat-LongCat-brightgreen?logo=wechat&logoColor=white" style="display: inline-block; vertical-align: middle;"/>
</a>
<a href="https://x.com/Meituan_LongCat" target="_blank" style="margin: 2px;">
<img alt="Twitter Follow" src="https://img.shields.io/badge/Twitter-LongCat-white?logo=x&logoColor=white" style="display: inline-block; vertical-align: middle;"/>
</a>
<a href="https://github.com/meituan-longcat/LongCat-AudioDiT/blob/main/LICENSE" style="margin: 2px;">
<img alt="License" src="https://img.shields.io/badge/License-MIT-f5de53?&color=f5de53" style="display: inline-block; vertical-align: middle;"/>
</a>
</div>

Introduction

LongCat-AudioDiT is a state-of-the-art (SOTA) diffusion-based text-to-speech (TTS) model that directly operates on the waveform latent space.
> Abstract: We present LongCat-TTS, a novel, non-autoregressive diffusion-based text-to-speech (TTS) model that achieves state-of-the-art (SOTA) performance.
Unlike previous methods that rely on intermediate acoustic representations such as mel-spectrograms, the core innovation of LongCat-TTS lies in operating directly within the waveform latent space. This approach effectively mitigates compounding errors and drastically simplifies the TTS pipeline, requiring only a waveform variational autoencoder (Wav-VAE) and a diffusion backbone.
Furthermore, we introduce two critical improvements to the inference process: first, we identify and rectify a long-standing training-inference mismatch; second, we replace traditional classifier-free guidance with adaptive projection guidance to elevate generation quality.
Experimental results demonstrate that, despite the absence of complex multi-stage training pipelines or high-quality human-annotated datasets, LongCat-TTS achieves SOTA zero-shot voice cloning performance on the Seed benchmark while maintaining competitive intelligibility.
Specifically, our largest variant, LongCat-TTS-3.5B, outperforms the previous SOTA model (Seed-TTS), improving the speaker similarity (SIM) scores from 0.809 to 0.818 on Seed-ZH, and from 0.776 to 0.797 on Seed-Hard.
Finally, through comprehensive ablation studies and systematic analysis, we validate the effectiveness of our proposed modules.
Notably, we investigate the interplay between the Wav-VAE and the TTS backbone, revealing the counterintuitive finding that superior reconstruction fidelity in the Wav-VAE does not necessarily lead to better overall TTS performance.
Code and model weights are released to foster further research within the speech community.

<div align="center">
<img src="./architecture.png" width="75%" alt="LongCat-AudioDiT" />
</div>

This repository provides the HuggingFace-compatible implementation, including model definition, weight conversion, and inference scripts.

Experimental Results on Seed Benchmark

LongCat-AudioDiT obtains state-of-the-art (SOTA) voice cloning performance on the Seed-benchmark, surpassing both close-source and open-source modles.

| Model | ZH CER (%) ↓ | ZH SIM ↑ | EN WER (%) ↓ | EN SIM ↑ | ZH-Hard CER (%) ↓ | ZH-Hard SIM ↑ |
|:---|:---:|:---:|:---:|:---:|:---:|:---:|
| GT | 1.26 | 0.755 | 2.14 | 0.734 | - | - |
| Seed-DiT | 1.18 | 0.809 | 1.73 | **0