MiniMax Music3
简介
核心亮点
- 支持自然语言驱动,零门槛创作高质量音乐
- Apache-2.0 协议,对开发者极其友好且灵活
- 极速生成商业级音频,适配短视频与游戏场景
- 精准捕捉情绪与风格,大幅提升音频素材产出率
使用方法
# 安装 Hugging Face transformers
pip install transformers torch
# 使用 transformers 加载模型
from transformers import AutoModel, AutoTokenizer
model = AutoModel.from_pretrained("MiniMaxAI/MiniMax-Music3")
tokenizer = AutoTokenizer.from_pretrained("MiniMaxAI/MiniMax-Music3")
Hugging Face 下载
我们推荐使用命令行或者 Hugging Face Hub SDK 来进行模型的下载。
操作指引:在下载前,请先通过如下命令安装 huggingface_hub:
pip install -U huggingface_hub
命令行下载
下载完整模型库
huggingface-cli download MiniMaxAI/MiniMax-Music3
下载单个文件到指定本地文件夹(以下载 config.json 到当前路径下 ./dir 目录为例)
huggingface-cli download MiniMaxAI/MiniMax-Music3 config.json --local-dir ./dir
SDK 下载
# 模型下载
from huggingface_hub import snapshot_download
model_dir = snapshot_download('MiniMaxAI/MiniMax-Music3')
Git 下载
请确保 lfs 已经被正确安装
git lfs install
git clone https://huggingface.co/MiniMaxAI/MiniMax-Music3
如果您希望跳过 lfs 大文件下载,可以使用如下命令
GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/MiniMaxAI/MiniMax-Music3
模型文件托管在 Hugging Face Hub,使用 HF CLI / SDK / Git 直接下载,不经过本站。
PyTorch / Transformers 使用
安装 Transformers
pip install -U transformers torch
模型加载和推理
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained('MiniMaxAI/MiniMax-Music3')
tokenizer = AutoTokenizer.from_pretrained('MiniMaxAI/MiniMax-Music3')
完整文档
---
library_name: diffusers
pipeline_tag: text-to-audio
tags:
- music-generation
- text-to-music
- pytorch
- sglang-omni
---
<div align="center">
<img width="100%" src="figures/Music3.png" alt="MiniMax">
</div>
<p align="center">
<a href="https://agent.minimax.io/" target="_blank"><img src="https://img.shields.io/badge/MiniMax%20Agent-FF6C37?logo=minimax&logoColor=white" alt="MiniMax Agent"></a>
<a href="https://platform.minimax.io/docs/guides/text-generation" target="_blank"><img src="https://img.shields.io/badge/API-FF6C37?logo=minimax&logoColor=white" alt="API"></a>
<a href="https://www.minimax.io" target="_blank"><img src="https://img.shields.io/badge/MiniMax%20Website-FF6C37?logo=minimax&logoColor=white" alt="MiniMax Website"></a>
<br>
<a href="https://modelscope.cn/organization/minimax" target="_blank" rel="noopener noreferrer"><img alt="ModelScope MiniMax AI" src="https://img.shields.io/badge/ModelScope-MiniMax%20AI-white?labelColor=%23EF3D5D"></a>
<a href="https://platform.minimaxi.com/docs/faq/contact-us" target="_blank"><img src="https://img.shields.io/badge/WeChat-07C160?logo=wechat&logoColor=white" alt="WeChat"></a>
<a href="https://discord.com/invite/DPC4AHFCBw" target="_blank"><img src="https://img.shields.io/badge/Discord-5865F2?logo=discord&logoColor=white" alt="Discord"></a>
<a href="https://huggingface.co/MiniMaxAI" target="_blank"><img src="https://img.shields.io/badge/Hugging%20Face-FFD21E?logo=huggingface&logoColor=black" alt="Hugging Face"></a>
<a href="https://github.com/MiniMax-AI/MiniMax-Music3" target="_blank"><img src="https://img.shields.io/badge/GitHub-181717?logo=github&logoColor=white" alt="GitHub"></a>
<a href="https://huggingface.co/MiniMaxAI/MiniMax-Music3/blob/main/LICENSE" target="_blank"><img src="https://img.shields.io/badge/LICENSE-4CAF50?logo=creativecommons&logoColor=white" alt="LICENSE"></a>
</p>
MiniMax Music 3
MiniMax Music 3 is a high-performance music generation model for creating complete songs up to five minutes long. Conditioned on lyrics and a detailed music description, it generates structurally coherent songs with expressive vocals, evolving arrangements, and stable long-form audio quality.
MiniMax Music 3 combines an 8B Global LLM for long-range musical structure, a 0.6B Local LLM for frame-level acoustic detail, and a continuous hidden-state synthesis system based on Flow Matching and Flow-VAE. The model produces 32 kHz, 16-bit stereo WAV audio.
Demo
Explore music generation examples on the MiniMax Music 3 Demo.
<p align="center">
<img width="100%" src="figures/music3.0-Architecture-Diagram.png">
</p>
Complete Songs with Long-Range Coherence
MiniMax Music 3 natively supports full-song generation up to five minutes. The model maintains musical themes, rhythm, vocal identity, and arrangement progression across long sequences, enabling complete structures such as intro, verse, pre-chorus, chorus, bridge, instrumental break, and outro.
Fine-Grained Music Control
The model accepts two complementary inputs:
- Lyrics define the words to be sung and may include explicit section tags such as
[Intro],[Verse],[Pre-Chorus],[Chorus],[Post-Chorus],[Bridge],[Instrumental],[Solo], and[Outro].
- Music description defines the musical style, emotional progression, vocal performance, instrumentation, arrangement, and production profile.
For precise control, we recommend using a Structured Caption with three sections:
- Global Metadata: genre, subgenre, BPM, key, scale, emotional progression, listening scenario, and production profile.
- Vocal Details: vocal gender, timbre, performance style, harmony, backing vocals, and vocal effects.
- Arrangement: primary and secondary instruments, section-level instrument evolution, groove, bass, percussion, textures, and spatial effects.
This representation allows the model to follow not only a global style, but also the musical development of the song over time.
Hybrid-LM
MiniMax Music 3 uses a hierarchical autoregressive architecture that separates global musical modeling from local acoustic modeling.
- The Global LLM (8B) predicts the first RVQ codebook frame by frame and models the song's long-range semantic and structural progression.
- The Local LLM (0.6B) predicts the remaining acoustic codebooks within each frame and restores fine-grained acoustic information.
The Global LLM is initialized from Qwen3-8B. During training, its embedding and output layers are first adapted to semantic music tokens. The Global and Local LLMs are then jointly trained to model all RVQ codebooks.
Continuous Hidden-State Synthesis
Instead of decoding only from discrete RVQ tokens, the synthesis module fuses the final hidden states of the Global and Local LLMs. These continuous representations preserve richer acoustic information for vocal articulation, instrumental texture, and temporal continuity.
The synthesis path is:
Global and Local LLM hidden states
↓
Hidden-state fusion
↓
Flow Matching (2.4B)
↓
Flow-VAE latent
↓
Flow-VAE Decoder (123M)
↓
32 kHz stereo audioThe Flow-VAE architecture is adapted from MiniMax Speech and retrained for the dynamic range and spectral characteristics of music.
Music Tokenizer
The training tokenizer uses eight layers of Residual Vector Quantization (RVQ):
- The first semantic codebook contains 16,384 entries and captures the core musical semantics and structure.
- The remaining seven acoustic codebooks contain 1,024 entries each and represent residual acoustic details.
Training first optimizes the semantic codebook, then jointly trains all eight codebooks. At inference time, waveform synthesis uses the fused LLM hidden states and does not require the discrete tokenizer decoder.
How to Use
MiniMax Music 3 is supported by SGLang-Omni. Follow the official installation guide to prepare the runtime environment.
Download the Model
hf download MiniMaxAI/MiniMax-Music3 --local-dir /path/to/minimax_ttmWe recommend the following inference frameworks to serve the model:
- diffusers \- see diffusers docs
Serve with SGLang-Omni
sgl-omni serve --model-path MiniMaxAI/MiniMax-Music3 --port 8000Generate Music
The service uses the shared speech API. Put the lyrics in input and the music description in instructions. Put lyric structure tags such as [Verse] and [Chorus] on their own lines.
curl http://127.0.0.1:8000/v1/audio/speech \
-H 'Content-Type: application/json' \
-d '{
"model": "MiniMaxAI/MiniMax-Music3",
"input": "[Verse]\nMorning light filtering through the pine\n[Chorus]\nSoftly the world begins to breathe",
"instructions": "A warm acoustic pop song with intimate female vocals, fingerpicked guitar, soft piano, and a gradual emotional build into a wide final chorus.",
"response_format": "wav",
"seed": 7,
"max_new_tokens": 750,
"stream": false
}' \
--output minimax_music3.wavmax_new_tokens sets the maximum number of audio frames at 25 frames per second. Generation may finish before this limit when the model emits an end-of-audio token. The response is a 32 kHz, 16-bit stereo WAV file.