MiniMax Music3

提供商MiniMaxAI
分类text-to-audio
许可证Apache-2.0
下载量8.6K
星标0

简介

MiniMax Music3 是一款由 MiniMax 推出的文本生成音频模型,专注于将文字描述直接转化为高质量的音乐片段。它打破了传统音乐创作的高门槛,用户无需掌握乐理或复杂的 DAW 软件,只需通过自然语言描述风格、情绪或场景,即可快速生成具有商业质感的音频。对于国内开发者而言,其 Apache-2.0 协议提供了极高的灵活度,非常适合集成到短视频配乐、游戏背景音或 AI 虚拟人等应用场景中,是目前实现“文本到音乐”快速原型的理想工具。

核心亮点

  • 支持自然语言驱动,零门槛创作高质量音乐
  • Apache-2.0 协议,对开发者极其友好且灵活
  • 极速生成商业级音频,适配短视频与游戏场景
  • 精准捕捉情绪与风格,大幅提升音频素材产出率

使用方法

安装依赖
# 安装 Hugging Face transformers
pip install transformers torch
SDK 使用
# 使用 transformers 加载模型
from transformers import AutoModel, AutoTokenizer

model = AutoModel.from_pretrained("MiniMaxAI/MiniMax-Music3")
tokenizer = AutoTokenizer.from_pretrained("MiniMaxAI/MiniMax-Music3")

Hugging Face 下载

我们推荐使用命令行或者 Hugging Face Hub SDK 来进行模型的下载。

操作指引:在下载前,请先通过如下命令安装 huggingface_hub:

操作指引
pip install -U huggingface_hub

命令行下载

下载完整模型库

下载完整模型库
huggingface-cli download MiniMaxAI/MiniMax-Music3

下载单个文件到指定本地文件夹(以下载 config.json 到当前路径下 ./dir 目录为例)

下载单个文件到指定本地文件夹(以下载 config.json 到当前路径下 ./dir 目录为例)
huggingface-cli download MiniMaxAI/MiniMax-Music3 config.json --local-dir ./dir

更多命令行下载选项,可参见官方文档

SDK 下载

SDK 下载
# 模型下载
from huggingface_hub import snapshot_download
model_dir = snapshot_download('MiniMaxAI/MiniMax-Music3')

Git 下载

请确保 lfs 已经被正确安装

Git 下载
git lfs install
git clone https://huggingface.co/MiniMaxAI/MiniMax-Music3

如果您希望跳过 lfs 大文件下载,可以使用如下命令

跳过 LFS
GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/MiniMaxAI/MiniMax-Music3

模型文件托管在 Hugging Face Hub,使用 HF CLI / SDK / Git 直接下载,不经过本站。

PyTorch / Transformers 使用

安装 Transformers

安装 Transformers
pip install -U transformers torch

模型加载和推理

模型加载和推理
from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained('MiniMaxAI/MiniMax-Music3')
tokenizer = AutoTokenizer.from_pretrained('MiniMaxAI/MiniMax-Music3')

完整文档

来源: HuggingFace

---
library_name: diffusers
pipeline_tag: text-to-audio
tags:
- music-generation
- text-to-music
- pytorch
- sglang-omni
---

<div align="center">
<img width="100%" src="figures/Music3.png" alt="MiniMax">
</div>
<p align="center">
<a href="https://agent.minimax.io/" target="_blank"><img src="https://img.shields.io/badge/MiniMax%20Agent-FF6C37?logo=minimax&logoColor=white" alt="MiniMax Agent"></a>
<a href="https://platform.minimax.io/docs/guides/text-generation" target="_blank"><img src="https://img.shields.io/badge/API-FF6C37?logo=minimax&logoColor=white" alt="API"></a>
<a href="https://www.minimax.io" target="_blank"><img src="https://img.shields.io/badge/MiniMax%20Website-FF6C37?logo=minimax&logoColor=white" alt="MiniMax Website"></a>
<br>
<a href="https://modelscope.cn/organization/minimax" target="_blank" rel="noopener noreferrer"><img alt="ModelScope MiniMax AI" src="https://img.shields.io/badge/ModelScope-MiniMax%20AI-white?labelColor=%23EF3D5D"></a>
<a href="https://platform.minimaxi.com/docs/faq/contact-us" target="_blank"><img src="https://img.shields.io/badge/WeChat-07C160?logo=wechat&logoColor=white" alt="WeChat"></a>
<a href="https://discord.com/invite/DPC4AHFCBw" target="_blank"><img src="https://img.shields.io/badge/Discord-5865F2?logo=discord&logoColor=white" alt="Discord"></a>
<a href="https://huggingface.co/MiniMaxAI" target="_blank"><img src="https://img.shields.io/badge/Hugging%20Face-FFD21E?logo=huggingface&logoColor=black" alt="Hugging Face"></a>
<a href="https://github.com/MiniMax-AI/MiniMax-Music3" target="_blank"><img src="https://img.shields.io/badge/GitHub-181717?logo=github&logoColor=white" alt="GitHub"></a>
<a href="https://huggingface.co/MiniMaxAI/MiniMax-Music3/blob/main/LICENSE" target="_blank"><img src="https://img.shields.io/badge/LICENSE-4CAF50?logo=creativecommons&logoColor=white" alt="LICENSE"></a>
</p>

MiniMax Music 3

MiniMax Music 3 is a high-performance music generation model for creating complete songs up to five minutes long. Conditioned on lyrics and a detailed music description, it generates structurally coherent songs with expressive vocals, evolving arrangements, and stable long-form audio quality.

MiniMax Music 3 combines an 8B Global LLM for long-range musical structure, a 0.6B Local LLM for frame-level acoustic detail, and a continuous hidden-state synthesis system based on Flow Matching and Flow-VAE. The model produces 32 kHz, 16-bit stereo WAV audio.

Demo

Explore music generation examples on the MiniMax Music 3 Demo.

<p align="center">
<img width="100%" src="figures/music3.0-Architecture-Diagram.png">
</p>

Complete Songs with Long-Range Coherence

MiniMax Music 3 natively supports full-song generation up to five minutes. The model maintains musical themes, rhythm, vocal identity, and arrangement progression across long sequences, enabling complete structures such as intro, verse, pre-chorus, chorus, bridge, instrumental break, and outro.

Fine-Grained Music Control

The model accepts two complementary inputs:

  • Lyrics define the words to be sung and may include explicit section tags such as [Intro], [Verse], [Pre-Chorus], [Chorus], [Post-Chorus], [Bridge], [Instrumental], [Solo], and [Outro].
  • Music description defines the musical style, emotional progression, vocal performance, instrumentation, arrangement, and production profile.

For precise control, we recommend using a Structured Caption with three sections:

  • Global Metadata: genre, subgenre, BPM, key, scale, emotional progression, listening scenario, and production profile.
  • Vocal Details: vocal gender, timbre, performance style, harmony, backing vocals, and vocal effects.
  • Arrangement: primary and secondary instruments, section-level instrument evolution, groove, bass, percussion, textures, and spatial effects.

This representation allows the model to follow not only a global style, but also the musical development of the song over time.

Hybrid-LM

MiniMax Music 3 uses a hierarchical autoregressive architecture that separates global musical modeling from local acoustic modeling.

  • The Global LLM (8B) predicts the first RVQ codebook frame by frame and models the song's long-range semantic and structural progression.
  • The Local LLM (0.6B) predicts the remaining acoustic codebooks within each frame and restores fine-grained acoustic information.

The Global LLM is initialized from Qwen3-8B. During training, its embedding and output layers are first adapted to semantic music tokens. The Global and Local LLMs are then jointly trained to model all RVQ codebooks.

Continuous Hidden-State Synthesis

Instead of decoding only from discrete RVQ tokens, the synthesis module fuses the final hidden states of the Global and Local LLMs. These continuous representations preserve richer acoustic information for vocal articulation, instrumental texture, and temporal continuity.

The synthesis path is:

text
Global and Local LLM hidden states
                ↓
       Hidden-state fusion
                ↓
     Flow Matching (2.4B)
                ↓
        Flow-VAE latent
                ↓
    Flow-VAE Decoder (123M)
                ↓
       32 kHz stereo audio

The Flow-VAE architecture is adapted from MiniMax Speech and retrained for the dynamic range and spectral characteristics of music.

Music Tokenizer

The training tokenizer uses eight layers of Residual Vector Quantization (RVQ):

  • The first semantic codebook contains 16,384 entries and captures the core musical semantics and structure.
  • The remaining seven acoustic codebooks contain 1,024 entries each and represent residual acoustic details.

Training first optimizes the semantic codebook, then jointly trains all eight codebooks. At inference time, waveform synthesis uses the fused LLM hidden states and does not require the discrete tokenizer decoder.

How to Use

MiniMax Music 3 is supported by SGLang-Omni. Follow the official installation guide to prepare the runtime environment.

Download the Model

bash
hf download MiniMaxAI/MiniMax-Music3 --local-dir /path/to/minimax_ttm

We recommend the following inference frameworks to serve the model:

Serve with SGLang-Omni

bash
sgl-omni serve --model-path MiniMaxAI/MiniMax-Music3 --port 8000

Generate Music

The service uses the shared speech API. Put the lyrics in input and the music description in instructions. Put lyric structure tags such as [Verse] and [Chorus] on their own lines.

bash
curl http://127.0.0.1:8000/v1/audio/speech \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "MiniMaxAI/MiniMax-Music3",
    "input": "[Verse]\nMorning light filtering through the pine\n[Chorus]\nSoftly the world begins to breathe",
    "instructions": "A warm acoustic pop song with intimate female vocals, fingerpicked guitar, soft piano, and a gradual emotional build into a wide final chorus.",
    "response_format": "wav",
    "seed": 7,
    "max_new_tokens": 750,
    "stream": false
  }' \
  --output minimax_music3.wav

max_new_tokens sets the maximum number of audio frames at 25 frames per second. Generation may finish before this limit when the model emits an end-of-audio token. The response is a 32 kHz, 16-bit stereo WAV file.

R