music generation

提供商DancingIguana
分类audio-generation
许可证apache-2.0
下载量100
星标0

简介

这是一个基于 Apache-2.0 开源协议的音频生成模型,旨在将文本描述转化为高质量的音乐片段。与复杂的专业编曲软件不同,它降低了音乐创作的门槛,开发者可以通过 API 快速将其集成到短视频背景音乐生成、游戏动态音效或个性化铃声等应用场景中。对于习惯使用 Suno 或 Udio 的用户来说,该模型提供了更灵活的部署可能,适合需要私有化部署或对生成成本有严格控制的开发者快速上手。

核心亮点

  • 开源协议灵活,支持商业化私有部署
  • 文本驱动生成,无需音乐理论基础
  • 适配短视频和游戏等轻量化音频场景
  • 上手难度低,可快速集成至现有工作流

使用方法

安装依赖
# 安装 Hugging Face transformers
pip install transformers torch
SDK 使用
# 使用 transformers 加载模型
from transformers import AutoModel, AutoTokenizer

model = AutoModel.from_pretrained("DancingIguana/music-generation")
tokenizer = AutoTokenizer.from_pretrained("DancingIguana/music-generation")

Hugging Face 下载

我们推荐使用命令行或者 Hugging Face Hub SDK 来进行模型的下载。

操作指引:在下载前,请先通过如下命令安装 huggingface_hub:

操作指引
pip install -U huggingface_hub

命令行下载

下载完整模型库

下载完整模型库
huggingface-cli download DancingIguana/music-generation

下载单个文件到指定本地文件夹(以下载 config.json 到当前路径下 ./dir 目录为例)

下载单个文件到指定本地文件夹(以下载 config.json 到当前路径下 ./dir 目录为例)
huggingface-cli download DancingIguana/music-generation config.json --local-dir ./dir

更多命令行下载选项,可参见官方文档

SDK 下载

SDK 下载
# 模型下载
from huggingface_hub import snapshot_download
model_dir = snapshot_download('DancingIguana/music-generation')

Git 下载

请确保 lfs 已经被正确安装

Git 下载
git lfs install
git clone https://huggingface.co/DancingIguana/music-generation

如果您希望跳过 lfs 大文件下载,可以使用如下命令

跳过 LFS
GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/DancingIguana/music-generation

模型文件托管在 Hugging Face Hub,使用 HF CLI / SDK / Git 直接下载,不经过本站。

PyTorch / Transformers 使用

安装 Transformers

安装 Transformers
pip install -U transformers torch

模型加载和推理

模型加载和推理
from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained('DancingIguana/music-generation')
tokenizer = AutoTokenizer.from_pretrained('DancingIguana/music-generation')

完整文档

来源: HuggingFace

---
license: apache-2.0
tags:

  • generated_from_trainer

model-index:
  • name: music-generation

results: []
---

<!-- This model card has been generated automatically according to the information the Trainer had access to. You
should probably proofread and complete it, then remove this comment. -->

music-generation

This model a trained from scratch version of distilgpt2 on a dataset where the text represents musical notes. The dataset consists of one stream of notes from MIDI files (the stream with most notes), where all of the melodies were transposed either to C major or A minor. Also, the BPM of the song is ignored, the duration of each note is based on its quarter length.

Each element in the melody is represented by a series of letters and numbers with the following structure.

  • For a note: ns[pitch of the note as a string]s[duration]

* Examples: nsC4s0p25, nsF7s1p0,
  • For a rest: rs[duration]:

* Examples: rs0p5, rs1q6
  • For a chord: cs[number of notes in chord]s[pitches of chords separated by "s"]s[duration]

* Examples: cs2sE7sF7s1q3, cs2sG3sGw3s0p25

The following special symbols are replaced in the strings by the following:
  • . = p

  • / = q

  • # =

  • - = t

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 0.0005

  • train_batch_size: 32

  • eval_batch_size: 32

  • seed: 42

  • gradient_accumulation_steps: 8

  • total_train_batch_size: 256

  • optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08

  • lr_scheduler_type: cosine

  • lr_scheduler_warmup_steps: 1000

  • num_epochs: 100

  • mixed_precision_training: Native AMP

Training results

Framework versions

  • Transformers 4.19.4
  • Pytorch 1.11.0+cu113
  • Datasets 2.2.2
  • Tokenizers 0.12.1