{"id":27974796,"url":"https://github.com/funaudiollm/inspiremusic","last_synced_at":"2025-05-15T01:04:07.352Z","repository":{"id":265357689,"uuid":"880111885","full_name":"FunAudioLLM/InspireMusic","owner":"FunAudioLLM","description":"InspireMusic: A Unified Framework for Music, Song, Audio Generation.","archived":false,"fork":false,"pushed_at":"2025-05-02T09:53:34.000Z","size":3813,"stargazers_count":1084,"open_issues_count":18,"forks_count":100,"subscribers_count":18,"default_branch":"main","last_synced_at":"2025-05-08T00:26:55.688Z","etag":null,"topics":["audio-generation","audio-processing","music-generation","pytorch"],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/FunAudioLLM.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE.txt","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null}},"created_at":"2024-10-29T06:13:25.000Z","updated_at":"2025-05-07T15:08:34.000Z","dependencies_parsed_at":null,"dependency_job_id":"7d154ac5-c919-4056-8352-2162509137b4","html_url":"https://github.com/FunAudioLLM/InspireMusic","commit_stats":null,"previous_names":["funaudiollm/inspiremusic"],"tags_count":1,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/FunAudioLLM%2FInspireMusic","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/FunAudioLLM%2FInspireMusic/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/FunAudioLLM%2FInspireMusic/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/FunAudioLLM%2FInspireMusic/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/FunAudioLLM","download_url":"https://codeload.github.com/FunAudioLLM/InspireMusic/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":254254040,"owners_count":22039792,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["audio-generation","audio-processing","music-generation","pytorch"],"created_at":"2025-05-08T00:22:50.060Z","updated_at":"2025-05-15T01:04:07.340Z","avatar_url":"https://github.com/FunAudioLLM.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"\u003cp\u003e \u003ca href=\"https://github.com/FunAudioLLM/InspireMusic\" target=\"_blank\"\u003e \u003cimg alt=\"logo\" src=\"./asset/logo.png\" width=\"100%\"\u003e\u003c/a\u003e\u003c/p\u003e\n\n\u003cp\u003e\n \u003ca href=\"https://funaudiollm.github.io/inspiremusic\" target=\"_blank\"\u003e\u003cimg alt=\"Demo\" src=\"https://img.shields.io/badge/Demo-InspireMusic?labelColor=%20%23FDB062\u0026label=InspireMusic\u0026color=%20%23f79009\"\u003e\u003c/a\u003e\n\u003ca href=\"https://github.com/FunAudioLLM/InspireMusic\" target=\"_blank\"\u003e\u003cimg alt=\"Code\" src=\"https://img.shields.io/badge/Code-InspireMusic?labelColor=%20%237372EB\u0026label=InspireMusic\u0026color=%20%235462eb\"\u003e\u003c/a\u003e\n\u003ca href=\"https://modelscope.cn/models/iic/InspireMusic\" target=\"_blank\"\u003e\u003cimg alt=\"Model\" src=\"https://img.shields.io/badge/InspireMusic-Model-green\"\u003e\u003c/a\u003e\n\u003ca href=\"https://modelscope.cn/studios/iic/InspireMusic/summary\" target=\"_blank\"\u003e\u003cimg alt=\"Space\" src=\"https://img.shields.io/badge/Spaces-ModelScope-pink?labelColor=%20%237b8afb\u0026label=Spaces\u0026color=%20%230a5af8\"\u003e\u003c/a\u003e\n\u003ca href=\"https://huggingface.co/spaces/FunAudioLLM/InspireMusic\" target=\"_blank\"\u003e\u003cimg alt=\"Space\" src=\"https://img.shields.io/badge/HuggingFace-Spaces?labelColor=%20%239b8afb\u0026label=Spaces\u0026color=%20%237a5af8\"\u003e\u003c/a\u003e\n\u003ca href=\"http://arxiv.org/abs/2503.00084\" target=\"_blank\"\u003e\u003cimg alt=\"Paper\" src=\"https://img.shields.io/badge/arXiv-Paper-green\"\u003e\u003c/a\u003e\n\u003c/p\u003e\n\n![GitHub Repo stars](https://img.shields.io/github/stars/FunAudioLLM/InspireMusic) Please support our community by starring it 感谢大家支持\n\n[**Highlights**](#highlights)\n| [**Introduction**](#introduction)\n| [**Installation**](#installation)\n| [**Quick Start**](#quick-start)\n| [**Tutorial**](https://github.com/FunAudioLLM/InspireMusic#tutorial)\n| [**Models**](#model-zoo)\n| [**Contact**](#contact)\n\n---\n\u003ca name=\"highlights\"\u003e\u003c/a\u003e\n## Highlights\n**InspireMusic** focuses on music generation, song generation, and audio generation.\n- A unified toolkit designed for music, song, and audio generation.\n- Music generation tasks with high audio quality. \n- Long-form music generation.\n\n\u003ca name=\"introduction\"\u003e\u003c/a\u003e\n## Introduction\n\u003e [!Note]\n\u003e This repo contains the algorithm infrastructure and some simple examples.\n\n\u003e [!Tip]\n\u003e To preview the performance, please refer to [InspireMusic Demo Page](https://funaudiollm.github.io/inspiremusic).\n\nInspireMusic is a toolkit for music, song, and audio generation. It consists of an autoregressive transformer with a flow-matching based model. This toolkit is for users to generate music, song, and audio. InspireMusic can generate high-quality music in long-form with text-to-music and music continuation. InspireMusic incorporates audio tokenizers with autoregressive transformer and flow-matching modeling to generate music, song, and audio with text and music prompts. The toolkit currently supports music generation.\n\n## InspireMusic\n\u003cp align=\"center\"\u003e\u003ctable\u003e\u003ctr\u003e\u003ctd style=\"text-align:center;\"\u003e\u003cimg alt=\"Light\" src=\"asset/InspireMusic.png\" width=\"100%\" /\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd style=\"text-align:center;\"\u003e\nFigure 1: An overview of the InspireMusic. We introduce InspireMusic, a toolkit designed for music, song, audio generation capable of producing high-quality long-form music. InspireMusic consists of the following three key components. Audio Tokenizers convert the raw audio waveform into discrete audio tokens that can be efficiently processed and trained by the autoregressive transformer model. Audio waveform of lower sampling rate has converted to discrete tokens via a high bitrate compression audio tokenizer\u003ca href=\"https://openreview.net/forum?id=yBlVlS2Fd9\" target=\"_blank\"\u003e\u003csup\u003e[1]\u003c/sup\u003e\u003c/a\u003e. Autoregressive Transformer model is based on Qwen2.5\u003ca href=\"https://arxiv.org/abs/2412.15115\" target=\"_blank\"\u003e\u003csup\u003e[2]\u003c/sup\u003e\u003c/a\u003e as the backbone model and is trained using a next-token prediction approach on both text and audio tokens, enabling it to generate coherent and contextually relevant token sequences. The audio and text tokens are the inputs of an autoregressive model with the next token prediction to generate tokens. Super-Resolution Flow-Matching Model based on flow modeling method, maps the generated tokens to latent features with high-resolution fine-grained acoustic details\u003ca href=\"https://arxiv.org/abs/2305.02765\" target=\"_blank\"\u003e\u003csup\u003e[3]\u003c/sup\u003e\u003c/a\u003e obtained from a higher sampling rate of audio to ensure the acoustic information flow connected with high fidelity through models. A vocoder then generates the final audio waveform from these enhanced latent features. InspireMusic supports a range of tasks including text-to-music, music continuation, music reconstruction and super resolution..\n\u003c/td\u003e\u003c/tr\u003e\u003c/table\u003e\u003c/p\u003e\n\n\u003ca name=\"installation\"\u003e\u003c/a\u003e\n## Installation\n### Clone\n- Clone the repo\n``` sh\ngit clone --recursive https://github.com/FunAudioLLM/InspireMusic.git\n# If you failed to clone submodule due to network failures, please run the following command until success\ncd InspireMusic\ngit submodule update --recursive\n# or you can download the third_party repo Matcha-TTS manually\ncd third_party \u0026\u0026 git clone https://github.com/shivammehta25/Matcha-TTS.git\n```\n\n### Install from Source\nInspireMusic requires Python\u003e=3.8, PyTorch\u003e=2.0.1, flash attention==2.6.2/2.6.3, CUDA\u003e=11.8. You can install the dependencies with the following commands:\n\n- Install Conda: please see https://docs.conda.io/en/latest/miniconda.html\n- Create Conda env:\n``` shell\nconda create -n inspiremusic python=3.8\nconda activate inspiremusic\ncd InspireMusic\n# pynini is required by WeTextProcessing, use conda to install it as it can be executed on all platforms.\nconda install -y -c conda-forge pynini==2.1.5\npip install -r requirements.txt -i https://mirrors.aliyun.com/pypi/simple/ --trusted-host=mirrors.aliyun.com\n# install flash attention to speedup training\npip install flash-attn --no-build-isolation\n```\n\n- Install within the package:\n```shell\ncd InspireMusic\n# You can run to install the packages\npython setup.py install\npip install flash-attn --no-build-isolation\n```\nWe also recommend having `sox` or `ffmpeg` installed, either through your system or Anaconda:\n```shell\n# # Install sox\n# ubuntu\nsudo apt-get install sox libsox-dev\n# centos\nsudo yum install sox sox-devel\n\n# Install ffmpeg\n# ubuntu\nsudo apt-get install ffmpeg\n# centos\nsudo yum install ffmpeg\n```\n\n### Use Docker\nRun the following command to build a docker image from Dockerfile provided.\n```shell\ndocker build -t inspiremusic .\n```\nRun the following command to start the docker container in interactive mode.\n```shell\ndocker run -ti --gpus all -v .:/workspace/InspireMusic inspiremusic\n```\n\n### Use Docker Compose\nRun the following command to build a docker compose environment and docker image from the docker-compose.yml file.\n```shell\ndocker compose up -d --build\n```\nRun the following command to attach to the docker container in interactive mode.\n```shell\ndocker exec -ti inspire-music bash\n```\n\n\u003ca name=\"quick-start\"\u003e\u003c/a\u003e\n### Quick Start\nHere is a quick example inference script for music generation. \n``` shell\ncd InspireMusic\nmkdir -p pretrained_models\n\n# Download models\n# ModelScope\ngit clone https://www.modelscope.cn/iic/InspireMusic-1.5B-Long.git pretrained_models/InspireMusic-1.5B-Long\n# HuggingFace\ngit clone https://huggingface.co/FunAudioLLM/InspireMusic-1.5B-Long.git pretrained_models/InspireMusic-1.5B-Long\n\ncd examples/music_generation\n# run a quick inference example\nsh infer_1.5b_long.sh\n```\n\nHere is a quick start running script to run music generation task including data preparation pipeline, model training, inference. \n``` shell\ncd InspireMusic/examples/music_generation/\nsh run.sh\n```\n\n### One-line Inference\n#### Text-to-music Task\nOne-line Shell script for text-to-music task.\n```shell\ncd examples/music_generation\n# with flow matching, use one-line command to get a quick try\npython -m inspiremusic.cli.inference\n\n# custom the config like the following one-line command\npython -m inspiremusic.cli.inference --task text-to-music -m \"InspireMusic-1.5B-Long\" -g 0 -t \"Experience soothing and sensual instrumental jazz with a touch of Bossa Nova, perfect for a relaxing restaurant or spa ambiance.\" -c intro -s 0.0 -e 30.0 -r \"exp/inspiremusic\" -o output -f wav \n\n# without flow matching, use one-line command to get a quick try\npython -m inspiremusic.cli.inference --task text-to-music -g 0 -t \"Experience soothing and sensual instrumental jazz with a touch of Bossa Nova, perfect for a relaxing restaurant or spa ambiance.\" --fast True\n```\n\nAlternatively, you can run the inference with just a few lines of Python code.\n```python\nfrom inspiremusic.cli.inference import InspireMusicModel\nfrom inspiremusic.cli.inference import env_variables\nif __name__ == \"__main__\":\n  env_variables()\n  model = InspireMusicModel(model_name = \"InspireMusic-Base\")\n  model.inference(\"text-to-music\", \"Experience soothing and sensual instrumental jazz with a touch of Bossa Nova, perfect for a relaxing restaurant or spa ambiance.\")\n```\n\n#### Music Continuation Task\nOne-line Shell script for music continuation task.\n```shell\ncd examples/music_generation\n# with flow matching\npython -m inspiremusic.cli.inference --task continuation -g 0 -a audio_prompt.wav\n# without flow matching\npython -m inspiremusic.cli.inference --task continuation -g 0 -a audio_prompt.wav --fast True\n```\n\nAlternatively, you can run the inference with just a few lines of Python code.\n```python\nfrom inspiremusic.cli.inference import InspireMusicModel\nfrom inspiremusic.cli.inference import env_variables\nif __name__ == \"__main__\":\n  env_variables()\n  model = InspireMusicModel(model_name = \"InspireMusic-Base\")\n  # just use audio prompt\n  model.inference(\"continuation\", None, \"audio_prompt.wav\")\n  # use both text prompt and audio prompt\n  model.inference(\"continuation\", \"Continue to generate jazz music.\", \"audio_prompt.wav\")\n```\n\u003ca name=\"model-zoo\"\u003e\u003c/a\u003e\n## Models\n### Download Models\nYou may download our pretrained InspireMusic models for music generation.\n```shell\n# use git to download models，please make sure git lfs is installed.\nmkdir -p pretrained_models\ngit clone https://www.modelscope.cn/iic/InspireMusic.git pretrained_models/InspireMusic\n```\n\n### Available Models\nCurrently, we open source the music generation models support 24KHz mono and 48KHz stereo audio. \nThe table below presents the links to the ModelScope and Huggingface model hub.\n\n| Model name                           | Model Links                                                                                                                                                                                                                                                                                                                                   | Remarks                                                                                                  |\n|--------------------------------------|-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------|\n| InspireMusic-Base-24kHz              | [![model](https://img.shields.io/badge/ModelScope-Model-green.svg)](https://modelscope.cn/models/iic/InspireMusic-Base-24kHz/summary) [![model](https://img.shields.io/badge/HuggingFace-Model-green.svg)](https://huggingface.co/FunAudioLLM/InspireMusic-Base-24kHz)                                                                        | Pre-trained Music Generation Model, 24kHz mono, 30s                                                      |\n| InspireMusic-Base                    | [![model](https://img.shields.io/badge/ModelScope-Model-green.svg)](https://modelscope.cn/models/iic/InspireMusic/summary) [![model](https://img.shields.io/badge/HuggingFace-Model-green.svg)](https://huggingface.co/FunAudioLLM/InspireMusic-Base)                                                                                         | Pre-trained Music Generation Model, 48kHz, 30s                                                           |\n| InspireMusic-1.5B-24kHz              | [![model](https://img.shields.io/badge/ModelScope-Model-green.svg)](https://modelscope.cn/models/iic/InspireMusic-1.5B-24kHz/summary) [![model](https://img.shields.io/badge/HuggingFace-Model-green.svg)](https://huggingface.co/FunAudioLLM/InspireMusic-1.5B-24kHz)                                                                        | Pre-trained Music Generation 1.5B Model, 24kHz mono, 30s                                                 |\n| InspireMusic-1.5B                    | [![model](https://img.shields.io/badge/ModelScope-Model-green.svg)](https://modelscope.cn/models/iic/InspireMusic-1.5B/summary) [![model](https://img.shields.io/badge/HuggingFace-Model-green.svg)](https://huggingface.co/FunAudioLLM/InspireMusic-1.5B)                                                                                    | Pre-trained Music Generation 1.5B Model, 48kHz, 30s                                                      |\n| InspireMusic-1.5B-Long               | [![model](https://img.shields.io/badge/ModelScope-Model-green.svg)](https://modelscope.cn/models/iic/InspireMusic-1.5B-Long/summary) [![model](https://img.shields.io/badge/HuggingFace-Model-green.svg)](https://huggingface.co/FunAudioLLM/InspireMusic-1.5B-Long)                                                                          | Pre-trained Music Generation 1.5B Model, 48kHz, support long-form music generation up to several minutes |\n| InspireSong-1.5B                     | [![model](https://img.shields.io/badge/ModelScope-Model-lightgrey.svg)]() [![model](https://img.shields.io/badge/HuggingFace-Model-lightgrey.svg)]()                                                                                                                                                                                          | Pre-trained Song Generation 1.5B Model, 48kHz stereo                                                     |\n| InspireAudio-1.5B                    | [![model](https://img.shields.io/badge/ModelScope-Model-lightgrey.svg)]() [![model](https://img.shields.io/badge/HuggingFace-Model-lightgrey.svg)]()                                                                                                                                                                                          | Pre-trained Audio Generation 1.5B Model, 48kHz stereo                                                    |\n| Wavtokenizer[\u003csup\u003e[1]\u003c/sup\u003e](https://openreview.net/forum?id=yBlVlS2Fd9) (75Hz) | [![model](https://img.shields.io/badge/ModelScope-Model-green.svg)](https://modelscope.cn/models/iic/InspireMusic-1.5B-Long/file/view/master?fileName=wavtokenizer%252Fmodel.pt) [![model](https://img.shields.io/badge/HuggingFace-Model-green.svg)](https://huggingface.co/FunAudioLLM/InspireMusic-1.5B-Long/tree/main/wavtokenizer)       | An extreme low bitrate audio tokenizer for music with one codebook at 24kHz audio.                       |\n| Music_tokenizer (75Hz)               | [![model](https://img.shields.io/badge/ModelScope-Model-green.svg)](https://modelscope.cn/models/iic/InspireMusic-1.5B-24kHz/file/view/master?fileName=music_tokenizer%252Fmodel.pt) [![model](https://img.shields.io/badge/HuggingFace-Model-green.svg)](https://huggingface.co/FunAudioLLM/InspireMusic-1.5B-24kHz/tree/main/music_tokenizer) | A music tokenizer based on HifiCodec\u003csup\u003e[3]\u003c/sup\u003e at 24kHz audio.                                       |\n| Music_tokenizer (150Hz)              | [![model](https://img.shields.io/badge/ModelScope-Model-green.svg)](https://modelscope.cn/models/iic/InspireMusic-1.5B-Long/file/view/master?fileName=music_tokenizer%252Fmodel.pt) [![model](https://img.shields.io/badge/HuggingFace-Model-green.svg)](https://huggingface.co/FunAudioLLM/InspireMusic-1.5B-Long/tree/main/music_tokenizer) | A music tokenizer based on HifiCodec\u003csup\u003e[3]\u003c/sup\u003e at 48kHz audio.                                       |\n\n\u003ca name=\"tutorial\"\u003e\u003c/a\u003e\n## Basic Usage\nAt the moment, InspireMusic contains the training and inference codes for [music generation](https://github.com/FunAudioLLM/InspireMusic/tree/main/examples/music_generation). \n\n### Training\nHere is an example to train LLM model, support BF16/FP16 training. \n```shell\ntorchrun --nnodes=1 --nproc_per_node=8 \\\n    --rdzv_id=1024 --rdzv_backend=\"c10d\" --rdzv_endpoint=\"localhost:0\" \\\n    inspiremusic/bin/train.py \\\n    --train_engine \"torch_ddp\" \\\n    --config conf/inspiremusic.yaml \\\n    --train_data data/train.data.list \\\n    --cv_data data/dev.data.list \\\n    --model llm \\\n    --model_dir `pwd`/exp/music_generation/llm/ \\\n    --tensorboard_dir `pwd`/tensorboard/music_generation/llm/ \\\n    --ddp.dist_backend \"nccl\" \\\n    --num_workers 8 \\\n    --prefetch 100 \\\n    --pin_memory \\\n    --deepspeed_config ./conf/ds_stage2.json \\\n    --deepspeed.save_states model+optimizer \\\n    --fp16\n```\n\nHere is an example code to train flow matching model, does not support FP16 training.\n```shell\ntorchrun --nnodes=1 --nproc_per_node=8 \\\n    --rdzv_id=1024 --rdzv_backend=\"c10d\" --rdzv_endpoint=\"localhost:0\" \\\n    inspiremusic/bin/train.py \\\n    --train_engine \"torch_ddp\" \\\n    --config conf/inspiremusic.yaml \\\n    --train_data data/train.data.list \\\n    --cv_data data/dev.data.list \\\n    --model flow \\\n    --model_dir `pwd`/exp/music_generation/flow/ \\\n    --tensorboard_dir `pwd`/tensorboard/music_generation/flow/ \\\n    --ddp.dist_backend \"nccl\" \\\n    --num_workers 8 \\\n    --prefetch 100 \\\n    --pin_memory \\\n    --deepspeed_config ./conf/ds_stage2.json \\\n    --deepspeed.save_states model+optimizer\n```\n\n### Inference\n\nHere is an example script to quickly do model inference.\n```shell\ncd InspireMusic/examples/music_generation/\nsh infer.sh\n```\nHere is an example code to run inference with normal mode, i.e., with flow matching model for text-to-music and music continuation tasks.\n```shell\npretrained_model_dir = \"pretrained_models/InspireMusic/\"\nfor task in 'text-to-music' 'continuation'; do\n  python inspiremusic/bin/inference.py --task $task \\\n      --gpu 0 \\\n      --config conf/inspiremusic.yaml \\\n      --prompt_data data/test/parquet/data.list \\\n      --flow_model $pretrained_model_dir/flow.pt \\\n      --llm_model $pretrained_model_dir/llm.pt \\\n      --music_tokenizer $pretrained_model_dir/music_tokenizer \\\n      --wavtokenizer $pretrained_model_dir/wavtokenizer \\\n      --result_dir `pwd`/exp/inspiremusic/${task}_test \\\n      --chorus verse \ndone\n```\nHere is an example code to run inference with fast mode, i.e., without flow matching model for text-to-music and music continuation tasks.\n```shell\npretrained_model_dir = \"pretrained_models/InspireMusic/\"\nfor task in 'text-to-music' 'continuation'; do\n  python inspiremusic/bin/inference.py --task $task \\\n      --gpu 0 \\\n      --config conf/inspiremusic.yaml \\\n      --prompt_data data/test/parquet/data.list \\\n      --flow_model $pretrained_model_dir/flow.pt \\\n      --llm_model $pretrained_model_dir/llm.pt \\\n      --music_tokenizer $pretrained_model_dir/music_tokenizer \\\n      --wavtokenizer $pretrained_model_dir/wavtokenizer \\\n      --result_dir `pwd`/exp/inspiremusic/${task}_test \\\n      --chorus verse \\\n      --fast \ndone\n```\n\n### Hardware Requirements\nPrevious test on H800 GPU, InspireMusic could generate 30 seconds audio with real-time factor (RTF) around 1.4~1.8. For normal mode, we recommend using hardware with at least 24GB of GPU memory for better experience. For fast mode, 12GB GPU memory is enough.\n\n## Roadmap\n- [x] 2024/12\n  - [x] 75Hz InspireMusic-Base model for music generation\n    \n- [x] 2025/01\n    - [x] Support to generate 48kHz\n    - [x] 75Hz InspireMusic-1.5B model for music generation\n    - [x] 75Hz InspireMusic-1.5B-Long model for long-form music generation\n\n- [x] 2025/02\n    - [x] Release technical report\n\n- [ ] Future work\n    - [ ] InspireAudio model for audio generation\n    - [ ] InspireSong model for song generation\n    - [ ] Support multilingual generation\n\n## Citation\n```bibtex\n@misc{InspireMusic2025,\n      title={InspireMusic: Integrating Super Resolution and Large Language Model for High-Fidelity Long-Form Music Generation}, \n      author={Chong Zhang and Yukun Ma and Qian Chen and Wen Wang and Shengkui Zhao and Zexu Pan and Hao Wang and Chongjia Ni and Trung Hieu Nguyen and Kun Zhou and Yidi Jiang and Chaohong Tan and Zhifu Gao and Zhihao Du and Bin Ma},\n      year={2025},\n      eprint={2503.00084},\n      archivePrefix={arXiv},\n      primaryClass={cs.SD},\n      url={https://arxiv.org/abs/2503.00084}, \n}\n```\n---\n\u003ca name=\"contact\"\u003e\u003c/a\u003e\n## Community \u0026 Discussion\n* Welcome to join our DingTalk and WeChat groups to share and discuss algorithms, technology, and user experience feedback. You may scan the following QR codes to join our official chat groups accordingly.\n\n\u003cp align=\"center\"\u003e\u003ctable\u003e\u003ctr\u003e\n\u003ctd style=\"text-align:center;\"\u003e\u003ca href=\"./asset/dingding.png\"\u003e\u003cimg alt=\"FunAudioLLM in DingTalk\" src=\"https://img.shields.io/badge/FunAudioLLM-DingTalk-green\"\u003e\u003c/a\u003e\u003c/td\u003e\n\u003ctd style=\"text-align:center;\"\u003e\u003ca href=\"./asset/QR.jpg\"\u003e\u003cimg alt=\"InspireMusic in WeChat\" src=\"https://img.shields.io/badge/InspireMusic-WeChat-green\"\u003e\u003c/a\u003e\u003c/td\u003e\n\u003c/tr\u003e\u003ctr\u003e\u003ctd style=\"text-align:center;\"\u003e\u003cimg alt=\"Light\" src=\"./asset/dingding.png\" width=\"100%\" /\u003e\n\u003ctd style=\"text-align:center;\"\u003e\u003cimg alt=\"Light\" src=\"./asset/QR.jpg\" width=\"80%\" /\u003e\u003c/td\u003e\n\u003c/tr\u003e\u003c/table\u003e\u003c/p\u003e\n\n## Acknowledgement\n1. codes from CosyVoice.\n3. codes from WavTokenizer.\n4. codes from AcademiCodec.\n5. codes from FunASR.\n6. codes from FunCodec.\n7. codes from Matcha-TTS.\n9. codes from WeNet.\n\n## Disclaimer\nThe content provided above is for research purposes only and is intended to demonstrate technical capabilities. Some examples are sourced from the internet. If any content infringes on your rights, please contact us to request its removal.\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Ffunaudiollm%2Finspiremusic","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Ffunaudiollm%2Finspiremusic","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Ffunaudiollm%2Finspiremusic/lists"}