{"id":13862094,"url":"https://github.com/webpolis/musai","last_synced_at":"2025-04-12T10:40:31.787Z","repository":{"id":176503539,"uuid":"646207511","full_name":"webpolis/musai","owner":"webpolis","description":"Machine learning-powered music generation. Full-featured tokenizer, customization options, and high-quality output files. Integration with music production tools.","archived":false,"fork":false,"pushed_at":"2025-02-03T22:11:23.000Z","size":15882,"stargazers_count":9,"open_issues_count":0,"forks_count":1,"subscribers_count":2,"default_branch":"main","last_synced_at":"2025-03-26T05:33:10.253Z","etag":null,"topics":["deep-learning","generative-art","large-language-models","llm","machine-learning","midi","music","music-generation","nlp","recurrent-neural-networks","rnn","text-generation","tokenizer","vae","variational-autoencoder"],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/webpolis.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2023-05-27T16:23:11.000Z","updated_at":"2025-02-27T13:17:41.000Z","dependencies_parsed_at":"2024-01-21T03:59:05.834Z","dependency_job_id":"36f7181d-4d9f-45f1-9067-46df48bccd39","html_url":"https://github.com/webpolis/musai","commit_stats":null,"previous_names":["webpolis/musai"],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/webpolis%2Fmusai","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/webpolis%2Fmusai/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/webpolis%2Fmusai/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/webpolis%2Fmusai/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/webpolis","download_url":"https://codeload.github.com/webpolis/musai/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":248556573,"owners_count":21124141,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["deep-learning","generative-art","large-language-models","llm","machine-learning","midi","music","music-generation","nlp","recurrent-neural-networks","rnn","text-generation","tokenizer","vae","variational-autoencoder"],"created_at":"2024-08-05T06:01:37.030Z","updated_at":"2025-04-12T10:40:31.778Z","avatar_url":"https://github.com/webpolis.png","language":"Python","funding_links":[],"categories":["Python"],"sub_categories":[],"readme":"# MusAI\n\nMusAI is an innovative project that leverages the power of machine learning to generate unique and creative MIDI music sequences. With MusAI, you can explore the intersection of art and technology, and unleash your creativity by generating original music compositions.\n\n## Features\n\n- Full-featured tokenizer using parallelization via [Ray](https://www.ray.io/)\n- MIDI music generation using a combination of architectures ([RWKV](https://github.com/BlinkDL/RWKV-LM), [VAE](https://en.wikipedia.org/wiki/Variational_autoencoder), etc.)\n- Fine-tune or generate a new model from scratch using a custom dataset \n- Instrument based sequence training and generation (drums, bass, etc.)\n- Pre-trained embeddings using a Variational Autoencoder ([Experimental](https://github.com/webpolis/musai/wiki/Experimental))\n- Adjustable parameters to customize the style and complexity of the generated music\n- High-quality output MIDI files for further refinement or direct use in your projects\n- Seamless integration with your favorite music production tools via VST bridge (@WIP)\n\n## Installation\n\n`pip install -U -r requirements.txt`\n\n## Usage\n\nThe typical workflow is:\n\n- Convert MIDI files into tokens\n- Train the model\n- Generate new sequences\n\n### [Tokenizer](src/tools/tokenizer.py)\n\n\n### Usage\n\n```sh\ntokenizer.py [-h] [-t TOKENS_PATH] [-m MIDIS_PATH] [-g MIDIS_GLOB] [-b] [-p] [-a {REMI,MMM}] [-c CLASSES]\n                    [-r CLASSES_REQ] [-l LENGTH] [-d]\n\noptions:\n  -h, --help            show this help message and exit\n  -t TOKENS_PATH, --tokens_path TOKENS_PATH\n                        The output path were tokens are saved\n  -m MIDIS_PATH, --midis_path MIDIS_PATH\n                        The path where MIDI files can be located or a file containing a list of paths\n  -g MIDIS_GLOB, --midis_glob MIDIS_GLOB\n                        The glob pattern used to locate MIDI files\n  -b, --bpe             Applies BPE to the corpora of tokens\n  -p PARAMS_PATH, --preload PARAMS_PATH\n                        Absolute path to existing token_params.cfg settings\n  -a {REMI,MMM}, --algo {REMI,MMM}\n                        Tokenization algorithm\n  -c CLASSES, --classes CLASSES\n                        Only extract these instruments classes (e.g. 1,14,16,3,4,10,11)\n  -r CLASSES_REQ, --classes_req CLASSES_REQ\n                        Minimum set of instruments classes required (e.g. 1,14,16)\n  -l LENGTH, --length LENGTH\n                        Minimum sequence length (in beats)\n  -n MAX_FILES, --num_limit MAX_FILES\n                        Limit number of files to process (random selection)\n  -d, --debug           Debug mode (disables Ray).\n\n```\n\n#### Instrument Classes\n\n```\nCLASS | NAME\n------+---------------------\n  0   | Piano\n  1   | Chromatic Percussion\n  2   | Organ\n  3   | Guitar\n  4   | Bass\n  5   | Strings\n  6   | Ensemble\n  7   | Brass\n  8   | Reed\n  9   | Pipe\n 10   | Synth Lead\n 11   | Synth Pad\n 12   | Synth Effects\n 13   | Ethnic\n 14   | Percussive\n 15   | Sound Effects  \u003c-- Effects are automatically removed because they don't introduce \n                           relevant information to the model.\n 16   | Drums\n```\n\n### [Trainer](src/tools/trainer.py)\n\n### (optional) Embeddings\n\n\u003e Additionally, you can choose to use a VAE model in replacement of the default architecture's embedding module ([read more](https://github.com/webpolis/musai/wiki/Experimental)).\n\n\u003e There is an extra cost on training performance if you choose to build the VAE embeddings from scratch (using `--vae_emb true`) while training the main model, so it is recommended to train the embeddings alone beforehand (`--vae_emb train` or see [example](notebooks/vae.ipynb)), but make sure you use the same values for the embeddings size (`--embed_num`) when building the final model.\n\n\u003e Training the final model using pre-trained embeddings (`--vae_emb path_to_pth_file`) will **save** significant _VRAM_.\n\nTo train the embeddings alone from scratch, change the arguments to match your needs and run:\n\n```sh\npython src/tools/trainer.py -t path_to_tokenized_dataset -o output_path -v train -e 768 -b 24 -p 20 -s 1000 -i 1e-5\n```\n\nThe saved embedding model will be stored in the output path with a name such as `embvae_#.pth` where `#` is the epoch number. Afterwards, you can use that file as the pre-trained embeddings for training the main and final model, using a command similar to:\n\n```sh\npython src/tools/trainer.py -t path_to_tokenized_dataset -o output_path -v path_to_pretrained_embeddings.pth -e 768 -c 2048 -n 12 -b 24 -p 100 -s 1000 -i 1e-5 -g -q\n```\n\nYou can avoid using VAE entirely and let the RWKV architecture build its own embeddings by removing the `-v` or `--vae_emb` option.\n\n### Usage\n\n```sh\ntrainer.py [-h] [-t TOKENS_PATH] [-o OUTPUT_PATH] [-m BASE_MODEL] [-r LORA_CKPT] [-c CTX_LEN]\n                  [-b BATCHES_NUM] [-e EMBED_NUM] [-n LAYERS_NUM] [-p EPOCHS_NUM] [-s STEPS_NUM] [-i LR_RATE]\n                  [-d LR_DECAY] [-a] [-l] [-g]\n\noptions:\n  -h, --help            show this help message and exit\n  -t DATASET_PATH, --dataset_path DATASET_PATH\n                        The path were tokens parameters were saved by the tokenizer\n  -x, --binidx          Dataset is in binidx format (Generated via https://github.com/Abel2076/json2binidx_tool) \n  -o OUTPUT_PATH, --output_path OUTPUT_PATH\n                        The output path were model binaries will be saved\n  -m BASE_MODEL, --base_model BASE_MODEL\n                        Full path for base model/checkpoint (*)\n  -r LORA_CKPT, --lora_ckpt LORA_CKPT\n                        Full path for LoRa checkpoint (*)\n  -v VAE_EMB, --vae_emb VAE_EMB\n                        The pre-trained VAE embeddings. Possible options: \n                        \"train\" for training alone, from scratch.\n                        \"train path_to_existing_embeddings.pth\" for training alone, from saved model.\n                        \"true\" for training from scratch together with the main model (slow).\n                        \"path_to_existing_embeddings.pth\" to use existing embeddings model while training main model (fast).\n  -c CTX_LEN, --ctx_len CTX_LEN\n                        The context length\n  -b BATCHES_NUM, --batches_num BATCHES_NUM\n                        Number of batches\n  -e EMBED_NUM, --embed_num EMBED_NUM\n                        Size of the embeddings dimension\n  -n LAYERS_NUM, --layers_num LAYERS_NUM\n                        Number of block layers (*)\n  -p EPOCHS_NUM, --epochs_num EPOCHS_NUM\n                        Number of epochs\n  -s STEPS_NUM, --steps_num STEPS_NUM\n                        Number of steps per epoch\n  -i LR_RATE, --lr_rate LR_RATE\n                        Learning rate. Initial \u0026 final derivates from it.\n  -d LR_DECAY, --lr_decay LR_DECAY\n                        Learning rate decay thru steps\n  -a, --attention       Enable tiny attention (*)\n  -l, --lora            Activate LoRa (Low-Rank Adaptation) (*)\n  -u, --offload         DeepSpeed offload (**)\n  -q, --head_qk         Enable head QK (*)\n\n```\n\n_* Only used when training the main model._\n_** Slower, but more VRAM room_\n\n### Runner\n\n(@WIP)\n\n\n## Examples\n\nCheck out the [examples](examples/) folder.\n\n## Text Generation\n\nThe trainer tool and model included in Musai can be easily reused for training a Large Language Model (LLM) or a Text Generation model of any size. Just provide the corresponding dataset processed with the [binidx](https://github.com/Abel2076/json2binidx_tool) and follow the given instructions (ignore the _Tokenizer_ section entirely).\n\n## Contributing\n\nContributions to MusAI are welcome! If you have any ideas, suggestions, or bug reports, please open an issue or submit a pull request.\n\n## License\n\nThis project is licensed under the [MIT License](LICENSE).\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fwebpolis%2Fmusai","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fwebpolis%2Fmusai","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fwebpolis%2Fmusai/lists"}