{"id":9982828,"url":"https://github.com/PlayVoice/whisper-vits-svc","last_synced_at":"2025-08-30T19:30:44.942Z","repository":{"id":140990730,"uuid":"539767354","full_name":"PlayVoice/whisper-vits-svc","owner":"PlayVoice","description":"Core Engine of Singing Voice Conversion \u0026 Singing Voice Clone","archived":false,"fork":false,"pushed_at":"2024-04-23T10:50:32.000Z","size":43324,"stargazers_count":2643,"open_issues_count":50,"forks_count":922,"subscribers_count":29,"default_branch":"bigvgan-mix-v2","last_synced_at":"2024-10-29T15:28:19.940Z","etag":null,"topics":["change","diff-svc","diffusion","diffusion-svc","singing-voice-conversion","sovits","svc","vits","vits2","voice"],"latest_commit_sha":null,"homepage":"https://huggingface.co/spaces/maxmax20160403/sovits5.0","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/PlayVoice.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2022-09-22T02:36:42.000Z","updated_at":"2024-10-29T02:57:33.000Z","dependencies_parsed_at":"2024-04-23T12:23:47.061Z","dependency_job_id":null,"html_url":"https://github.com/PlayVoice/whisper-vits-svc","commit_stats":null,"previous_names":["playvoice/vits-svc","playvoice/so-vits-svc-5.0","playvoice/whisper-vits-svc"],"tags_count":15,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/PlayVoice%2Fwhisper-vits-svc","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/PlayVoice%2Fwhisper-vits-svc/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/PlayVoice%2Fwhisper-vits-svc/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/PlayVoice%2Fwhisper-vits-svc/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/PlayVoice","download_url":"https://codeload.github.com/PlayVoice/whisper-vits-svc/tar.gz/refs/heads/bigvgan-mix-v2","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":231443882,"owners_count":18377608,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["change","diff-svc","diffusion","diffusion-svc","singing-voice-conversion","sovits","svc","vits","vits2","voice"],"created_at":"2024-05-20T09:04:06.516Z","updated_at":"2024-12-27T17:30:42.577Z","avatar_url":"https://github.com/PlayVoice.png","language":"Python","funding_links":[],"categories":["Python","语音识别与合成_其他"],"sub_categories":["网络服务_其他"],"readme":"\u003cdiv align=\"center\"\u003e\n\u003ch1\u003e Variational Inference with adversarial learning for end-to-end Singing Voice Conversion based on VITS \u003c/h1\u003e\n    \n[![Hugging Face Spaces](https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-Spaces-blue)](https://huggingface.co/spaces/maxmax20160403/sovits5.0)\n\u003cimg alt=\"GitHub Repo stars\" src=\"https://img.shields.io/github/stars/PlayVoice/so-vits-svc-5.0\"\u003e\n\u003cimg alt=\"GitHub forks\" src=\"https://img.shields.io/github/forks/PlayVoice/so-vits-svc-5.0\"\u003e\n\u003cimg alt=\"GitHub issues\" src=\"https://img.shields.io/github/issues/PlayVoice/so-vits-svc-5.0\"\u003e\n\u003cimg alt=\"GitHub\" src=\"https://img.shields.io/github/license/PlayVoice/so-vits-svc-5.0\"\u003e\n\n[中文文档](./README_ZH.md)\n\nThe tree [bigvgan-mix-v2](https://github.com/PlayVoice/whisper-vits-svc/tree/bigvgan-mix-v2) has good audio quality\n\nThe tree [RoFormer-HiFTNet](https://github.com/PlayVoice/whisper-vits-svc/tree/RoFormer-HiFTNet) has fast infer speed\n\nNo More Upgrade\n\n\u003c/div\u003e\n\n- This project targets deep learning beginners, basic knowledge of Python and PyTorch are the prerequisites for this project;\n- This project aims to help deep learning beginners get rid of boring pure theoretical learning, and master the basic knowledge of deep learning by combining it with practices;\n- This project does not support real-time voice converting; (need to replace whisper if real-time voice converting is what you are looking for)\n- This project will not develop one-click packages for other purposes;\n\n![vits-5.0-frame](https://github.com/PlayVoice/so-vits-svc-5.0/assets/16432329/3854b281-8f97-4016-875b-6eb663c92466)\n\n- A minimum VRAM requirement of 6GB for training\n\n- Support for multiple speakers\n\n- Create unique speakers through speaker mixing\n\n- It can even convert voices with light accompaniment\n\n- You can edit F0 using Excel\n\nhttps://github.com/PlayVoice/so-vits-svc-5.0/assets/16432329/6a09805e-ab93-47fe-9a14-9cbc1e0e7c3a\n\nPowered by [@ShadowVap](https://space.bilibili.com/491283091)\n\n## Model properties\n\n| Feature | From | Status | Function |\n| :--- | :--- | :--- | :--- |\n| whisper | OpenAI | ✅ | strong noise immunity |\n| bigvgan  | NVIDA | ✅ | alias and snake | The formant is clearer and the sound quality is obviously improved |\n| natural speech | Microsoft | ✅ | reduce mispronunciation |\n| neural source-filter | Xin Wang | ✅ | solve the problem of audio F0 discontinuity |\n| pitch quantization | Xin Wang | ✅ | quantize the F0 for embedding |\n| speaker encoder | Google | ✅ | Timbre Encoding and Clustering |\n| GRL for speaker | Ubisoft |✅ | Preventing Encoder Leakage Timbre |\n| SNAC |  Samsung | ✅ | One Shot Clone of VITS |\n| SCLN |  Microsoft | ✅ | Improve Clone |\n| Diffusion |  HuaWei | ✅ | Improve sound quality |\n| PPG perturbation | this project | ✅ | Improved noise immunity and de-timbre |\n| HuBERT perturbation | this project | ✅ | Improved noise immunity and de-timbre |\n| VAE perturbation | this project | ✅ | Improve sound quality |\n| MIX encoder | this project | ✅ | Improve conversion stability |\n| USP infer | this project | ✅ | Improve conversion stability |\n| HiFTNet | Columbia University | ✅ | NSF-iSTFTNet for speed up |\n| RoFormer | Zhuiyi Technology | ✅ | Rotary Positional Embeddings |\n\ndue to the use of data perturbation, it takes longer to train than other projects.\n\n**USP : Unvoice and Silence with Pitch when infer**\n![vits_svc_usp](https://github.com/PlayVoice/so-vits-svc-5.0/assets/16432329/ba733b48-8a89-4612-83e0-a0745587d150)\n\n## Why mix\n\n![mix_frame](https://github.com/PlayVoice/whisper-vits-svc/assets/16432329/3ffa1be0-1a21-4752-96b5-6220f98f2313)\n\n## Plug-In-Diffusion\n\n![plug-in-diffusion](https://github.com/PlayVoice/so-vits-svc-5.0/assets/16432329/54a61c90-a97b-404d-9cc9-a2151b2db28f)\n\n## Setup Environment\n\n1. Install [PyTorch](https://pytorch.org/get-started/locally/).\n\n2. Install project dependencies\n    ```shell\n    pip install -i https://pypi.tuna.tsinghua.edu.cn/simple -r requirements.txt\n    ```\n    **Note: whisper is already built-in, do not install it again otherwise it will cuase conflict and error**\n3. Download the Timbre Encoder: [Speaker-Encoder by @mueller91](https://drive.google.com/drive/folders/15oeBYf6Qn1edONkVLXe82MzdIi3O_9m3), put `best_model.pth.tar`  into `speaker_pretrain/`.\n\n4. Download whisper model [whisper-large-v2](https://openaipublic.azureedge.net/main/whisper/models/81f7c96c852ee8fc832187b0132e569d6c3065a3252ed18e56effd0b6a73e524/large-v2.pt). Make sure to download `large-v2.pt`，put it into `whisper_pretrain/`.\n\n5. Download [hubert_soft model](https://github.com/bshall/hubert/releases/tag/v0.1)，put `hubert-soft-0d54a1f4.pt` into `hubert_pretrain/`.\n\n6. Download pitch extractor [crepe full](https://github.com/maxrmorrison/torchcrepe/tree/master/torchcrepe/assets)，put `full.pth` into `crepe/assets`.\n\n   **Note: crepe full.pth is 84.9 MB, not 6kb**\n   \n7. Download pretrain model [sovits5.0.pretrain.pth](https://github.com/PlayVoice/so-vits-svc-5.0/releases/tag/5.0/), and put it into `vits_pretrain/`.\n    ```shell\n    python svc_inference.py --config configs/base.yaml --model ./vits_pretrain/sovits5.0.pretrain.pth --spk ./configs/singers/singer0001.npy --wave test.wav\n    ```\n\n## Dataset preparation\n\nNecessary pre-processing:\n1. Separate voice and accompaniment with [UVR](https://github.com/Anjok07/ultimatevocalremovergui) (skip if no accompaniment)\n2. Cut audio input to shorter length with [slicer](https://github.com/flutydeer/audio-slicer), whisper takes input less than 30 seconds.\n3. Manually check generated audio input, remove inputs shorter than 2 seconds or with obivous noise.\n4. Adjust loudness if necessary, recommend Adobe Audiiton.\n5. Put the dataset into the `dataset_raw` directory following the structure below.\n```\ndataset_raw\n├───speaker0\n│   ├───000001.wav\n│   ├───...\n│   └───000xxx.wav\n└───speaker1\n    ├───000001.wav\n    ├───...\n    └───000xxx.wav\n```\n\n## Data preprocessing\n```shell\npython svc_preprocessing.py -t 2\n```\n`-t`: threading, max number should not exceed CPU core count, usually 2 is enough.\nAfter preprocessing you will get an output with following structure.\n```\ndata_svc/\n└── waves-16k\n│    └── speaker0\n│    │      ├── 000001.wav\n│    │      └── 000xxx.wav\n│    └── speaker1\n│           ├── 000001.wav\n│           └── 000xxx.wav\n└── waves-32k\n│    └── speaker0\n│    │      ├── 000001.wav\n│    │      └── 000xxx.wav\n│    └── speaker1\n│           ├── 000001.wav\n│           └── 000xxx.wav\n└── pitch\n│    └── speaker0\n│    │      ├── 000001.pit.npy\n│    │      └── 000xxx.pit.npy\n│    └── speaker1\n│           ├── 000001.pit.npy\n│           └── 000xxx.pit.npy\n└── hubert\n│    └── speaker0\n│    │      ├── 000001.vec.npy\n│    │      └── 000xxx.vec.npy\n│    └── speaker1\n│           ├── 000001.vec.npy\n│           └── 000xxx.vec.npy\n└── whisper\n│    └── speaker0\n│    │      ├── 000001.ppg.npy\n│    │      └── 000xxx.ppg.npy\n│    └── speaker1\n│           ├── 000001.ppg.npy\n│           └── 000xxx.ppg.npy\n└── speaker\n│    └── speaker0\n│    │      ├── 000001.spk.npy\n│    │      └── 000xxx.spk.npy\n│    └── speaker1\n│           ├── 000001.spk.npy\n│           └── 000xxx.spk.npy\n└── singer\n│   ├── speaker0.spk.npy\n│   └── speaker1.spk.npy\n|\n└── indexes\n    ├── speaker0\n    │   ├── some_prefix_hubert.index\n    │   └── some_prefix_whisper.index\n    └── speaker1\n        ├── hubert.index\n        └── whisper.index\n```\n\n1.  Re-sampling\n    - Generate audio with a sampling rate of 16000Hz in `./data_svc/waves-16k` \n    ```\n    python prepare/preprocess_a.py -w ./dataset_raw -o ./data_svc/waves-16k -s 16000\n    ```\n    \n    - Generate audio with a sampling rate of 32000Hz in `./data_svc/waves-32k`\n    ```\n    python prepare/preprocess_a.py -w ./dataset_raw -o ./data_svc/waves-32k -s 32000\n    ```\n2. Use 16K audio to extract pitch\n    ```\n    python prepare/preprocess_crepe.py -w data_svc/waves-16k/ -p data_svc/pitch\n    ```\n3. Use 16K audio to extract ppg\n    ```\n    python prepare/preprocess_ppg.py -w data_svc/waves-16k/ -p data_svc/whisper\n    ```\n4. Use 16K audio to extract hubert\n    ```\n    python prepare/preprocess_hubert.py -w data_svc/waves-16k/ -v data_svc/hubert\n    ```\n5. Use 16k audio to extract timbre code\n    ```\n    python prepare/preprocess_speaker.py data_svc/waves-16k/ data_svc/speaker\n    ```\n6. Extract the average value of the timbre code for inference; it can also replace a single audio timbre in generating the training index, and use it as the unified timbre of the speaker for training \n    ```\n    python prepare/preprocess_speaker_ave.py data_svc/speaker/ data_svc/singer\n    ``` \n7. Use 32k audio to extract the linear spectrum\n    ```\n    python prepare/preprocess_spec.py -w data_svc/waves-32k/ -s data_svc/specs\n    ``` \n8. Use 32k audio to generate training index\n    ```\n    python prepare/preprocess_train.py\n    ```\n11. Training file debugging\n    ```\n    python prepare/preprocess_zzz.py\n    ```\n\n## Train\n1. If fine-tuning is based on the pre-trained model, you need to download the pre-trained model: [sovits5.0.pretrain.pth](https://github.com/PlayVoice/so-vits-svc-5.0/releases/tag/5.0). Put pretrained model under project root, change this line\n    ```\n    pretrain: \"./vits_pretrain/sovits5.0.pretrain.pth\"\n    ```\n    in `configs/base.yaml`，and adjust the learning rate appropriately, eg 5e-5.\n   \n   `batch_size`: for GPU with 6G VRAM, 6 is the recommended value, 8 will work but step speed will be much slower.\n2. Start training\n   ```\n   python svc_trainer.py -c configs/base.yaml -n sovits5.0\n   ``` \n3. Resume training\n   ```\n   python svc_trainer.py -c configs/base.yaml -n sovits5.0 -p chkpt/sovits5.0/sovits5.0_***.pt\n   ```\n4. Log visualization\n   ```\n   tensorboard --logdir logs/\n   ```\n\n![sovits5 0_base](https://github.com/PlayVoice/so-vits-svc-5.0/assets/16432329/1628e775-5888-4eac-b173-a28dca978faa)\n\n![sovits_spec](https://github.com/PlayVoice/so-vits-svc-5.0/assets/16432329/c4223cf3-b4a0-4325-bec0-6d46d195a1fc)\n\n## Inference\n\n1. Export inference model: text encoder, Flow network, Decoder network\n   ```\n   python svc_export.py --config configs/base.yaml --checkpoint_path chkpt/sovits5.0/***.pt\n   ```\n2. Inference\n   - if there is no need to adjust `f0`, just run the following command.\n   ```\n   python svc_inference.py --config configs/base.yaml --model sovits5.0.pth --spk ./data_svc/singer/your_singer.spk.npy --wave test.wav --shift 0\n   ```\n   - if `f0` will be adjusted manually, follow the steps:\n     1. use whisper to extract content encoding, generate `test.vec.npy`.\n       ```\n       python whisper/inference.py -w test.wav -p test.ppg.npy\n       ```\n     2. use hubert to extract content vector, without using one-click reasoning, in order to reduce GPU memory usage\n       ```\n       python hubert/inference.py -w test.wav -v test.vec.npy\n       ```\n     3. extract the F0 parameter to the csv text format, open the csv file in Excel, and manually modify the wrong F0 according to Audition or SonicVisualiser\n       ```\n       python pitch/inference.py -w test.wav -p test.csv\n       ```\n     4. final inference\n       ```\n       python svc_inference.py --config configs/base.yaml --model sovits5.0.pth --spk ./data_svc/singer/your_singer.spk.npy --wave test.wav --ppg test.ppg.npy --vec test.vec.npy --pit test.csv --shift 0\n       ```\n3. Notes\n\n    - when `--ppg` is specified, when the same audio is reasoned multiple times, it can avoid repeated extraction of audio content codes; if it is not specified, it will be automatically extracted;\n\n    - when `--vec` is specified, when the same audio is reasoned multiple times, it can avoid repeated extraction of audio content codes; if it is not specified, it will be automatically extracted;\n\n    - when `--pit` is specified, the manually tuned F0 parameter can be loaded; if not specified, it will be automatically extracted;\n\n    - generate files in the current directory:svc_out.wav\n\n4. Arguments ref\n\n    | args |--config | --model | --spk | --wave | --ppg | --vec | --pit | --shift |\n    | :---:  | :---: | :---: | :---: | :---: | :---: | :---: | :---: | :---: |\n    | name | config path | model path | speaker | wave input | wave ppg | wave hubert | wave pitch | pitch shift |\n\n5. post by vad\n```\npython svc_inference_post.py --ref test.wav --svc svc_out.wav --out svc_out_post.wav\n```\n\n## Train Feature Retrieval Index (Optional)\n\nTo increase the stability of the generated timbre, you can use the method described in the \n[Retrieval-based-Voice-Conversion](https://github.com/RVC-Project/Retrieval-based-Voice-Conversion-WebUI/blob/main/docs/en/README.en.md) \nrepository. This method consists of 2 steps: \n\n1. Training the retrieval index on hubert and whisper features\n    Run training with default settings:\n    ```\n    python svc_train_retrieval.py\n    ```\n   \n    If the number of vectors is more than 200_000 they will be compressed to 10_000 using the MiniBatchKMeans algorithm.\n    You can change these settings using command line options:\n    ```\n    usage: crate faiss indexes for feature retrieval [-h] [--debug] [--prefix PREFIX] [--speakers SPEAKERS [SPEAKERS ...]] [--compress-features-after COMPRESS_FEATURES_AFTER]\n                                                     [--n-clusters N_CLUSTERS] [--n-parallel N_PARALLEL]\n\n    options:\n      -h, --help            show this help message and exit\n      --debug\n      --prefix PREFIX       add prefix to index filename\n      --speakers SPEAKERS [SPEAKERS ...]\n                            speaker names to create an index. By default all speakers are from data_svc\n      --compress-features-after COMPRESS_FEATURES_AFTER\n                            If the number of features is greater than the value compress feature vectors using MiniBatchKMeans.\n      --n-clusters N_CLUSTERS\n                            Number of centroids to which features will be compressed\n      --n-parallel N_PARALLEL\n                            Nuber of parallel job of MinibatchKmeans. Default is cpus-1\n    ``` \n    Compression of training vectors can speed up index inference, but reduces the quality of the retrieve.\n    Use vector count compression if you really have a lot of them.\n \n    The resulting indexes will be stored in the \"indexes\" folder as:\n    ``` \n    data_svc\n    ...\n    └── indexes\n        ├── speaker0\n        │   ├── some_prefix_hubert.index\n        │   └── some_prefix_whisper.index\n        └── speaker1\n            ├── hubert.index\n            └── whisper.index\n    ```\n2. At the inference stage adding the n closest features in a certain proportion of the vits model\n    Enable Feature Retrieval with settings:\n    ```\n    python svc_inference.py --config configs/base.yaml --model sovits5.0.pth --spk ./data_svc/singer/your_singer.spk.npy --wave test.wav --shift 0 \\\n    --enable-retrieval \\\n    --retrieval-ratio 0.5 \\\n    --n-retrieval-vectors 3\n    ``` \n    For a better retrieval effect, you can try to cycle through different parameters: `--retrieval-ratio` and `--n-retrieval-vectors`\n \n    If you have multiple sets of indexes, you can specify a specific set via the parameter: `--retrieval-index-prefix`\n \n    You can explicitly specify the paths to the hubert and whisper indexes using the parameters: `--hubert-index-path` and `--whisper-index-path`\n    \n\n## Create singer\nnamed by pure coincidence：average -\u003e ave -\u003e eva，eve(eva) represents conception and reproduction\n\n```\npython svc_eva.py\n```\n\n```python\neva_conf = {\n    './configs/singers/singer0022.npy': 0,\n    './configs/singers/singer0030.npy': 0,\n    './configs/singers/singer0047.npy': 0.5,\n    './configs/singers/singer0051.npy': 0.5,\n}\n```\n\nthe generated singer file will be `eva.spk.npy`.\n\n## Data set\n\n| Name | URL |\n| :--- | :--- |\n|KiSing         |http://shijt.site/index.php/2021/05/16/kising-the-first-open-source-mandarin-singing-voice-synthesis-corpus/|\n|PopCS          |https://github.com/MoonInTheRiver/DiffSinger/blob/master/resources/apply_form.md|\n|opencpop       |https://wenet.org.cn/opencpop/download/|\n|Multi-Singer   |https://github.com/Multi-Singer/Multi-Singer.github.io|\n|M4Singer       |https://github.com/M4Singer/M4Singer/blob/master/apply_form.md|\n|CSD            |https://zenodo.org/record/4785016#.YxqrTbaOMU4|\n|KSS            |https://www.kaggle.com/datasets/bryanpark/korean-single-speaker-speech-dataset|\n|JVS MuSic      |https://sites.google.com/site/shinnosuketakamichi/research-topics/jvs_music|\n|PJS            |https://sites.google.com/site/shinnosuketakamichi/research-topics/pjs_corpus|\n|JUST Song      |https://sites.google.com/site/shinnosuketakamichi/publication/jsut-song|\n|MUSDB18        |https://sigsep.github.io/datasets/musdb.html#musdb18-compressed-stems|\n|DSD100         |https://sigsep.github.io/datasets/dsd100.html|\n|Aishell-3      |http://www.aishelltech.com/aishell_3|\n|VCTK           |https://datashare.ed.ac.uk/handle/10283/2651|\n|Korean Songs   |http://urisori.co.kr/urisori-en/doku.php/|\n\n## Code sources and references\n\nhttps://github.com/facebookresearch/speech-resynthesis [paper](https://arxiv.org/abs/2104.00355)\n\nhttps://github.com/jaywalnut310/vits [paper](https://arxiv.org/abs/2106.06103)\n\nhttps://github.com/openai/whisper/ [paper](https://arxiv.org/abs/2212.04356)\n\nhttps://github.com/NVIDIA/BigVGAN [paper](https://arxiv.org/abs/2206.04658)\n\nhttps://github.com/mindslab-ai/univnet [paper](https://arxiv.org/abs/2106.07889)\n\nhttps://github.com/nii-yamagishilab/project-NN-Pytorch-scripts/tree/master/project/01-nsf\n\nhttps://github.com/huawei-noah/Speech-Backbones/tree/main/Grad-TTS\n\nhttps://github.com/brentspell/hifi-gan-bwe\n\nhttps://github.com/mozilla/TTS\n\nhttps://github.com/bshall/soft-vc\n\nhttps://github.com/maxrmorrison/torchcrepe\n\nhttps://github.com/MoonInTheRiver/DiffSinger\n\nhttps://github.com/OlaWod/FreeVC [paper](https://arxiv.org/abs/2210.15418)\n\nhttps://github.com/yl4579/HiFTNet [paper](https://arxiv.org/abs/2309.09493)\n\n[Autoregressive neural f0 model for statistical parametric speech synthesis](https://web.archive.org/web/20210718024752id_/https://ieeexplore.ieee.org/ielx7/6570655/8356719/08341752.pdf)\n\n[One-shot Voice Conversion by Separating Speaker and Content Representations with Instance Normalization](https://arxiv.org/abs/1904.05742)\n\n[SNAC : Speaker-normalized Affine Coupling Layer in Flow-based Architecture for Zero-Shot Multi-Speaker Text-to-Speech](https://github.com/hcy71o/SNAC)\n\n[Adapter-Based Extension of Multi-Speaker Text-to-Speech Model for New Speakers](https://arxiv.org/abs/2211.00585)\n\n[AdaSpeech: Adaptive Text to Speech for Custom Voice](https://arxiv.org/pdf/2103.00993.pdf)\n\n[AdaVITS: Tiny VITS for Low Computing Resource Speaker Adaptation](https://arxiv.org/pdf/2206.00208.pdf)\n\n[Cross-Speaker Prosody Transfer on Any Text for Expressive Speech Synthesis](https://github.com/ubisoft/ubisoft-laforge-daft-exprt)\n\n[Learn to Sing by Listening: Building Controllable Virtual Singer by Unsupervised Learning from Voice Recordings](https://arxiv.org/abs/2305.05401)\n\n[Adversarial Speaker Disentanglement Using Unannotated External Data for Self-supervised Representation Based Voice Conversion](https://arxiv.org/pdf/2305.09167.pdf)\n\n[Multilingual Speech Synthesis and Cross-Language Voice Cloning: GRL](https://arxiv.org/abs/1907.04448)\n\n[RoFormer: Enhanced Transformer with rotary position embedding](https://arxiv.org/abs/2104.09864)\n\n## Method of Preventing Timbre Leakage Based on Data Perturbation\n\nhttps://github.com/auspicious3000/contentvec/blob/main/contentvec/data/audio/audio_utils_1.py\n\nhttps://github.com/revsic/torch-nansy/blob/main/utils/augment/praat.py\n\nhttps://github.com/revsic/torch-nansy/blob/main/utils/augment/peq.py\n\nhttps://github.com/biggytruck/SpeechSplit2/blob/main/utils.py\n\nhttps://github.com/OlaWod/FreeVC/blob/main/preprocess_sr.py\n\n## Contributors\n\n\u003ca href=\"https://github.com/PlayVoice/so-vits-svc/graphs/contributors\"\u003e\n  \u003cimg src=\"https://contrib.rocks/image?repo=PlayVoice/so-vits-svc\" /\u003e\n\u003c/a\u003e\n\n## Thanks to\n\nhttps://github.com/Francis-Komizu/Sovits\n\n## Relevant Projects\n- [LoRA-SVC](https://github.com/PlayVoice/lora-svc): decoder only svc\n- [Grad-SVC](https://github.com/PlayVoice/Grad-SVC): diffusion based svc\n\n## Original evidence\n2022.04.12 https://mp.weixin.qq.com/s/autNBYCsG4_SvWt2-Ll_zA\n\n2022.04.22 https://github.com/PlayVoice/VI-SVS\n\n2022.07.26 https://mp.weixin.qq.com/s/qC4TJy-4EVdbpvK2cQb1TA\n\n2022.09.08 https://github.com/PlayVoice/VI-SVC\n\n## Be copied by svc-develop-team/so-vits-svc\n![coarse_f0_1](https://github.com/PlayVoice/so-vits-svc-5.0/assets/16432329/e2f5e5d3-d169-42c1-953f-4e1648b6da37)\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FPlayVoice%2Fwhisper-vits-svc","html_url":"https://awesome.ecosyste.ms/projects/github.com%2FPlayVoice%2Fwhisper-vits-svc","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FPlayVoice%2Fwhisper-vits-svc/lists"}