{"id":18838125,"url":"https://github.com/isabelleysseric/voice-cloning","last_synced_at":"2026-08-05T06:31:11.644Z","repository":{"id":214125011,"uuid":"733809992","full_name":"isabelleysseric/voice-cloning","owner":"isabelleysseric","description":"Speech synthesis with conditioning on very small dataset. Using Nvidia's Tacotron2 and WaveGlow models with Pytorch.","archived":false,"fork":false,"pushed_at":"2024-09-10T21:29:29.000Z","size":29372,"stargazers_count":0,"open_issues_count":0,"forks_count":0,"subscribers_count":1,"default_branch":"main","last_synced_at":"2025-05-29T19:53:53.836Z","etag":null,"topics":["nvidia","signal-processing","speech-recognition","speech-synthesis","speechbrain","tacotron2","text-to-speech","tts","waveglow"],"latest_commit_sha":null,"homepage":"","language":"Jupyter Notebook","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/isabelleysseric.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2023-12-20T07:14:42.000Z","updated_at":"2024-09-10T21:29:32.000Z","dependencies_parsed_at":"2024-09-10T23:07:43.157Z","dependency_job_id":"a989eb38-2272-473f-aef9-fe0b043782b9","html_url":"https://github.com/isabelleysseric/voice-cloning","commit_stats":null,"previous_names":["isabelleysseric/voice-cloning"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/isabelleysseric/voice-cloning","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/isabelleysseric%2Fvoice-cloning","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/isabelleysseric%2Fvoice-cloning/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/isabelleysseric%2Fvoice-cloning/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/isabelleysseric%2Fvoice-cloning/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/isabelleysseric","download_url":"https://codeload.github.com/isabelleysseric/voice-cloning/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/isabelleysseric%2Fvoice-cloning/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":278903014,"owners_count":26065785,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","status":"online","status_checked_at":"2025-10-08T02:00:06.501Z","response_time":56,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["nvidia","signal-processing","speech-recognition","speech-synthesis","speechbrain","tacotron2","text-to-speech","tts","waveglow"],"created_at":"2024-11-08T02:38:01.294Z","updated_at":"2025-10-08T06:38:02.227Z","avatar_url":"https://github.com/isabelleysseric.png","language":"Jupyter Notebook","funding_links":[],"categories":[],"sub_categories":[],"readme":"\u003ch1 align=\"center\"\u003e Voice Cloning\u003c/h1\u003e\n\u003cbr/\u003e\n\u003cbr/\u003e\n\n\u003cp align=\"center\"\u003e\n  \u003cimg src=\"https://github.com/isabelleysseric/voice-cloning/blob/main/output/test_model_image.png\" /\u003e\n\u003c/p\u003e  \n\n\u003ch2 align=\"center\"\u003e    \n\n  \u003c!-- GitHub --\u003e\n  \u003ca href=\"https://github.com/isabelleysseric/\"\u003e\n    \u003cimg src=\"https://img.shields.io/badge/GitHub-100000?style=for-the-badge\u0026logo=github\u0026logoColor=white\" \u003e\n  \u003c/a\u003e  \n\n  \u003c!-- Project Repo --\u003e\n  \u003ca href=\"https://github.com/isabelleysseric/Voice-Cloning/\"\u003e\n    \u003cimg src=\"https://img.shields.io/badge/Repo-Voice_Cloning-green?style=for-the-badge\u0026logo={Voice-Cloning}\u0026logoColor=white\" \u003e\n  \u003c/a\u003e\n\n  \u003c!-- Wiki Project --\u003e\n  \u003ca href=\"https://github.com/isabelleysseric/Voice-Cloning/wiki/\"\u003e\n    \u003cimg src=\"https://img.shields.io/badge/Wiki-Voice_Cloning-green?style=for-the-badge\u0026logo={Voice-Cloning}\u0026logoColor=white\" \u003e\n  \u003c/a\u003e\u003cbr\u003e \n  \n\u003c/h2\u003e\n\u003cbr/\u003e\n\n**Note**: The `Voice_cloning_Training_with_Tacotron2_and_WaveGlow.ipynb` notebook is to be run in *Google Colab*. Once in Colab, you need to import the *data_cleaned.zip* dataset into the current folder `/content/`.\nReplace the files in the folder `/content/TTS-TT2/filelists/` with my files that have the same name after installing Tacotron2. LThe rest of the code will take care of unzipping it and putting it in the new folder `/content/TTS-TT2/wavs/`.\nThe program will then ask you to load your transcription file. You will give it the `list.txt` file.\n\nThe files in the `input` folder are needed to give input to the speech synthesis model. They are also found at the root of the project. The wav files correspond to the zip file: `data_cleaned.zip` and the `list.txt`, `ljs_audio_text_val_filelists.txt`, `ljs_audio_text_val_filelists.txt` and `ljs_audio_text_val_filelists.txt`files are also found at the root of the project.\nThe files in the `output` folder are the results of the model, during and after training.\n\n**TREE**:\n\n[input](https://github.com/isabelleysseric/voice-cloning/tree/main/input)\n\n  - [filelists](https://github.com/isabelleysseric/voice-cloning/tree/main/input/wavs)\n      - `list.txt`\n      - `ljs_audio_text_test_filelists.txt`\n      - `ljs_audio_text_train_filelists.txt`\n      - `ljs_audio_text_val_filelists.txt`\n        \n  - [wavs](https://github.com/isabelleysseric/voice-cloning/tree/main/input/audio)\n      - `1.npy`\n      - `1.wav`\n      - ...\n      - `60.npy`\n      - `60.wav`\n   \n[output](https://github.com/isabelleysseric/voice-cloning/tree/main/output)\n\n  - [audio](https://github.com/isabelleysseric/voice-cloning/tree/main/output/audio)\n      - `model_BS_6_0.00003_350epoch_0_original_audio.wav`\n      - `model_BS_6_0.00003_350epoch_0_predicted_audio.wav`\n      - ...\n      - `model_BS_6_0.00003_350epoch_20_original_audio.wav`\n      - `model_BS_6_0.00003_350epoch_20_predicted_audio.wav`\n      \n      - `model_BS_6_0.00003_350signals_epoch_0.png`\n      - ...\n      - `model_BS_6_0.00003_350signals_epoch_20.png`\n      \n  - [images](https://github.com/isabelleysseric/voice-cloning/tree/main/output/images)\n      - `model_BS_6_0.00003_350_Alignment_Epoch_0_Iteration_9_Validation_Loss_1.7767614126205444.png`\n      - ...\n      - `model_BS_6_0.00003_350_Alignment_Epoch_20_Iteration_189_Validation_Loss_1.0240533351898193.png`\n\n  - [logs](https://github.com/isabelleysseric/voice-cloning/tree/main/output/logs)\n      - `events.out.tfevents.1703405636.c8a2ca7defbc.1806.11`\n        \n  - [loss](https://github.com/isabelleysseric/voice-cloning/tree/main/output/loss)\n      - `model_BS_6_0.00003_350loss_curve_epoch_0.png`\n      - ...\n      - `model_BS_6_0.00003_350loss_curve_epoch_22.png`\n        \n  - [spectrogram](https://github.com/isabelleysseric/voice-cloning/tree/main/output/spectrogram)\n      - `model_BS_6_0.00003_350spectrograms_epoch_0.png`\n      - ...\n      - `model_BS_6_0.00003_350spectrograms_epoch_20.png`\n\n\n[Voice_cloning_Training_with_Tacotron2_and_WaveGlow.ipynb](https://github.com/isabelleysseric/voice-cloning/blob/main/Voice_cloning_Training_with_Tacotron2_and_WaveGlow.ipynb)  \n[MLSP Presentation_Clonage_de_la_voix.pdf](https://github.com/isabelleysseric/voice-cloning/blob/main/MLSP%20Presentation_Clonage_de_la_voix.pdf)  \n[MLSP_Rapport_Clonage_de_la_voix.pdf](https://github.com/isabelleysseric/voice-cloning/blob/main/MLSP_Rapport_Clonage_de_la_voix.pdf)  \n[README.md](https://github.com/isabelleysseric/voice-cloning/blob/main/README.md)  \n[data_cleaned.zip](https://github.com/isabelleysseric/voice-cloning/blob/main/data_cleaned.zip)  \n[list.txt](https://github.com/isabelleysseric/voice-cloning/blob/main/list.txt)  \n[ljs_audio_text_test_filelist.txt](https://github.com/isabelleysseric/voice-cloning/blob/main/ljs_audio_text_test_filelist.txt)  \n[ljs_audio_text_train_filelist.txt](https://github.com/isabelleysseric/voice-cloning/blob/main/ljs_audio_text_train_filelist.txt)  \n[ljs_audio_text_val_filelist.txt](https://github.com/isabelleysseric/voice-cloning/blob/main/ljs_audio_text_val_filelist.txt)  \n\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fisabelleysseric%2Fvoice-cloning","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fisabelleysseric%2Fvoice-cloning","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fisabelleysseric%2Fvoice-cloning/lists"}