{"id":28042418,"url":"https://github.com/egorsmkv/tts_uk","last_synced_at":"2025-05-11T14:47:01.762Z","repository":{"id":280305177,"uuid":"941568450","full_name":"egorsmkv/tts_uk","owner":"egorsmkv","description":"High-fidelity speech synthesis for Ukrainian using modern neural networks.","archived":false,"fork":false,"pushed_at":"2025-04-28T09:10:37.000Z","size":1849,"stargazers_count":8,"open_issues_count":5,"forks_count":0,"subscribers_count":1,"default_branch":"main","last_synced_at":"2025-04-28T10:29:22.791Z","etag":null,"topics":["audio","sound","speech-uk","synthesis","text-to-speech","tts","ukrainian","vocos","wav","wave"],"latest_commit_sha":null,"homepage":"https://huggingface.co/spaces/Yehor/radtts-uk-vocos-demo","language":"Jupyter Notebook","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/egorsmkv.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":".github/FUNDING.yml","license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":"CITATION.cff","codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null},"funding":{"custom":["https://send.monobank.ua/jar/3Saxixsdua"]}},"created_at":"2025-03-02T15:51:24.000Z","updated_at":"2025-04-28T09:10:44.000Z","dependencies_parsed_at":"2025-03-02T16:34:35.166Z","dependency_job_id":"851d30be-efac-47fd-b890-46d1f2dd3906","html_url":"https://github.com/egorsmkv/tts_uk","commit_stats":null,"previous_names":["egorsmkv/tts_uk"],"tags_count":7,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/egorsmkv%2Ftts_uk","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/egorsmkv%2Ftts_uk/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/egorsmkv%2Ftts_uk/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/egorsmkv%2Ftts_uk/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/egorsmkv","download_url":"https://codeload.github.com/egorsmkv/tts_uk/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":253584235,"owners_count":21931540,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["audio","sound","speech-uk","synthesis","text-to-speech","tts","ukrainian","vocos","wav","wave"],"created_at":"2025-05-11T14:47:01.069Z","updated_at":"2025-05-11T14:47:01.754Z","avatar_url":"https://github.com/egorsmkv.png","language":"Jupyter Notebook","funding_links":["https://send.monobank.ua/jar/3Saxixsdua"],"categories":[],"sub_categories":[],"readme":"\u003cp align=\"center\"\u003e\n  \u003ca href=\"https://github.com/egorsmkv/tts_uk\"\u003e\n    \u003cimg loading=\"lazy\" alt=\"tts_uk\" src=\"https://github.com/egorsmkv/tts_uk/raw/main/assets/logo_github.jpg\" width=\"100%\"/\u003e\n  \u003c/a\u003e\n\u003c/p\u003e\n\n# Text-to-Speech for Ukrainian\n\n[![PyPI Version](https://img.shields.io/pypi/v/tts_uk)](https://pypi.org/project/tts_uk/)\n[![License MIT](https://img.shields.io/github/license/egorsmkv/tts_uk)](https://opensource.org/licenses/MIT)\n[![PyPI Downloads](https://static.pepy.tech/badge/tts_uk/month)](https://pepy.tech/projects/tts_uk)\n[![DOI](https://zenodo.org/badge/DOI/10.5281/zenodo.14966501.svg)](https://doi.org/10.5281/zenodo.14966501)\n[![FOSSA Status](https://app.fossa.com/api/projects/git%2Bgithub.com%2Fegorsmkv%2Ftts_uk.svg?type=small)](https://app.fossa.com/projects/git%2Bgithub.com%2Fegorsmkv%2Ftts_uk?ref=badge_small)\n\nHigh-fidelity speech synthesis for Ukrainian using modern neural networks.\n\n## Statuses\n\n[![CI Pipeline](https://github.com/egorsmkv/tts_uk/actions/workflows/ci.yml/badge.svg)](https://github.com/egorsmkv/tts_uk/actions/workflows/ci.yml)\n[![Dependabot Updates](https://github.com/egorsmkv/tts_uk/actions/workflows/dependabot/dependabot-updates/badge.svg)](https://github.com/egorsmkv/tts_uk/actions/workflows/dependabot/dependabot-updates)\n[![Snyk Security](https://github.com/egorsmkv/tts_uk/actions/workflows/snyk-python.yml/badge.svg)](https://github.com/egorsmkv/tts_uk/actions/workflows/snyk-python.yml)\n\n## Demo\n\n[![HF Space](https://img.shields.io/badge/HF-%F0%9F%A4%97%20Space-yellow)](https://huggingface.co/spaces/Yehor/radtts-uk-vocos-demo)\n[![Google Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/drive/1sdCPnZJRNAf12PhPut4gu6T_o6lYaUdo?usp=sharing)\n\nCheck out our demo on [Hugging Face space](https://huggingface.co/spaces/Yehor/radtts-uk-vocos-demo) or just [listen to samples here](https://huggingface.co/spaces/speech-uk/listen-tts-voices).\n\n## Features\n\n- Multi-speaker model: 2 **female** (Tetiana, Lada) + 1 **male** (Mykyta) voices;\n- Fine-grained control over speech parameters, including duration, fundamental frequency (F0), and energy;\n- High-fidelity speech generation using the [RAD-TTS++](https://github.com/egorsmkv/radtts-uk) acoustic model;\n- Fast vocoding using [Vocos](https://github.com/gemelo-ai/vocos);\n- Synthesizes long sentences effectively;\n- Supports a sampling rate of 44.1 kHz;\n- Tested on Linux environments and **Windows**/**WSL**;\n- Python API (requires Python 3.9 or later);\n- CUDA-enabled for GPU acceleration.\n\n## Installation\n\n```shell\n# Install from PyPI\npip install tts-uk\n\n# OR, for the latest development version:\npip install git+https://github.com/egorsmkv/tts_uk\n\n# OR, use git and local setup\ngit clone https://github.com/egorsmkv/tts_uk\ncd tts_uk\nuv sync # uv will handle the virtual environment\n```\n\nRead [uv's installation](https://github.com/astral-sh/uv?tab=readme-ov-file#installation) section.\n\nAlso, you can [download the repository](https://github.com/egorsmkv/tts_uk/archive/refs/heads/main.zip) as a ZIP archive.\n\n## Getting started\n\nCode example:\n\n```python\nimport torchaudio\n\nfrom tts_uk.inference import synthesis\n\nsampling_rate = 44_100\n\n# Perform the synthesis, `synthesis` function returns:\n# - mels: Mel spectrograms of the generated audio.\n# - wave: The synthesized waveform by a Vocoder as a PyTorch tensor.\n# - stats: A dictionary containing synthesis statistics (processing time, duration, speech rate, etc).\nmels, wave, stats = synthesis(\n    text=\"Ви можете протестувати синтез мовлення українською мовою. Просто введіть текст, який ви хочете прослухати.\",\n    voice=\"tetiana\",  # tetiana, mykyta, lada\n    n_takes=1,\n    use_latest_take=False,\n    token_dur_scaling=1,\n    f0_mean=0,\n    f0_std=0,\n    energy_mean=0,\n    energy_std=0,\n    sigma_decoder=0.8,\n    sigma_token_duration=0.666,\n    sigma_f0=1,\n    sigma_energy=1,\n)\n\nprint(stats)\n\n# Save the generated audio to a WAV file.\ntorchaudio.save(\"audio.wav\", wave.cpu(), sampling_rate, encoding=\"PCM_S\")\n```\n\nUse these Google colabs:\n\n- [CPU inference](https://colab.research.google.com/drive/1dsQiVhTaNw5lRfUiCZeECMuEbtEEYqbZ?usp=sharing)\n- [GPU inference](https://colab.research.google.com/drive/1sdCPnZJRNAf12PhPut4gu6T_o6lYaUdo?usp=sharing) on T4 card (long document to synthesize)\n\nOr run synthesis in a terminal:\n\n```shell\nuv run example.py\n```\n\nIf you need to synthesize articles we recommend consider [wtpsplit](https://github.com/segment-any-text/wtpsplit).\n\n## Get help and support\n\nPlease feel free to connect with us using [the Issues section](https://github.com/egorsmkv/tts_uk/issues).\n\n## License\n\nCode has the MIT license.\n\n## Model authors\n\n### Acoustic\n\n- [Yehor Smoliakov](https://github.com/egorsmkv), [HF profile](https://huggingface.co/Yehor) \n\n### Vocoder\n\n- [Serhiy Stetskovych](https://github.com/patriotyk), [HF profile](https://huggingface.co/patriotyk) \n\n## Community\n\n[![Discord](https://img.shields.io/discord/1199769227192713226?label=\u0026logo=discord\u0026logoColor=white\u0026color=7289DA\u0026style=flat-square)](https://bit.ly/discord-uds)\n\n- Discord: https://bit.ly/discord-uds\n- Speech Recognition: https://t.me/speech_recognition_uk\n- Speech Synthesis: https://t.me/speech_synthesis_uk\n\nAlso, follow [our Speech-UK initiative](https://huggingface.co/speech-uk) on Hugging Face!\n\n## Acknowledgements\n\n- [RAD-TTS by NVIDIA](https://github.com/NVIDIA/radtts)\n- [Vocos fork by langtech-bsc](https://github.com/langtech-bsc/vocos/tree/matcha)\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fegorsmkv%2Ftts_uk","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fegorsmkv%2Ftts_uk","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fegorsmkv%2Ftts_uk/lists"}