{"id":24251185,"url":"https://github.com/nazdridoy/kokoro-tts","last_synced_at":"2025-05-16T08:03:20.564Z","repository":{"id":272490426,"uuid":"916766950","full_name":"nazdridoy/kokoro-tts","owner":"nazdridoy","description":"A CLI text-to-speech tool using the Kokoro model, supporting multiple languages, voices (with blending), and various input formats including EPUB books and PDF documents.","archived":false,"fork":false,"pushed_at":"2025-05-03T09:00:59.000Z","size":1373,"stargazers_count":388,"open_issues_count":6,"forks_count":60,"subscribers_count":4,"default_branch":"main","last_synced_at":"2025-05-03T09:40:22.312Z","etag":null,"topics":["audiobook","epub","kokoro","kokoro-tts","pdf","podcast","python","tts"],"latest_commit_sha":null,"homepage":"https://huggingface.co/hexgrad/Kokoro-82M","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/nazdridoy.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":"CONTRIBUTING.md","funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null}},"created_at":"2025-01-14T18:12:50.000Z","updated_at":"2025-05-03T09:01:02.000Z","dependencies_parsed_at":"2025-04-12T03:52:41.840Z","dependency_job_id":"39295aab-2e07-485b-b35a-356d8a4d2f6f","html_url":"https://github.com/nazdridoy/kokoro-tts","commit_stats":null,"previous_names":["nazdridoy/kokoro-tts"],"tags_count":4,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/nazdridoy%2Fkokoro-tts","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/nazdridoy%2Fkokoro-tts/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/nazdridoy%2Fkokoro-tts/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/nazdridoy%2Fkokoro-tts/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/nazdridoy","download_url":"https://codeload.github.com/nazdridoy/kokoro-tts/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":254493381,"owners_count":22080126,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["audiobook","epub","kokoro","kokoro-tts","pdf","podcast","python","tts"],"created_at":"2025-01-15T02:01:36.206Z","updated_at":"2025-05-16T08:03:20.557Z","avatar_url":"https://github.com/nazdridoy.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"# Kokoro TTS\n\nA CLI text-to-speech tool using the Kokoro model, supporting multiple languages, voices (with blending), and various input formats including EPUB books and PDF documents.\n\n![ngpt-s-c](https://raw.githubusercontent.com/nazdridoy/kokoro-tts/main/previews/kokoro-tts-h.png)\n\n## Features\n\n- Multiple language and voice support\n- Voice blending with customizable weights\n- EPUB, PDF and TXT file input support\n- Standard input (stdin) and `|` piping from other programs\n- Streaming audio playback\n- Split output into chapters\n- Adjustable speech speed\n- WAV and MP3 output formats\n- Chapter merging capability\n- Detailed debug output option\n- GPU Support\n\n## Demo\n\nKokoro TTS is an open-source CLI tool that delivers high-quality text-to-speech right from your terminal. Think of it as your personal voice studio, capable of transforming any text into natural-sounding speech with minimal effort.\n\nhttps://github.com/user-attachments/assets/8413e640-59e9-490e-861d-49187e967526\n\n[Demo Audio (MP3)](https://github.com/nazdridoy/kokoro-tts/raw/main/previews/demo.mp3) | [Demo Audio (WAV)](https://github.com/nazdridoy/kokoro-tts/raw/main/previews/demo.wav)\n\n## TODO\n\n- [x] Add GPU support\n- [x] Add PDF support\n- [ ] Add GUI\n\n## Prerequisites\n\n- Python 3.12\n\n## Installation\n\n1. Clone the repository:\n```bash\ngit clone https://github.com/nazdridoy/kokoro-tts.git\ncd kokoro-tts\n```\n\n2. Install required packages:\n```bash\npip install -r requirements.txt\n```\nor\n```bash\nuv sync\n```\nNote: You can also use `uv` as a faster alternative to pip for package installation. (This is a uv project)\nNote: Python\u003e=3.13 is not currently supported.\n\n3. Download the required model files:\n```bash\n# Download either voices.json or voices.bin (bin is preferred)\nwget https://github.com/nazdridoy/kokoro-tts/releases/download/v1.0.0/voices-v1.0.bin\n\n# Download the model\nwget https://github.com/nazdridoy/kokoro-tts/releases/download/v1.0.0/kokoro-v1.0.onnx\n```\nNote: The script will automatically use voices.bin if present, falling back to voices.json if bin is not available.\n\n\n## Supported voices:\n\n| **Category** | **Voices** | **Language Code** |\n| --- | --- | --- |\n| 🇺🇸 👩 | af\\_alloy, af\\_aoede, af\\_bella, af\\_heart, af\\_jessica, af\\_kore, af\\_nicole, af\\_nova, af\\_river, af\\_sarah, af\\_sky | **en-us** |\n| 🇺🇸 👨 | am\\_adam, am\\_echo, am\\_eric, am\\_fenrir, am\\_liam, am\\_michael, am\\_onyx, am\\_puck | **en-us** |\n| 🇬🇧 | bf\\_alice, bf\\_emma, bf\\_isabella, bf\\_lily, bm\\_daniel, bm\\_fable, bm\\_george, bm\\_lewis | **en-gb** |\n| 🇫🇷 | ff\\_siwis | **fr-fr** |\n| 🇮🇹 | if\\_sara, im\\_nicola | **it** |\n| 🇯🇵 | jf\\_alpha, jf\\_gongitsune, jf\\_nezumi, jf\\_tebukuro, jm\\_kumo | **ja** |\n| 🇨🇳 | zf\\_xiaobei, zf\\_xiaoni, zf\\_xiaoxiao, zf\\_xiaoyi, zm\\_yunjian, zm\\_yunxi, zm\\_yunxia, zm\\_yunyang | **cmn** |\n\n\n## Usage\n\nBasic usage:\n```bash\n./kokoro-tts \u003cinput_text_file\u003e [\u003coutput_audio_file\u003e] [options]\n```\n\n### Commands\n\n- `-h, --help`: Show help message\n- `--help-languages`: List supported languages\n- `--help-voices`: List available voices\n- `--merge-chunks`: Merge existing chunks into chapter files\n\n### Options\n\n- `--stream`: Stream audio instead of saving to file\n- `--speed \u003cfloat\u003e`: Set speech speed (default: 1.0)\n- `--lang \u003cstr\u003e`: Set language (default: en-us)\n- `--voice \u003cstr\u003e`: Set voice or blend voices (default: interactive selection)\n  - Single voice: Use voice name (e.g., \"af_sarah\")\n  - Blended voices: Use \"voice1:weight,voice2:weight\" format\n- `--split-output \u003cdir\u003e`: Save each chunk as separate file in directory\n- `--format \u003cstr\u003e`: Audio format: wav or mp3 (default: wav)\n- `--debug`: Show detailed debug information during processing\n\n### Input Formats\n\n- `.txt`: Text file input\n- `.epub`: EPUB book input (will process chapters)\n- `.pdf`: PDF document input (extracts chapters from TOC or content)\n\n### Examples\n\n```bash\n# Basic usage with output file\nkokoro-tts input.txt output.wav --speed 1.2 --lang en-us --voice af_sarah\n\n# Read from standard input (stdin)\necho \"Hello World\" | kokoro-tts /dev/stdin --stream\ncat input.txt | kokoro-tts /dev/stdin output.wav\n\n# Use voice blending (60-40 mix)\nkokoro-tts input.txt output.wav --voice \"af_sarah:60,am_adam:40\"\n\n# Use equal voice blend (50-50)\nkokoro-tts input.txt --stream --voice \"am_adam,af_sarah\"\n\n# Process EPUB and split into chunks\nkokoro-tts input.epub --split-output ./chunks/ --format mp3\n\n# Stream audio directly\nkokoro-tts input.txt --stream --speed 0.8\n\n# Merge existing chunks\nkokoro-tts --merge-chunks --split-output ./chunks/ --format wav\n\n# Process EPUB with detailed debug output\nkokoro-tts input.epub --split-output ./chunks/ --debug\n\n# Process PDF and split into chapters\nkokoro-tts input.pdf --split-output ./chunks/ --format mp3\n# List available voices\nkokoro-tts --help-voices\n\n# List supported languages\nkokoro-tts --help-languages\n```\n\n## Features in Detail\n\n### EPUB Processing\n- Automatically extracts chapters from EPUB files\n- Preserves chapter titles and structure\n- Creates organized output for each chapter\n- Detailed debug output available for troubleshooting\n\n### Audio Processing\n- Chunks long text into manageable segments\n- Supports streaming for immediate playback\n- Voice blending with customizable mix ratios\n- Progress indicators for long processes\n- Handles interruptions gracefully\n\n### Output Options\n- Single file output\n- Split output with chapter organization\n- Chunk merging capability\n- Multiple audio format support\n\n### Debug Mode\n- Shows detailed information about file processing\n- Displays NCX parsing details for EPUB files\n- Lists all found chapters and their metadata\n- Helps troubleshoot processing issues\n\n### Input Options\n- Text file input (.txt)\n- EPUB book input (.epub)\n- Standard input (stdin)\n- Supports piping from other programs\n\n## Contributing\n\nThis is a personal project. But if you want to contribute, please feel free to submit a Pull Request.\n\n## License\n\nThis project is licensed under the MIT License. See the [LICENSE](LICENSE) file for details.\n\n## Acknowledgments\n\n- [Kokoro-ONNX](https://github.com/thewh1teagle/kokoro-onnx)\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fnazdridoy%2Fkokoro-tts","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fnazdridoy%2Fkokoro-tts","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fnazdridoy%2Fkokoro-tts/lists"}