Projects in Awesome Lists tagged with multi-speaker
A curated list of projects in awesome lists tagged with multi-speaker .
https://github.com/netease-youdao/emotivoice
EmotiVoice 😊: a Multi-Voice and Prompt-Controlled TTS Engine
ai deep-learning emotion emotivoice multi-speaker prompt python pytorch speech speech-synthesis style text-to-speech tts
Last synced: 13 May 2025
https://github.com/mikebrady/shairport-sync
AirPlay and AirPlay 2 audio player
airplay airplay-2 audio audio-player audio-streaming embedded-systems multi-room-audio multi-speaker synchronized-audio
Last synced: 07 Oct 2025
https://github.com/netease-youdao/EmotiVoice
EmotiVoice 😊: a Multi-Voice and Prompt-Controlled TTS Engine
ai deep-learning emotion emotivoice multi-speaker prompt python pytorch speech speech-synthesis style text-to-speech tts
Last synced: 24 Mar 2025
https://github.com/r9y9/deepvoice3_pytorch
PyTorch implementation of convolutional neural networks-based text-to-speech synthesis models
end-to-end machine-learning multi-speaker python pytorch speech-processing speech-synthesis tts
Last synced: 14 May 2025
https://github.com/aishoot/LSTM_PIT_Speech_Separation
Two-talker Speech Separation with LSTM/BLSTM by Permutation Invariant Training method.
audio-separation multi-speaker permutation-invariant-training robust-speech-recognition speech-enhancement speech-separation
Last synced: 01 Apr 2025
https://github.com/keonlee9420/comprehensive-e2e-tts
A Non-Autoregressive End-to-End Text-to-Speech (text-to-wav), supporting a family of SOTA unsupervised duration modelings. This project grows with the research community, aiming to achieve the ultimate E2E-TTS
deep-learning end-to-end fastspeech2 hifi-gan jets multi-speaker neural-tts non-ar non-autoregressive pytorch single-speaker sota speech-synthesis text-to-speech text-to-wav tts ultimate-tts unsupervised
Last synced: 01 Jun 2026
https://github.com/jayspiffy/draft-to-take
Draft to Take beta: local-first AI audio production studio powered by IndexTTS2, Docker, Qwen, OmniVoice, SFX, ambience, and music sidecars.
ai-audio docker draft-to-take fastapi gpu index-tts indextts indextts2 local-ai multi-speaker self-hosted speaker-prep speech-synthesis text-to-speech timeline-editor tts voice-cloning
Last synced: 14 Jun 2026
https://github.com/theseraphim/scribe-forge-ai
🎵 Complete offline audio transcription system with speaker diarization using OpenAI Whisper and PyAnnote. Features automatic audio cleaning, precise timestamps, multiple output formats (JSON/TXT/Markdown), and support for 20+ audio formats. No external APIs required - works entirely offline.
audio-analysis audio-cleaning audio-processing audio-transcription diarization ffmpeg huggingface machine-learning multi-speaker nlp offline-transcription openai-whisper pyannote python speaker-diarization speech-recognition speech-to-text timestamps transcription-tool whisper
Last synced: 05 May 2026
https://github.com/charles-forsyth/generate-tts
A professional CLI for Google Gemini's Native 2.5 TTS. Generate multi-speaker podcasts ('Deep Dive'), audio summaries, and expressive speech from text/files.
ai-tools audio-generation cli gemini-api google-cloud multi-speaker podcast-generator python text-to-speech tts
Last synced: 13 Jan 2026
https://github.com/drajabr/audio-matrix-router
Grid style audio router
audio audio-routing multi-speaker speaker-array virtual-audio
Last synced: 26 Apr 2026