Projects in Awesome Lists tagged with audio-generation
A curated list of projects in awesome lists tagged with audio-generation .
https://github.com/mudler/localai
:robot: The free, Open Source alternative to OpenAI, Claude and others. Self-hosted and local-first. Drop-in replacement for OpenAI, running on consumer-grade hardware. No GPU required. Runs gguf, transformers, diffusers and many more models architectures. Features: Generate Text, Audio, Video, Images, Voice Cloning, Distributed, P2P inference
ai api audio-generation distributed gemma gpt4all image-generation kubernetes libp2p llama llama3 llm mamba mistral musicgen rerank rwkv stable-diffusion text-generation tts
Last synced: 14 May 2026
https://github.com/go-skynet/LocalAI
:robot: The free, Open Source alternative to OpenAI, Claude and others. Self-hosted and local-first. Drop-in replacement for OpenAI, running on consumer-grade hardware. No GPU required. Runs gguf, transformers, diffusers and many more models architectures. Features: Generate Text, Audio, Video, Images, Voice Cloning, Distributed, P2P inference
ai api audio-generation distributed gemma gpt4all image-generation kubernetes libp2p llama llama3 llm mamba mistral musicgen rerank rwkv stable-diffusion text-generation tts
Last synced: 03 May 2025
https://github.com/mudler/LocalAI
:robot: The free, Open Source alternative to OpenAI, Claude and others. Self-hosted and local-first. Drop-in replacement for OpenAI, running on consumer-grade hardware. No GPU required. Runs gguf, transformers, diffusers and many more models architectures. Features: Generate Text, Audio, Video, Images, Voice Cloning, Distributed inference
ai api audio-generation distributed gemma gpt4all image-generation kubernetes llama llama3 llm mamba mistral musicgen p2p rerank rwkv stable-diffusion text-generation tts
Last synced: 14 Mar 2025
https://github.com/QwenAudio/CosyVoice
Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.
audio-generation cantonese chatbot chatgpt chinese cosyvoice cross-lingual english fine-grained fine-tuning gpt-4o japanese korean multi-lingual natural-language-generation python text-to-speech tts voice-cloning
Last synced: 29 Jul 2026
https://github.com/funaudiollm/cosyvoice
Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.
audio-generation cantonese chatbot chatgpt chinese cosyvoice cross-lingual english fine-grained fine-tuning gpt-4o japanese korean multi-lingual natural-language-generation python text-to-speech tts voice-cloning
Last synced: 20 Oct 2025
https://github.com/open-mmlab/amphion
Amphion (/æmˈfaɪən/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support reproducible research and help junior researchers and engineers get started in the field of audio, music, and speech generation research and development.
audio-generation audio-synthesis audioldm audit emilia fastspeech2 maskgct music-generation naturalspeech2 singing-voice-conversion speech-synthesis text-to-audio text-to-speech vall-e vits vocoder voice-conversion
Last synced: 12 May 2025
https://github.com/open-mmlab/Amphion
Amphion (/æmˈfaɪən/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support reproducible research and help junior researchers and engineers get started in the field of audio, music, and speech generation research and development.
audio-generation audio-synthesis audioldm audit emilia fastspeech2 maskgct music-generation naturalspeech2 singing-voice-conversion speech-synthesis text-to-audio text-to-speech vall-e vits vocoder voice-conversion
Last synced: 28 Mar 2025
https://github.com/FunAudioLLM/CosyVoice
Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.
audio-generation cantonese chatbot chatgpt chinese cosyvoice cross-lingual english fine-grained fine-tuning gpt-4o japanese korean multi-lingual natural-language-generation python text-to-speech tts voice-cloning
Last synced: 24 Mar 2025
https://github.com/multimodal-art-projection/YuE
YuE: Open Full-song Music Generation Foundation Model, something similar to Suno.ai but open
ai audio-generation deep-learning foundation-models gpt huggingface llama llms music-generation style-transfers voice-cloning
Last synced: 16 Oct 2025
https://github.com/multimodal-art-projection/yue
YuE: Open Full-song Music Generation Foundation Model, something similar to Suno.ai but open
ai audio-generation deep-learning foundation-models gpt huggingface llama llms music-generation style-transfers voice-cloning
Last synced: 13 May 2025
https://github.com/rsxdalv/tts-webui
A single Gradio + React WebUI with extensions for ACE-Step, Kimi Audio, Piper TTS, GPT-SoVITS, CosyVoice, XTTSv2, DIA, Kokoro, OpenVoice, ParlerTTS, Stable Audio, MMS, StyleTTS2, MAGNet, AudioGen, MusicGen, Tortoise, RVC, Vocos, Demucs, SeamlessM4T, and Bark!
ace-step ai audio-generation cosyvoice generative-ai generator gradio music musicgen openai-api openvoice rvc styletts2 text-to-speech tortoise-tts tts vocos
Last synced: 05 Apr 2026
https://github.com/haoheliu/audioldm
AudioLDM: Generate speech, sound effects, music and beyond, with text.
Last synced: 13 May 2025
https://github.com/haoheliu/AudioLDM
AudioLDM: Generate speech, sound effects, music and beyond, with text.
Last synced: 27 Mar 2025
https://github.com/rsxdalv/TTS-WebUI
A single Gradio + React WebUI with extensions for ACE-Step, Kimi Audio, Piper TTS, GPT-SoVITS, CosyVoice, XTTSv2, DIA, Kokoro, OpenVoice, ParlerTTS, Stable Audio, MMS, StyleTTS2, MAGNet, AudioGen, MusicGen, Tortoise, RVC, Vocos, Demucs, SeamlessM4T, and Bark!
ai audio-generation generative-ai generator gradio magnet music musicgen openai-api rvc styletts2 text-to-speech tortoise-tts tts vocos
Last synced: 10 Jun 2025
https://github.com/rsxdalv/tts-generation-webui
TTS Generation Web UI (Bark, MusicGen + AudioGen, Tortoise, RVC, Vocos, Demucs, SeamlessM4T, MAGNet, StyleTTS2, MMS, Stable Audio, Mars5, F5-TTS, ParlerTTS)
ai audio-generation audiogen bark deep-learning generator gradio machine-learning magnet music musicgen rvc seamlessm4t styletts2 text-to-speech torch tortoise-tts tts vocos web
Last synced: 04 Apr 2025
https://github.com/archinetai/audio-diffusion-pytorch
Audio generation using diffusion models, in PyTorch.
artificial-intelligence audio-generation deep-learning denoising-diffusion
Last synced: 14 May 2025
https://github.com/archinetai/audio-ai-timeline
A timeline of the latest AI models for audio generation, starting in 2023!
artificial-intelligence audio-generation machine-learning
Last synced: 05 Feb 2026
https://github.com/lucidrains/soundstorm-pytorch
Implementation of SoundStorm, Efficient Parallel Audio Generation from Google Deepmind, in Pytorch
artificial-intelligence attention-mechanism audio-generation deep-learning non-autoregressive transformers
Last synced: 14 May 2025
https://github.com/declare-lab/tango
A family of diffusion models for text-to-audio generation.
audio-generation diffusion diffusion-models language-models large-language-models text-to-audio
Last synced: 16 May 2025
https://github.com/funaudiollm/inspiremusic
InspireMusic: A Unified Framework for Music, Song, Audio Generation.
audio-generation audio-processing music-generation pytorch
Last synced: 15 May 2025
https://github.com/nvidia/bigvgan
Official PyTorch implementation of BigVGAN (ICLR 2023)
audio-generation audio-synthesis music-synthesis neural-vocoder singing-voice-synthesis speech-synthesis
Last synced: 19 Oct 2025
https://binwang28.github.io/audio-ai-hub/
The hub for audio AI research: papers, open models, benchmarks & datasets across audio LLMs, speech recognition, TTS, music & audio generation.
audio-ai audio-generation audio-llm music-generation paper-list speech-recognition tts
Last synced: 22 Jun 2026
https://github.com/Yuan-ManX/ai-audio-datasets
AI Audio Datasets (AI-ADS) 🎵, including Speech, Music, and Sound Effects, which can provide training data for Generative AI, AIGC, AI model training, intelligent audio tool development, and audio applications.
aigc artificial-intelligence audio audio-effect audio-generation datasets deep-learning machine-learning music-generation
Last synced: 17 Mar 2025
https://github.com/modelscope/funcodec
FunCodec is a research-oriented toolkit for audio quantization and downstream applications, such as text-to-speech synthesis, music generation et.al.
audio-generation audio-quantization codec encodec speech-synthesis speech-to-text tts voicecloning
Last synced: 05 Apr 2025
https://github.com/modelscope/FunCodec
FunCodec is a research-oriented toolkit for audio quantization and downstream applications, such as text-to-speech synthesis, music generation et.al.
audio-generation audio-quantization codec encodec speech-synthesis speech-to-text tts voicecloning
Last synced: 26 Oct 2025
https://github.com/v-iashin/SpecVQGAN
Source code for "Taming Visually Guided Sound Generation" (Oral at the BMVC 2021)
audio audio-generation bmvc evaluation-metrics gan melgan multi-modal pytorch transformer vas vggsound video video-features video-understanding vqvae
Last synced: 09 Apr 2025
https://github.com/Yuan-ManX/audio-development-tools
This is a list of sound, audio and music development tools which contains machine learning, audio generation, audio signal processing, sound synthesis, spatial audio, music information retrieval, music generation, speech recognition, speech synthesis, singing voice synthesis and more.
artificial-intelligence audio audio-generation audio-processing deep-learning dsp machine-learning music music-generation signal-processing speech speech-processing speech-synthesis
Last synced: 17 Mar 2025
https://github.com/cabralpinto/modular-diffusion
Python library for designing and training your own Diffusion Models with PyTorch
audio-generation deep-learning diffusion-models image-generation machine-learning modular-design python pytorch text-generation transformer u-net
Last synced: 23 Aug 2025
https://github.com/sony/bigvsan
Pytorch implementation of BigVSAN
audio-generation audio-synthesis gan neural-vocoder pytorch speech-synthesis
Last synced: 03 Apr 2025
https://github.com/happylittlecat2333/Auffusion
Official codes and models of the paper "Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation"
audio-generation diffusion diffusion-models large-language-models text-to-audio
Last synced: 07 May 2025
https://github.com/archinetai/audio-data-pytorch
A collection of useful audio datasets and transforms for PyTorch.
artifical-intelligense audio-generation datasets deep-learning pytorch
Last synced: 30 Oct 2025
https://github.com/devnen/dia-tts-server
Self-host the powerful Dia TTS model. This server offers a user-friendly Web UI, flexible API endpoints (incl. OpenAI compatible), support for SafeTensors/BF16, voice cloning, dialogue generation, and GPU/CPU execution.
ai api-server audio-generation cuda dia dia-tts dialogue-tts fastapi huggingface openai-api python pytorch speech-synthesis speech-synthesis-api text-to-speech tts tts-api voice-cloning web-ui
Last synced: 06 May 2025
https://github.com/archinetai/audio-diffusion-pytorch-trainer
Trainer for audio-diffusion-pytorch
artificial-intelligence audio-generation deep-learning denoising-diffusion
Last synced: 05 Oct 2025
https://github.com/ilaria-manco/word2wave
Word2Wave: a framework for generating short audio samples from a text prompt using WaveGAN and COALA.
ai-music audio-generation music-generation text-to-audio
Last synced: 14 Jul 2025
https://github.com/sony/soundctm
Pytorch implementation of SoundCTM
audio-generation diffusion-models pytorch text-to-audio
Last synced: 07 Apr 2025
https://github.com/SamurAIGPT/n8n-nodes-muapi
n8n community nodes for MuAPI — generate images, videos & audio with 60+ AI models (FLUX, Midjourney V7, Veo 3, Suno, Kling, Runway) in your n8n workflows
ai-nodes ai-workflow audio-generation automation flux generative-ai image-generation image-to-video midjourney muapi n8n n8n-community-node n8n-nodes stable-diffusion suno text-to-image text-to-video typescript video-generation workflow-automation
Last synced: 04 Apr 2026
https://github.com/openmoss/omnivae
OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation
audio audio-generation audio-video-generation dit multimodal-generation vaes video video-generation
Last synced: 02 Aug 2026
https://github.com/rsxdalv/musicgen-prompts
Site for sharing MusicGen + AudioGen Prompts and Creations
ai audio-generation audiogen generator machine-learning musicgen
Last synced: 06 Feb 2026
https://github.com/olaviinha/neuraltexttoaudio
Text prompt steered synthetic audio generators
audio audio-generation audio-processing audio-synthesis audioldm colab colab-notebook mubert mubertai music-generation text2audio text2music voice-cloning voice-synthesis
Last synced: 22 Sep 2025
https://github.com/Bai-YT/ConsistencyTTA
ConsistencyTTA: Accelerating Diffusion-Based Text-to-Audio Generation with Consistency Distillation
audio-generation audio-processing consistency-models diffusion-models ldm
Last synced: 29 Aug 2025
https://github.com/josefalbers/aggressor
Ultra-minimal autoregressive diffusion model for image generation
artificial-intelligence audio-generation autoregressive autoregressive-diffusion autoregressive-models deep-learning diffusion diffusion-models image-generation mlx tts
Last synced: 11 Mar 2026
https://github.com/warma10032/easytts
打造最简单的TTS前端集合,最简单的有声小说制作工作流。基于正则规则对小说进行分句,基于RoBERTa对小说中的对话进行说话人识别,从而实现一键式生成多人有声小说。多说话人的语音合成,高质量的有声小说制作。
ai audio-generation nlp pyqt speaker-identification tts
Last synced: 13 Jun 2025
https://github.com/bean980310/stable-diffusion-docker-project
Stable Diffusion WebUI and KohyaSS, ComfyUI, InvokeAI, Fooocus, and more Generative AI on Docker
ai audio-generation chatbotai container-image deep-learning docker generative-ai gradio image-generation img2img machine-learning python pytorch stable-diffusion text-generation torch txt2img video-generation
Last synced: 22 Apr 2025
https://github.com/gianpaj/sexyvoice
Voice cloning and Text to Speech platform. Perfect for content creators, developers, and storytellers.
ai audio-generation generate-audio generative-ai text-to-speech
Last synced: 03 Feb 2026
https://github.com/fsecada01/midi-drums
🥁 Comprehensive MIDI drum generation system with 4 genres (Metal, Rock, Jazz, Funk), 28 styles, and 7 authentic drummer personalities. Plugin-based architecture for professional EZDrummer-compatible output.
audio-generation drummer-styles drums ezdrummer midi multi-genre music-generation music-production plugin-architecture python
Last synced: 17 Feb 2026
https://github.com/voltsygm/openvoice
🔊 Clone voices accurately while controlling style and tone in multiple languages seamlessly with OpenVoice.
audio-generation conversational-ai converter docker generative-ai generator onnxruntime openai-api openvoiceos rvc-voices styletts2 tensorflow tortoise-tts tts voice-cloning voice-conversion voice-conversion-gan zero-shot-tts
Last synced: 12 May 2026
https://github.com/hddevteam/speechify
🎧 Text-to-speech VS Code extension with 200+ Azure voices, TypeScript architecture, multilingual support (EN/CN), and advanced voice customization. Professional solution for accessibility, content creation, and language learning powered by Azure Speech Services.
accessibility audio-generation azure-speech-services content-creation language-learning multilingual speech-synthesis text-to-speech typescript vscode-extension
Last synced: 14 Feb 2026
https://github.com/merekat/children-stories
OhanashiGPT is an application that generates personalized children's stories based on parameters like age and preferences. It narrates these stories using an AI-generated voice that mimics a parent, trained on their audio samples. The app also creates illustrations to accompany each story, providing a unique and engaging experience for children.
ai audio-generation data-science image-generation large-language-models llama lora lux neural-networks stable-diffusion story text-generation tts xtts
Last synced: 04 Jul 2025
https://github.com/justmalhar/tts-studio
Text to Speech Studio to convert text into natural-sounding speech using advanced AI models from leading providers like Replicate, OpenAI, and ElevenLabs.
ai audio-generation elevenlabs kokoro nextjs openai reactjs replicate speech-synthesis tailwindcss text-to-speech tts tts-api vercel vercel-deploy
Last synced: 23 Apr 2025
https://github.com/socaity/socaity
SDK for generative AI.
3d-generation api artificial-intelligence audio-generation bark clip deepseek-r1 flux generative-ai hosting image-captioning image-generation inference-api llama3 llm runpod speech-synthesis stable-diffusion text-to-speech video-generation
Last synced: 15 Apr 2025
https://github.com/0x7o/deepmozart
Audio generation using diffusion models
audio-generation diffusion-models neural-network
Last synced: 06 Mar 2026
https://github.com/radoslawregula/voxg
Singing voice synthesizer using GANs
audio audio-generation audio-processing deep-learning gan generative-adversarial-network machine-learning music-programming music-technology python singing-voice-synthesis tensorflow voice-synthesis
Last synced: 18 May 2026
https://github.com/simonbernarding/ohanashigpt-children-story-generation
OhanashiGPT is an application that generates personalized children's stories based on parameters like age and preferences. It narrates these stories using an AI-generated voice that mimics a parent, trained on their audio samples. The app also creates illustrations to accompany each story, providing a unique and engaging experience for children.
ai audio-generation data-science image-generation large-language-model text-generation
Last synced: 08 Apr 2025
https://github.com/runapi-ai/suno-sdk
RunAPI Suno SDK for text-to-music, lyric generation and blending, cover audio, music extension, stem separation, voice validation phrase, custom voice, and related audio workflows in JavaScript, Python, Ruby, Go, Java, and PHP
api audio-generation golang gradle java maven music-api music-generation python ruby runapi runapi-ai sdk suno suno-ai-api typescript
Last synced: 28 Jul 2026
https://github.com/lucadellalib/bigvgan
A single-file implementation of BigVGAN generator
audio-generation audio-synthesis bigvgan music-synthesis neural-vocoder pytorch singing-voice-synthesis speech-synthesis
Last synced: 05 May 2026
https://github.com/bocaletto-luca/cw-generator
The CW (Morse) Generator is a versatile and powerful application for Morse communication enthusiasts and anyone interested in learning or practicing this classic language of communication. This software allows you to convert text to Morse code and vice versa, providing a complete suite of tools to create, interpret, and reproduce ...
audio-generation cw-trasmission desktop-application education gui ham-radio morse-code morse-code-translator morse-decoding open-source python signal-processing text-to-morse
Last synced: 18 Jun 2025
https://github.com/farhann-saleem/curssed-resonance
Universal Audio & SFX Engine (共鳴り). Dynamically swaps ControlFoley and ACE-Step models in VRAM on a single cheap GPU to generate music and sound effects on-the-fly. Designed for the Kamui pipeline.
audio-generation audio-generation-ai automation cloudflare-r2 music-generation runpod serverless sound-effects
Last synced: 15 Aug 2026
https://github.com/koppalexander/ohanashi-childgpt
OhanashiGPT is an application that generates personalized children's stories based on parameters like age and preferences. It narrates these stories using an AI-generated voice that mimics a parent, trained on their audio samples. The app also creates illustrations to accompany each story, providing a unique and engaging experience for children.
ai audio-generation data-science flux generative-ai image-generation large-language-models llama lora neural-networks stable-diffusion story text-generation tts xtts
Last synced: 09 Mar 2026
https://github.com/iris2c/inspiremusic
InspireMusic: A Unified Framework for Music, Song and Audio Generation
audio-generation music-generation pytorch text-to-music
Last synced: 29 Jul 2025
https://github.com/jxoesneon/gemini-audio-mcp
A high-performance Model Context Protocol (MCP) server in Rust that generates infinite, context-aware environmental soundscapes and professional audio using Gemini 2.0 Multimodal Live API.
ai-agents ai-audio audio audio-generation ffmpeg gemini gemini-api google-ai lyria mcp mcp-server model-context-protocol modelcontextprotocol music-generation npx rust soundscape
Last synced: 03 Apr 2026
https://github.com/criadacasa/podcastfy-saas
SaaS platform for generating AI podcasts from multimodal content - Built with Hono and Cloudflare Pages
ai audio-generation cloudflare-pages cloudflare-workers hono podcast podcastfy saas text-to-speech typescript
Last synced: 15 Apr 2026
https://github.com/polaroteam/moltdj-skill
SoundCloud for AI agents — skill files for the first music and podcast platform built for autonomous bots
agent-skill ai-agent ai-agents ai-api ai-music ai-music-generation ai-music-platform ai-podcast audio-generation claude-code-skill clawdbot mcp moltbot music-api music-generation music-platform openclaw podcast-generation soundcloud text-to-music
Last synced: 07 Mar 2026
https://github.com/mahshid1378/tts-generation-webui
TTS Generation Web UI (Bark, MusicGen + AudioGen, Tortoise, RVC, Vocos, Demucs, SeamlessM4T, MAGNet, StyleTTS2, MMS, Stable Audio, Mars5, F5-TTS, ParlerTTS)
ai audio-generation audiogen bark deep-learning generator gradio machine-learning magnet music musicgen rvc seamlessm4t styletts2 text-to-speech torch tortoise-tts vocos web
Last synced: 14 Jul 2025
https://github.com/work-nobu/ohanashigpt
OhanashiGPT is an application that generates personalized children's stories based on parameters like age and preferences. It narrates these stories using an AI-generated voice that mimics a parent, trained on their audio samples. The app also creates illustrations to accompany each story, providing a unique and engaging experience for children.
ai audio-generation data-science image-generation large-language-models llama3 llamacpp lora low-rank-adaptation stable-diffusion text-generation xtts
Last synced: 09 Feb 2026
https://github.com/jet-logic/vocal_vse
Generates voiceover audio from text strips in the Blender Video Sequence Editor
audio-generation blender blender-addon blender-python text-to-speech tts video-editing vse
Last synced: 14 May 2026
https://github.com/deepgram-starters/rust-text-to-speech
Get started using Deepgram's Text-to-Speech with this Rust demo app
audio-generation axum deepgram demo quickstart rust text-to-speech tts
Last synced: 05 Apr 2026
https://github.com/runapi-ai/elevenlabs-sdk
RunAPI ElevenLabs SDK for text-to-speech, dialogue generation, sound effects, speech transcription, and audio isolation workflows in JavaScript, Python, Ruby, Go, Java, and PHP
api audio-generation elevenlabs elevenlabs-api golang gradle java maven python ruby runapi runapi-ai sdk speech-to-text text-to-speech typescript
Last synced: 28 Jul 2026
https://github.com/anas436/image-to-audio-app
Image Captioning and Text-to-Speech
audio-generation image-processing text-to-speech
Last synced: 27 Mar 2025
https://github.com/aidayang/inspiremusic-oneclick
InspireMusic文本转音乐软件免安装一键启动整合包
audio-generation audio-processing inspiremusic music-generation python pytorch
Last synced: 15 May 2026
https://github.com/charles-forsyth/generate-tts
A professional CLI for Google Gemini's Native 2.5 TTS. Generate multi-speaker podcasts ('Deep Dive'), audio summaries, and expressive speech from text/files.
ai-tools audio-generation cli gemini-api google-cloud multi-speaker podcast-generator python text-to-speech tts
Last synced: 13 Jan 2026