An open API service indexing awesome lists of open source software.

Projects in Awesome Lists tagged with audio-generation

A curated list of projects in awesome lists tagged with audio-generation .

https://github.com/mudler/localai

:robot: The free, Open Source alternative to OpenAI, Claude and others. Self-hosted and local-first. Drop-in replacement for OpenAI, running on consumer-grade hardware. No GPU required. Runs gguf, transformers, diffusers and many more models architectures. Features: Generate Text, Audio, Video, Images, Voice Cloning, Distributed, P2P inference

ai api audio-generation distributed gemma gpt4all image-generation kubernetes libp2p llama llama3 llm mamba mistral musicgen rerank rwkv stable-diffusion text-generation tts

Last synced: 14 May 2026

https://github.com/go-skynet/LocalAI

:robot: The free, Open Source alternative to OpenAI, Claude and others. Self-hosted and local-first. Drop-in replacement for OpenAI, running on consumer-grade hardware. No GPU required. Runs gguf, transformers, diffusers and many more models architectures. Features: Generate Text, Audio, Video, Images, Voice Cloning, Distributed, P2P inference

ai api audio-generation distributed gemma gpt4all image-generation kubernetes libp2p llama llama3 llm mamba mistral musicgen rerank rwkv stable-diffusion text-generation tts

Last synced: 03 May 2025

https://github.com/mudler/LocalAI

:robot: The free, Open Source alternative to OpenAI, Claude and others. Self-hosted and local-first. Drop-in replacement for OpenAI, running on consumer-grade hardware. No GPU required. Runs gguf, transformers, diffusers and many more models architectures. Features: Generate Text, Audio, Video, Images, Voice Cloning, Distributed inference

ai api audio-generation distributed gemma gpt4all image-generation kubernetes llama llama3 llm mamba mistral musicgen p2p rerank rwkv stable-diffusion text-generation tts

Last synced: 14 Mar 2025

https://github.com/open-mmlab/amphion

Amphion (/æmˈfaɪən/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support reproducible research and help junior researchers and engineers get started in the field of audio, music, and speech generation research and development.

audio-generation audio-synthesis audioldm audit emilia fastspeech2 maskgct music-generation naturalspeech2 singing-voice-conversion speech-synthesis text-to-audio text-to-speech vall-e vits vocoder voice-conversion

Last synced: 12 May 2025

https://github.com/open-mmlab/Amphion

Amphion (/æmˈfaɪən/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support reproducible research and help junior researchers and engineers get started in the field of audio, music, and speech generation research and development.

audio-generation audio-synthesis audioldm audit emilia fastspeech2 maskgct music-generation naturalspeech2 singing-voice-conversion speech-synthesis text-to-audio text-to-speech vall-e vits vocoder voice-conversion

Last synced: 28 Mar 2025

https://github.com/multimodal-art-projection/YuE

YuE: Open Full-song Music Generation Foundation Model, something similar to Suno.ai but open

ai audio-generation deep-learning foundation-models gpt huggingface llama llms music-generation style-transfers voice-cloning

Last synced: 16 Oct 2025

https://github.com/multimodal-art-projection/yue

YuE: Open Full-song Music Generation Foundation Model, something similar to Suno.ai but open

ai audio-generation deep-learning foundation-models gpt huggingface llama llms music-generation style-transfers voice-cloning

Last synced: 13 May 2025

https://github.com/rsxdalv/tts-webui

A single Gradio + React WebUI with extensions for ACE-Step, Kimi Audio, Piper TTS, GPT-SoVITS, CosyVoice, XTTSv2, DIA, Kokoro, OpenVoice, ParlerTTS, Stable Audio, MMS, StyleTTS2, MAGNet, AudioGen, MusicGen, Tortoise, RVC, Vocos, Demucs, SeamlessM4T, and Bark!

ace-step ai audio-generation cosyvoice generative-ai generator gradio music musicgen openai-api openvoice rvc styletts2 text-to-speech tortoise-tts tts vocos

Last synced: 05 Apr 2026

https://github.com/haoheliu/audioldm

AudioLDM: Generate speech, sound effects, music and beyond, with text.

audio-generation

Last synced: 13 May 2025

https://github.com/haoheliu/AudioLDM

AudioLDM: Generate speech, sound effects, music and beyond, with text.

audio-generation

Last synced: 27 Mar 2025

https://github.com/haoheliu/audioldm2

Text-to-Audio/Music Generation

audio-generation

Last synced: 14 May 2025

https://github.com/haoheliu/AudioLDM2

Text-to-Audio/Music Generation

audio-generation

Last synced: 24 Mar 2025

https://github.com/rsxdalv/TTS-WebUI

A single Gradio + React WebUI with extensions for ACE-Step, Kimi Audio, Piper TTS, GPT-SoVITS, CosyVoice, XTTSv2, DIA, Kokoro, OpenVoice, ParlerTTS, Stable Audio, MMS, StyleTTS2, MAGNet, AudioGen, MusicGen, Tortoise, RVC, Vocos, Demucs, SeamlessM4T, and Bark!

ai audio-generation generative-ai generator gradio magnet music musicgen openai-api rvc styletts2 text-to-speech tortoise-tts tts vocos

Last synced: 10 Jun 2025

https://github.com/rsxdalv/tts-generation-webui

TTS Generation Web UI (Bark, MusicGen + AudioGen, Tortoise, RVC, Vocos, Demucs, SeamlessM4T, MAGNet, StyleTTS2, MMS, Stable Audio, Mars5, F5-TTS, ParlerTTS)

ai audio-generation audiogen bark deep-learning generator gradio machine-learning magnet music musicgen rvc seamlessm4t styletts2 text-to-speech torch tortoise-tts tts vocos web

Last synced: 04 Apr 2025

https://github.com/archinetai/audio-ai-timeline

A timeline of the latest AI models for audio generation, starting in 2023!

artificial-intelligence audio-generation machine-learning

Last synced: 05 Feb 2026

https://github.com/lucidrains/soundstorm-pytorch

Implementation of SoundStorm, Efficient Parallel Audio Generation from Google Deepmind, in Pytorch

artificial-intelligence attention-mechanism audio-generation deep-learning non-autoregressive transformers

Last synced: 14 May 2025

https://github.com/declare-lab/tango

A family of diffusion models for text-to-audio generation.

audio-generation diffusion diffusion-models language-models large-language-models text-to-audio

Last synced: 16 May 2025

https://github.com/funaudiollm/inspiremusic

InspireMusic: A Unified Framework for Music, Song, Audio Generation.

audio-generation audio-processing music-generation pytorch

Last synced: 15 May 2025

https://binwang28.github.io/audio-ai-hub/

The hub for audio AI research: papers, open models, benchmarks & datasets across audio LLMs, speech recognition, TTS, music & audio generation.

audio-ai audio-generation audio-llm music-generation paper-list speech-recognition tts

Last synced: 22 Jun 2026

https://github.com/Yuan-ManX/ai-audio-datasets

AI Audio Datasets (AI-ADS) 🎵, including Speech, Music, and Sound Effects, which can provide training data for Generative AI, AIGC, AI model training, intelligent audio tool development, and audio applications.

aigc artificial-intelligence audio audio-effect audio-generation datasets deep-learning machine-learning music-generation

Last synced: 17 Mar 2025

https://github.com/modelscope/funcodec

FunCodec is a research-oriented toolkit for audio quantization and downstream applications, such as text-to-speech synthesis, music generation et.al.

audio-generation audio-quantization codec encodec speech-synthesis speech-to-text tts voicecloning

Last synced: 05 Apr 2025

https://github.com/modelscope/FunCodec

FunCodec is a research-oriented toolkit for audio quantization and downstream applications, such as text-to-speech synthesis, music generation et.al.

audio-generation audio-quantization codec encodec speech-synthesis speech-to-text tts voicecloning

Last synced: 26 Oct 2025

https://github.com/v-iashin/SpecVQGAN

Source code for "Taming Visually Guided Sound Generation" (Oral at the BMVC 2021)

audio audio-generation bmvc evaluation-metrics gan melgan multi-modal pytorch transformer vas vggsound video video-features video-understanding vqvae

Last synced: 09 Apr 2025

https://github.com/Yuan-ManX/audio-development-tools

This is a list of sound, audio and music development tools which contains machine learning, audio generation, audio signal processing, sound synthesis, spatial audio, music information retrieval, music generation, speech recognition, speech synthesis, singing voice synthesis and more.

artificial-intelligence audio audio-generation audio-processing deep-learning dsp machine-learning music music-generation signal-processing speech speech-processing speech-synthesis

Last synced: 17 Mar 2025

https://github.com/happylittlecat2333/Auffusion

Official codes and models of the paper "Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation"

audio-generation diffusion diffusion-models large-language-models text-to-audio

Last synced: 07 May 2025

https://github.com/archinetai/audio-data-pytorch

A collection of useful audio datasets and transforms for PyTorch.

artifical-intelligense audio-generation datasets deep-learning pytorch

Last synced: 30 Oct 2025

https://github.com/devnen/dia-tts-server

Self-host the powerful Dia TTS model. This server offers a user-friendly Web UI, flexible API endpoints (incl. OpenAI compatible), support for SafeTensors/BF16, voice cloning, dialogue generation, and GPU/CPU execution.

ai api-server audio-generation cuda dia dia-tts dialogue-tts fastapi huggingface openai-api python pytorch speech-synthesis speech-synthesis-api text-to-speech tts tts-api voice-cloning web-ui

Last synced: 06 May 2025

https://github.com/ilaria-manco/word2wave

Word2Wave: a framework for generating short audio samples from a text prompt using WaveGAN and COALA.

ai-music audio-generation music-generation text-to-audio

Last synced: 14 Jul 2025

https://github.com/sony/soundctm

Pytorch implementation of SoundCTM

audio-generation diffusion-models pytorch text-to-audio

Last synced: 07 Apr 2025

https://github.com/SamurAIGPT/n8n-nodes-muapi

n8n community nodes for MuAPI — generate images, videos & audio with 60+ AI models (FLUX, Midjourney V7, Veo 3, Suno, Kling, Runway) in your n8n workflows

ai-nodes ai-workflow audio-generation automation flux generative-ai image-generation image-to-video midjourney muapi n8n n8n-community-node n8n-nodes stable-diffusion suno text-to-image text-to-video typescript video-generation workflow-automation

Last synced: 04 Apr 2026

https://github.com/openmoss/omnivae

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation

audio audio-generation audio-video-generation dit multimodal-generation vaes video video-generation

Last synced: 02 Aug 2026

https://github.com/rsxdalv/musicgen-prompts

Site for sharing MusicGen + AudioGen Prompts and Creations

ai audio-generation audiogen generator machine-learning musicgen

Last synced: 06 Feb 2026

https://github.com/Bai-YT/ConsistencyTTA

ConsistencyTTA: Accelerating Diffusion-Based Text-to-Audio Generation with Consistency Distillation

audio-generation audio-processing consistency-models diffusion-models ldm

Last synced: 29 Aug 2025

https://github.com/warma10032/easytts

打造最简单的TTS前端集合,最简单的有声小说制作工作流。基于正则规则对小说进行分句,基于RoBERTa对小说中的对话进行说话人识别,从而实现一键式生成多人有声小说。多说话人的语音合成,高质量的有声小说制作。

ai audio-generation nlp pyqt speaker-identification tts

Last synced: 13 Jun 2025

https://github.com/gianpaj/sexyvoice

Voice cloning and Text to Speech platform. Perfect for content creators, developers, and storytellers.

ai audio-generation generate-audio generative-ai text-to-speech

Last synced: 03 Feb 2026

https://github.com/fsecada01/midi-drums

🥁 Comprehensive MIDI drum generation system with 4 genres (Metal, Rock, Jazz, Funk), 28 styles, and 7 authentic drummer personalities. Plugin-based architecture for professional EZDrummer-compatible output.

audio-generation drummer-styles drums ezdrummer midi multi-genre music-generation music-production plugin-architecture python

Last synced: 17 Feb 2026

https://github.com/hddevteam/speechify

🎧 Text-to-speech VS Code extension with 200+ Azure voices, TypeScript architecture, multilingual support (EN/CN), and advanced voice customization. Professional solution for accessibility, content creation, and language learning powered by Azure Speech Services.

accessibility audio-generation azure-speech-services content-creation language-learning multilingual speech-synthesis text-to-speech typescript vscode-extension

Last synced: 14 Feb 2026

https://github.com/merekat/children-stories

OhanashiGPT is an application that generates personalized children's stories based on parameters like age and preferences. It narrates these stories using an AI-generated voice that mimics a parent, trained on their audio samples. The app also creates illustrations to accompany each story, providing a unique and engaging experience for children.

ai audio-generation data-science image-generation large-language-models llama lora lux neural-networks stable-diffusion story text-generation tts xtts

Last synced: 04 Jul 2025

https://github.com/justmalhar/tts-studio

Text to Speech Studio to convert text into natural-sounding speech using advanced AI models from leading providers like Replicate, OpenAI, and ElevenLabs.

ai audio-generation elevenlabs kokoro nextjs openai reactjs replicate speech-synthesis tailwindcss text-to-speech tts tts-api vercel vercel-deploy

Last synced: 23 Apr 2025

https://github.com/0x7o/deepmozart

Audio generation using diffusion models

audio-generation diffusion-models neural-network

Last synced: 06 Mar 2026

https://github.com/simonbernarding/ohanashigpt-children-story-generation

OhanashiGPT is an application that generates personalized children's stories based on parameters like age and preferences. It narrates these stories using an AI-generated voice that mimics a parent, trained on their audio samples. The app also creates illustrations to accompany each story, providing a unique and engaging experience for children.

ai audio-generation data-science image-generation large-language-model text-generation

Last synced: 08 Apr 2025

https://github.com/runapi-ai/suno-sdk

RunAPI Suno SDK for text-to-music, lyric generation and blending, cover audio, music extension, stem separation, voice validation phrase, custom voice, and related audio workflows in JavaScript, Python, Ruby, Go, Java, and PHP

api audio-generation golang gradle java maven music-api music-generation python ruby runapi runapi-ai sdk suno suno-ai-api typescript

Last synced: 28 Jul 2026

https://github.com/bocaletto-luca/cw-generator

The CW (Morse) Generator is a versatile and powerful application for Morse communication enthusiasts and anyone interested in learning or practicing this classic language of communication. This software allows you to convert text to Morse code and vice versa, providing a complete suite of tools to create, interpret, and reproduce ...

audio-generation cw-trasmission desktop-application education gui ham-radio morse-code morse-code-translator morse-decoding open-source python signal-processing text-to-morse

Last synced: 18 Jun 2025

https://github.com/farhann-saleem/curssed-resonance

Universal Audio & SFX Engine (共鳴り). Dynamically swaps ControlFoley and ACE-Step models in VRAM on a single cheap GPU to generate music and sound effects on-the-fly. Designed for the Kamui pipeline.

audio-generation audio-generation-ai automation cloudflare-r2 music-generation runpod serverless sound-effects

Last synced: 15 Aug 2026

https://github.com/koppalexander/ohanashi-childgpt

OhanashiGPT is an application that generates personalized children's stories based on parameters like age and preferences. It narrates these stories using an AI-generated voice that mimics a parent, trained on their audio samples. The app also creates illustrations to accompany each story, providing a unique and engaging experience for children.

ai audio-generation data-science flux generative-ai image-generation large-language-models llama lora neural-networks stable-diffusion story text-generation tts xtts

Last synced: 09 Mar 2026

https://github.com/iris2c/inspiremusic

InspireMusic: A Unified Framework for Music, Song and Audio Generation

audio-generation music-generation pytorch text-to-music

Last synced: 29 Jul 2025

https://github.com/jxoesneon/gemini-audio-mcp

A high-performance Model Context Protocol (MCP) server in Rust that generates infinite, context-aware environmental soundscapes and professional audio using Gemini 2.0 Multimodal Live API.

ai-agents ai-audio audio audio-generation ffmpeg gemini gemini-api google-ai lyria mcp mcp-server model-context-protocol modelcontextprotocol music-generation npx rust soundscape

Last synced: 03 Apr 2026

https://github.com/criadacasa/podcastfy-saas

SaaS platform for generating AI podcasts from multimodal content - Built with Hono and Cloudflare Pages

ai audio-generation cloudflare-pages cloudflare-workers hono podcast podcastfy saas text-to-speech typescript

Last synced: 15 Apr 2026

https://github.com/mahshid1378/tts-generation-webui

TTS Generation Web UI (Bark, MusicGen + AudioGen, Tortoise, RVC, Vocos, Demucs, SeamlessM4T, MAGNet, StyleTTS2, MMS, Stable Audio, Mars5, F5-TTS, ParlerTTS)

ai audio-generation audiogen bark deep-learning generator gradio machine-learning magnet music musicgen rvc seamlessm4t styletts2 text-to-speech torch tortoise-tts vocos web

Last synced: 14 Jul 2025

https://github.com/work-nobu/ohanashigpt

OhanashiGPT is an application that generates personalized children's stories based on parameters like age and preferences. It narrates these stories using an AI-generated voice that mimics a parent, trained on their audio samples. The app also creates illustrations to accompany each story, providing a unique and engaging experience for children.

ai audio-generation data-science image-generation large-language-models llama3 llamacpp lora low-rank-adaptation stable-diffusion text-generation xtts

Last synced: 09 Feb 2026

https://github.com/jet-logic/vocal_vse

Generates voiceover audio from text strips in the Blender Video Sequence Editor

audio-generation blender blender-addon blender-python text-to-speech tts video-editing vse

Last synced: 14 May 2026

https://github.com/deepgram-starters/rust-text-to-speech

Get started using Deepgram's Text-to-Speech with this Rust demo app

audio-generation axum deepgram demo quickstart rust text-to-speech tts

Last synced: 05 Apr 2026

https://github.com/runapi-ai/elevenlabs-sdk

RunAPI ElevenLabs SDK for text-to-speech, dialogue generation, sound effects, speech transcription, and audio isolation workflows in JavaScript, Python, Ruby, Go, Java, and PHP

api audio-generation elevenlabs elevenlabs-api golang gradle java maven python ruby runapi runapi-ai sdk speech-to-text text-to-speech typescript

Last synced: 28 Jul 2026

https://github.com/anas436/image-to-audio-app

Image Captioning and Text-to-Speech

audio-generation image-processing text-to-speech

Last synced: 27 Mar 2025

https://github.com/aidayang/inspiremusic-oneclick

InspireMusic文本转音乐软件免安装一键启动整合包

audio-generation audio-processing inspiremusic music-generation python pytorch

Last synced: 15 May 2026

https://github.com/charles-forsyth/generate-tts

A professional CLI for Google Gemini's Native 2.5 TTS. Generate multi-speaker podcasts ('Deep Dive'), audio summaries, and expressive speech from text/files.

ai-tools audio-generation cli gemini-api google-cloud multi-speaker podcast-generator python text-to-speech tts

Last synced: 13 Jan 2026