{"id":13434181,"url":"https://github.com/kennethleungty/Text-to-Audio-with-Bark","last_synced_at":"2025-03-17T14:30:52.631Z","repository":{"id":197480805,"uuid":"697602603","full_name":"kennethleungty/Text-to-Audio-with-Bark","owner":"kennethleungty","description":"Exploring Bark, the Open-Source Text-to-Audio Generative Model","archived":false,"fork":false,"pushed_at":"2023-10-10T16:55:13.000Z","size":2804,"stargazers_count":15,"open_issues_count":0,"forks_count":4,"subscribers_count":3,"default_branch":"main","last_synced_at":"2025-03-04T08:51:18.874Z","etag":null,"topics":["ai","artificial-intelligence","bark","data-science","deep-learning","gen-ai","generative-ai","machine-learning","prompt-engineering","speech","text-prompt","text-to-audio","text-to-music","text-to-sound","text-to-speech"],"latest_commit_sha":null,"homepage":"https://betterprogramming.pub/text-to-audio-generation-with-bark-clearly-explained-4ee300a3713a","language":"Jupyter Notebook","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/kennethleungty.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null}},"created_at":"2023-09-28T04:57:01.000Z","updated_at":"2024-07-04T18:59:10.000Z","dependencies_parsed_at":null,"dependency_job_id":"1f911441-bb4d-44af-956f-550dbee03e51","html_url":"https://github.com/kennethleungty/Text-to-Audio-with-Bark","commit_stats":null,"previous_names":["kennethleungty/text-to-audio-with-bark"],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/kennethleungty%2FText-to-Audio-with-Bark","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/kennethleungty%2FText-to-Audio-with-Bark/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/kennethleungty%2FText-to-Audio-with-Bark/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/kennethleungty%2FText-to-Audio-with-Bark/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/kennethleungty","download_url":"https://codeload.github.com/kennethleungty/Text-to-Audio-with-Bark/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":244050061,"owners_count":20389632,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["ai","artificial-intelligence","bark","data-science","deep-learning","gen-ai","generative-ai","machine-learning","prompt-engineering","speech","text-prompt","text-to-audio","text-to-music","text-to-sound","text-to-speech"],"created_at":"2024-07-31T02:01:48.826Z","updated_at":"2025-03-17T14:30:52.168Z","avatar_url":"https://github.com/kennethleungty.png","language":"Jupyter Notebook","funding_links":[],"categories":["Jupyter Notebook"],"sub_categories":[],"readme":"# Exploring Text-to-Audio with Bark\n\nLink to article: https://betterprogramming.pub/text-to-audio-generation-with-bark-clearly-explained-4ee300a3713a\n\n## Context\n- Amidst the transformative surge of generative AI, text-to-audio models are emerging as one of the most promising frontiers. \n- These advances are not just about converting text to speech, but also about crafting audio experiences that are indistinguishable from human-produced content.\n- From audiobooks narrated in any voice imaginable to dynamic music compositions prompted by mere sentences, the potential applications are vast and captivating.\n- In this article, we delve into the capabilities and technical intricacies of Bark, an open-source text-prompted audio generation model in Python.\n\n___\n\n## Introducing Bark\nBark is a transformer-based text-to-audio model capable of generating realistic multilingual speech, music, and sound effects. It is created by Suno, a research-driven company that develops cutting-edge audio AI.\nAs Bark was developed for research purposes, its pre-trained model checkpoints have been made open-source and available for commercial use, which is a valuable contribution to the generative AI community.\n\n___\n\n### References\n- https://github.com/suno-ai/bark\n- https://audiocraft.metademolab.com/encodec.html\n- https://www.streamingmedia.com/Articles/ReadArticle.aspx?ArticleID=74487 \n- https://towardsdatascience.com/optimizing-vector-quantization-methods-by-machine-learning-algorithms-77c436d0749d\n- https://www.assemblyai.com/blog/what-is-residual-vector-quantization/\n- https://github.com/facebookresearch/encodec\n- https://ai.meta.com/blog/ai-powered-audio-compression-technique/\n- https://arxiv.org/abs/2210.13438\n- https://github.com/facebookresearch/encodec#extracting-discrete-representations \n- https://paperswithcode.com/paper/speaker-anonymization-using-neural-audio\n- https://huggingface.co/suno/bark/tree/main/speaker_embeddings/v2\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fkennethleungty%2FText-to-Audio-with-Bark","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fkennethleungty%2FText-to-Audio-with-Bark","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fkennethleungty%2FText-to-Audio-with-Bark/lists"}