{"id":14400006,"url":"https://github.com/jianchang512/vocal-separate","last_synced_at":"2025-05-14T18:03:24.802Z","repository":{"id":214141107,"uuid":"735808717","full_name":"jianchang512/vocal-separate","owner":"jianchang512","description":"an extremely simple tool for separating vocals and background music, completely localized for web operation,  using 2stems/4stems/5stems models  这是一个极简的人声和背景音乐分离工具，本地化网页操作，无需连接外网","archived":false,"fork":false,"pushed_at":"2024-11-26T06:47:03.000Z","size":141377,"stargazers_count":1512,"open_issues_count":11,"forks_count":174,"subscribers_count":9,"default_branch":"main","last_synced_at":"2025-04-13T13:15:53.445Z","etag":null,"topics":["music-separation","spleeter","vocal-separation","voice-separation"],"latest_commit_sha":null,"homepage":"https://pyvideotrans.com","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"gpl-3.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/jianchang512.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2023-12-26T06:20:35.000Z","updated_at":"2025-04-11T06:48:14.000Z","dependencies_parsed_at":"2024-12-05T20:11:03.727Z","dependency_job_id":null,"html_url":"https://github.com/jianchang512/vocal-separate","commit_stats":{"total_commits":25,"total_committers":1,"mean_commits":25.0,"dds":0.0,"last_synced_commit":"52df271986b3b974eee38b42abcf6fe16b2acc9f"},"previous_names":["jianchang512/vocal-separate"],"tags_count":5,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/jianchang512%2Fvocal-separate","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/jianchang512%2Fvocal-separate/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/jianchang512%2Fvocal-separate/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/jianchang512%2Fvocal-separate/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/jianchang512","download_url":"https://codeload.github.com/jianchang512/vocal-separate/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":254198452,"owners_count":22030964,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["music-separation","spleeter","vocal-separation","voice-separation"],"created_at":"2024-08-29T07:01:04.528Z","updated_at":"2025-05-14T18:03:19.793Z","avatar_url":"https://github.com/jianchang512.png","language":"Python","funding_links":[],"categories":["AI 绘画 / 音频视频创作 \u003ca name=\"index--art\"\u003e\u0026nbsp;\u003c/a\u003e","语音识别与合成_其他"],"sub_categories":["音频视频处理","网络服务_其他"],"readme":"[English README](./README_EN.md) / [👑捐助该项目](https://github.com/jianchang512/pyvideotrans/blob/main/about.md) / [Discord](https://discord.gg/TMCM2PfHzQ) \n\n# 音乐人声分离工具\n\n这是一个极简的人声和背景音乐分离工具，本地化网页操作，无需连接外网，使用 2stems/4stems/5stems 模型。\n\n\n将一首歌曲或者含有背景音乐的音视频文件，拖拽到本地网页中，即可将其中的人声和音乐声分离为单独的音频wav文件，可选单独分离“钢琴声”、“贝斯声”、“鼓声”等\n\n自动调用本地浏览器打开本地网页，模型已内置，无需连接外网下载。\n\n支持视频(mp4/mov/mkv/avi/mpeg)和音频(mp3/wav)格式\n\n只需点两下鼠标，一选择音视频文件，二启动处理。\n\n\n\u003e **[赞助商]**\n\u003e \n\u003e [![](https://github.com/user-attachments/assets/5348c86e-2d5f-44c7-bc1b-3cc5f077e710)](https://gpt302.saaslink.net/teRK8Y)\n\u003e  [302.AI](https://gpt302.saaslink.net/teRK8Y)是一个按需付费的一站式AI应用平台，开放平台，开源生态, [302.AI开源地址](https://github.com/302ai)\n\u003e \n\u003e 集合了最新最全的AI模型和品牌/按需付费零月费/管理和使用分离/所有AI能力均提供API/每周推出2-3个新应用\n\n\n# 视频演示\n\nhttps://github.com/jianchang512/vocal-separate/assets/3378335/8e6b1b20-70d4-45e3-b106-268888fc0240\n\n\n\n![image](./images/1.png)\n\n\n\n# 预编译Win版使用方法/Linux和Mac源码部署\n\n1. [点击此处打开Releases页面下载](https://github.com/jianchang512/vocal-separate/releases)预编译文件\n\n2. 下载后解压到某处，比如 E:/vocal-separate\n\n3. 双击 start.exe ，等待自动打开浏览器窗口即可\n\n4. 点击页面中的上传区域，在弹窗中找到想分离的音视频文件，或直接拖拽音频文件到上传区域，然后点击“立即分离”，稍等片刻，底部会显示每个分离文件以及播放控件，点击播放。\n\n5. 如果机器拥有英伟达GPU，并正确配置了CUDA环境，将自动使用CUDA加速\n\n\n# 源码部署(Linux/Mac/Window)\n\n0. 要求 python 3.9-\u003e3.11\n\n1. 创建空目录，比如 E:/vocal-separate, 在这个目录下打开 cmd 窗口，方法是地址栏中输入 `cmd`, 然后回车。\n\n\t使用git拉取源码到当前目录 ` git clone git@github.com:jianchang512/vocal-separate.git . `\n\n2. 创建虚拟环境 `python -m venv venv`\n\n3. 激活环境，win下命令 `%cd%/venv/scripts/activate`，linux和Mac下命令 `source ./venv/bin/activate`\n\n4. 安装依赖: `pip install -r requirements.txt`\n\n5. win下解压 ffmpeg.7z，将其中的`ffmpeg.exe`和`ffprobe.exe`放在项目目录下, linux和mac 到 [ffmpeg官网](https://ffmpeg.org/download.html)下载对应版本ffmpeg，解压其中的`ffmpeg`和`ffprobe`二进制程序放到项目根目录下\n\n6. [下载模型压缩包](https://github.com/jianchang512/vocal-separate/releases/download/0.0/models-all.7z)，在项目根目录下的 `pretrained_models` 文件夹中解压，解压后，`pretrained_models`中将有3个文件夹，分别是`2stems`/`3stems`/`5stems`\n\n7. 执行  `python  start.py `，等待自动打开本地浏览器窗口。\n\n\n# API 接口\n\n接口地址: http://127.0.0.1:9999/api\n\n请求方法: POST\n\n请求参数:\n\n    file: 要分离的音视频文件\n\n    model: 模型名称 2stems,4stems,5stems\n\n返回响应: json\n    code:int, 0 处理成功完成，\u003e0 出错\n\n    msg:str,  出错时填充错误信息\n\n    data: List[str], 每个分离后的wav url地址，例如 ['http://127.0.0.1:9999/static/files/2/accompaniment.wav']\n\n    status_text: dict[str,str], 每个分离后wav文件的包含信息,{'accompaniment': '伴奏', 'bass': '低音', 'drums': '鼓', 'other': '其他', 'piano': '琴', 'vocals': '人声'}\n\n```\nimport requests\n# 请求地址\nurl = \"http://127.0.0.1:9999/api\"\nfiles = {\"file\": open(\"C:\\\\Users\\\\c1\\\\Videos\\\\2.wav\", \"rb\")}\ndata={\"model\":\"2stems\"}\nresponse = requests.request(\"POST\", url, timeout=600, data=data,files=files)\nprint(response.json())\n\n{'code': 0, 'data': ['http://127.0.0.1:9999/static/files/2/accompaniment.wav', 'http://127.0.0.1:9999/static/files/2/vocals.wav'], 'msg': '分离成功\n', 'status_text': {'accompaniment': '伴奏', 'bass': '低音', 'drums': '鼓', 'other': '其他', 'piano': '琴', 'vocals': '人声'}}\n\n\n```\n\n\n\n# CUDA 加速支持\n\n**安装CUDA工具** [详细安装方法](https://juejin.cn/post/7318704408727519270)\n\n如果你的电脑拥有 Nvidia 显卡，先升级显卡驱动到最新，然后去安装对应的 \n   [CUDA Toolkit 11.8](https://developer.nvidia.com/cuda-downloads)  和  [cudnn for CUDA11.X](https://developer.nvidia.com/rdp/cudnn-archive)。\n   \n   安装完成成，按`Win + R`,输入 `cmd`然后回车，在弹出的窗口中输入`nvcc --version`,确认有版本信息显示，类似该图\n   ![image](https://github.com/jianchang512/pyvideotrans/assets/3378335/e68de07f-4bb1-4fc9-bccd-8f841825915a)\n\n   然后继续输入`nvidia-smi`,确认有输出信息，并且能看到cuda版本号，类似该图\n   ![image](https://github.com/jianchang512/pyvideotrans/assets/3378335/71f1d7d3-07f9-4579-b310-39284734006b)\n\n\n\n# 注意事项\n\n0. 中文音乐或中式乐器，建议选择使用`2stems`模型，其他模型对“钢琴、贝斯、鼓”可单独分离出文件\n1. 如果电脑没有NVIDIA显卡或未配置cuda环境，不要选择 4stems和5stems模型，尤其是处理较长时长的音频时, 否则很可能耗尽内存\n\n\n\n# 致谢\n\n本项目主要依赖的其他项目\n\n1. https://github.com/deezer/spleeter\n2. https://github.com/pallets/flask\n3. https://ffmpeg.org/\n4. https://layui.dev\n\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fjianchang512%2Fvocal-separate","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fjianchang512%2Fvocal-separate","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fjianchang512%2Fvocal-separate/lists"}