{"id":13545284,"url":"https://github.com/snakers4/silero-vad","last_synced_at":"2025-05-13T20:02:45.208Z","repository":{"id":37398935,"uuid":"315270108","full_name":"snakers4/silero-vad","owner":"snakers4","description":"Silero VAD: pre-trained enterprise-grade Voice Activity Detector","archived":false,"fork":false,"pushed_at":"2025-03-24T16:02:57.000Z","size":105166,"stargazers_count":5711,"open_issues_count":16,"forks_count":545,"subscribers_count":55,"default_branch":"master","last_synced_at":"2025-05-06T19:51:52.044Z","etag":null,"topics":["onnx","onnx-runtime","onnxruntime","pytorch","speech","speech-processing","vad","voice-activity-detection","voice-commands","voice-control","voice-detection","voice-recognition"],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/snakers4.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":"CODE_OF_CONDUCT.md","threat_model":null,"audit":null,"citation":"CITATION.cff","codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null}},"created_at":"2020-11-23T09:54:16.000Z","updated_at":"2025-05-06T13:57:12.000Z","dependencies_parsed_at":"2023-02-08T09:45:57.614Z","dependency_job_id":"974cdf1e-1869-424e-bebc-131e6a037a63","html_url":"https://github.com/snakers4/silero-vad","commit_stats":{"total_commits":287,"total_committers":37,"mean_commits":7.756756756756757,"dds":0.554006968641115,"last_synced_commit":"9060f664f20eabb66328e4002a41479ff288f14c"},"previous_names":[],"tags_count":9,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/snakers4%2Fsilero-vad","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/snakers4%2Fsilero-vad/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/snakers4%2Fsilero-vad/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/snakers4%2Fsilero-vad/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/snakers4","download_url":"https://codeload.github.com/snakers4/silero-vad/tar.gz/refs/heads/master","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":254020468,"owners_count":22000749,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["onnx","onnx-runtime","onnxruntime","pytorch","speech","speech-processing","vad","voice-activity-detection","voice-commands","voice-control","voice-detection","voice-recognition"],"created_at":"2024-08-01T11:01:00.229Z","updated_at":"2025-05-13T20:02:45.132Z","avatar_url":"https://github.com/snakers4.png","language":"Python","funding_links":[],"categories":["Python","Uncategorized","Speech Enhancement \u0026 Audio Processing","VAD (Voice Activity Detection) | 语音活动检测","Repos","Tools\u003ca id=\"tool\"\u003e\u003c/a\u003e","🤖 AI \u0026 Machine Learning","6. Voice activity detection and turn-taking"],"sub_categories":["Uncategorized","Voice Activity Detection (VAD)","Core VAD Models | 核心 VAD 模型","Others\u003ca id=\"paper11\"\u003e\u003c/a\u003e","Voice-specific prompting and tools"],"readme":"[![Mailing list : test](http://img.shields.io/badge/Email-gray.svg?style=for-the-badge\u0026logo=gmail)](mailto:hello@silero.ai) [![Mailing list : test](http://img.shields.io/badge/Telegram-blue.svg?style=for-the-badge\u0026logo=telegram)](https://t.me/silero_speech) [![License: CC BY-NC 4.0](https://img.shields.io/badge/License-MIT-lightgrey.svg?style=for-the-badge)](https://github.com/snakers4/silero-vad/blob/master/LICENSE) [![downloads](https://img.shields.io/pypi/dm/silero-vad?style=for-the-badge)](https://pypi.org/project/silero-vad/)\n\n[![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/snakers4/silero-vad/blob/master/silero-vad.ipynb)\n\n![header](https://user-images.githubusercontent.com/12515440/89997349-b3523080-dc94-11ea-9906-ca2e8bc50535.png)\n\n\u003cbr/\u003e\n\u003ch1 align=\"center\"\u003eSilero VAD\u003c/h1\u003e\n\u003cbr/\u003e\n\n**Silero VAD** - pre-trained enterprise-grade [Voice Activity Detector](https://en.wikipedia.org/wiki/Voice_activity_detection) (also see our [STT models](https://github.com/snakers4/silero-models)).\n\n\u003cbr/\u003e\n\n\u003cp align=\"center\"\u003e\n  \u003cimg src=\"https://github.com/snakers4/silero-vad/assets/36505480/300bd062-4da5-4f19-9736-9c144a45d7a7\" /\u003e\n\u003c/p\u003e\n\n\n\u003cdetails\u003e\n\u003csummary\u003eReal Time Example\u003c/summary\u003e\n\nhttps://user-images.githubusercontent.com/36505480/144874384-95f80f6d-a4f1-42cc-9be7-004c891dd481.mp4\n\nPlease note, that video loads only if you are logged in your GitHub account. \n\n\u003c/details\u003e\n\n\u003cbr/\u003e\n\n\u003ch2 align=\"center\"\u003eFast start\u003c/h2\u003e\n\u003cbr/\u003e\n\n\u003cdetails\u003e\n\u003csummary\u003eDependencies\u003c/summary\u003e\n\n  System requirements to run python examples on `x86-64` systems:\n  \n  - `python 3.8+`;\n  - 1G+ RAM;\n  - A modern CPU with AVX, AVX2, AVX-512 or AMX instruction sets.\n\n  Dependencies:\n  \n  - `torch\u003e=1.12.0`;\n  - `torchaudio\u003e=0.12.0` (for I/O only);\n  - `onnxruntime\u003e=1.16.1` (for ONNX model usage).\n  \n  Silero VAD uses torchaudio library for audio I/O (`torchaudio.info`, `torchaudio.load`, and `torchaudio.save`), so a proper audio backend is required:\n  \n  - Option №1 - [**FFmpeg**](https://www.ffmpeg.org/) backend. `conda install -c conda-forge 'ffmpeg\u003c7'`;\n  - Option №2 - [**sox_io**](https://pypi.org/project/sox/) backend. `apt-get install sox`, TorchAudio is tested on libsox 14.4.2;\n  - Option №3 - [**soundfile**](https://pypi.org/project/soundfile/) backend. `pip install soundfile`.\n\nIf you are planning to run the VAD using solely the `onnx-runtime`, it will run on any other system architectures where onnx-runtume is [supported](https://onnxruntime.ai/getting-started). In this case please note that:\n\n- You will have to implement the I/O;\n- You will have to adapt the existing wrappers / examples / post-processing for your use-case.\n\n\u003c/details\u003e\n\n**Using pip**:\n`pip install silero-vad`\n\n```python3\nfrom silero_vad import load_silero_vad, read_audio, get_speech_timestamps\nmodel = load_silero_vad()\nwav = read_audio('path_to_audio_file')\nspeech_timestamps = get_speech_timestamps(\n  wav,\n  model,\n  return_seconds=True,  # Return speech timestamps in seconds (default is samples)\n)\n```\n\n**Using torch.hub**:\n```python3\nimport torch\ntorch.set_num_threads(1)\n\nmodel, utils = torch.hub.load(repo_or_dir='snakers4/silero-vad', model='silero_vad')\n(get_speech_timestamps, _, read_audio, _, _) = utils\n\nwav = read_audio('path_to_audio_file')\nspeech_timestamps = get_speech_timestamps(\n  wav,\n  model,\n  return_seconds=True,  # Return speech timestamps in seconds (default is samples)\n)\n```\n\n\u003cbr/\u003e\n\n\u003ch2 align=\"center\"\u003eKey Features\u003c/h2\u003e\n\u003cbr/\u003e\n\n- **Stellar accuracy**\n\n  Silero VAD has [excellent results](https://github.com/snakers4/silero-vad/wiki/Quality-Metrics#vs-other-available-solutions) on speech detection tasks.\n  \n- **Fast**\n\n  One audio chunk (30+ ms) [takes](https://github.com/snakers4/silero-vad/wiki/Performance-Metrics#silero-vad-performance-metrics) less than **1ms** to be processed on a single CPU thread. Using batching or GPU can also improve performance considerably. Under certain conditions ONNX may even run up to 4-5x faster. \n\n- **Lightweight**\n\n  JIT model is around two megabytes in size.\n\n- **General**\n\n  Silero VAD was trained on huge corpora that include over **6000** languages and it performs well on audios from different domains with various background noise and quality levels.\n\n- **Flexible sampling rate**\n\n  Silero VAD [supports](https://github.com/snakers4/silero-vad/wiki/Quality-Metrics#sample-rate-comparison)  **8000 Hz** and **16000 Hz** [sampling rates](https://en.wikipedia.org/wiki/Sampling_(signal_processing)#Sampling_rate).\n\n- **Highly Portable**\n\n  Silero VAD reaps benefits from the rich ecosystems built around **PyTorch** and **ONNX** running everywhere where these runtimes are available.\n\n- **No Strings Attached**\n\n   Published under permissive license (MIT) Silero VAD has zero strings attached - no telemetry, no keys, no registration, no built-in expiration, no keys or vendor lock.\n\n\u003cbr/\u003e\n\n\u003ch2 align=\"center\"\u003eTypical Use Cases\u003c/h2\u003e\n\u003cbr/\u003e\n\n- Voice activity detection for IOT / edge / mobile use cases\n- Data cleaning and preparation, voice detection in general\n- Telephony and call-center automation, voice bots\n- Voice interfaces\n\n\u003cbr/\u003e\n\u003ch2 align=\"center\"\u003eLinks\u003c/h2\u003e\n\u003cbr/\u003e\n\n\n- [Examples and Dependencies](https://github.com/snakers4/silero-vad/wiki/Examples-and-Dependencies#dependencies)\n- [Quality Metrics](https://github.com/snakers4/silero-vad/wiki/Quality-Metrics)\n- [Performance Metrics](https://github.com/snakers4/silero-vad/wiki/Performance-Metrics)\n- [Versions and Available Models](https://github.com/snakers4/silero-vad/wiki/Version-history-and-Available-Models)\n- [Further reading](https://github.com/snakers4/silero-models#further-reading)\n- [FAQ](https://github.com/snakers4/silero-vad/wiki/FAQ)\n\n\u003cbr/\u003e\n\u003ch2 align=\"center\"\u003eGet In Touch\u003c/h2\u003e\n\u003cbr/\u003e\n\nTry our models, create an [issue](https://github.com/snakers4/silero-vad/issues/new), start a [discussion](https://github.com/snakers4/silero-vad/discussions/new), join our telegram [chat](https://t.me/silero_speech), [email](mailto:hello@silero.ai) us, read our [news](https://t.me/silero_news).\n\nPlease see our [wiki](https://github.com/snakers4/silero-models/wiki) for relevant information and [email](mailto:hello@silero.ai) us directly.\n\n**Citations**\n\n```\n@misc{Silero VAD,\n  author = {Silero Team},\n  title = {Silero VAD: pre-trained enterprise-grade Voice Activity Detector (VAD), Number Detector and Language Classifier},\n  year = {2024},\n  publisher = {GitHub},\n  journal = {GitHub repository},\n  howpublished = {\\url{https://github.com/snakers4/silero-vad}},\n  commit = {insert_some_commit_here},\n  email = {hello@silero.ai}\n}\n```\n\n\u003cbr/\u003e\n\u003ch2 align=\"center\"\u003eExamples and VAD-based Community Apps\u003c/h2\u003e\n\u003cbr/\u003e\n\n- Example of VAD ONNX Runtime model usage in [C++](https://github.com/snakers4/silero-vad/tree/master/examples/cpp) \n\n- Voice activity detection for the [browser](https://github.com/ricky0123/vad) using ONNX Runtime Web\n\n- [Rust](https://github.com/snakers4/silero-vad/tree/master/examples/rust-example), [Go](https://github.com/snakers4/silero-vad/tree/master/examples/go), [Java](https://github.com/snakers4/silero-vad/tree/master/examples/java-example), [C++](https://github.com/snakers4/silero-vad/tree/master/examples/cpp), [C#](https://github.com/snakers4/silero-vad/tree/master/examples/csharp) and [other](https://github.com/snakers4/silero-vad/tree/master/examples) community examples\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fsnakers4%2Fsilero-vad","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fsnakers4%2Fsilero-vad","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fsnakers4%2Fsilero-vad/lists"}