An open API service indexing awesome lists of open source software.

Projects in Awesome Lists tagged with vad

A curated list of projects in awesome lists tagged with vad .

https://github.com/modelscope/funasr

A Fundamental End-to-End Speech Recognition Toolkit and Open Source SOTA Pretrained Models, Supporting Speech Recognition, Voice Activity Detection, Text Post-processing etc.

audio-visual-speech-recognition conformer dfsmn paraformer pretrained-model punctuation pytorch rnnt speaker-diarization speech-recognition speechgpt speechllm vad voice-activity-detection whisper

Last synced: 16 May 2025

https://github.com/modelscope/FunASR

A Fundamental End-to-End Speech Recognition Toolkit and Open Source SOTA Pretrained Models, Supporting Speech Recognition, Voice Activity Detection, Text Post-processing etc.

audio-visual-speech-recognition conformer dfsmn paraformer pretrained-model punctuation pytorch rnnt speaker-diarization speech-recognition speechgpt speechllm vad voice-activity-detection whisper

Last synced: 24 Mar 2025

https://github.com/k2-fsa/sherpa-ncnn

Real-time speech recognition and voice activity detection (VAD) using next-gen Kaldi with ncnn without Internet connection. Support iOS, Android, Linux, macOS, Windows, Raspberry Pi, VisionFive2, LicheePi4A etc.

asr c cpp csharp go kotlin python speech-recognition vad voice-activity-detection

Last synced: 13 May 2025

https://github.com/jtkim-kaist/VAD

Voice activity detection (VAD) toolkit including DNN, bDNN, LSTM and ACAM based VAD. We also provide our directly recorded dataset.

acam attention bdnn data dnn lstm speech speech-activity-detection speech-recognition vad voice-activity-detection voice-detection

Last synced: 07 May 2025

https://github.com/amsehili/auditok

An audio/acoustic activity detection and audio segmentation tool

audio-activities audio-data audio-segmentation vad voice-activity-detection voice-detection

Last synced: 22 Mar 2025

https://github.com/FluidInference/FluidAudio

Native Swift and CoreML SDK for local speaker diarization, VAD, and speech-to-text for real-time workloads. Works on iOS and macOS.

ane asr audio automatic-speech-recognition avfoundation coreml ios macos nvidia parakeet real-time speaker-diarization speaker-embedding speaker-identification speaker-recognition speech-to-text swift vad voice-activity-detection

Last synced: 31 Aug 2025

https://github.com/FireRedTeam/FireRedASR2S

A SOTA Industrial-Grade All-in-One ASR system with ASR, VAD, LID, and Punc modules. FireRedASR2 supports Chinese (Mandarin, 20+ dialects/accents), English, code-switching, and both speech and singing ASR. FireRedVAD supports speech/singing/music in 100+ langs. FireRedLID supports 100+ langs and 20+ zh dialects. FireRedPunc supports zh and en.

asr asr-pipeline audio-event-classification audio-event-detection automatic-speech-recognition industrial-grade language-identification lid llm multimodal-llm open-source punctuation-prediction punctuation-restoration sota speech-recognition speechllm vad voice-activity-detection

Last synced: 06 May 2026

https://github.com/DmitryRyumin/ICASSP-2023-24-Papers

ICASSP 2023-2024 Papers: A complete collection of influential and exciting research papers from the ICASSP 2023-24 conferences. Explore the latest advancements in acoustics, speech and signal processing. Code included. Star the repository to support the advancement of audio and signal processing!

asr denoising domain-adaptation face-recognition generative-models icassp icassp2023 icassp2024 image-generation keyword-spotting language-modeling multimodal-learning music-generation self-supervised-learning semantic-segmentation signal-processing signal-restoration speech-recognition spoken-language-understanding vad

Last synced: 14 Jul 2025

https://github.com/dmitryryumin/icassp-2023-24-papers

ICASSP 2023-2024 Papers: A complete collection of influential and exciting research papers from the ICASSP 2023-24 conferences. Explore the latest advancements in acoustics, speech and signal processing. Code included. Star the repository to support the advancement of audio and signal processing!

asr denoising domain-adaptation face-recognition generative-models icassp icassp2023 icassp2024 image-generation keyword-spotting language-modeling multimodal-learning music-generation self-supervised-learning semantic-segmentation signal-processing signal-restoration speech-recognition spoken-language-understanding vad

Last synced: 08 Apr 2025

https://github.com/shashikg/whispers2t

An Optimized Speech-to-Text Pipeline for the Whisper Model Supporting Multiple Inference Engine

asr deep-learning speech-recognition speech-to-text tensorrt tensorrt-llm vad voice-activity-detection whisper

Last synced: 12 Apr 2025

https://github.com/Baidu-AIP/speech-vad-demo

集成Webrtc的VAD,用于切分音频文件

speech vad webrtc webrtc-vad

Last synced: 04 May 2025

https://github.com/baidu-aip/speech-vad-demo

集成Webrtc的VAD,用于切分音频文件

speech vad webrtc webrtc-vad

Last synced: 06 Apr 2025

https://github.com/etienneab3d/whisperhallu

Experimental code: sound file preprocessing to optimize Whisper transcriptions without hallucinated texts

asr audio-processing noise-removal sound-processing text-to-speech vad vocals whisper

Last synced: 16 May 2025

https://github.com/shashikg/WhisperS2T

An Optimized Speech-to-Text Pipeline for the Whisper Model Supporting Multiple Inference Engine

asr deep-learning speech-recognition speech-to-text tensorrt tensorrt-llm vad voice-activity-detection whisper

Last synced: 08 May 2025

https://github.com/picovoice/cobra

On-device voice activity detection (VAD) powered by deep learning

on-device speech-recognition vad voice-activity voice-activity-detection voice-activity-detector

Last synced: 15 May 2025

https://github.com/Picovoice/cobra

On-device voice activity detection (VAD) powered by deep learning

on-device speech-recognition vad voice-activity voice-activity-detection voice-activity-detector

Last synced: 07 May 2025

https://github.com/eesungkim/Voice_Activity_Detector

A statistical model-based Voice Activity Detection

vad voice-activity-detection voice-detection

Last synced: 07 May 2025

https://github.com/xiongyihui/python-webrtc-audio-processing

Python bindings of WebRTC Audio Processing

agc ns python vad webrtc-audio-processing

Last synced: 05 Apr 2025

https://github.com/0vercl0k/sic

Enumerate user mode shared memory mappings on Windows.

driver ntoskrnl prototype-pte shared-memory shm vad windows-10 windows-kernel

Last synced: 14 Apr 2025

https://github.com/kigner/audio.cpp-webui

audio.cpp with a full-task WebUI - pure C++ audio-model inference engine powered by ggml. TTS, ASR/STT, VAD, voice conversion, speaker diarization, music generation. No Python dependency.

asr audio cpp ggml inference-engine music-generation speaker-diarization speech-recognition speech-synthesis speech-to-text text-to-speech tts vad voice-conversion webui

Last synced: 25 Aug 2026

https://github.com/xia-chu/webrtc_apm

webrtc中apm相关代码的提取,包括AEC/NS/AGC/VAD ,另外还包括mp3/aac编码器、SoundTouch

aac aec agc jni mp3 ns soundtouch vad webrtc

Last synced: 23 Apr 2025

https://github.com/mgonzs13/whisper_ros

Speech-to-Text based on SileroVAD + whisper.cpp (GGML Whisper) for ROS 2

asr automatic-speech-recognition ggml ros2 speech-recognition speech-to-text vad voice-activity-detection whisper whisper-cpp

Last synced: 30 Aug 2025

https://github.com/etienneab3d/karaok-ai

Karaoke Player / Editor with automatic clip creation from any song file using vocals and lyrics extraction (Speech-to-Text)

djing karaoke karaoke-maker lyrics mp3-player music party-apps sound-processing speech-to-text srt-subtitles subtitles vad whisper

Last synced: 26 Oct 2025

https://github.com/EtienneAb3d/karaok-AI

Karaoke Player / Editor with automatic clip creation from any song file using vocals and lyrics extraction (Speech-to-Text)

djing karaoke karaoke-maker lyrics mp3-player music party-apps sound-processing speech-to-text srt-subtitles subtitles vad whisper

Last synced: 15 Apr 2025

https://github.com/mounalab/LSTM-RNN-VAD

Voice Activity Detection LSTM-RNN learning model

lstm lstm-neural-network nlp-machine-learning rnn rnn-tensorflow tensorflow vad

Last synced: 07 May 2025

https://github.com/baabaaox/go-webrtcvad

WebRTC Voice Activity Detection for Golang

cgo go golang vad webrtc webrtcvad

Last synced: 20 Jan 2026

https://github.com/pguso/voice-agents-from-scratch

From-scratch voice agents in Python: end-to-end speech pipelines, runnable chapters, and a small shared library. Local models, explicit streaming behavior.

agents edge-ai faster-whisper kokoro-82m llm local-ai onnx python speech-to-text streaming text-to-speech tool-calling tool-calling-agent tutorial uv vad voice-activity-detection voice-agent voice-agents whisper

Last synced: 15 Aug 2026

https://github.com/mochi-neko/voice-activity-detection-unity

A voice activity detection (VAD) library for Unity.

unity vad

Last synced: 08 Apr 2026

https://github.com/lgrammel/whisperwriter

Local & private voice controlled notepad using whisper.cpp

nextjs stt transcription vad whisper-cpp

Last synced: 05 May 2025

https://github.com/asiff00/on-device-speech-to-speech-conversational-ai

This is an on-CPU real-time conversational system for two-way speech communication with AI models, utilizing a continuous streaming architecture for fluid conversations with immediate responses and natural interruption handling.

asr audio-processing conversational-ai kokoro-tts ollama speech-to-speech tts vad voice-assistant

Last synced: 24 Feb 2026

https://github.com/xulihang/Silhouette

An open source computer-aided translation tool for audios and videos

computer-aided-translation forced-alignment mac speech-recognition subtitle vad whisper

Last synced: 11 Jul 2026

https://github.com/thewh1teagle/vad-rs

Speech detection using silero vad in Rust

onnxruntime rust speech-recognition vad

Last synced: 18 Mar 2025

https://github.com/pykeio/earshot

Ridiculously fast voice activity detection in pure #[no_std] Rust

rust vad voice-activity-detection

Last synced: 11 May 2025

https://github.com/sshh12/conv-vad

A packaged convolutional voice activity detector for noisy environments.

convolutional-neural-networks keras melspectrogram vad voice-activity-detection

Last synced: 19 Mar 2025

https://github.com/thurti/vad-audio-worklet

Voice Activity Detection (VAD) AudioWorklet

audioworklet audioworkletprocessor speech vad voice-activity-detection

Last synced: 07 Apr 2025

https://github.com/picovoice/voice-activity-benchmark

Voice activity engine benchmark framework

benchmark benchmark-framework vad voice-activity

Last synced: 14 Jul 2025

https://github.com/panmasuo/voice-activity-detection

Voice activity detection algorithm written in C

alsa c language paho-mqtt vad voice-activity-detection

Last synced: 10 Apr 2025

https://github.com/zygotecode/vadsharp

Enterprise VAD (Voice Activity Detection) in C#.NET (.NET 6.0+) with Microsoft.ML.Net, ONNXRuntime and DirectML. The easiest, efficient, and performant Silero VAD implementation! Always open for PRs.

csharp dotnet onnx onnx-runtime onnxruntime silero-vad speech speech-processing vad voice-activity-detection voice-control voice-recognition

Last synced: 10 Feb 2026

https://github.com/sangmin7648/tacit

Harness tacit knowledge into the context of AI agent with on device always-on transcription

agent agent-skills autodetection harness harness-engineering local-ai ondevice-ai vad

Last synced: 14 Apr 2026

https://github.com/chenqianhe/vad-addon

This repo provides an addon that can perform VAD model reasoning in nodes and electric environments, based on cmake-js and Fastdeploy. Silero VAD is a pre-trained enterprise-grade Voice Activity Detector.

addon electron electron-addon node-addon nodejs silero-vad vad

Last synced: 11 Apr 2025

https://github.com/daanzu/py-silero-vad-lite

Lightweight wrapper for Silero VAD using internal ONNX Runtime and with no python package dependencies

python speech speech-processing vad voice voice-activity-detection

Last synced: 19 Apr 2025

https://github.com/4players/odin-sdk

Reliable cross-platform SDK enabling developers to integrate real-time VoIP chat technology into games, apps and websites

apm c chat client console-client cross-platform http3 network opus-codec pre-compiled proximity-chat rust sdk transport vad voice voip webtransport

Last synced: 18 Jan 2026

https://github.com/dbklim/webrtcvad_wrapper

A simple Python wrapper to simplify working with WebRTC VAD and its rougher analogue based on RMS and ZCR (useful for processing audio recordings before using them with neural networks).

audio audio-processing dsp forced-alignment python silence-suppression vad vad-detection voice-activity-detection webrtc webrtc-tools webrtc-vad webrtcvad-wrapper

Last synced: 19 Jul 2025

https://github.com/helloooideeeeea/realtimecutvadlibrary

A real-time Voice Activity Detection (VAD) library for iOS and macOS using Silero models powered by ONNX Runtime. Includes advanced noise suppression and audio preprocessing with WebRTC APM, supporting seamless WAV data output with header metadata.

ios macos onnxruntime silero-vad vad webrtc-audio-processing

Last synced: 12 Apr 2025

https://github.com/nico-martin/vad-recorder

A browser-focused TypeScript library that combines voice activity detection (VAD) with automatic audio segment recording.

silero-vad vad webml

Last synced: 28 May 2026

https://github.com/decibri/decibri

Audio capture, playback, and processing for Node.js and the browser. One Rust codebase, zero system dependencies.

audio audio-capture audio-playback browser cross-platform microphone microphone-capture nodejs pcm rust speech-recognition vad voice voice-activity-detection voice-ai

Last synced: 07 May 2026

https://github.com/bincrafters/conan-libfvad

Conan.io package for libfvad project

conan libfvad vad voice voice-activity-detection

Last synced: 29 Apr 2026

https://jacoblincool.github.io/awesome-recorder/

Effortless audio recording with built-in Voice Activity Detection and optimized MP3/WAV/PCM outputs in modern browsers.

audio-recorder browser mp3 vad

Last synced: 22 May 2026

https://github.com/sirmews/textcast

Browser audio recorder with Whisper transcription

silero transformers-js vad webaudio webgpu whisper

Last synced: 28 Apr 2026

https://github.com/chicogong/realtime-ai

Real-time AI voice conversation platform with WebSocket, supporting streaming STT/LLM/TTS

azure-speech fastapi openai python real-time realtime-voice streaming vad voice-assistant

Last synced: 08 Feb 2026

https://github.com/eja/wav2vad

A command line tool for voice activity detection.

silero vad wav

Last synced: 28 Apr 2025

https://github.com/monhi/vad_recorder

a mfc program to capture audio into file by finding activity on the captured audio. speex open source library is used in this project.

michrophone online vad

Last synced: 11 Jun 2026

https://github.com/nico-byte/whisper-web

The Whisper Web Transcription Server is a Python-based real-time speech-to-text transcription system powered by OpenAI's Whisper models. It leverages state-of-the-art models like Distil-Whisper to transcribe audio input in real-time.

ai asr automatic-speech-recognition distil-whisper distil-whisper-large-v3 huggingface huggingface-transformers server vad voice web websockets whisper

Last synced: 26 Apr 2026

https://github.com/chicogong/ffvoice-engine

🎙️ 高性能 C++ 语音引擎 - 实时音频处理 + AI 语音识别 + 边录边转写 | High-performance C++ voice engine with real-time ASR and RNNoise

ai audio-processing audio-recording cmake cpp20 ffmpeg flac machine-learning noise-suppression offline-asr portaudio real-time real-time-transcription rnnoise speech-recognition speech-to-text subtitle-generation vad voice-activity-detection whisper

Last synced: 29 Apr 2026

https://github.com/sultanfariz/go-vad

Voice Activity Detection package for Golang. Might be useful for speech-to-speech or speech-to-text system.

speech speech-processing vad voice-activity-detection

Last synced: 13 Jul 2026

https://github.com/rogerchappel/bargekit

Local-first VAD, barge-in, and turn-taking primitives for interruptible voice agents.

agents barge-in duplex echo-guard local-first microphone speech-detection turn-taking vad voice-agent voice-ui

Last synced: 05 Jun 2026

https://github.com/kazuhito00/silero-vad-onnx-sample

Silero VADのONNX推論(PyTorch依存処理無し)サンプル

colaboratory onnx onnxruntime python silero-vad vad

Last synced: 02 Jul 2026

https://github.com/bubustack/livekit-turn-detector-engram

Turn detector Engram for bobrapet — buffers PCM audio, runs WebRTC VAD, and emits speech turn payloads.

batch bubustack engram go kubernetes livekit streaming turn-detection vad

Last synced: 20 May 2026

https://github.com/egorsmkv/marblenet-inference

Inference code for Frame MarbleNet (VAD from NeMo)

marblenet ml nemo nvidia speech vad voice-activity-detection

Last synced: 07 Oct 2025

https://github.com/prkbuilds/pablos_therapy

Pablo's Therapy - A multilingual therapy voice session made with generative AI chatbot

dart elevenlabs flutter genai ios portfolio vad

Last synced: 27 Jan 2026

https://github.com/xulihang/silhouette

A computer-aided translation tool for audios and videos

computer-aided-translation speech-recognition subtitle vad whisper

Last synced: 08 Mar 2026

https://github.com/ganymedenil/go-webrtcvad

cgo interface to WebRTC Voice Activity Dectection

cgo vad webrtc webrtc-vad webrtcvad

Last synced: 22 Apr 2026

https://github.com/steinathan/telephony-server

Telephony Server is a powerful bridge that connects telephony providers (Twilio, Vonage, Plivo, etc.) with real-time communication platforms (LiveKit, Jay.so, Pipecat, etc.). It enables seamless call routing, robust metrics collection, and observability features for enhanced telephony operations.

agent jay langchain llm multimodal pilvo pipecat pydantic telephony telephonymanager twilio vad vocode voice voiceai vonage

Last synced: 11 Jul 2025

https://github.com/nipponjo/silero-vad-wasm

WebAssembly build of Silero VAD with a small browser demo

speech vad voice voice-activity-detection wasm webassembly

Last synced: 14 Aug 2026

https://github.com/OpenVoiceOS/ovos-vad-plugin-webrtcvad

ovos plugin for voice activity detection using webrtcvad

openvoiceos ovos vad voice-activity-detection

Last synced: 14 May 2025

https://github.com/nestarz/vad

ES6 Voice Activity Detection using Silero models

browser deno vad voice-activity-detection

Last synced: 30 Apr 2026

https://github.com/OpenVoiceOS/ovos-vad-plugin-silero

ovos plugin for voice activity detection using silero vad

openvoiceos ovos plugin vad voice-activity-detection

Last synced: 14 May 2025

https://github.com/openvoiceos/ovos-vad-plugin-webrtcvad

ovos plugin for voice activity detection using webrtcvad

openvoiceos ovos vad voice-activity-detection

Last synced: 14 Mar 2025

https://github.com/bubustack/silero-vad-engram

Silero VAD Engram for bobrapet — ONNX-based voice activity detection for streaming audio pipelines.

batch bubustack engram go kubernetes onnx silero streaming vad voice-activity-detection

Last synced: 20 May 2026

https://github.com/seriouslysean/transcript-precombobulator

A tool to preprocess audio files for transcription by removing silence and preparing them for whisper.cpp

audio-processing python transcription vad whisper

Last synced: 28 Mar 2025

https://github.com/numq/voice-activity-detection

JVM library for voice activity detection written in Kotlin based on C library fvad and Silero

cpp fvad java jni jvm kotlin libfvad ml onnx silero silero-vad vad voice-activity-detection

Last synced: 07 Apr 2026

https://github.com/xdcobra/react-native-sherpa-onnx

Offline Speech Processing SDK for React Native using sherpa-onnx. Supports Speech-to-Text, Text-to-Speech, Speaker Diarization, Speech Enhancement, Source Separation & VAD.

android ctc data-privacy funasr ios paraformer react-native sdk sensevoice sherpa-onnx source-separation speaker-diarization speech-enhancement speech-to-text text-to-speech vad voice-activity-detection whisper zipformer

Last synced: 02 Apr 2026

https://github.com/gbibbo/vad_benchmark

Privacy‑preserving VAD benchmark on domestic audio (CHiME‑Home): 8 models, accuracy vs efficiency.

audio-processing benchmark cnn edge-ai f1-score panns passt privacy privacy-preserving real-time roc-auc silero-vad transformer vad voice-activity-detection webrtc whisper

Last synced: 16 Apr 2026

https://github.com/wavekat/wavekat-lab

Developer experimentation tools for the WaveKat libraries. Includes vad-lab, a web-based tool for testing and comparing VAD backends side by side.

audio audio-processing developer-tools rust speech-detection vad voice voice-ai wavekat

Last synced: 15 May 2026

https://github.com/ekhodzitsky/polyvoice

Speaker diarization for Rust — who spoke when, without Python. Silero VAD + WeSpeaker + AHC in a single Pipeline::run() call.

audio diarization machine-learning onnx python-bindings rust speaker-diarization speech vad voice

Last synced: 16 May 2026

https://github.com/seriouslysean/transcript-combobulator

A tool to preprocess audio files for transcription by removing silence and preparing them for whisper.cpp

audio-processing python transcription vad whisper

Last synced: 22 Jun 2026

https://github.com/sheldonix/silero-vad-rust

Rust port of Silero VAD for ultra-fast inference with bundled ONNX models

audio onnx onnxruntime silero speech vad voice voice-activity-detection voice-assistant voice-detection voice-recognition

Last synced: 16 Dec 2025

https://github.com/clg3227783168/paper

paper readings

control heartfailure rl vad

Last synced: 12 Aug 2026