Projects in Awesome Lists tagged with speaker-diarization
A curated list of projects in awesome lists tagged with speaker-diarization .
https://github.com/modelscope/funasr
A Fundamental End-to-End Speech Recognition Toolkit and Open Source SOTA Pretrained Models, Supporting Speech Recognition, Voice Activity Detection, Text Post-processing etc.
audio-visual-speech-recognition conformer dfsmn paraformer pretrained-model punctuation pytorch rnnt speaker-diarization speech-recognition speechgpt speechllm vad voice-activity-detection whisper
Last synced: 16 May 2025
https://github.com/espnet/espnet
End-to-End Speech Processing Toolkit
chainer deep-learning end-to-end kaldi machine-translation pytorch singing-voice-synthesis speaker-diarization speech-enhancement speech-recognition speech-separation speech-synthesis speech-translation spoken-language-understanding text-to-speech voice-conversion
Last synced: 08 Apr 2026
https://github.com/speechbrain/speechbrain
A PyTorch-based Speech Toolkit
asr audio audio-processing deep-learning huggingface language-model pytorch speaker-diarization speaker-recognition speaker-verification speech-enhancement speech-processing speech-recognition speech-separation speech-to-text speech-toolkit speechrecognition spoken-language-understanding transformers voice-recognition
Last synced: 13 May 2025
https://github.com/pyannote/pyannote-audio
Neural building blocks for speaker diarization: speech activity detection, speaker change detection, overlapped speech detection, speaker embedding
overlapped-speech-detection pretrained-models pytorch speaker-change-detection speaker-diarization speaker-embedding speaker-recognition speaker-verification speech-activity-detection speech-processing voice-activity-detection
Last synced: 13 May 2025
https://github.com/modelscope/FunASR
A Fundamental End-to-End Speech Recognition Toolkit and Open Source SOTA Pretrained Models, Supporting Speech Recognition, Voice Activity Detection, Text Post-processing etc.
audio-visual-speech-recognition conformer dfsmn paraformer pretrained-model punctuation pytorch rnnt speaker-diarization speech-recognition speechgpt speechllm vad voice-activity-detection whisper
Last synced: 24 Mar 2025
https://github.com/mahmoudashraf97/whisper-diarization
Automatic Speech Recognition with Speaker Diarization based on OpenAI Whisper
asr speaker-diarization speech speech-recognition speech-to-text whisper
Last synced: 13 May 2025
https://github.com/MahmoudAshraf97/whisper-diarization
Automatic Speech Recognition with Speaker Diarization based on OpenAI Whisper
asr speaker-diarization speech speech-recognition speech-to-text whisper
Last synced: 28 Mar 2025
https://github.com/linto-ai/whisper-timestamped
Multilingual Automatic Speech Recognition with word-level timestamps and confidence
asr attention-is-all-you-need attention-mechanism attention-model attention-network attention-seq2seq attention-visualization deep-learning machine-learning multilingual-models python python3 pytorch speaker-diarization speech speech-processing speech-recognition speech-to-text transformers whisper
Last synced: 13 May 2025
https://github.com/purfview/whisper-standalone-win
Whisper & Faster-Whisper standalone executables for those who don't want to bother with Python.
asr ctranslate2 diarization faster-whisper openai speaker-diarization speech-recognition speech-to-text subtitles transcriber uvr vocal-extractor whisper whisper-faster whisperx
Last synced: 14 May 2025
https://github.com/Purfview/whisper-standalone-win
Whisper & Faster-Whisper standalone executables for those who don't want to bother with Python.
asr ctranslate2 diarization faster-whisper openai speaker-diarization speech-recognition speech-to-text subtitles transcriber uvr vocal-extractor whisper whisper-faster whisperx
Last synced: 28 Mar 2025
https://github.com/google/uis-rnn
This is the library for the Unbounded Interleaved-State Recurrent Neural Network (UIS-RNN) algorithm, corresponding to the paper Fully Supervised Speaker Diarization.
clustering machine-learning speaker-diarization speaker-recognition supervised-clustering supervised-learning uis-rnn
Last synced: 14 May 2025
https://github.com/juanmc2005/diart
A python package to build AI-powered real-time audio applications
deep-learning real-time speaker-diarization speaker-embedding streaming-audio transcription voice-activity-detection
Last synced: 14 May 2025
https://github.com/FunAudioLLM/Fun-ASR
End-to-end speech recognition large model: 31 languages, dialects, accents, lyrics, hotwords, timestamps, speaker diarization. Trained on tens of millions of hours.
31-languages asr audio-language-model chinese-dialects fun-asr llm-asr multilingual-asr pytorch real-time-asr speaker-diarization speech-recognition speech-to-text transcription whisper-alternative
Last synced: 13 Jun 2026
https://github.com/modelscope/3d-speaker
A Repository for Single- and Multi-modal Speaker Verification, Speaker Recognition and Speaker Diarization
3d-speaker campplus cnceleb eres2net language-identification modelscope rdino speaker-diarization speaker-verification voxceleb
Last synced: 14 May 2025
https://github.com/wenet-e2e/wespeaker
Research and Production Oriented Speaker Verification, Recognition and Diarization Toolkit
asv campplus cnceleb dino ecapa-tdnn eres2net nist-sre plda production-ready pytorch redimnet repvgg resnet speaker-diarization speaker-recognition speaker-verification ssl voxceleb wavlm xvector
Last synced: 16 May 2025
https://github.com/transcriptionstream/transcriptionstream
turnkey self-hosted offline transcription and diarization service with llm summary
automation diarization llm mistral-7b ollama speaker-diarization speech-recognition transcription whisper whisperx
Last synced: 07 Apr 2025
https://github.com/soniqo/speech-swift
AI speech toolkit for Apple Silicon — ASR, TTS, speech-to-speech, VAD, and diarization powered by MLX and CoreML
apple-silicon asr coreml ios macos mlx neural-engine on-device speaker-diarization speech-enhancement speech-recognition speech-to-speech swift text-to-speech tts voice-activity-detection
Last synced: 24 May 2026
https://github.com/corvo007/MioSub
一站式全自动字幕生成软件,下载、转录、翻译、压制全流程覆盖,无需人工介入 / One-stop automated subtitle generator. Handles downloading, transcription, translation, and hardcoding—zero human intervention required.
alignment ass-subtitles captions diarization ffmpeg forced-alignment gemini-api gemini-subtitle-pro i18n speaker-diarization speech-to-text srt-subtitles substation-alpha subtitle-generator subtitle-translation subtitles subtitles-generator transcription whisper
Last synced: 13 Aug 2026
https://github.com/yinruiqing/pyannote-whisper
asr chatgpt meeting-summarization pyannote speaker-diarization whisper
Last synced: 04 Apr 2025
https://github.com/FluidInference/FluidAudio
Native Swift and CoreML SDK for local speaker diarization, VAD, and speech-to-text for real-time workloads. Works on iOS and macOS.
ane asr audio automatic-speech-recognition avfoundation coreml ios macos nvidia parakeet real-time speaker-diarization speaker-embedding speaker-identification speaker-recognition speech-to-text swift vad voice-activity-detection
Last synced: 31 Aug 2025
https://github.com/wq2012/spectralcluster
Python re-implementation of the (constrained) spectral clustering algorithms used in Google's speaker diarization papers.
auto-tune clustering constrained-clustering machine-learning python speaker-diarization spectral-clustering unsupervised-clustering unsupervised-learning
Last synced: 16 May 2025
https://github.com/revdotcom/reverb
Open source inference code for Rev's model
asr asr-model canary deeplearning diarization docker huggingface neural-network open-source opensource pyannote rev revai speaker-diarization speech-recognition speech-to-text speechrecognition wenet whisper
Last synced: 15 May 2025
https://github.com/manojpamk/pytorch_xvectors
Deep speaker embeddings in PyTorch, including x-vectors. Code used in this work: https://arxiv.org/abs/2007.16196
speaker-diarization speaker-embeddings speaker-recognition speaker-verification
Last synced: 26 Apr 2025
https://github.com/Frikallo/parakeet.cpp
Ultra fast and portable Parakeet implementation for on-device inference in C++ using Axiom with MPS+Unified Memory
asr automatic-speech-recognition axiom nvidia parakeet speaker-diarization speech speech-recognition speech-to-text
Last synced: 07 Jul 2026
https://github.com/IBM-Cloud/chatbot-watson-android
An Android ChatBot powered by Watson Services - Assistant, Speech-to-Text and Text-to-Speech on IBM Cloud.
android android-studio chatbot conversation conversation-service dialog entity ibm-cloud ibm-cloud-solutions ibm-watson ibm-watson-services intent java speaker-diarization speaker-recognition speech watson watson-services workspace
Last synced: 13 May 2025
https://github.com/agentem-ai/izwi
On-device Voice AI engine for transcription, TTS, and voice workflows.
asr audio-inference local-first openai-compatible-api self-hosted-ai speaker-diarization speech-to-text text-to-speech tts voice-cloning
Last synced: 10 Mar 2026
https://github.com/altunenes/parakeet-rs
very fast speech-to-text, diarization, streaming (even in CPU) with NVIDIA Parakeet in Rust
asr automatic-speech-recognition onnx parakeet speaker-diarization speaker-identification speech speech-recognition speech-to-text
Last synced: 06 Feb 2026
https://github.com/narcotic-sh/senko
Very fast, accurate speaker diarization
audio-ai diarization fbank pyannote rapids silero-vad speaker-diarization zanshin
Last synced: 02 Oct 2025
https://github.com/yufan-aslp/AliMeeting
The project is associated with the recently-launched ICASSP 2022 Multi-channel Multi-party Meeting Transcription Challenge (M2MeT) to provide participants with baseline systems for speech recognition and speaker diarization in conference scenario.
aishell-4 alimeeting asr challenge m2met multi-speaker-asr speaker-diarization
Last synced: 21 Jul 2025
https://github.com/kigner/audio.cpp-webui
audio.cpp with a full-task WebUI - pure C++ audio-model inference engine powered by ggml. TTS, ASR/STT, VAD, voice conversion, speaker diarization, music generation. No Python dependency.
asr audio cpp ggml inference-engine music-generation speaker-diarization speech-recognition speech-synthesis speech-to-text text-to-speech tts vad voice-conversion webui
Last synced: 25 Aug 2026
https://github.com/nezhar/speech-condenser
A tool for summarizing dialogues from videos or audio
asr speach-recognition speaker-diarization speaker-identification summarization
Last synced: 09 Jul 2025
https://github.com/vidyasagarmsc/watbot
An Android ChatBot powered by IBM Watson Services (Assistant V1, Text-to-Speech, and Speech-to-Text with Speaker Recognition) on IBM Cloud.
android android-studio assistant chatbot cognitive-services conversation conversation-service dialog entity ibm-cloud intent speaker-diarization speaker-labels speaker-recognition speech speech-to-text text-to-speech watson watson-assistant-service workspace
Last synced: 12 Aug 2025
https://github.com/clement-pages/gryannote
Provide Gradio custom components to make the diarization-based audio labeling process easier and faster.
annotation-processing annotation-tool audio gradio gradio-custom-component interspeech2024 pyannote speaker-diarization speech-processing
Last synced: 05 Apr 2025
https://github.com/wq2012/simpleder
A lightweight library to compute Diarization Error Rate (DER).
diarization machine-learning metrics speaker-diarization speech-processing speech-recognition
Last synced: 30 Aug 2025
https://github.com/narcotic-sh/zanshin
A novel media player that allows you to navigate by speaker
local macos media-player senko speaker-diarization visualization youtube
Last synced: 13 Oct 2025
https://github.com/juanmc2005/rttm-viewer
Application for viewing Rich Transcription Time Marked (RTTM) files in an interactive way
plotly rttm speaker-diarization visualization
Last synced: 11 Apr 2025
https://github.com/picovoice/falcon
On-device speaker diarization powered by deep learning
deep-learning diarization on-device speaker-diarization speaker-recognition
Last synced: 31 Mar 2025
https://github.com/nttcslab-sp/mamba-diarization
Official repository for Mamba-based Segmentation Model for Speaker Diarization
mamba-state-space-models pyannote speaker-diarization state-space-models
Last synced: 22 Jun 2025
https://github.com/ubclaunchpad/minutes
:telescope: Speaker diarization via transfer learning
library machine-learning python speaker-diarization speech transfer-learning ubc
Last synced: 15 May 2025
https://github.com/linto-ai/linto-diarization
Speaker diarization service
asr linto speaker-diarization speaker-identification
Last synced: 10 Jun 2025
https://github.com/MSKazemi/yazses
Free, open-source, fully-offline-by-default voice dictation for Linux (X11 & Wayland), macOS & Windows. Hold a key, speak, release — on-device faster-whisper types it into any app. Also transcribes recordings & captures meetings with speaker labels. No cloud, no account, no subscription.
accessibility assistive-technology dictation faster-whisper linux macos meeting-notes meeting-transcription offline privacy speaker-diarization speech-recognition speech-to-text transcription voice-commands voice-dictation voice-typing wayland whisper windows
Last synced: 01 Sep 2026
https://github.com/cadia-lvl/kaldi-speaker-diarization
This repository creates speaker diarization recipes to be used within the egs folder of kaldi.
ahc audio-files diarization icelandic kaldi mfccs plda speaker-diarization wav
Last synced: 11 Mar 2026
https://github.com/wq2012/vb_diarization
VB Diarization with Eigenvoice and HMM Priors, refactored
machine-learning speaker-diarization speech-processing speech-recognition
Last synced: 12 Apr 2025
https://github.com/Gr122lyBr/voicetag
Speaker identification powered by pyannote and resemblyzer
audio-transcription deep-learning deepgram diarization groq machine-learning nlp pyannote python resemblyzer speaker-diarization speaker-identification speaker-recognition speech-processing speech-to-text transcription voice-recognition whisper whisper-ai
Last synced: 03 Apr 2026
https://github.com/shashikg/x-vector-based-speaker-diarization
Course project for EE698R (2020-21 Sem 2). An X-Vector Based Speaker Diarization System with AutoEncoder based clustering method. Also supports spectral and KMeans clustering method.
deep-clustering deep-learning python pytroch speaker-diarization
Last synced: 18 Mar 2025
https://github.com/elmiraghorbani/gpt-speaker-diarization
Conversational Speaker Diarization using OpenAI AI Language Models(gpt-4) and OpenAI Whisper.
asr diarization gpt-4 openai speaker-diarization speech-recognition speech-to-text voice-activity-detection whisper youtube-dl
Last synced: 10 Oct 2025
https://github.com/juanmc2005/csda
Companion repository for the paper "Continual Self-supervised Domain Adaptation for End-to-end Speaker Diarization"
continual-learning domain-adaptation end-to-end pytorch pytorch-lightning self-supervised-learning speaker-diarization
Last synced: 11 Apr 2025
https://github.com/bunyaminergen/wavlmmsdd
This repository combines `WavLM`, a powerful speech representation model from Microsoft, with `MSDD` (Multi-Scale Diarization Decoder), a state-of-the-art approach for speaker diarization from Nvidia.
diarization embedding microsoft nvidia-nemo speaker-diarization speech speech-embedding wavlm
Last synced: 15 Aug 2025
https://github.com/luongndcoder/scribble
Ứng dụng ghi chú cuộc họp thông minh — ghi âm, phiên dịch realtime, realtime dịch đa ngôn ngữ và tạo biên bản tự động bằng AI. Smart meeting notes app — record, real-time transcription, multi-language cabin translation, and AI-powered meeting minutes generation.
ai-meeting-assistant ai-summarization cabin-translation cross-platform desktop-app fastapi llm meeting-minutes meeting-notes meeting-transcription nvidia-riva python react realtime-transcription speaker-diarization speech-to-text tauri tauri-app transcription typescript
Last synced: 28 May 2026
https://github.com/gorkemkaramolla/whisper-run
Faster Whisper with Speaker Diarization
distil-whisper faster-whisper openai pyannote speaker-diarization speech-recognition transcription whisper whisper-large
Last synced: 23 Oct 2025
https://github.com/scionoftech/speaker_diarization
speaker diarization using spectralcluster and Deeplearning
clustering speaker-diarization speech-recognition
Last synced: 14 Jun 2025
https://github.com/parva101/speaker_diarization_identification
A Streamlit web app for speaker diarization and identification in audio files. Upload or record audio, transcribe conversations, and automatically segment and label speakers using reference samples. This app makes it easy to analyze multi-speaker audio, export transcripts, and identify "who spoke when" for meetings, interviews, and more.
assembly-ai speaker-diarization speaker-identification speaker-verification speechbrain transcription voiceai
Last synced: 04 Apr 2026
https://github.com/mmxgn/smooth-convex-kl-nmf
Repository holding various implementation of specific NMF methods for speaker diarization
nmf nonnegative-matrix-factorization smoothness sparsity speaker-diarization
Last synced: 06 Apr 2025
https://github.com/maxhollmann/lium-diarization-editor
A very simple viewer/editor for LIUM speaker diarizations.
Last synced: 26 Jul 2025
https://github.com/neuralwork/audio2chat
Convert multi-speaker audio files to structured chat data for LLMs
chat llm llm-datasets speaker-diarization transcription whisper
Last synced: 04 Mar 2025
https://github.com/mikeesto/gemini-transcribe
Transcribe audio and video files with speaker diarization and logically grouped timestamps
gemini-flash speaker-diarization speech-to-text sveltekit transcription
Last synced: 27 Oct 2025
https://github.com/jpzinn654/speaker-diarization-portuguese
This project implements speaker diarization for Portuguese audio using WhisperX for transcription and PyAnotAudio's Speaker-Diarization 3.1 for speaker separation. It includes a Flask UI for easy file upload, transcription, and speaker identification.
flask gender-detection portuguese-language speaker-diarization speaker-recognition speech-recognition transcription whisper
Last synced: 26 Feb 2026
https://github.com/mafiatun/un-webcast-analyzer
AI-powered platform for analyzing UN WebTV sessions with automated transcription, speaker diarization, entity extraction (speakers, countries, SDGs), semantic search, and RAG-based chat interface. Built with Azure OpenAI, Cosmos DB, and Streamlit.
azure azureopenai entity-extraction international-relations rag semantic-search speaker-diarization streamlit un unwebtv
Last synced: 12 Apr 2026
https://github.com/martossien/transcria
Self-hosted meeting transcription portal — speech-to-text, speaker diarization, LLM-corrected transcripts, structured summaries and Word minutes, on your own GPUs. Flask + PostgreSQL, GDPR audit trail, distributed GPU topologies
asr flask gdrp gpu llm meeting-minutes on-premise postgresql pyannote self-hosted speaker-diarization speech-to-text srt transcription whisper
Last synced: 27 Jun 2026
https://github.com/nicknaskida/insanely-fast-whisper
Incredibly fast Whisper-large-v3 with speaker diarization
diarization speaker-diarization transfromers whisper whisper-ai whisper-faster whisper-large
Last synced: 29 Sep 2025
https://github.com/ztxtech/meeting-auto-summary
Local meeting audio/video transcription skill with speaker diarization, subtitles, summaries, reports, and optional translation.
ai-skills claude-code codex local-ai meeting-summary mlx mlx-audio opencode qwen3-asr speaker-diarization speech-to-text subtitles
Last synced: 23 May 2026
https://github.com/zhima-mochi/whisper-v3-server
A robust backend server for audio processing, delivering high-accuracy transcription and speaker diarization. Powered by Whisper for speech-to-text and Pyannote for speaker segmentation, wrapped in a clean, maintainable architecture based on Domain-Driven Design (DDD) and Hexagonal Architecture.
audio-processing domain-driven-design fastapi ports-and-adapters-architecture pyannote speaker-diarization speech-recognition speech-to-text whisper
Last synced: 13 Apr 2026
https://github.com/aeronjl/transcribe
Python package for accurate audio transcription with speaker diarisation
audio-transcription gpt speaker-diarization whisper
Last synced: 04 Jul 2025
https://github.com/theseraphim/scribe-forge-ai
🎵 Complete offline audio transcription system with speaker diarization using OpenAI Whisper and PyAnnote. Features automatic audio cleaning, precise timestamps, multiple output formats (JSON/TXT/Markdown), and support for 20+ audio formats. No external APIs required - works entirely offline.
audio-analysis audio-cleaning audio-processing audio-transcription diarization ffmpeg huggingface machine-learning multi-speaker nlp offline-transcription openai-whisper pyannote python speaker-diarization speech-recognition speech-to-text timestamps transcription-tool whisper
Last synced: 05 May 2026
https://github.com/strcoder4007/voice-sentiment-analysis
Customer call analysis app with: Speech-to-Text + speaker diarization via ElevenLabs Structured conversation analysis via OpenAI React frontend for multi-file upload and rich results display. Takes into account not only the text but tone, rythym etc.
call-analysis speaker-diarization speech-to-text
Last synced: 14 Jul 2026
https://github.com/bochengyang/mac-mlx-meeting-minutes
Local macOS meeting minutes: record mic + system audio and transcribe on-device with MLX Whisper
apple-silicon macos meeting-minutes mlx on-device privacy speaker-diarization transcription whisper
Last synced: 17 Jul 2026
https://github.com/ogwata/kikoyu
町内会用:DGX Spark上でローカル文字起こし・話者分離(whisper-large-v3 + pyannote)
dgx-spark pyannote speaker-diarization whisper whisperx
Last synced: 17 Jul 2026
https://github.com/aidayang/funasr-oneclick
FunASR实时语音识别版,识别麦克风和电脑内播放的声音,电脑语音打字软件
audio-visual-speech-recognition conformer dfsmn funasr paraformer pretrained-models punctuation pytorch rnnt speaker-diarization speech-recognition speechgpt speechllm vad voice-activity-detection whisper
Last synced: 24 Oct 2025
https://github.com/davidamacey/opentranscribe
Self-hosted AI-powered transcription platform with speaker diarization, search, and collaboration features. Built with Svelte, FastAPI, and Docker for easy deployment.
ai audio-processing docker fastapi machine-learning nlp open-source self-hosted speaker-diarization speech-to-text svelte transcription video-transcription whisper
Last synced: 17 Jan 2026
https://github.com/katagaki/firesidesubtitles
Video transcription, speaker diarization, and face detection in Python.
audio dnn face-detection openai openai-whisper opencv python speaker-diarization transcription video
Last synced: 15 Apr 2026
https://github.com/nicknaskida/cog-whisper-diarization
Cog implementation of transcribing + diarization pipeline with Whisper & Pyannote
diarization openai-whisper pyannote replicate speaker-diarization whisper whisper-faster whisperx
Last synced: 01 Oct 2025
https://github.com/xdcobra/react-native-sherpa-onnx
Offline Speech Processing SDK for React Native using sherpa-onnx. Supports Speech-to-Text, Text-to-Speech, Speaker Diarization, Speech Enhancement, Source Separation & VAD.
android ctc data-privacy funasr ios paraformer react-native sdk sensevoice sherpa-onnx source-separation speaker-diarization speech-enhancement speech-to-text text-to-speech vad voice-activity-detection whisper zipformer
Last synced: 02 Apr 2026
https://github.com/flo-bit/youtube-speaker-separation
simple python script that outputs separate audio files for each speaker in a youtube video, using whisper on replicate
speaker-diarization speech-to-text text-to-speech voice-cloning whisper youtube
Last synced: 11 Feb 2026
https://github.com/konhi/elevenlabs-speech-to-text-api-ui
elevenlabs speaker-diarization speech-to-text transcription
Last synced: 15 Feb 2026
https://github.com/mathusanm6/amaze-voice-lab
The goal of this research project was to be able to control the movements of characters in a Maze game using real-time voice commands such as saying out loud Up, Down, Left or Right.
asr automatic-speech-recognition game java maze research speaker-diarization speaker-recognition voice-recognition
Last synced: 03 Aug 2025
https://github.com/cervantesvive/vidnotes-tools
Local-first CLI pipeline that turns recorded videos into speaker-labeled transcripts, summaries, and action items using WhisperX and Ollama.
cli llm local-first meeting-notes ollama python speaker-diarization transcription video whisperx
Last synced: 16 Aug 2026
https://github.com/collectiveai-team/coro
OpenAI-compatible ASR + speaker-diarization HTTP server with pluggable backends (Faster-Whisper, onnx-asr Parakeet, onnx-genai Nemotron) and NeMo Sortformer diarization.
asr fastapi faster-whisper nemo onnxruntime openai-api parakeet python sortformer speaker-diarization speech-recognition speech-to-text transcription whisper
Last synced: 22 Jul 2026
https://github.com/yinruiqing/annotation_generator
annotation generator for diarization task
e2e-speaker-diarization speaker-diarization target-speaker-vad
Last synced: 06 Apr 2025
https://github.com/aeronjl/transcribe-streamlit
Streamlit user interface for transcribing conversations with speaker diarisation
audio-transcription speaker-diarization streamlit
Last synced: 19 Apr 2026
https://github.com/nlink-jp/gem-transcribe
Audio transcription CLI built on Vertex AI Gemini — speaker name inference, multi-language output, structured JSON
audio-transcription cli gemini multilingual python speaker-diarization srt subtitles vertex-ai webvtt
Last synced: 04 Jun 2026
https://github.com/navopw/runpod-worker-whisper-diarization
RunPod serverless worker for audio transcription with speaker diarization using Whisper and pyannote.
asr runpod speaker-diarization speech-to-text transcription whisper
Last synced: 25 Jul 2026
https://github.com/zxkane/audio-transcriber-funasr
Agent skill for multi-speaker meeting & podcast transcription with FunASR speaker diarization and LLM cleanup. Supports 99 languages (zh/en/ja/ko/yue + Whisper). GPU & CPU. Packaged as a Claude Code plugin.
agent-skill asr chinese-asr claude-code claude-code-skill funasr interview-transcription meeting-transcription multilingual paraformer podcast-transcription skills-sh speaker-diarization speech-to-text whisper
Last synced: 01 May 2026
https://github.com/koradripless624/un-webcast-analyzer
📊 Transform UN WebTV sessions into structured insights with AI-driven transcription, entity extraction, and interactive analytics for research-ready knowledge.
azure azureopenai entity-extraction international-relations rag semantic-search speaker-diarization streamlit unwebtv
Last synced: 01 May 2026
https://github.com/jithinolickal/meeting-recorder
macOS menu bar app that records meetings and generates transcripts with speaker identification - fully local via Whisper (MLX), no cloud or API costs
blackhole macos menu-bar-app mlx python speaker-diarization transcription whisper
Last synced: 16 Aug 2026
https://github.com/ekhodzitsky/polyvoice
Speaker diarization for Rust — who spoke when, without Python. Silero VAD + WeSpeaker + AHC in a single Pipeline::run() call.
audio diarization machine-learning onnx python-bindings rust speaker-diarization speech vad voice
Last synced: 16 May 2026
https://github.com/lenik/wav2chat
Convert phone/meeting audio into speaker-segmented chat transcripts
asr funasr python speaker-diarization transcription
Last synced: 13 Jun 2026
https://github.com/vipul-sharma20/diarization-service
HTTP service for pyannote speaker diarization
Last synced: 03 Aug 2026
https://github.com/mtwn105/audio-intel
AudioIntel - Audio/Video Intelligence, Transcripts, Summary, and much more
ai assemblyai audio audio-processing diarization lemur sonet speaker-diarization speaker-recognition speech-recognition speech-to-text transcript
Last synced: 04 Apr 2025
https://github.com/elien666/diarize
On-device speaker diarization and transcription for macOS — CLI, SwiftUI app, and Swift library powered by FluidAudio and GRDB.
audio cli coreml fluidaudio grdb macos speaker-diarization swift swiftui transcription
Last synced: 26 Jun 2026
https://github.com/biyachuev/yt-transcriber
AI-powered audio/video processing: transcription, speaker diarization, LLM refinement, translation | AI-обработка аудио/видео: транскрибация, распознавание спикеров, перевод
ai-pipeline audio-intelligence audio-processing batch-processing document-generation llm media-processing nllb nlp ollama openai-api speaker-diarization speech-to-text translation video-transcription voice-activity-detection whisper youtube-downloader
Last synced: 16 May 2026
https://github.com/ericrihm/yt-whisper
Fast, local YouTube transcription with speaker diarization and a keyboard-first Textual TUI. YouTube-subs fast path, faster-whisper on CUDA, opt-in pyannote diarization, prompt profile auto-detection.
cli cuda faster-whisper pyannote python speaker-diarization textual transcription tui whisper youtube
Last synced: 31 May 2026
https://github.com/werserk/techstormhack-1st-place
Решение соревнования ТехШторм от корпорации ТатНефть по анализу активности членов команды на ВКС
pyannote speaker-diarization speech-recognition streamlit whisper
Last synced: 19 Feb 2026