An open API service indexing awesome lists of open source software.

Projects in Awesome Lists tagged with speaker-diarization

A curated list of projects in awesome lists tagged with speaker-diarization .

https://github.com/modelscope/funasr

A Fundamental End-to-End Speech Recognition Toolkit and Open Source SOTA Pretrained Models, Supporting Speech Recognition, Voice Activity Detection, Text Post-processing etc.

audio-visual-speech-recognition conformer dfsmn paraformer pretrained-model punctuation pytorch rnnt speaker-diarization speech-recognition speechgpt speechllm vad voice-activity-detection whisper

Last synced: 16 May 2025

https://github.com/pyannote/pyannote-audio

Neural building blocks for speaker diarization: speech activity detection, speaker change detection, overlapped speech detection, speaker embedding

overlapped-speech-detection pretrained-models pytorch speaker-change-detection speaker-diarization speaker-embedding speaker-recognition speaker-verification speech-activity-detection speech-processing voice-activity-detection

Last synced: 13 May 2025

https://github.com/modelscope/FunASR

A Fundamental End-to-End Speech Recognition Toolkit and Open Source SOTA Pretrained Models, Supporting Speech Recognition, Voice Activity Detection, Text Post-processing etc.

audio-visual-speech-recognition conformer dfsmn paraformer pretrained-model punctuation pytorch rnnt speaker-diarization speech-recognition speechgpt speechllm vad voice-activity-detection whisper

Last synced: 24 Mar 2025

https://github.com/mahmoudashraf97/whisper-diarization

Automatic Speech Recognition with Speaker Diarization based on OpenAI Whisper

asr speaker-diarization speech speech-recognition speech-to-text whisper

Last synced: 13 May 2025

https://github.com/MahmoudAshraf97/whisper-diarization

Automatic Speech Recognition with Speaker Diarization based on OpenAI Whisper

asr speaker-diarization speech speech-recognition speech-to-text whisper

Last synced: 28 Mar 2025

https://github.com/google/uis-rnn

This is the library for the Unbounded Interleaved-State Recurrent Neural Network (UIS-RNN) algorithm, corresponding to the paper Fully Supervised Speaker Diarization.

clustering machine-learning speaker-diarization speaker-recognition supervised-clustering supervised-learning uis-rnn

Last synced: 14 May 2025

https://github.com/juanmc2005/diart

A python package to build AI-powered real-time audio applications

deep-learning real-time speaker-diarization speaker-embedding streaming-audio transcription voice-activity-detection

Last synced: 14 May 2025

https://github.com/FunAudioLLM/Fun-ASR

End-to-end speech recognition large model: 31 languages, dialects, accents, lyrics, hotwords, timestamps, speaker diarization. Trained on tens of millions of hours.

31-languages asr audio-language-model chinese-dialects fun-asr llm-asr multilingual-asr pytorch real-time-asr speaker-diarization speech-recognition speech-to-text transcription whisper-alternative

Last synced: 13 Jun 2026

https://github.com/modelscope/3d-speaker

A Repository for Single- and Multi-modal Speaker Verification, Speaker Recognition and Speaker Diarization

3d-speaker campplus cnceleb eres2net language-identification modelscope rdino speaker-diarization speaker-verification voxceleb

Last synced: 14 May 2025

https://github.com/transcriptionstream/transcriptionstream

turnkey self-hosted offline transcription and diarization service with llm summary

automation diarization llm mistral-7b ollama speaker-diarization speech-recognition transcription whisper whisperx

Last synced: 07 Apr 2025

https://github.com/soniqo/speech-swift

AI speech toolkit for Apple Silicon — ASR, TTS, speech-to-speech, VAD, and diarization powered by MLX and CoreML

apple-silicon asr coreml ios macos mlx neural-engine on-device speaker-diarization speech-enhancement speech-recognition speech-to-speech swift text-to-speech tts voice-activity-detection

Last synced: 24 May 2026

https://github.com/corvo007/MioSub

一站式全自动字幕生成软件,下载、转录、翻译、压制全流程覆盖,无需人工介入 / One-stop automated subtitle generator. Handles downloading, transcription, translation, and hardcoding—zero human intervention required.

alignment ass-subtitles captions diarization ffmpeg forced-alignment gemini-api gemini-subtitle-pro i18n speaker-diarization speech-to-text srt-subtitles substation-alpha subtitle-generator subtitle-translation subtitles subtitles-generator transcription whisper

Last synced: 13 Aug 2026

https://github.com/FluidInference/FluidAudio

Native Swift and CoreML SDK for local speaker diarization, VAD, and speech-to-text for real-time workloads. Works on iOS and macOS.

ane asr audio automatic-speech-recognition avfoundation coreml ios macos nvidia parakeet real-time speaker-diarization speaker-embedding speaker-identification speaker-recognition speech-to-text swift vad voice-activity-detection

Last synced: 31 Aug 2025

https://github.com/wq2012/spectralcluster

Python re-implementation of the (constrained) spectral clustering algorithms used in Google's speaker diarization papers.

auto-tune clustering constrained-clustering machine-learning python speaker-diarization spectral-clustering unsupervised-clustering unsupervised-learning

Last synced: 16 May 2025

https://github.com/manojpamk/pytorch_xvectors

Deep speaker embeddings in PyTorch, including x-vectors. Code used in this work: https://arxiv.org/abs/2007.16196

speaker-diarization speaker-embeddings speaker-recognition speaker-verification

Last synced: 26 Apr 2025

https://github.com/Frikallo/parakeet.cpp

Ultra fast and portable Parakeet implementation for on-device inference in C++ using Axiom with MPS+Unified Memory

asr automatic-speech-recognition axiom nvidia parakeet speaker-diarization speech speech-recognition speech-to-text

Last synced: 07 Jul 2026

https://github.com/altunenes/parakeet-rs

very fast speech-to-text, diarization, streaming (even in CPU) with NVIDIA Parakeet in Rust

asr automatic-speech-recognition onnx parakeet speaker-diarization speaker-identification speech speech-recognition speech-to-text

Last synced: 06 Feb 2026

https://github.com/narcotic-sh/senko

Very fast, accurate speaker diarization

audio-ai diarization fbank pyannote rapids silero-vad speaker-diarization zanshin

Last synced: 02 Oct 2025

https://github.com/yufan-aslp/AliMeeting

The project is associated with the recently-launched ICASSP 2022 Multi-channel Multi-party Meeting Transcription Challenge (M2MeT) to provide participants with baseline systems for speech recognition and speaker diarization in conference scenario.

aishell-4 alimeeting asr challenge m2met multi-speaker-asr speaker-diarization

Last synced: 21 Jul 2025

https://github.com/kigner/audio.cpp-webui

audio.cpp with a full-task WebUI - pure C++ audio-model inference engine powered by ggml. TTS, ASR/STT, VAD, voice conversion, speaker diarization, music generation. No Python dependency.

asr audio cpp ggml inference-engine music-generation speaker-diarization speech-recognition speech-synthesis speech-to-text text-to-speech tts vad voice-conversion webui

Last synced: 25 Aug 2026

https://github.com/nezhar/speech-condenser

A tool for summarizing dialogues from videos or audio

asr speach-recognition speaker-diarization speaker-identification summarization

Last synced: 09 Jul 2025

https://github.com/vidyasagarmsc/watbot

An Android ChatBot powered by IBM Watson Services (Assistant V1, Text-to-Speech, and Speech-to-Text with Speaker Recognition) on IBM Cloud.

android android-studio assistant chatbot cognitive-services conversation conversation-service dialog entity ibm-cloud intent speaker-diarization speaker-labels speaker-recognition speech speech-to-text text-to-speech watson watson-assistant-service workspace

Last synced: 12 Aug 2025

https://github.com/clement-pages/gryannote

Provide Gradio custom components to make the diarization-based audio labeling process easier and faster.

annotation-processing annotation-tool audio gradio gradio-custom-component interspeech2024 pyannote speaker-diarization speech-processing

Last synced: 05 Apr 2025

https://github.com/wq2012/simpleder

A lightweight library to compute Diarization Error Rate (DER).

diarization machine-learning metrics speaker-diarization speech-processing speech-recognition

Last synced: 30 Aug 2025

https://github.com/narcotic-sh/zanshin

A novel media player that allows you to navigate by speaker

local macos media-player senko speaker-diarization visualization youtube

Last synced: 13 Oct 2025

https://github.com/juanmc2005/rttm-viewer

Application for viewing Rich Transcription Time Marked (RTTM) files in an interactive way

plotly rttm speaker-diarization visualization

Last synced: 11 Apr 2025

https://github.com/picovoice/falcon

On-device speaker diarization powered by deep learning

deep-learning diarization on-device speaker-diarization speaker-recognition

Last synced: 31 Mar 2025

https://github.com/nttcslab-sp/mamba-diarization

Official repository for Mamba-based Segmentation Model for Speaker Diarization

mamba-state-space-models pyannote speaker-diarization state-space-models

Last synced: 22 Jun 2025

https://github.com/ubclaunchpad/minutes

:telescope: Speaker diarization via transfer learning

library machine-learning python speaker-diarization speech transfer-learning ubc

Last synced: 15 May 2025

https://github.com/MSKazemi/yazses

Free, open-source, fully-offline-by-default voice dictation for Linux (X11 & Wayland), macOS & Windows. Hold a key, speak, release — on-device faster-whisper types it into any app. Also transcribes recordings & captures meetings with speaker labels. No cloud, no account, no subscription.

accessibility assistive-technology dictation faster-whisper linux macos meeting-notes meeting-transcription offline privacy speaker-diarization speech-recognition speech-to-text transcription voice-commands voice-dictation voice-typing wayland whisper windows

Last synced: 01 Sep 2026

https://github.com/cadia-lvl/kaldi-speaker-diarization

This repository creates speaker diarization recipes to be used within the egs folder of kaldi.

ahc audio-files diarization icelandic kaldi mfccs plda speaker-diarization wav

Last synced: 11 Mar 2026

https://github.com/wq2012/vb_diarization

VB Diarization with Eigenvoice and HMM Priors, refactored

machine-learning speaker-diarization speech-processing speech-recognition

Last synced: 12 Apr 2025

https://github.com/shashikg/x-vector-based-speaker-diarization

Course project for EE698R (2020-21 Sem 2). An X-Vector Based Speaker Diarization System with AutoEncoder based clustering method. Also supports spectral and KMeans clustering method.

deep-clustering deep-learning python pytroch speaker-diarization

Last synced: 18 Mar 2025

https://github.com/elmiraghorbani/gpt-speaker-diarization

Conversational Speaker Diarization using OpenAI AI Language Models(gpt-4) and OpenAI Whisper.

asr diarization gpt-4 openai speaker-diarization speech-recognition speech-to-text voice-activity-detection whisper youtube-dl

Last synced: 10 Oct 2025

https://github.com/juanmc2005/csda

Companion repository for the paper "Continual Self-supervised Domain Adaptation for End-to-end Speaker Diarization"

continual-learning domain-adaptation end-to-end pytorch pytorch-lightning self-supervised-learning speaker-diarization

Last synced: 11 Apr 2025

https://github.com/bunyaminergen/wavlmmsdd

This repository combines `WavLM`, a powerful speech representation model from Microsoft, with `MSDD` (Multi-Scale Diarization Decoder), a state-of-the-art approach for speaker diarization from Nvidia.

diarization embedding microsoft nvidia-nemo speaker-diarization speech speech-embedding wavlm

Last synced: 15 Aug 2025

https://github.com/luongndcoder/scribble

Ứng dụng ghi chú cuộc họp thông minh — ghi âm, phiên dịch realtime, realtime dịch đa ngôn ngữ và tạo biên bản tự động bằng AI. Smart meeting notes app — record, real-time transcription, multi-language cabin translation, and AI-powered meeting minutes generation.

ai-meeting-assistant ai-summarization cabin-translation cross-platform desktop-app fastapi llm meeting-minutes meeting-notes meeting-transcription nvidia-riva python react realtime-transcription speaker-diarization speech-to-text tauri tauri-app transcription typescript

Last synced: 28 May 2026

https://github.com/scionoftech/speaker_diarization

speaker diarization using spectralcluster and Deeplearning

clustering speaker-diarization speech-recognition

Last synced: 14 Jun 2025

https://github.com/parva101/speaker_diarization_identification

A Streamlit web app for speaker diarization and identification in audio files. Upload or record audio, transcribe conversations, and automatically segment and label speakers using reference samples. This app makes it easy to analyze multi-speaker audio, export transcripts, and identify "who spoke when" for meetings, interviews, and more.

assembly-ai speaker-diarization speaker-identification speaker-verification speechbrain transcription voiceai

Last synced: 04 Apr 2026

https://github.com/mmxgn/smooth-convex-kl-nmf

Repository holding various implementation of specific NMF methods for speaker diarization

nmf nonnegative-matrix-factorization smoothness sparsity speaker-diarization

Last synced: 06 Apr 2025

https://github.com/maxhollmann/lium-diarization-editor

A very simple viewer/editor for LIUM speaker diarizations.

lium speaker-diarization

Last synced: 26 Jul 2025

https://github.com/neuralwork/audio2chat

Convert multi-speaker audio files to structured chat data for LLMs

chat llm llm-datasets speaker-diarization transcription whisper

Last synced: 04 Mar 2025

https://github.com/mikeesto/gemini-transcribe

Transcribe audio and video files with speaker diarization and logically grouped timestamps

gemini-flash speaker-diarization speech-to-text sveltekit transcription

Last synced: 27 Oct 2025

https://github.com/jpzinn654/speaker-diarization-portuguese

This project implements speaker diarization for Portuguese audio using WhisperX for transcription and PyAnotAudio's Speaker-Diarization 3.1 for speaker separation. It includes a Flask UI for easy file upload, transcription, and speaker identification.

flask gender-detection portuguese-language speaker-diarization speaker-recognition speech-recognition transcription whisper

Last synced: 26 Feb 2026

https://github.com/mafiatun/un-webcast-analyzer

AI-powered platform for analyzing UN WebTV sessions with automated transcription, speaker diarization, entity extraction (speakers, countries, SDGs), semantic search, and RAG-based chat interface. Built with Azure OpenAI, Cosmos DB, and Streamlit.

azure azureopenai entity-extraction international-relations rag semantic-search speaker-diarization streamlit un unwebtv

Last synced: 12 Apr 2026

https://github.com/martossien/transcria

Self-hosted meeting transcription portal — speech-to-text, speaker diarization, LLM-corrected transcripts, structured summaries and Word minutes, on your own GPUs. Flask + PostgreSQL, GDPR audit trail, distributed GPU topologies

asr flask gdrp gpu llm meeting-minutes on-premise postgresql pyannote self-hosted speaker-diarization speech-to-text srt transcription whisper

Last synced: 27 Jun 2026

https://github.com/ztxtech/meeting-auto-summary

Local meeting audio/video transcription skill with speaker diarization, subtitles, summaries, reports, and optional translation.

ai-skills claude-code codex local-ai meeting-summary mlx mlx-audio opencode qwen3-asr speaker-diarization speech-to-text subtitles

Last synced: 23 May 2026

https://github.com/zhima-mochi/whisper-v3-server

A robust backend server for audio processing, delivering high-accuracy transcription and speaker diarization. Powered by Whisper for speech-to-text and Pyannote for speaker segmentation, wrapped in a clean, maintainable architecture based on Domain-Driven Design (DDD) and Hexagonal Architecture.

audio-processing domain-driven-design fastapi ports-and-adapters-architecture pyannote speaker-diarization speech-recognition speech-to-text whisper

Last synced: 13 Apr 2026

https://github.com/aeronjl/transcribe

Python package for accurate audio transcription with speaker diarisation

audio-transcription gpt speaker-diarization whisper

Last synced: 04 Jul 2025

https://github.com/nikitalpopov/master

research for master degree

python3 speaker-diarization

Last synced: 18 Apr 2026

https://github.com/theseraphim/scribe-forge-ai

🎵 Complete offline audio transcription system with speaker diarization using OpenAI Whisper and PyAnnote. Features automatic audio cleaning, precise timestamps, multiple output formats (JSON/TXT/Markdown), and support for 20+ audio formats. No external APIs required - works entirely offline.

audio-analysis audio-cleaning audio-processing audio-transcription diarization ffmpeg huggingface machine-learning multi-speaker nlp offline-transcription openai-whisper pyannote python speaker-diarization speech-recognition speech-to-text timestamps transcription-tool whisper

Last synced: 05 May 2026

https://github.com/strcoder4007/voice-sentiment-analysis

Customer call analysis app with: Speech-to-Text + speaker diarization via ElevenLabs Structured conversation analysis via OpenAI React frontend for multi-file upload and rich results display. Takes into account not only the text but tone, rythym etc.

call-analysis speaker-diarization speech-to-text

Last synced: 14 Jul 2026

https://github.com/bochengyang/mac-mlx-meeting-minutes

Local macOS meeting minutes: record mic + system audio and transcribe on-device with MLX Whisper

apple-silicon macos meeting-minutes mlx on-device privacy speaker-diarization transcription whisper

Last synced: 17 Jul 2026

https://github.com/ogwata/kikoyu

町内会用:DGX Spark上でローカル文字起こし・話者分離(whisper-large-v3 + pyannote)

dgx-spark pyannote speaker-diarization whisper whisperx

Last synced: 17 Jul 2026

https://github.com/davidamacey/opentranscribe

Self-hosted AI-powered transcription platform with speaker diarization, search, and collaboration features. Built with Svelte, FastAPI, and Docker for easy deployment.

ai audio-processing docker fastapi machine-learning nlp open-source self-hosted speaker-diarization speech-to-text svelte transcription video-transcription whisper

Last synced: 17 Jan 2026

https://github.com/katagaki/firesidesubtitles

Video transcription, speaker diarization, and face detection in Python.

audio dnn face-detection openai openai-whisper opencv python speaker-diarization transcription video

Last synced: 15 Apr 2026

https://github.com/nicknaskida/cog-whisper-diarization

Cog implementation of transcribing + diarization pipeline with Whisper & Pyannote

diarization openai-whisper pyannote replicate speaker-diarization whisper whisper-faster whisperx

Last synced: 01 Oct 2025

https://github.com/xdcobra/react-native-sherpa-onnx

Offline Speech Processing SDK for React Native using sherpa-onnx. Supports Speech-to-Text, Text-to-Speech, Speaker Diarization, Speech Enhancement, Source Separation & VAD.

android ctc data-privacy funasr ios paraformer react-native sdk sensevoice sherpa-onnx source-separation speaker-diarization speech-enhancement speech-to-text text-to-speech vad voice-activity-detection whisper zipformer

Last synced: 02 Apr 2026

https://github.com/flo-bit/youtube-speaker-separation

simple python script that outputs separate audio files for each speaker in a youtube video, using whisper on replicate

speaker-diarization speech-to-text text-to-speech voice-cloning whisper youtube

Last synced: 11 Feb 2026

https://github.com/mathusanm6/amaze-voice-lab

The goal of this research project was to be able to control the movements of characters in a Maze game using real-time voice commands such as saying out loud Up, Down, Left or Right.

asr automatic-speech-recognition game java maze research speaker-diarization speaker-recognition voice-recognition

Last synced: 03 Aug 2025

https://github.com/cervantesvive/vidnotes-tools

Local-first CLI pipeline that turns recorded videos into speaker-labeled transcripts, summaries, and action items using WhisperX and Ollama.

cli llm local-first meeting-notes ollama python speaker-diarization transcription video whisperx

Last synced: 16 Aug 2026

https://github.com/collectiveai-team/coro

OpenAI-compatible ASR + speaker-diarization HTTP server with pluggable backends (Faster-Whisper, onnx-asr Parakeet, onnx-genai Nemotron) and NeMo Sortformer diarization.

asr fastapi faster-whisper nemo onnxruntime openai-api parakeet python sortformer speaker-diarization speech-recognition speech-to-text transcription whisper

Last synced: 22 Jul 2026

https://github.com/aeronjl/transcribe-streamlit

Streamlit user interface for transcribing conversations with speaker diarisation

audio-transcription speaker-diarization streamlit

Last synced: 19 Apr 2026

https://github.com/nlink-jp/gem-transcribe

Audio transcription CLI built on Vertex AI Gemini — speaker name inference, multi-language output, structured JSON

audio-transcription cli gemini multilingual python speaker-diarization srt subtitles vertex-ai webvtt

Last synced: 04 Jun 2026

https://github.com/navopw/runpod-worker-whisper-diarization

RunPod serverless worker for audio transcription with speaker diarization using Whisper and pyannote.

asr runpod speaker-diarization speech-to-text transcription whisper

Last synced: 25 Jul 2026

https://github.com/zxkane/audio-transcriber-funasr

Agent skill for multi-speaker meeting & podcast transcription with FunASR speaker diarization and LLM cleanup. Supports 99 languages (zh/en/ja/ko/yue + Whisper). GPU & CPU. Packaged as a Claude Code plugin.

agent-skill asr chinese-asr claude-code claude-code-skill funasr interview-transcription meeting-transcription multilingual paraformer podcast-transcription skills-sh speaker-diarization speech-to-text whisper

Last synced: 01 May 2026

https://github.com/koradripless624/un-webcast-analyzer

📊 Transform UN WebTV sessions into structured insights with AI-driven transcription, entity extraction, and interactive analytics for research-ready knowledge.

azure azureopenai entity-extraction international-relations rag semantic-search speaker-diarization streamlit unwebtv

Last synced: 01 May 2026

https://github.com/jithinolickal/meeting-recorder

macOS menu bar app that records meetings and generates transcripts with speaker identification - fully local via Whisper (MLX), no cloud or API costs

blackhole macos menu-bar-app mlx python speaker-diarization transcription whisper

Last synced: 16 Aug 2026

https://github.com/ekhodzitsky/polyvoice

Speaker diarization for Rust — who spoke when, without Python. Silero VAD + WeSpeaker + AHC in a single Pipeline::run() call.

audio diarization machine-learning onnx python-bindings rust speaker-diarization speech vad voice

Last synced: 16 May 2026

https://github.com/lenik/wav2chat

Convert phone/meeting audio into speaker-segmented chat transcripts

asr funasr python speaker-diarization transcription

Last synced: 13 Jun 2026

https://github.com/vipul-sharma20/diarization-service

HTTP service for pyannote speaker diarization

speaker-diarization

Last synced: 03 Aug 2026

https://github.com/elien666/diarize

On-device speaker diarization and transcription for macOS — CLI, SwiftUI app, and Swift library powered by FluidAudio and GRDB.

audio cli coreml fluidaudio grdb macos speaker-diarization swift swiftui transcription

Last synced: 26 Jun 2026

https://github.com/biyachuev/yt-transcriber

AI-powered audio/video processing: transcription, speaker diarization, LLM refinement, translation | AI-обработка аудио/видео: транскрибация, распознавание спикеров, перевод

ai-pipeline audio-intelligence audio-processing batch-processing document-generation llm media-processing nllb nlp ollama openai-api speaker-diarization speech-to-text translation video-transcription voice-activity-detection whisper youtube-downloader

Last synced: 16 May 2026

https://github.com/ericrihm/yt-whisper

Fast, local YouTube transcription with speaker diarization and a keyboard-first Textual TUI. YouTube-subs fast path, faster-whisper on CUDA, opt-in pyannote diarization, prompt profile auto-detection.

cli cuda faster-whisper pyannote python speaker-diarization textual transcription tui whisper youtube

Last synced: 31 May 2026

https://github.com/werserk/techstormhack-1st-place

Решение соревнования ТехШторм от корпорации ТатНефть по анализу активности членов команды на ВКС

pyannote speaker-diarization speech-recognition streamlit whisper

Last synced: 19 Feb 2026