Ecosyste.ms: Awesome

An open API service indexing awesome lists of open source software.

Awesome Lists | Featured Topics | Projects

Whisper

Whisper is an autoregressive language model developed by OpenAI. It is trained on a large corpus of text using a transformer architecture and is capable of generating high-quality natural language text. Whisper can be used for tasks such as language modeling, text completion, and text generation. It has shown impressive performance on various benchmarks and has been released by OpenAI to encourage research in the field of language modeling. Whisper is not yet available for public use, but it has the potential to transform the field of natural language processing and generate new opportunities for language-based applications.

GitHub: https://github.com/topics/whisper
Repo: https://github.com/openai/whisper
Created by: OpenAI
Released: August 2021
Related Topics: machine-learning, artificial-intelligence, language-modeling,
Last updated: 2025-02-01 00:33:26 UTC
JSON Representation

https://github.com/fer14/videoseek

Intelligent video search tool powered by AI

bert timestamp video whisper youtube-api

Last synced: 14 Jan 2025

https://github.com/JoSuru/speeka

Speeaka is an open-source project that uses the Whisper model of OpenAI to transcribe audio into text. Its intuitive web interface makes it easy to use. Contributions are welcome.

open-source python python3 speech-to-text streamlit whisper

Last synced: 24 Oct 2024

https://github.com/datarabbit-ai/transcription_service

System/service with REST API for extracting text transcriptions from movies and audio recordings in most popular video formats.

containers datarabbit rest-api speech-to-text stt transcription transcription-services whisper

Last synced: 09 Oct 2024

https://github.com/limdongjin/ignkafasr

Real-Time In-memory Speaker Verification and Speech Recognition Project using apache ignite, apache kafka, speechbrain, whisper, stomp, spring webflux, kubernetes(k8s)

apache-ignite apache-kafka asr audio-recorder google-kubernetes-engine k8s kubernetes speaker-recognition speaker-verification speech-recognition speechbrain springframework stomp stompwebsocket webflux whisper

Last synced: 24 Oct 2024

https://github.com/imsanjoykb/speech-nlp-bootcamp

Speech NLP Bootcamp

asr audio-analysis audio-applications bangla-nlp huggingface-transformers seq2seq speech speech-recognition tts wav2vec2 whisper

Last synced: 18 Jan 2025

https://github.com/astrologos/py-speakeasy

Speakeasy GPT is a Jupyter notebook that utilizes several natural language processing utilities to provide a seamless and low-latency speech interface to ChatGPT and other large language models.

automatic-speech-recognition chat-gpt coqui-ai coqui-tts elevenlabs-api mimic mycroftai text-to-speech whisper

Last synced: 24 Oct 2024

https://github.com/Lord-Haji/ChatAudio

chatbot gpt-3-5-turbo gpt-4 langchain langchain-python speech-recognition whisper whisper-api

Last synced: 24 Oct 2024

https://github.com/otonomee/mic2transcript

CLI tool that continuously transcribes audio from the device's built-in microphone to a text file. Runs in the background, providing an ongoing log of ambient audio as text.

audio cli cli-tool openai speech speech-transcription transcription whisper

Last synced: 09 Oct 2024

https://github.com/firefly55lm/bisbigliatorev2

Automatic audio transcriber notebook based on Whisper

colab-notebook speech-to-text whisper

Last synced: 25 Jan 2025

https://github.com/seitzquest/RavenWhisperer

Listens to your voice and queries a language model for answers when a question is detected

rwkv whisper

Last synced: 22 Nov 2024

https://github.com/marketcalls/openalgo-voice-based-orders

OpenAlgo Voice Based Orders

flask groq openai python speech-to-text whisper

Last synced: 19 Dec 2024

https://github.com/romiconez/konspecto-llm

LLM agent that provides tools for convenient work with personal documents using voice or text.

agent backend docker docx frontend google langchain llamaindex llm managment nlp rag whisper

Last synced: 05 Jan 2025

https://github.com/vimwei/whispertranscriber

Whisper Transcribe and srt Resegment

speech-to-text subtitle whisper

Last synced: 17 Oct 2024

https://github.com/sakurajimamai-1202/stream-translator-gpt-webui

A web ui application that utilizes the stream-translator-gpt

faster-whisper gemini gpt transcribe translate translation translator webui whisper yt-dlp

Last synced: 11 Oct 2024

https://github.com/tensoraws/yuisub

Auto translation of new anime episodes based on Yui-MHCP001

anime chatgpt llm openai pysubs2 subtitle translation whisper

Last synced: 09 Oct 2024

https://github.com/ashot72/speech-to-text-to-image

Generating texts from your voice then images form the texts

chatgpt large-language-models llm replicate speech-to-text speechtotext stability-ai text-to-image texttoimage whisper whisper-ai

Last synced: 30 Dec 2024

https://github.com/alancunningham/chatgpt-assistant

A ChatGPT assistant with voice activation and image generation, connected to a Raspberry Pi display.

chatgpt chatgpt-api dall-e dall-e-api porcupine python raspberry-pi whisper

Last synced: 06 Jan 2025

https://github.com/saadkh1/docqa-textsummarization-app

A Streamlit app for document question answering and text summarization.

langchain llama-2 llamacpp pytesseract question-answering streamlit summarization whisper

Last synced: 07 Jan 2025

https://github.com/i4ds/whisper-prep

Data preparation utility for the finetuning of OpenAI's Whisper model.

fine-tuning nlp speech-to-text whisper

Last synced: 09 Nov 2024

https://github.com/stefanasandei/youtube-to-text

Speech to text for any YouTube video.

ai api flask openai python server speech-to-text web-server whisper youtube youtube-dl

Last synced: 04 Jan 2025

https://github.com/sonhm3029/realtime-vietnamese-asr-react-native-and-whisper

This project implement end to end realtime vietnamese speech recognition with PhoWhisper in Backend and frontend in React Native

asr phowhiper react-native realtime realtime-speech-recognition speech-recognition speech-to-text vietnamese whisper

Last synced: 16 Nov 2024

https://github.com/ksylvest/omniai-openai

An implementation of the OmniAI interface for OpenAI.

chatgpt omniai openai ruby whisper

Last synced: 10 Jan 2025

https://github.com/flyingfathead/youwhisper-cli

A streamlined CLI tool combining `yt-dlp` and `whisperx` (or `openai-whisper`) for quick and efficient audio transcription from various video platforms.

cli cli-app python transcribe transcriber transcription whisper whisper-ai whisperx youtube-downloader yt-dlp yt-dlp-wrapper

Last synced: 11 Jan 2025

https://github.com/knot-inc/john

John is a web app that records video, analyzes audio with AI, and identifies the speaker's native language from their English accent, simplifying language assessment.

audio-analysis machine-learning whisper

Last synced: 17 Nov 2024

https://github.com/abhishtagatya/polly

☎️ Language Learning Chatbot

chatbot chatgpt python telegram whisper

Last synced: 17 Nov 2024

https://github.com/t-h-chung/note-taker

Note-taking app for online/local video/audio using Whisper transcription, ChatGPT, and Notion

chatgpt notes notion transcription whisper youtube

Last synced: 09 Oct 2024

https://github.com/daisyyedda/whisper-large-v2-atcosim_corpus

A fine-tuned Whisper model (whisper-large-v2) for aviation audio transcription. WER < 5%.

asr-model nlp whisper whisper-ai

Last synced: 09 Oct 2024

https://github.com/jemtaly/whispering

A real-time transcription and translation tool implemented in Python based on the fast-whisper library.

live-caption python real-time-transcription real-time-translation tkinter transcription translation whisper

Last synced: 09 Jan 2025

https://github.com/sanket-poojary-03/fine-tuning-whisper

Fine tuning Whisper-Small LLM for Hinglish Audio dataset

audio-dataset audio-to-text deep-learning fine-tuning huggingface-transformers python speech-recognition speech-to-text whisper whisper-ai

Last synced: 09 Oct 2024

https://github.com/gurpreetkaurjethra/multimodal-ai-app-using-llava-7b

Multimodal AI App using Llava 7B and Gradio

ai generative-ai gradio large-language-models llava llavacpp llm multimodal voice-assistant whisper

Last synced: 22 Nov 2024

https://github.com/bharathajjarapu/voicecipher

Local Speech transcription

transformerjs whisper

Last synced: 09 Oct 2024

https://github.com/kazkozdev/video-analyser

⚡ The YouTube Video Analyzer Pro brings AI-powered analysis capabilities to your fingertips, offering deep insights for content creators and marketers.

ai content-analytics fastapi llama3 llm ollama-api python3 video-analysis video-analysis-client whisper youtube youtube-analytics youtube-api youtube-subscribers

Last synced: 13 Jan 2025

https://github.com/williamwa/mssmith

A Telegram bot that utilizes the ChatGPT API and can communicate through voice.

chatpgt-api telegram-bot tts whisper

Last synced: 31 Dec 2024

https://github.com/sovit-123/sam_molmo_whisper

An integration of Segment Anything Model, Molmo, and, Whisper to segment objects using voice and natural language.

molmo segment-anything-model segmentanythingmodel vlm whisper

Last synced: 18 Oct 2024

https://github.com/ndjenkins85/afkode

Personal voice command interface for iPhone on pythonista powered by Whisper and ChatGPT.

chatgpt openai python-packaging quick-start whisper

Last synced: 12 Oct 2024

https://github.com/benitomartin/youtube-llm

LLM Q&A and Summarization App

chromadb langchain python streamlit whisper

Last synced: 31 Dec 2024

https://github.com/alessioborgi/stylealigned_multireference-multimodal

Novel framework for Zero-Shot Style Alignment in Text-to-Image generation, incorporating Multi-Modal Context-Awareness and Multi-Reference Style Alignment, using minimal attention sharing, ensuring consistent style transfer without fine-tuning.

adain blip clap context-awareness multi-modal multi-style-transfer no-fine-tuning shared-attention-heads style-aligned text-to-image-generation whisper zero-shot-learning

Last synced: 18 Oct 2024

https://github.com/rokbenko/arctic-meet

ArcticMeet is an AI meeting assistant using Streamlit for the GUI and the Snowflake Arctic LLM via the Snowflake Cortex for the AI features

ffmpeg pandas plotly python pytorch snowflake snowflake-arctic snowflake-cortex snowpark streamlit transformers whisper

Last synced: 11 Jan 2025

https://github.com/kunesj/holo-subs-search

Tool for searching transcriptions of vtuber videos.

holodex pyannote transcription vtuber whisper youtube

Last synced: 19 Jan 2025

https://github.com/ahmetoner/master-whisper

Master Whisper transcription with CTranslate2

deep-learning inference openai quantization speech-recognition speech-to-text transformer whisper

Last synced: 08 Jan 2025

https://github.com/mikeesto/whispercpp-android

An Android app using whisper.cpp to do voice-to-text transcriptions

android kotlin speech-to-text whisper whisper-cpp

Last synced: 17 Dec 2024

https://github.com/bigyaa/transcription-system

This versatile tool is designed for anyone in need of a robust solution for transcribing and diarizing large volumes of audio files. Whether you are dealing with terabytes or even larger quantities, our tool ensures efficient and accurate processing. Ideal for researchers, content creators, and businesses.

accessibility diarization speech-to-text storytelling-with-data transcription whisper

Last synced: 19 Dec 2024

https://github.com/adisol07/sharpspeech

SharpSpeech is free, local and open source way to speech and wake word recognition.

audio speech speech-recognition speech-to-text wake-word-detection wakeword whisper whisper-ai

Last synced: 19 Dec 2024

https://github.com/maawad/luna

Personal assistant

bot openai personal-assistant whisper

Last synced: 17 Dec 2024

https://github.com/huuquyet/phowhisper-next

Demo using PhoWhisper models of VinAI built with Transformers.js + Next.js

nextjs onnx-models phowhisper speech-recognition transformersjs vietnamese vinai whisper

Last synced: 19 Dec 2024

https://github.com/toLSC/tolsc-speech-to-text

Speech to text service for toLSC app implemented with OpenAI Whisper model

fastapi python speech-recognition speech-to-text tts whisper

Last synced: 24 Oct 2024

https://github.com/bbc-esq/batch-openai-whisper-ctranslate2

Batch process multiple files using the fasted ctranslate2 implementation of Open AI's Whisper

batch-processing batch-script openai openai-whisper pyside6 transcription translation whisper whisperx

Last synced: 11 Jan 2025

https://github.com/carlosulisesochoa/whisper-ai-transcription-audio-to-text-file

A Python tool that uses OpenAI's Whisper model to batch transcribe audio files with GPU acceleration. Features include multi-language support, timestamp-based output, automatic file status checking, and CUDA support for faster processing. Perfect for transcribing lectures, interviews, or any audio content with high accuracy.

ai audio-to-text transcription whisper

Last synced: 28 Jan 2025

https://github.com/chaoticbyte/audio-summarize

An audio summarizer (faster-whisper and BART glued together)

ai ai-summarizer audio bart ctranslate2 faster-whisper nlp speech-to-text summarization whisper

Last synced: 09 Oct 2024

https://github.com/jpzinn654/speaker-diarization-portuguese

This project implements speaker diarization for Portuguese audio using WhisperX for transcription and PyAnotAudio's Speaker-Diarization 3.1 for speaker separation. It includes a Flask UI for easy file upload, transcription, and speaker identification.

flask gender-detection portuguese-language speaker-diarization speaker-recognition speech-recognition transcription whisper

Last synced: 28 Jan 2025

https://github.com/etienneab3d/srt-sync

Synchronize SRT timestamps over an existing accurate transcription

aligner asr nlp subtitles text-to-speech whisper

Last synced: 19 Dec 2024

https://github.com/wtlow003/auto-subtitles

CLI tool to transcribe (+ translate) videos and embed subtitles automatically.

faster-whisper nllb subtitles subtitles-generator translation whisper whisper-cpp

Last synced: 15 Nov 2024

https://github.com/team-mansumugang/mansumugang-backend

만수무강 서비스의 스프링 부트 어플리케이션입니다.

aws github-actions jpa jpa-hibernate spring-boot whisper

Last synced: 09 Oct 2024

https://github.com/xawos/owt

🦙🗣️ Ollama and Whisper Telegram bot, with advanced configuration

ai-bots local-ai ollama telegram-aichatbot telegram-bots whisper

Last synced: 28 Jan 2025

https://github.com/gamut73/quizinator

Generating quizzes, on Android, from YouTube videos.

kotlin-android llm python whisper

Last synced: 19 Dec 2024

https://github.com/mikeesto/subber

A small CLI tool for converting video & audio to a text transcription

audio cli ffmpeg golang transcribe video whisper

Last synced: 19 Dec 2024

https://github.com/nerdimite/meetsy-backend

AI Backend for the Workshop on Building an End-to-End AI Meeting Assistant

gpt-3 nextjs sentence-transformers tailwindcss whisper

Last synced: 24 Oct 2024

https://github.com/aeronjl/transcribe

Python package for accurate audio transcription with speaker diarisation

audio-transcription gpt speaker-diarization whisper

Last synced: 09 Oct 2024

https://github.com/roman01la/sub-deep

Transcribe and translate audio with AI

deepl transcribe translate whisper

Last synced: 30 Dec 2024

https://github.com/sumitesh9/localizedwhisper

An initiative to make OpenAI Whisper more localized by adding support for more languages.

albanian albanian-language huggingface openai speech speech-to-text whisper

Last synced: 02 Jan 2025

https://github.com/sbadulin/obsidian-dictation-plugin

Obsidian dictation plugin

dictation gpt-35-turbo obsidian obsidian-plugin openai speech-to-text whisper

Last synced: 02 Feb 2025

https://github.com/niqifan007/openai-tts-stt-streamlit

A gui interface for tts (text-to-speech) and stt (speech-to-text) interfaces using the openai api developed by Streamlit, with a history function一个使用Streamlit开发的openai的api接口的tts（文字转语音）和stt（语音转文字）接口的gui界面，带有历史记录功能

openai openai-api streamlit stt-gui tts tts-gui whisper whisper-api

Last synced: 09 Oct 2024

https://github.com/mickekring/top-of-mind-clara

Clara är en prototyp som möjliggör att anonymt kunna göra sin röst hörd. Medarbetaren kan prata eller skriva in det du vill säga och AI anonymiserar det. Medarbetaren har dessutom tillgång till en chatbot att rådfråga. Därefter analyseras och sammanställs alla medarbetares tankar i en dashboard.

ai chatbot feedback openai python streamlit transcription whisper

Last synced: 22 Dec 2024

https://github.com/gabriellopesdesouza2002/funcspy

Functions to help you develop any program or script you want

automation chatbot dall-e email email-library ocr openai-api openai-chatgpt openai-whisper pdf pdf-tools python regex selenium selenium-webdriver whisper

Last synced: 30 Oct 2024

https://github.com/slinusc/speaker_identification_evaluation

Evaluating the Effectiveness of Transformer Layers in Wav2Vec 2.0, XLS-R, and Whisper for Speaker Identification Tasks

wav2vec2 whisper xls-r

Last synced: 09 Oct 2024

https://github.com/jgw96/speech-to-text-web-toolkit

Making Speech-To-Text on the web easy, both local and in the cloud

ai lit transformersjs webcomponents whisper

Last synced: 01 Feb 2025

https://github.com/antoniosbarotsis/telegram-transcriber

A Telegram bot for transcribing voice messages

telegram transcribe voice whisper

Last synced: 26 Dec 2024

https://github.com/tposcic/audio-to-srt-transcriber

Audio to srt transcriber in Python using whisper for transcription and Tcl/Tk for GUI

audio python3 srt transcription whisper

Last synced: 05 Jan 2025

https://github.com/utrechtuniversity/transcription-d-lucea

python utrecht-university whisper

Last synced: 23 Jan 2025

https://github.com/aitor-alvarez/large-speech-models

Fine-tuning Multilingual Large Speech Recognition Models: Wav2vec and Whisper

arabic-speech-recognition asr asr-model finetuning-wav2vec finetuning-whisper large-speech-models speech-recognition-model wav2vec2 whisper

Last synced: 25 Jan 2025

https://github.com/aspadax/subtitlegenerator

Automatically generate a subtitle for your video.

gpt machine-learning openai rust streamlit subtitles-generator whisper

Last synced: 09 Oct 2024

https://github.com/marquesafonso/multilang-asr-captioner

A multilingual automatic speech recognition and video captioning tool using faster whisper. Supports real-time translation to english. Runs on consumer grade cpu.

automatic-speech-recognition captioning-videos faster-whisper whisper

Last synced: 24 Oct 2024

https://github.com/stnderror/robotron

🤖 A personal robot assistant for Telegram

assistant bot dall-e gpt-35-turbo openai telegram-bot whisper

Last synced: 25 Jan 2025

https://github.com/drankush/voxrad

VOXRAD is a voice transcription application for radiologists leveraging locally deployed ASR and LLM models.

desktop-app ffmpeg gemini gpt llm macos medical-informatics multimodal natural-language-processing nlp openai openai-api productivity python radiology reporting transcription voice-recognition whisper windows

Last synced: 31 Jan 2025

https://github.com/ma2za/telegram-voice-journal

telegram telegram-bot whisper

Last synced: 24 Oct 2024

https://github.com/thewh1teagle/whisper.zig

Transcribe audio with whisper in zig

asr openai whisper zig

Last synced: 24 Jan 2025

https://github.com/tracywong117/ai-learning-material-from-video

Support subtitling, translating, RAG to generate language learning material from video.

ai auto-subtitle gpt-translate groq groq-api rag subtitles-generator translate whisper

Last synced: 19 Jan 2025

https://github.com/shtirmann/v2t

Telegram bot which automatically transcribes all voice and video messages to text.

ai aiogram faster-whisper python telegram-bot telegram-bot-python voice-to-text whisper

Last synced: 09 Oct 2024

https://github.com/csuoc/breaking_vlad_audio_AI

ai audio aws openai polly python pytube selenium transcription whisper

Last synced: 24 Oct 2024

https://github.com/i4ds/whisper-finetune

This repository contains code for fine-tuning the Whisper speech-to-text model.

fine-tuning nlp speech-to-text whisper

Last synced: 09 Oct 2024

https://github.com/jesse-c/local-audio-toolkit

Some handy tools to do with audio locally.

large-language-models lm-studio macos side-project whisper

Last synced: 29 Jan 2025

https://github.com/jlcarveth/skreech

An HTTP API wrapper around Whisper for transcribing audio files.

speech-recognition speech-to-text whisper whisper-ai

Last synced: 19 Jan 2025

https://github.com/nicknaskida/insanely-fast-whisper

Incredibly fast Whisper-large-v3 with speaker diarization

diarization speaker-diarization transfromers whisper whisper-ai whisper-faster whisper-large

Last synced: 19 Jan 2025

https://github.com/h3yn3s/tl-dl

A selfhostable webapp which helps you read those uselessly long (by nature) voice messages with the power of AI.

sveltekit tailwind whisper

Last synced: 24 Oct 2024

https://github.com/juanestban/whisper-tnode

cli ts typescript whisper whisper-cpp whisper-ia whisper-node whisper-node-ts

Last synced: 21 Dec 2024

https://github.com/adamelkholyy/whisper-yt

Toolkit for using Whisper to transcribe YouTube videos. Includes Whisper transcription of YouTube videos, conversion of YouTube video into HuggingFace dataset (using audio and subtitles) and evaluation of Whisper transcription against YouTube subtitles

asr diarization huggingface-datasets pyannote transcription whisper word-error-rate youtube

Last synced: 10 Dec 2024

https://github.com/Shtirmann/V2T

Telegram bot which automatically transcribes all voice and video messages to text.

ai aiogram faster-whisper python telegram-bot telegram-bot-python voice-to-text whisper

Last synced: 24 Oct 2024

https://github.com/lazauk/aoai-entraidauth-sdkv1

Authenticating with Entra ID (former Azure AD) to access Azure OpenAI models in Python SDK v1.x

ai authentication azure azure-active-directory dall-e embeddings entra-id gpt openai whisper

Last synced: 12 Jan 2025

https://github.com/amanpriyanshu/medtranslate-360

MedTranslate 360 redefines medical documentation by providing an AI-powered assistant designed specifically for healthcare professionals.

ai gemini gemini-api hackathon llm llms medical ml privacy streamlit whisper

Last synced: 15 Dec 2024

https://github.com/crone-ai/force-align-wordstamps

Takes audio (mp3) and text input (string) and force aligns the text to the audio. Uses stable-ts and whisperx.

captions faster-whisper force-alignment stable-ts whisper

Last synced: 17 Jan 2025

https://github.com/jojasadventure/whisper-client

Very simple Python based client for Whisper compatible endpoint

desktop-app dictation faster-whisper macos productivity python speech-to-text stt whisper

Last synced: 09 Oct 2024

https://github.com/maylad31/colab-codes

some useful colab files

clip colab-notebook speech-recognition whisper zero-shot-classification

Last synced: 11 Jan 2025

https://github.com/extrange/transcription-benchmarks

Speech to text model benchmarks

transcription whisper

Last synced: 08 Dec 2024

https://github.com/bbc-esq/whisper-solo-with-gui

OpenAI's Whisper program with a simple lightweight GUI.

pyqt pyqt6 pyqt6-gui transcribe transcribe-audio-files translate whisper

Last synced: 11 Jan 2025

https://github.com/valiantlynx/custom-whisper-api

This project provides a custom API wrapper for the open-source Whisper model using FastAPI. It allows you to integrate Whisper into your applications for automatic speech recognition (ASR) tasks.

ai docker-compose fastapi python whisper

Last synced: 10 Jan 2025

https://github.com/fukuro-kun/wortweber

Wortweber ist ein sich in der Entwicklung befindendes Open-Source-Projekt, das Echtzeit-Sprachtranskription mit KI-Technologie erforscht. Es dient als Lern- und Experimentierplattform für Spracherkennung in Deutsch und Englisch.

speech-to-text whisper

Last synced: 17 Jan 2025

https://github.com/ayeshaaaaaaaaa/ai-powered-video-analysis-with-object-detection-and-detailed-scene-narratives

AI-driven video analysis system that extracts and transcribes audio with Whisper, detects objects using YOLO, and generates comprehensive scene descriptions with GPT-2. The project combines transcriptions and object detections to produce detailed, context-aware video narratives.

bart gpt2 video-analysis whisper yolov8

Last synced: 02 Jan 2025