Ecosyste.ms: Awesome

An open API service indexing awesome lists of open source software.

Awesome Lists | Featured Topics | Projects

Whisper

Whisper is an autoregressive language model developed by OpenAI. It is trained on a large corpus of text using a transformer architecture and is capable of generating high-quality natural language text. Whisper can be used for tasks such as language modeling, text completion, and text generation. It has shown impressive performance on various benchmarks and has been released by OpenAI to encourage research in the field of language modeling. Whisper is not yet available for public use, but it has the potential to transform the field of natural language processing and generate new opportunities for language-based applications.

https://github.com/sovit-123/sam_molmo_whisper

An integration of Segment Anything Model, Molmo, and, Whisper to segment objects using voice and natural language.

molmo segment-anything-model segmentanythingmodel vlm whisper

Last synced: 18 Oct 2024

https://github.com/ndjenkins85/afkode

Personal voice command interface for iPhone on pythonista powered by Whisper and ChatGPT.

chatgpt openai python-packaging quick-start whisper

Last synced: 12 Oct 2024

https://github.com/JoSuru/speeka

Speeaka is an open-source project that uses the Whisper model of OpenAI to transcribe audio into text. Its intuitive web interface makes it easy to use. Contributions are welcome.

open-source python python3 speech-to-text streamlit whisper

Last synced: 24 Oct 2024

https://github.com/t-h-chung/note-taker

Note-taking app for online/local video/audio using Whisper transcription, ChatGPT, and Notion

chatgpt notes notion transcription whisper youtube

Last synced: 08 Feb 2025

https://github.com/tonywu71/distilling-and-forgetting-in-large-pre-trained-models

Code for my dissertation on "Distilling and Forgetting in Large Pre-Trained Models" for the MPhil in Machine Learning and Machine Intelligence (MLMI) at the University of Cambridge.

continual-learning distillation speech-recognition whisper

Last synced: 04 Dec 2024

https://github.com/daisyyedda/whisper-large-v2-atcosim_corpus

A fine-tuned Whisper model (whisper-large-v2) for aviation audio transcription. WER < 5%.

asr-model nlp whisper whisper-ai

Last synced: 08 Feb 2025

https://github.com/abhishtagatya/polly

☎️ Language Learning Chatbot

chatbot chatgpt python telegram whisper

Last synced: 17 Nov 2024

https://github.com/amir-mohseni/voicebridge

This repository provides a dockerized Speech-to-Speech application that supports text-to-audio conversion, audio-to-text transcription, and interactive voice-based conversations. It is easy to set up and use, offering a versatile platform for speech and text processing.

docker huggingface python transformer tts whisper

Last synced: 17 Jan 2025

https://github.com/saadkh1/docqa-textsummarization-app

A Streamlit app for document question answering and text summarization.

langchain llama-2 llamacpp pytesseract question-answering streamlit summarization whisper

Last synced: 07 Jan 2025

https://github.com/amgawishx/voiceworker

A Web App UI for OpenAI's Whisper model for audio transcription and translation.

ai audio-processing python streamlit transcription translation webapp whisper

Last synced: 17 Jan 2025

https://github.com/phidlarkson/whisper-stt-api

Easy setup for the whisper speech to text

api flask speech-to-text whisper

Last synced: 01 Jan 2025

https://github.com/water25234/ChatREP

Summary on Youtube By ChatGPT & whisper

chatgpt-api openai python python3 video whisper youtube

Last synced: 24 Oct 2024

https://github.com/limdongjin/ignkafasr

Real-Time In-memory Speaker Verification and Speech Recognition Project using apache ignite, apache kafka, speechbrain, whisper, stomp, spring webflux, kubernetes(k8s)

apache-ignite apache-kafka asr audio-recorder google-kubernetes-engine k8s kubernetes speaker-recognition speaker-verification speech-recognition speechbrain springframework stomp stompwebsocket webflux whisper

Last synced: 24 Oct 2024

https://github.com/adamelkholyy/whisper-yt

Toolkit for using Whisper to transcribe YouTube videos. Includes Whisper transcription of YouTube videos, conversion of YouTube video into HuggingFace dataset (using audio and subtitles) and evaluation of Whisper transcription against YouTube subtitles

asr diarization huggingface-datasets pyannote transcription whisper word-error-rate youtube

Last synced: 05 Feb 2025

https://github.com/jacoblincool/wft

Run Whisper fine-tuning with ease—it works on MPS, CUDA, and CPU without code changes.

fine-tuning whisper

Last synced: 11 Dec 2024

https://github.com/ribartra/call-listener_bot

A bot that downloads, transcribes and analyzes calls to find insights for sales advisors.

api audio-analyser call-bot call-listener drive gcp openai python whisper

Last synced: 08 Feb 2025

https://github.com/alancunningham/chatgpt-assistant

A ChatGPT assistant with voice activation and image generation, connected to a Raspberry Pi display.

chatgpt chatgpt-api dall-e dall-e-api porcupine python raspberry-pi whisper

Last synced: 06 Jan 2025

https://github.com/williamwa/mssmith

A Telegram bot that utilizes the ChatGPT API and can communicate through voice.

chatpgt-api telegram-bot tts whisper

Last synced: 31 Dec 2024

https://github.com/datarabbit-ai/transcription_service

System/service with REST API for extracting text transcriptions from movies and audio recordings in most popular video formats.

containers datarabbit rest-api speech-to-text stt transcription transcription-services whisper

Last synced: 08 Feb 2025

https://github.com/bhattbhavesh91/neo4j-palm2-makersuite

Explore how to build a Q&A system on Neo4j using Google's Palm2 model with MakerSuite in this repository.

google google-api google-palm maker-suite neo4j-driver neo4j-python-scripts palm2 python table-qa voice-assistant whisper

Last synced: 17 Jan 2025

https://github.com/romiconez/konspecto-llm

LLM agent that provides tools for convenient work with personal documents using voice or text.

agent backend docker docx frontend google langchain llamaindex llm managment nlp rag whisper

Last synced: 05 Jan 2025

https://github.com/jemtaly/whispering

A real-time transcription and translation tool implemented in Python based on the fast-whisper library.

live-caption python real-time-transcription real-time-translation tkinter transcription translation whisper

Last synced: 09 Jan 2025

https://github.com/benitomartin/youtube-llm

LLM Q&A and Summarization App

chromadb langchain python streamlit whisper

Last synced: 31 Dec 2024

https://github.com/aspadax/subtitlegenerator

Automatically generate a subtitle for your video.

gpt machine-learning openai rust streamlit subtitles-generator whisper

Last synced: 08 Feb 2025

https://github.com/firefly55lm/bisbigliatorev2

Automatic audio transcriber notebook based on Whisper

colab-notebook speech-to-text whisper

Last synced: 25 Jan 2025

https://github.com/flyingfathead/youwhisper-cli

A streamlined CLI tool combining `yt-dlp` and `whisperx` (or `openai-whisper`) for quick and efficient audio transcription from various video platforms.

cli cli-app python transcribe transcriber transcription whisper whisper-ai whisperx youtube-downloader yt-dlp yt-dlp-wrapper

Last synced: 11 Jan 2025

https://github.com/kazkozdev/video-analyser

⚡ The YouTube Video Analyzer Pro brings AI-powered analysis capabilities to your fingertips, offering deep insights for content creators and marketers.

ai content-analytics fastapi llama3 llm ollama-api python3 video-analysis video-analysis-client whisper youtube youtube-analytics youtube-api youtube-subscribers

Last synced: 13 Jan 2025

https://github.com/astrologos/py-speakeasy

Speakeasy GPT is a Jupyter notebook that utilizes several natural language processing utilities to provide a seamless and low-latency speech interface to ChatGPT and other large language models.

automatic-speech-recognition chat-gpt coqui-ai coqui-tts elevenlabs-api mimic mycroftai text-to-speech whisper

Last synced: 24 Oct 2024

https://github.com/szilvia-csernus/openai-audio-api-calls

Speech-to-text and text-to-speech API call examples, using OpenAI's whisper-1 and tts-1 models.

jupyter-notebook openai openai-api tts-1 whisper

Last synced: 08 Feb 2025

https://github.com/TheGuysBrushes/Whisper

Secured chat application

android chat socket whisper

Last synced: 24 Oct 2024

https://github.com/seitzquest/RavenWhisperer

Listens to your voice and queries a language model for answers when a question is detected

rwkv whisper

Last synced: 22 Nov 2024

https://github.com/vimwei/whispertranscriber

Whisper Transcribe and srt Resegment

speech-to-text subtitle whisper

Last synced: 17 Oct 2024

https://github.com/i4ds/whisper-prep

Data preparation utility for the finetuning of OpenAI's Whisper model.

fine-tuning nlp speech-to-text whisper

Last synced: 09 Nov 2024

https://github.com/ksylvest/omniai-openai

An implementation of the OmniAI interface for OpenAI.

chatgpt omniai openai ruby whisper

Last synced: 10 Jan 2025

https://github.com/andreabak/whispersubs

Generate subtitles for your video or audio files using the power of AI

ai cuda deep-learning gpu-acceleration machine-learning srt subtitles transcribe transcription translate whisper

Last synced: 16 Nov 2024

https://github.com/otonomee/mic2transcript

CLI tool that continuously transcribes audio from the device's built-in microphone to a text file. Runs in the background, providing an ongoing log of ambient audio as text.

audio cli cli-tool openai speech speech-transcription transcription whisper

Last synced: 08 Feb 2025

https://github.com/awaisoem/interview-lingo

(Aug 2024) AI assistant which help with interviews, hiring, personality development and communication skills

ai ai71 drizzle-orm falcon neondb nextjs postgresql tailwindcss whisper

Last synced: 08 Feb 2025

https://github.com/my-north-ai/semantic_audio_filtering

Synthetic data augmentation technique via LLM for Automatic Speech Recognition fine tuning.

automatic-speech-recognition fine-tuning synthetic-dataset-generation text-to-speech whisper

Last synced: 24 Oct 2024

https://github.com/knot-inc/john

John is a web app that records video, analyzes audio with AI, and identifies the speaker's native language from their English accent, simplifying language assessment.

audio-analysis machine-learning whisper

Last synced: 17 Nov 2024

https://github.com/fer14/videoseek

Intelligent video search tool powered by AI

bert timestamp video whisper youtube-api

Last synced: 14 Jan 2025

https://github.com/gamut73/quizinator

Generating quizzes, on Android, from YouTube videos.

kotlin-android llm python whisper

Last synced: 19 Dec 2024

https://github.com/chinese-soup/cbot-telegram-whisper

Simple bot that transcribes Telegram voice messages. Powered by go-telegram-bot-api & whisper.cpp Go bindings.

bot cpu-inference golang openai speech-recognition speech-to-text whisper whisper-cpp whispercpp

Last synced: 17 Jan 2025

https://github.com/schnoddelbotz/whisper-ui

Transcribe audio/video to text, locally on macOS, Linux and Windows. A simple whisper.cpp wrapper/UI built with Go/Fyne.

ffmpeg ffmpeg-wrapper fyne gui local privacy speech-to-text transcription whisper whisper-cpp

Last synced: 27 Jan 2025

https://github.com/egorsmkv/star-adapt-uk

Fork of https://github.com/YUCHEN005/STAR-Adapt with some modifications for Ukrainian.

asr speech-recognition ukrainian whisper

Last synced: 19 Dec 2024

https://github.com/nri12/filter_voice

Dự án lọc và tắt tiếng video những từ khóa mong muốn

python tools whisper

Last synced: 19 Dec 2024

https://github.com/oov/aviutl_subtitler

AviUtl+拡張編集の環境で Whisper による文字起こしをするためのプラグイン

aviutl aviutl-plugin whisper

Last synced: 19 Dec 2024

https://github.com/canaxs/whisper-core

An application where users can make rumor-based news and earn money in return.

mysql panel spring spring-boot whisper

Last synced: 19 Dec 2024

https://github.com/breadrock1/audio-to-text

There is simple backend project to use whisper-rs.

actix-web audio-to-text rust swagger-ui whisper

Last synced: 10 Jan 2025

https://github.com/tranbavinhson/eth-decentralized-chat

Decentralized chat app by Ethereum Whisper protocol + Vuejs

ethereum vue vuejs whisper whisper-protocol

Last synced: 26 Dec 2024

https://github.com/ayeshaaaaaaaaa/ai-powered-video-analysis-with-object-detection-and-detailed-scene-narratives

AI-driven video analysis system that extracts and transcribes audio with Whisper, detects objects using YOLO, and generates comprehensive scene descriptions with GPT-2. The project combines transcriptions and object detections to produce detailed, context-aware video narratives.

bart gpt2 video-analysis whisper yolov8

Last synced: 02 Jan 2025

https://github.com/toomore/whisper

🔐📦📜🔑🍞 Write some notes by using the GPG encrypts.

gpg notes pgp quickstart whisper

Last synced: 23 Jan 2025

https://github.com/slinusc/speaker_identification_evaluation

Evaluating the Effectiveness of Transformer Layers in Wav2Vec 2.0, XLS-R, and Whisper for Speaker Identification Tasks

wav2vec2 whisper xls-r

Last synced: 08 Feb 2025

https://github.com/nerdimite/meetsy-backend

AI Backend for the Workshop on Building an End-to-End AI Meeting Assistant

gpt-3 nextjs sentence-transformers tailwindcss whisper

Last synced: 24 Oct 2024

https://github.com/fukuro-kun/wortweber

Wortweber ist ein sich in der Entwicklung befindendes Open-Source-Projekt, das Echtzeit-Sprachtranskription mit KI-Technologie erforscht. Es dient als Lern- und Experimentierplattform für Spracherkennung in Deutsch und Englisch.

speech-to-text whisper

Last synced: 17 Jan 2025

https://github.com/sumitesh9/localizedwhisper

An initiative to make OpenAI Whisper more localized by adding support for more languages.

albanian albanian-language huggingface openai speech speech-to-text whisper

Last synced: 02 Jan 2025

https://github.com/mikeesto/whispercpp-android

An Android app using whisper.cpp to do voice-to-text transcriptions

android kotlin speech-to-text whisper whisper-cpp

Last synced: 09 Feb 2025

https://github.com/noxs1d/speech-to-text

🤖ML project which record audio and converts it to text

ai fastapi ml torch whisper

Last synced: 07 Feb 2025

https://github.com/oussemabenhassena5/notegen-with-llama-and-whisper

AI-powered YouTube video notes generator

ai llama3 python whisper youtube-api

Last synced: 07 Feb 2025

https://github.com/valiantlynx/custom-whisper-api

This project provides a custom API wrapper for the open-source Whisper model using FastAPI. It allows you to integrate Whisper into your applications for automatic speech recognition (ASR) tasks.

ai docker-compose fastapi python whisper

Last synced: 10 Jan 2025

https://github.com/bbc-esq/whisper-solo-with-gui

OpenAI's Whisper program with a simple lightweight GUI.

pyqt pyqt6 pyqt6-gui transcribe transcribe-audio-files translate whisper

Last synced: 11 Jan 2025

https://github.com/gangula-karthik/memo-mate

🚀 Discord meetings redefined with Memo Mate: Transcribe, summarize, and automate minutes seamlessly! ✨

discord-bot huggingface mistral py-cord speech-to-text transcribe whisper

Last synced: 22 Dec 2024

https://github.com/lazauk/aoai-entraidauth-sdkv1

Authenticating with Entra ID (former Azure AD) to access Azure OpenAI models in Python SDK v1.x

ai authentication azure azure-active-directory dall-e embeddings entra-id gpt openai whisper

Last synced: 12 Jan 2025

https://github.com/winstxnhdw/capgen

A fast CPU-first video/audio transcriber for generating caption files with Whisper and CTranslate2, hosted on Hugging Face Spaces.

asr automatic-speech-recognition caddy ctranslate2 docker fastapi huggingface huggingface-spaces uvicorn-gunicorn whisper

Last synced: 23 Oct 2024

https://github.com/toLSC/tolsc-speech-to-text

Speech to text service for toLSC app implemented with OpenAI Whisper model

fastapi python speech-recognition speech-to-text tts whisper

Last synced: 24 Oct 2024

https://github.com/adisol07/sharpspeech

SharpSpeech is free, local and open source way to speech and wake word recognition.

audio speech speech-recognition speech-to-text wake-word-detection wakeword whisper whisper-ai

Last synced: 19 Dec 2024

https://github.com/wtlow003/auto-subtitles

CLI tool to transcribe (+ translate) videos and embed subtitles automatically.

faster-whisper nllb subtitles subtitles-generator translation whisper whisper-cpp

Last synced: 15 Nov 2024

https://github.com/platput/pysubs

api to get audio transcription for video files from youtube, aws s3 and such. using OpenAI Whisper

openai whisper

Last synced: 24 Oct 2024

https://github.com/tracywong117/ai-learning-material-from-video

Support subtitling, translating, RAG to generate language learning material from video.

ai auto-subtitle gpt-translate groq groq-api rag subtitles-generator translate whisper

Last synced: 19 Jan 2025

https://github.com/thewh1teagle/whisper.zig

Transcribe audio with whisper in zig

asr openai whisper zig

Last synced: 24 Jan 2025

https://github.com/marquesafonso/multilang-asr-captioner

A multilingual automatic speech recognition and video captioning tool using faster whisper. Supports real-time translation to english. Runs on consumer grade cpu.

automatic-speech-recognition captioning-videos faster-whisper whisper

Last synced: 24 Oct 2024

https://github.com/notyusheng/transcribe-translate

Local web app for transcription and translation services for audio and video using Whisper models

docker full-stack nodejs react reactjs self-hosted speech-to-text transcribe translate whisper

Last synced: 11 Oct 2024

https://github.com/antoniosbarotsis/telegram-transcriber

A Telegram bot for transcribing voice messages

telegram transcribe voice whisper

Last synced: 26 Dec 2024

https://github.com/Op27/meeting_minutes_generator

This Python application automates the process of generating meeting minutes from an audio recording. It uses the Whisper library for transcription and the OpenAI GPT models for summarizing content, then outputs the result in a Word document.

ai audio-processing document-automation meeting-minutes openai python speech-recognition text-summarization transcription whisper

Last synced: 24 Oct 2024

https://github.com/drankush/voxrad

VOXRAD is a voice transcription application for radiologists leveraging locally deployed ASR and LLM models.

desktop-app ffmpeg gemini gpt llm macos medical-informatics multimodal natural-language-processing nlp openai openai-api productivity python radiology reporting transcription voice-recognition whisper windows

Last synced: 31 Jan 2025

https://github.com/stnderror/robotron

🤖 A personal robot assistant for Telegram

assistant bot dall-e gpt-35-turbo openai telegram-bot whisper

Last synced: 25 Jan 2025

https://github.com/TranBaVinhSon/eth-decentralized-chat

Decentralized chat app by Ethereum Whisper protocol + Vuejs

ethereum vue vuejs whisper whisper-protocol

Last synced: 24 Oct 2024

https://github.com/pdcalado/waste

Whisper Audio Service for Transcription and Ergonomics

productivity rofi transcription tts whisper

Last synced: 21 Jan 2025

https://github.com/nerdimite/meetsy-app

Frontend for the Workshop on Building an End-to-End AI Meeting Assistant

gpt-3 nextjs sentence-transformers tailwindcss whisper

Last synced: 24 Oct 2024

https://github.com/bhattbhavesh91/openai-whisper-benchmarking

Comparing the performance of OpenAI's Whisper model on a GPU vs OpenAI's API

gpu openai speech-to-text whisper

Last synced: 16 Nov 2024

https://github.com/bigyaa/transcription-system

This versatile tool is designed for anyone in need of a robust solution for transcribing and diarizing large volumes of audio files. Whether you are dealing with terabytes or even larger quantities, our tool ensures efficient and accurate processing. Ideal for researchers, content creators, and businesses.

accessibility diarization speech-to-text storytelling-with-data transcription whisper

Last synced: 19 Dec 2024

https://github.com/jowadev/interview

Interview is an interactive application crafted to empower both students and professionals in honing their skills for job interviews.

interview-preparation job-interviews nextjs professional students whisper

Last synced: 07 Feb 2025

https://github.com/jlcarveth/skreech

An HTTP API wrapper around Whisper for transcribing audio files.

speech-recognition speech-to-text whisper whisper-ai

Last synced: 19 Jan 2025

https://github.com/xaionaro-go/speech

A Speech-To-Text (with translation) library for Go; currently uses Whisper (runs locally if needed; no need in any API keys)

ai converter go golang library module package speech speech-recognition speech-to-text text whisper

Last synced: 13 Jan 2025

https://github.com/Shtirmann/V2T

Telegram bot which automatically transcribes all voice and video messages to text.

ai aiogram faster-whisper python telegram-bot telegram-bot-python voice-to-text whisper

Last synced: 24 Oct 2024