{"id":16283356,"url":"https://github.com/liamdugan/speech-to-speech","last_synced_at":"2025-03-20T02:30:44.968Z","repository":{"id":188502102,"uuid":"595804868","full_name":"liamdugan/speech-to-speech","owner":"liamdugan","description":"Code for the INTERSPEECH 2023 paper \"Learning When to Speak: Latency and Quality Trade-offs for Simultaneous Speech-to-Speech Translation with Offline Models\"","archived":false,"fork":false,"pushed_at":"2025-01-14T21:49:06.000Z","size":1712,"stargazers_count":30,"open_issues_count":0,"forks_count":6,"subscribers_count":2,"default_branch":"main","last_synced_at":"2025-03-17T13:35:14.525Z","etag":null,"topics":["simultaneous-translation","speech","speech-processing","speech-to-speech","speech-translation"],"latest_commit_sha":null,"homepage":"https://arxiv.org/abs/2306.01201","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/liamdugan.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2023-01-31T21:01:26.000Z","updated_at":"2025-02-08T06:22:09.000Z","dependencies_parsed_at":null,"dependency_job_id":"d70c58df-ad8f-4aef-ac8e-6af4f72ab532","html_url":"https://github.com/liamdugan/speech-to-speech","commit_stats":null,"previous_names":["liamdugan/speech-to-speech"],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/liamdugan%2Fspeech-to-speech","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/liamdugan%2Fspeech-to-speech/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/liamdugan%2Fspeech-to-speech/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/liamdugan%2Fspeech-to-speech/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/liamdugan","download_url":"https://codeload.github.com/liamdugan/speech-to-speech/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":244538540,"owners_count":20468731,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["simultaneous-translation","speech","speech-processing","speech-to-speech","speech-translation"],"created_at":"2024-10-10T19:13:15.389Z","updated_at":"2025-03-20T02:30:44.950Z","avatar_url":"https://github.com/liamdugan.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"# Simultaneous Speech to Speech Translation w/ Whisper\n![/assets/demo.gif](/assets/demo.gif)\n\nThis repository contains code for the INTERSPEECH 2023 paper \"Learning When to Speak: Latency and Quality Trade-offs for Simultaneous Speech-to-Speech Translation with Offline Models\" ([link](https://www.isca-archive.org/interspeech_2023/dugan23_interspeech.pdf)). In our paper we introduce simple techniques for converting offline speech to text translation systems (like whisper) into real-time simultaneous translation systems.\n\n## Installation\n### Step 1: Install portaudio (required for PyAudio)\n**Windows**\n\nPyaudio seems to compile portaudio upon installation automatically. No manual installation is needed.\n\n**Linux**\n\nDownload the portaudio source from https://files.portaudio.com/download.html.\n\nExtract the source into its own directory and compile portaudio using `./configure \u0026\u0026 make install`.\n\n**Mac**\n```\nbrew install portaudio\n```\n### Step 2: Clone Repository\nClone the repo and cd into the main directory\n``` \ngit clone https://github.com/liamdugan/speech-to-speech.git\ncd speech-to-speech\n```\n\n### Step 3: Create Environment and install dependencies \n**Conda** (must be Python 3.8 or higher)\n```\nconda env create -n s2st python=3.8\nconda activate s2st\npip install -r requirements.txt\n```\n**Venv** (must be Python 3.8 or higher)\n```\npython -m venv env\nsource env/bin/activate\npip install -r requirements.txt\n```\nand you're good to go!\n\n## Usage\n\nFirst populate the `api_keys.yml` file with your API keys for both OpenAI and ElevenLabs. For OpenAI go to [https://platform.openai.com/account/api-keys](https://platform.openai.com/account/api-keys) and for ElevenLabs go to [https://api.elevenlabs.io/docs](https://api.elevenlabs.io/docs).\n\nAfter you are finished use `speech-to-speech.py` to run the system. Add the `--use_local` flag to use the local version of whisper (model size and input language will be taken from the `config.yml` file). The full options are listed below\n```\n$ python speech-to-speech.py -h\n  -h, --help           show this help message and exit\n  --mic MIC            Integer ID for the input microphone (default: 0)\n  --use_local          Whether to use local models instead of APIs\n  --api_keys API_KEYS  The path to the api keys file (default: api_keys.yml)\n  --config CONFIG      The path to the config file (default: keys.yml)\n```\n\nThis script will default to using the Whisper API if `--use_local` is not specified. If it is specified, it will default to using the GPU for the whisper model if it is available. It is highly recommended to use the API if your machine does not have access to a large CUDA-capable GPU.\n\n### Note: Microphone Selection\nIf you run the `speech-to-speech.py` script with an invalid microphone ID then the following message will be printed\n```\n$ python speech-to-speech.py --mic 10\nMicrophone Error: please select a different mic. Available devices:\nID: 0 -- Liam’s AirPods\nID: 1 -- Sceptre Z27\nID: 2 -- MacBook Pro Microphone\nID: 3 -- MacBook Pro Speakers\nID: 4 -- ZoomAudioDevice\nPlease quit with Control-C and try again\n```\nThis will allow you to pick the desired microphone from the list.\n\n### Note: Configuration\nTo edit the policy used for speaking output phrases, directly edit (or make copies of) the `config.yml` file. Fields of particular importance are the policy settings (`policy`, `consensus_threshold`, `confidence_threshold`, and `frame_width`), which can have a large impact on the final performance. \n\n## Evaluation\nBelow is a table showing the evaluation results (in BLEU and Average Lagging) for the four policies `greedy`, `offline`, `confidence`, and `consensus` using Whisper Medium on the CoVoST2 dataset - see the paper for more details on the evaluation setup.\n\n![assets/table.png](assets/table.png)\n\nTo reproduce these results run the `evaluation/pipeline.py` script. This script takes in elements from the CoVoST2 dataset and simulates a simultaneous environment by feeding the model portions of the audio in increments of `frame_width`.\n\n## Contribution\nWe appreciate any and all contributions but especially those containing implementations of new translators and vocalizers. \n\nTo implement a new translator simply add a new subfolder under `/translators` and create a class that implements the `Translator` interface. Likewise for vocalizers, simply add a new subfolder under `/vocalizers` and create a class that implements the `Vocalizer` interface. In particular we would love to have implementations for `whisper.cpp` and a local TTS system (possibly `balacoon` or some other fine-tuned `tacotron`).\n\n## Citation\nIf you use our code or findings in your research, please cite us as:\n```\n@inproceedings{dugan23_interspeech,\n  author={Liam Dugan and \n          Anshul Wadhawan and \n          Kyle Spence and \n          Chris Callison-Burch and \n          Morgan McGuire and \n          Victor Zordan},\n  title={{Learning When to Speak: Latency and Quality Trade-offs for Simultaneous Speech-to-Speech Translation with Offline Models}},\n  year=2023,\n  booktitle={Proc. INTERSPEECH 2023},\n  pages={5265--5266}\n}\n```\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fliamdugan%2Fspeech-to-speech","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fliamdugan%2Fspeech-to-speech","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fliamdugan%2Fspeech-to-speech/lists"}