{"id":31034633,"url":"https://github.com/freedomintelligence/echox","last_synced_at":"2025-09-14T02:46:33.069Z","repository":{"id":313952463,"uuid":"963288603","full_name":"FreedomIntelligence/EchoX","owner":"FreedomIntelligence","description":"EchoX: Towards Mitigating Acoustic-Semantic Gap via Echo Training for Speech-to-Speech LLMs","archived":false,"fork":false,"pushed_at":"2025-09-09T16:36:33.000Z","size":5736,"stargazers_count":3,"open_issues_count":0,"forks_count":0,"subscribers_count":12,"default_branch":"main","last_synced_at":"2025-09-09T19:55:25.303Z","etag":null,"topics":["ai","artificial-intelligence","dialogue-systems","llm","s2s","speech-to-speech"],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/FreedomIntelligence.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2025-04-09T13:00:12.000Z","updated_at":"2025-09-09T16:36:36.000Z","dependencies_parsed_at":"2025-09-09T19:55:28.207Z","dependency_job_id":"4d2bc455-8850-4ba9-8077-9b58a0fb0f70","html_url":"https://github.com/FreedomIntelligence/EchoX","commit_stats":null,"previous_names":["freedomintelligence/echox"],"tags_count":null,"template":false,"template_full_name":null,"purl":"pkg:github/FreedomIntelligence/EchoX","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/FreedomIntelligence%2FEchoX","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/FreedomIntelligence%2FEchoX/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/FreedomIntelligence%2FEchoX/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/FreedomIntelligence%2FEchoX/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/FreedomIntelligence","download_url":"https://codeload.github.com/FreedomIntelligence/EchoX/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/FreedomIntelligence%2FEchoX/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":275054971,"owners_count":25397576,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","status":"online","status_checked_at":"2025-09-14T02:00:10.474Z","response_time":75,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["ai","artificial-intelligence","dialogue-systems","llm","s2s","speech-to-speech"],"created_at":"2025-09-14T02:46:32.040Z","updated_at":"2025-09-14T02:46:33.060Z","avatar_url":"https://github.com/FreedomIntelligence.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"# EchoX: Towards Mitigating Acoustic-Semantic Gap via Echo Training for Speech-to-Speech LLMs\n\n\u003c!-- \u003e **EchoX: Towards Mitigating Acoustic-Semantic Gap via Echo Training for Speech-to-Speech LLMs**  \n\u003e Yuhao Zhang, Yuhao Du, Zhanchen Dai, et al. — *Under review at ICLR 2026*  \n\u003e 📄 [Paper (to be added)](https://arxiv.org/abs/XXXX.XXXX) --\u003e\n\n\u003cp align=\"center\"\u003e\n  \u003cimg src=\"https://img.shields.io/badge/License-Apache2.0-blue.svg\" alt=\"License\"\u003e\n  \u003cimg src=\"https://img.shields.io/badge/Python-3.10+-green.svg\" alt=\"Python\"\u003e\n  \u003cimg src=\"https://img.shields.io/badge/Model-8B%7C3B-orange.svg\" alt=\"Model Size\"\u003e\n\u003c/p\u003e\n\n\u003cp align=\"center\"\u003e\n   📄 \u003ca href=\"https://arxiv.org/abs/2509.09174\"\u003ePaper\u003c/a\u003e |\n   📦 \u003ca href=\"https://huggingface.co/FreedomIntelligence/EchoX-8B\"\u003eModel\u003c/a\u003e | \n   🚀 \u003ca href=\"https://huggingface.co/spaces/FreedomIntelligence/EchoX\"\u003eHF Space\u003c/a\u003e | \n   🌐 \u003ca href=\"https://freedomintelligence.github.io/EchoX\"\u003eWeb Demo\u003c/a\u003e | \n   📊 \u003ca href=\"https://huggingface.co/datasets/FreedomIntelligence/EchoX-Dialougues\"\u003eEchoX-Dialougues\u003c/a\u003e | \n   📊 \u003ca href=\"https://huggingface.co/datasets/KurtDu/EchoX-Dialogues-Plus\"\u003eEchoX-Dialogues-Plus\u003c/a\u003e\n\u003c/p\u003e\n\n\n\n## Contents\n- [EchoX: Towards Mitigating Acoustic-Semantic Gap via Echo Training for Speech-to-Speech LLMs](#echox-towards-mitigating-acoustic-semantic-gap-via-echo-training-for-speech-to-speech-llms)\n  - [Contents](#contents)\n  - [Key Features](#key-features)\n  - [Performance](#performance)\n  - [Datasets and Models](#datasets-and-models)\n    - [Dataset](#dataset)\n    - [Model](#model)\n  - [Quickstart](#quickstart)\n    - [Environment Setup](#environment-setup)\n    - [Model Download](#model-download)\n    - [Inference](#inference)\n  - [Citation](#citation)\n  - [License](#license)\n\n## Key Features\n- Mitigates Acoustic-Semantic Gap in Speech-to-Speech LLMs\n- Introduces Echo Training with a Novel Three-Stage Pipeline (S2T, T2C, Echo)\n- Trained on Only 6k Hours of Curated Data, Ensuring Efficiency\n- Achieves State-of-the-Art Performance in Knowledge-Based QA Benchmarks\n- Preserves Reasoning and Knowledge Abilities for Interactive Speech Tasks\n\n## Performance\n\n\u003cp align=\"center\"\u003e\n  \u003cimg src=\"asset/performance.png\" alt=\"Performance\" style=\"width:50%;\"\u003e\n\u003c/p\u003e\n\nEchoX demonstrates exceptional performance on knowledge-based question-answering tasks. The model achieves superior results with minimal training data, establishing a new benchmark for efficiency in speech-to-speech language models.\n\n## Datasets and Models\n\n### Dataset\nEchoX is trained on carefully curated datasets for each stage of the pipeline, ensuring optimal performance across ASR, TTS, and SQA tasks. The datasets used are as follows:\n\n| Task      | Data                | Size        | Duration(H) | Stage  | Download                                                                    |\n| :-------- | :------------------ | :---------- | :---------- | :----- | :-------------------------------------------------------------------------- |\n| ASR       | LibriSpeech         | 281,241     | 960         | I      | -                                                                           |\n| ASR       | MLS                 | 723,636     | 3,000       | I      | -                                                                           |\n| TTS       | AudioQA-1M          | 178,576     | 989         | II     | -                                                                           |\n| TTS       | SpeechInstruct      | 31,563      | 84          | II     | -                                                                           |\n| TTS       | HH-RLHF-Speech      | 124,945     | 656         | II     | -                                                                           |\n| SQA       | sharechatx          | 43,223      | 178         | I, III | [Link](https://huggingface.co/datasets/KurtDu/EchoX-Dialogues) |\n| SQA       | Magpie-Pro-Speech+   | 117,000     | 327         | I, III | [Link](https://huggingface.co/datasets/KurtDu/EchoX-Dialogues) |\n| **Total** |                     | **1,500,184** | **6,194**   |        |                                                                             |\n\n### Model\nThe following pre-trained models are available for download:\n\n| Model        | Parameters | Training Data | Download Link                                      |\n| ------------ | ---------- | ------------- | -------------------------------------------------- |\n| **EchoX-3B** | 3 billion  | 6k hours     | [EchoX-3B Model](https://huggingface.co/FreedomIntelligence/EchoX-3B) |\n| **EchoX-8B** | 8 billion  | 6k hours     | [EchoX-8B Model](https://huggingface.co/FreedomIntelligence/EchoX-8B) |\n\n## Quickstart\n\n### Environment Setup\nTo set up your environment, follow these steps:\n```bash\ngit clone https://github.com/FreedomIntelligence/EchoX.git\ncd EchoX\nconda create -n echox python=3.10 pip=24.0\nconda activate echox\npip install -r requirements.txt\n```\n\n### Model Download\nDownload the models to this repository directory using the following commands:\n\n```bash\npip install -U huggingface_hub\nhf download --resume-download FreedomIntelligence/EchoX-8B --local-dir EchoX-8B\nhf download --resume-download openai/whisper-large-v3 --local-dir whisper-large-v3\n```\n\n**Note**: If the models are downloaded to a different location, please update the model directory paths in [inference/echox_stream.py](inference/echox_stream.py) accordingly.\n\n### Inference\nRun inference on a test case:\n```bash\npython demo.py\n```\n\nAlternatively, start the Gradio web interface:\n```bash\npython app.py\n```\n\nTo use a specific GPU:\n```bash\nCUDA_VISIBLE_DEVICES=1 python app.py\n```\n\n## Citation\nIf you use EchoX in your research or projects, please cite our paper:\n\n```bibtex\n@misc{zhang2025echoxmitigatingacousticsemanticgap,\n      title={EchoX: Towards Mitigating Acoustic-Semantic Gap via Echo Training for Speech-to-Speech LLMs}, \n      author={Yuhao Zhang and Yuhao Du and Zhanchen Dai and Xiangnan Ma and Kaiqi Kou and Benyou Wang and Haizhou Li},\n      year={2025},\n      eprint={2509.09174},\n      archivePrefix={arXiv},\n      primaryClass={cs.CL},\n      url={https://arxiv.org/abs/2509.09174}, \n}\n```\n\n## License\nThis project is licensed under the Apache 2.0 License. See the [LICENSE](LICENSE) file for details.\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Ffreedomintelligence%2Fechox","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Ffreedomintelligence%2Fechox","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Ffreedomintelligence%2Fechox/lists"}