{"id":29826008,"url":"https://github.com/antgroup/echomimic_v3","last_synced_at":"2026-02-08T14:31:32.221Z","repository":{"id":306222936,"uuid":"1025268363","full_name":"antgroup/echomimic_v3","owner":"antgroup","description":"[AAAI 2026] EchoMimicV3: 1.3B Parameters are All You Need for Unified Multi-Modal and Multi-Task Human Animation","archived":false,"fork":false,"pushed_at":"2026-01-30T05:22:16.000Z","size":53312,"stargazers_count":739,"open_issues_count":21,"forks_count":83,"subscribers_count":8,"default_branch":"main","last_synced_at":"2026-01-30T21:42:23.220Z","etag":null,"topics":["audio-driven-body-animation","audio-driven-portrait-animations","human-animation","video-generation"],"latest_commit_sha":null,"homepage":"https://antgroup.github.io/ai/echomimic_v3/","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/antgroup.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE.txt","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2025-07-24T02:20:55.000Z","updated_at":"2026-01-30T05:33:41.000Z","dependencies_parsed_at":"2025-07-24T12:27:19.432Z","dependency_job_id":"f480e045-f7db-4346-98e4-4c9d7eac0657","html_url":"https://github.com/antgroup/echomimic_v3","commit_stats":null,"previous_names":["antgroup/echomimic_v3"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/antgroup/echomimic_v3","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/antgroup%2Fechomimic_v3","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/antgroup%2Fechomimic_v3/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/antgroup%2Fechomimic_v3/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/antgroup%2Fechomimic_v3/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/antgroup","download_url":"https://codeload.github.com/antgroup/echomimic_v3/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/antgroup%2Fechomimic_v3/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":29233163,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-02-08T14:18:14.570Z","status":"ssl_error","status_checked_at":"2026-02-08T14:18:14.071Z","response_time":57,"last_error":"SSL_read: unexpected eof while reading","robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":false,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["audio-driven-body-animation","audio-driven-portrait-animations","human-animation","video-generation"],"created_at":"2025-07-29T04:34:28.444Z","updated_at":"2026-02-08T14:31:32.206Z","avatar_url":"https://github.com/antgroup.png","language":"Python","funding_links":[],"categories":["开源数字人"],"sub_categories":[],"readme":"[简体中文](https://github.com/antgroup/echomimic_v3/blob/main/README_zh.md) | English \n\u003cp align=\"center\"\u003e\n  \u003cimg src=\"asset/EchoMimicV3_logo.png.jpg\"  height=60\u003e\n\u003c/p\u003e\n\n\u003ch1 align='center'\u003eEchoMimicV3: 1.3B Parameters are All You Need for Unified Multi-Modal and Multi-Task Human Animation\u003c/h1\u003e\n\n\u003cdiv align='center'\u003e\n    \u003ca href='https://github.com/mengrang' target='_blank'\u003eRang Meng\u003c/a\u003e\u003csup\u003e1\u003c/sup\u003e\u0026emsp;\n    \u003ca href='https://github.com/' target='_blank'\u003eYan Wang\u003c/a\u003e\u0026emsp;\n    \u003ca href='https://github.com/' target='_blank'\u003eWeipeng Wu\u003c/a\u003e\u0026emsp;\n    \u003ca href='https://github.com/' target='_blank'\u003eRuobing Zheng\u003c/a\u003e\u0026emsp;\n    \u003ca href='https://lymhust.github.io/' target='_blank'\u003eYuming Li\u003c/a\u003e\u003csup\u003e2\u003c/sup\u003e\u0026emsp;\n    \u003ca href='https://openreview.net/profile?id=~Chenguang_Ma3' target='_blank'\u003eChenguang Ma\u003c/a\u003e\u003csup\u003e2\u003c/sup\u003e\n\u003c/div\u003e\n\u003cdiv align='center'\u003e\nTerminal Technology Department, Alipay, Ant Group.\n\u003c/div\u003e\n\u003cp align='center'\u003e\n    \u003csup\u003e1\u003c/sup\u003eCore Contributor\u0026emsp;\n    \u003csup\u003e2\u003c/sup\u003eCorresponding Authors\n\u003c/p\u003e\n\u003cdiv align='center'\u003e\n    \u003ca href='https://github.com/antgroup/echomimic_v3'\u003e\u003cimg src='https://img.shields.io/github/stars/antgroup/echomimic_v3?style=social'\u003e\u003c/a\u003e\n    \u003ca href='https://antgroup.github.io/ai/echomimic_v3/'\u003e\u003cimg src='https://img.shields.io/badge/Project-Page-blue'\u003e\u003c/a\u003e\n    \u003ca href='https://arxiv.org/abs/2507.03905'\u003e\u003cimg src='https://img.shields.io/badge/Paper-Arxiv-red'\u003e\u003c/a\u003e\n    \u003ca href='https://huggingface.co/BadToBest/EchoMimicV3'\u003e\u003cimg src='https://img.shields.io/badge/%F0%9F%A4%97%20HuggingFace-Model-yellow'\u003e\u003c/a\u003e\n    \u003ca href='https://modelscope.cn/models/BadToBest/EchoMimicV3'\u003e\u003cimg src='https://img.shields.io/badge/ModelScope-Model-purple'\u003e\u003c/a\u003e\n    \u003ca href='https://github.com/antgroup/echomimic_v3/blob/main/asset/wechat_group.png'\u003e\u003cimg src='https://badges.aleen42.com/src/wechat.svg'\u003e\u003c/a\u003e\n    \u003ca href='https://github.com/antgroup/echomimic_v3/discussions/18'\u003e\u003cimg src='https://img.shields.io/badge/中文版-常见问题汇总-orange'\u003e\u003c/a\u003e\n    \u003c!--\u003ca href='https://antgroup.github.io/ai/echomimic_v2/'\u003e\u003cimg src='https://img.shields.io/badge/%F0%9F%A4%97%20HuggingFace-Demo-yellow'\u003e\u003c/a\u003e--\u003e\n    \u003c!--\u003ca href='https://antgroup.github.io/ai/echomimic_v2/'\u003e\u003cimg src='https://img.shields.io/badge/ModelScope-Demo-purple'\u003e\u003c/a\u003e--\u003e\n    \u003c!-- \u003ca href='https://openaccess.thecvf.com/content/CVPR2025/papers/Meng_EchoMimicV2_Towards_Striking_Simplified_and_Semi-Body_Human_Animation_CVPR_2025_paper.pdf'\u003e\u003cimg src='https://img.shields.io/badge/Paper-CVPR2025-blue'\u003e\u003c/a\u003e --\u003e\n  \n\u003c/div\u003e\n\u003c!-- \u003cdiv align='center'\u003e\n    \u003ca href='https://github.com/antgroup/echomimic_v3/discussions/0'\u003e\u003cimg src='https://img.shields.io/badge/English-Common Problems-orange'\u003e\u003c/a\u003e\n    \u003ca href='https://github.com/antgroup/echomimic_v3/discussions/1'\u003e\u003cimg src='https://img.shields.io/badge/中文版-常见问题汇总-orange'\u003e\u003c/a\u003e\n\u003c/div\u003e --\u003e\n\n\u003cp align=\"center\"\u003e\n  \u003cimg src=\"asset/algo_framework.png\"  height=700\u003e\n\u003c/p\u003e\n\n## \u0026#x1F680; EchoMimic Series\n* EchoMimicV1: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditioning. [GitHub](https://github.com/antgroup/echomimic)\n* EchoMimicV2: Towards Striking, Simplified, and Semi-Body Human Animation. [GitHub](https://github.com/antgroup/echomimic_v2)\n* EchoMimicV3: 1.3B Parameters are All You Need for Unified Multi-Modal and Multi-Task Human Animation. [GitHub](https://github.com/antgroup/echomimic_v3)\n\n## \u0026#x1F4E3; Updates\n* [2026.01.22] 🔥 We update our EchoMimicV3-Flash on [Huggingface](https://huggingface.co/BadToBest/EchoMimicV3/tree/main/echomimicv3-flash-pro).\n  - 🚀 8-step High-quality Generation.\n  - 🧩 No Face Mask required.\n  - 💾 12G VRAM Requirement.\n  - ✅ Supports up to 768×768 Resolution.\n* [2025.11.09] 🔥 EchoMimicV3 is accepted by AAAI 2026.\n* [2025.08.21] 🔥 EchoMimicV3 gradio demo on [modelscope](https://modelscope.cn/studios/BadToBest/EchoMimicV3) is ready.\n* [2025.08.12] 🔥🚀 **12G VRAM is All YOU NEED to Generate Video**. Please use this [GradioUI](https://github.com/antgroup/echomimic_v3/blob/main/app_mm.py). Check the [tutorial](https://www.bilibili.com/video/BV1W8tdzEEVN) from @[gluttony-10](https://github.com/gluttony-10). Thanks for the contribution.\n* [2025.08.12] 🔥 EchoMimicV3 can run on **16G VRAM** using [ComfyUI](https://github.com/smthemex/ComfyUI_EchoMimic). Thanks @[smthemex](https://github.com/smthemex) for the contribution.\n* [2025.08.09] 🔥 We release our [models](https://modelscope.cn/models/BadToBest/EchoMimicV3) on ModelScope.\n* [2025.08.08] 🔥 We release our [codes](https://github.com/antgroup/echomimic_v3) on GitHub and [models](https://huggingface.co/BadToBest/EchoMimicV3) on Huggingface.\n* [2025.07.08] 🔥 Our [paper](https://arxiv.org/abs/2507.03905) is in public on arxiv.\n\n## \u0026#x1F305; Gallery\n\u003cp align=\"center\"\u003e\n  \u003cimg src=\"asset/echomimicv3.jpg\"  height=1000\u003e\n\u003c/p\u003e\n\u003ctable class=\"center\"\u003e\n\u003ctr\u003e\n    \u003ctd width=100% style=\"border: none\"\u003e\n        \u003cvideo controls loop src=\"https://github.com/user-attachments/assets/f33edb30-66b1-484b-8be0-a5df20a44f3b\" muted=\"false\"\u003e\u003c/video\u003e\n    \u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n    \u003ctd width=100% style=\"border: none\"\u003e\n        \u003cvideo controls loop src=\"https://github.com/user-attachments/assets/056105d8-47cd-4a78-8ec2-328ceaf95a5a\" muted=\"false\"\u003e\u003c/video\u003e\n    \u003c/td\u003e\n\u003c/tr\u003e\n\u003c/table\u003e\n\n### Chinese Driven Audio\n\u003ctable class=\"center\"\u003e\n\u003ctr\u003e\n    \u003ctd width=25% style=\"border: none\"\u003e\n        \u003cvideo controls loop src=\"https://github.com/user-attachments/assets/fc1ebae4-b571-43eb-a13a-7d6d05b74082\" muted=\"false\"\u003e\u003c/video\u003e\n    \u003c/td\u003e\n    \u003ctd width=25% style=\"border: none\"\u003e\n        \u003cvideo controls loop src=\"https://github.com/user-attachments/assets/54607cc7-944c-4529-9bef-715862ba330d\" muted=\"false\"\u003e\u003c/video\u003e\n    \u003c/td\u003e\n    \u003ctd width=25% style=\"border: none\"\u003e\n        \u003cvideo controls loop src=\"https://github.com/user-attachments/assets/4d1de999-cce2-47ab-89ed-f2fa11c838fe\" muted=\"false\"\u003e\u003c/video\u003e\n    \u003c/td\u003e\n    \u003ctd width=25% style=\"border: none\"\u003e\n        \u003cvideo controls loop src=\"https://github.com/user-attachments/assets/41e701cc-ac3e-4dd8-b94c-859261f17344\" muted=\"false\"\u003e\u003c/video\u003e\n    \u003c/td\u003e\n\u003c/tr\u003e\n\u003c/table\u003e\n\nFor more demo videos, please refer to the [project page](https://antgroup.github.io/ai/echomimic_v3/)\n\n## Quick Start\n### Environment Setup\n- Tested System Environment: Centos 7.2/Ubuntu 22.04, Cuda \u003e= 12.1\n- Tested GPUs: A100(80G) / RTX4090D (24G) / V100(16G)\n- Tested Python Version: 3.10 / 3.11\n  \n### 🛠️Installation for Windows\n\n##### Please use the [one-click installation package](https://pan.baidu.com/share/init?surl=cV7i2V0wF4exDtKjJrAUeA) (passport: glut) to get started quickly for Quantified version.\n\n### 🛠️Installation for Linux\n#### 1. Create a conda environment\n```\nconda create -n echomimic_v3 python=3.10\nconda activate echomimic_v3\n```\n\n#### 2. Other dependencies\n```\npip install -r requirements.txt\n```\n### 🧱Model Preparation\n\n| Models        |                       Download Link                                           |    Notes                      |\n| --------------|-------------------------------------------------------------------------------|-------------------------------|\n| Wan2.1-Fun-V1.1-1.3B-InP  |      🤗 [Huggingface](https://huggingface.co/alibaba-pai/Wan2.1-Fun-V1.1-1.3B-InP)       | Base model\n| wav2vec2-base |      🤗 [Huggingface](https://huggingface.co/facebook/wav2vec2-base-960h)          | Audio encoder for preview\n| chinese-wav2vec2-base |      🤗 [Huggingface](https://modelscope.cn/models/TencentGameMate/chinese-wav2vec2-base)          | Audio encoder for Flash\n| EchoMimicV3-preview      |      🤗 [Huggingface](https://huggingface.co/BadToBest/EchoMimicV3)              | preview weights\n| EchoMimicV3-preview      |      🤗 [ModelScope](https://modelscope.cn/models/BadToBest/EchoMimicV3)              | preview weights\n| EchoMimicV3-Flash      |      🤗 [Huggingface](https://huggingface.co/BadToBest/EchoMimicV3/tree/main/echomimicv3-flash-pro)              | Flash weights\n\n-- The **weights** of EchoMimicV3-flash-pro is organized as follows.\n\n```\n./flash-pro/\n├── Wan2.1-Fun-V1.1-1.3B-InP\n├── chinese-wav2vec2-base\n└── transformer\n    └── diffusion_pytorch_model.safetensors\n```\n\n-- The **weights** is of EchoMimicV3-preview organized as follows.\n\n```\n./preview/\n├── Wan2.1-Fun-V1.1-1.3B-InP\n├── wav2vec2-base-960h\n└── transformer\n    └── diffusion_pytorch_model.safetensors\n``` \n### 🔑 Quick Inference for EchoMimicV3-flash-pro\n```\nbash run_flash_pro.sh\n```\n### 🔑 Quick Inference for EchoMimicV3-preview\n```\npython infer_preview.py\n```\nFor Quantified GradioUI version for EchoMimicV3-preview:\n```\npython app_mm.py\n```\n**images, audios, masks and prompts are provided in `datasets/echomimicv3_demos`**\n\n#### Tips\n- Audio CFG: Audio CFG `audio_guidance_scale` works optimally between 1.8~2. Increase the audio CFG value for better lip synchronization, while decreasing the audio CFG value can improve the visual quality.\n- Text CFG: Text CFG `guidance_scale` works optimally between 3~6. Increase the text CFG value for better prompt following, while decreasing the text CFG value can improve the visual quality.\n- TeaCache: The optimal range for `teacache_threshold` is between 0~0.1.\n- Sampling steps: 5 steps for talking head, 15~25 steps for talking body. \n- ​Long video generation: If you want to generate a video longer than 138 frames, you can use Long Video CFG.\n- Try setting `partial_video_length` to 81, 65 or smaller to reduce VRAM usage.\n\n## \u0026#x1F4D2; Citation\n\nIf you find our work useful for your research, please consider citing the paper :\n\n```\n@misc{meng2025echomimicv3,\n  title={EchoMimicV3: 1.3B Parameters are All You Need for Unified Multi-Modal and Multi-Task Human Animation},\n  author={Rang Meng, Yan Wang, Weipeng Wu, Ruobing Zheng, Yuming Li, Chenguang Ma},\n  year={2025},\n  eprint={2507.03905},\n  archivePrefix={arXiv}\n}\n```\n## Reference\n- Wan2.1: https://github.com/Wan-Video/Wan2.1/\n- VideoX-Fun: https://github.com/aigc-apps/VideoX-Fun/\n## 📜 License\nThe models in this repository are licensed under the Apache 2.0 License. We claim no rights over the your generated contents, \ngranting you the freedom to use them while ensuring that your usage complies with the provisions of this license. \nYou are fully accountable for your use of the models, which must not involve sharing any content that violates applicable laws, \ncauses harm to individuals or groups, disseminates personal information intended for harm, spreads misinformation, or targets vulnerable populations. \n\n\n## \u0026#x1F31F; Star History\n[![Star History Chart](https://api.star-history.com/svg?repos=antgroup/echomimic_v3\u0026type=Date)](https://www.star-history.com/#antgroup/echomimic_v3\u0026Date)\n\n\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fantgroup%2Fechomimic_v3","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fantgroup%2Fechomimic_v3","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fantgroup%2Fechomimic_v3/lists"}