{"id":28409091,"url":"https://github.com/harlanhong/actalker","last_synced_at":"2026-01-29T16:36:47.926Z","repository":{"id":285569622,"uuid":"951636984","full_name":"harlanhong/ACTalker","owner":"harlanhong","description":"ICCV 2025 ACTalker: an end-to-end video diffusion framework for talking head synthesis that supports both single and multi-signal control (e.g., audio, expression).","archived":false,"fork":false,"pushed_at":"2025-08-20T22:16:56.000Z","size":131284,"stargazers_count":374,"open_issues_count":7,"forks_count":39,"subscribers_count":44,"default_branch":"master","last_synced_at":"2025-08-21T00:22:29.919Z","etag":null,"topics":["avatar","diffusion-models","digitalhuman","face-animation","multi-modal","stablevideodiffusion","talking-head"],"latest_commit_sha":null,"homepage":"https://harlanhong.github.io/publications/actalker/index.html","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/harlanhong.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null}},"created_at":"2025-03-20T02:08:54.000Z","updated_at":"2025-08-20T22:17:00.000Z","dependencies_parsed_at":null,"dependency_job_id":"b79dff11-aced-4d3a-8485-5ae835ac0c73","html_url":"https://github.com/harlanhong/ACTalker","commit_stats":null,"previous_names":["harlanhong/actalker"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/harlanhong/ACTalker","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/harlanhong%2FACTalker","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/harlanhong%2FACTalker/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/harlanhong%2FACTalker/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/harlanhong%2FACTalker/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/harlanhong","download_url":"https://codeload.github.com/harlanhong/ACTalker/tar.gz/refs/heads/master","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/harlanhong%2FACTalker/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":28880980,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-01-29T10:31:27.438Z","status":"ssl_error","status_checked_at":"2026-01-29T10:31:01.017Z","response_time":59,"last_error":"SSL_connect returned=1 errno=0 peeraddr=140.82.121.6:443 state=error: unexpected eof while reading","robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":false,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["avatar","diffusion-models","digitalhuman","face-animation","multi-modal","stablevideodiffusion","talking-head"],"created_at":"2025-06-02T05:19:44.884Z","updated_at":"2026-01-29T16:36:47.906Z","avatar_url":"https://github.com/harlanhong.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"\n## :book: Audio-visual Controlled Video Diffusion with Masked Selective State Spaces Modeling for Natural Talking Head Generation\n\n\n\u003e [[Paper](https://arxiv.org/abs/2504.02542)] \u0026emsp; [[Project Page](https://harlanhong.github.io/publications/actalker/index.html)]  \u0026emsp; [[HuggingFace](https://huggingface.co/papers/2504.02542)]\u003cbr\u003e\n\u003c!-- \u003e [Fa-Ting Hong](https://harlanhong.github.io), [Longhao Zhang](https://dblp.org/pid/236/7382.html), [Li Shen](https://scholar.google.co.uk/citations?user=ABbCaxsAAAAJ\u0026hl=en), [Dan Xu](https://www.danxurgb.net) \u003cbr\u003e --\u003e\n\u003c!-- \u003e The Hong Kong University of Science and Technology, Alibaba Cloud --\u003e\n\u003e [Fa-Ting Hong](https://harlanhong.github.io)\u003csup\u003e1,2\u003c/sup\u003e, Zunnan Xu\u003csup\u003e2,3\u003c/sup\u003e, Zixiang Zhou\u003csup\u003e2\u003c/sup\u003e, Jun Zhou\u003csup\u003e2\u003c/sup\u003e, Xiu Li\u003csup\u003e3\u003c/sup\u003e, Qin Lin\u003csup\u003e2\u003c/sup\u003e, Qinglin Lu\u003csup\u003e2\u003c/sup\u003e, [Dan Xu](https://www.danxurgb.net)\u003csup\u003e1\u003c/sup\u003e \u003cbr\u003e\n\u003e \u003csup\u003e1\u003c/sup\u003eThe Hong Kong University of Science and Technology\u003cbr\u003e\n\u003e \u003csup\u003e2\u003c/sup\u003eTencent\u003cbr\u003e\n\u003e \u003csup\u003e3\u003c/sup\u003eTsinghua University\n\n\u003cimg src=\"assets/teaser_compressed.jpg\"\u003e\n:triangular_flag_on_post: **Updates**  \n\n\u0026#9745; arXiv paper is released [here](https://arxiv.org/abs/2504.02542) !\n\n## Framework \n\u003cimg src=\"assets/framework.png\"\u003e\n\n## TL;DR:\nWe propose ACTalker, an end-to-end video diffusion framework for talking head synthesis that supports both single and multi-signal control (e.g., audio, pose, expression). ACTalker uses a parallel mamba-based architecture with a gating mechanism to assign different control signals to specific facial regions, ensuring fine-grained and conflict-free generation. A mask-drop strategy further enhances regional independence and control stability. Experiments show that ACTalker produces natural, synchronized talking head videos under various control combinations.\n\n## Expression Driven Samples\nhttps://github.com/user-attachments/assets/fc46c7cd-d1b4-44a6-8649-2ef973107637\n\n## Audio Dirven Samples\nhttps://github.com/user-attachments/assets/8f9e18a0-6fff-4a31-bbf4-c21702d4da38\n\n## Audio-Visual Driven Samples\nhttps://github.com/user-attachments/assets/3d8af4ef-edc7-4971-87b6-7a9c77ee0cb2\n\nhttps://github.com/user-attachments/assets/2d12defd-de3d-4a33-8178-b5af30d7f0c2\n\n\n### :e-mail: Contact\n\nIf you have any question or collaboration need (research purpose or commercial purpose), please email `fhongac@connect.ust.hk`.\n\n# 📍Citation \nPlease feel free to leave a star⭐️⭐️⭐️ and cite our paper:\n```bibtex\n@article{hong2025audio,\n  title={Audio-visual controlled video diffusion with masked selective state spaces modeling for natural talking head generation},\n  author={Hong, Fa-Ting and Xu, Zunnan and Zhou, Zixiang and Zhou, Jun and Li, Xiu and Lin, Qin and Lu, Qinglin and Xu, Dan},\n  journal={arXiv preprint arXiv:2504.02542},\n  year={2025}\n}\n``` \n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fharlanhong%2Factalker","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fharlanhong%2Factalker","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fharlanhong%2Factalker/lists"}