{"id":13603566,"url":"https://ku-cvlab.github.io/MoDiTalker/","last_synced_at":"2025-04-11T22:31:37.612Z","repository":{"id":230027385,"uuid":"774830319","full_name":"cvlab-kaist/MoDiTalker","owner":"cvlab-kaist","description":null,"archived":false,"fork":false,"pushed_at":"2024-04-04T08:48:47.000Z","size":4693,"stargazers_count":159,"open_issues_count":11,"forks_count":11,"subscribers_count":8,"default_branch":"master","last_synced_at":"2024-10-28T16:39:47.490Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":null,"language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/cvlab-kaist.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null}},"created_at":"2024-03-20T09:15:01.000Z","updated_at":"2024-10-25T10:31:39.000Z","dependencies_parsed_at":"2024-04-04T09:36:42.961Z","dependency_job_id":null,"html_url":"https://github.com/cvlab-kaist/MoDiTalker","commit_stats":null,"previous_names":["ku-cvlab/moditalker","cvlab-kaist/moditalker"],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/cvlab-kaist%2FMoDiTalker","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/cvlab-kaist%2FMoDiTalker/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/cvlab-kaist%2FMoDiTalker/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/cvlab-kaist%2FMoDiTalker/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/cvlab-kaist","download_url":"https://codeload.github.com/cvlab-kaist/MoDiTalker/tar.gz/refs/heads/master","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":223483552,"owners_count":17152794,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2024-08-01T19:00:25.402Z","updated_at":"2025-04-11T22:31:37.582Z","avatar_url":"https://github.com/cvlab-kaist.png","language":"Python","funding_links":[],"categories":["Papers"],"sub_categories":["2D Video - Person independent"],"readme":"# MoDiTalker\n\nOfficial PyTorch implementation of **[\"MoDiTalker: Motion-Disentangled Diffusion\nModel for High-Fidelity Talking Head Generation\"](https://arxiv.org/abs/2403.19144)**.   \n\u003c!-- [Seyeon Kim](https://sihyun.me/)\u003csup\u003e*1\u003c/sup\u003e, \n[Siyoon Jin](https://sites.google.com/site/kihyuksml/)\u003csup\u003e*1\u003c/sup\u003e, \n[Jihye Park](https://subin-kim-cv.github.io/)\u003csup\u003e1\u003c/sup\u003e, \n[Kihong Kim](https://alinlab.kaist.ac.kr/shin.html)\u003csup\u003e2\u003c/sup\u003e,\n[Jiyoung Kim]()\u003csup\u003e1\u003c/sup\u003e,\n[Jisu Nam]()\u003csup\u003e1\u003c/sup\u003e and\n[Seungryong Kim]()\u003csup\u003e1\u003c/sup\u003e. --\u003e\nSeyeon Kim\u003csup\u003e\u0026#8727;1,2\u003c/sup\u003e, \nSiyoon Jin\u003csup\u003e\u0026#8727;1\u003c/sup\u003e, \nJihye Park\u003csup\u003e\u0026#8727;1,2\u003c/sup\u003e, \nKihong Kim\u003csup\u003e3\u003c/sup\u003e,\nJiyoung Kim\u003csup\u003e1\u003c/sup\u003e,\nJisu Nam\u003csup\u003e4\u003c/sup\u003e and\nSeungryong Kim\u003csup\u003e\u0026dagger;4\u003c/sup\u003e.\n\u003cbr\u003e\n\u003csup\u003e\u0026#8727;\u003c/sup\u003e Equal contribution ,\u003csup\u003e\u0026dagger;\u003c/sup\u003eCorresponding author\n\u003cbr\u003e\n\u003csup\u003e1\u003c/sup\u003eKorea University, \u003csup\u003e2\u003c/sup\u003eSamsung Electronics, \u003csup\u003e3\u003c/sup\u003eVIVE STUDIOS, \u003csup\u003e4\u003c/sup\u003eKAIST\n\n[paper](https://arxiv.org/abs/2403.19144) | [project page](https://cvlab-kaist.github.io/MoDiTalker/)\n\n\n## 1. Environment setup\n\n```bash\nconda create -n MoDiTalker python=3.8 -y\nconda activate MoDiTalker\npython -m pip install torch==1.12.1+cu116 torchvision==0.13.1+cu116 torchaudio==0.12.1 --extra-index-url https://download.pytorch.org/whl/cu116\npython -m pip install natsort tqdm gdown omegaconf einops lpips pyspng tensorboard imageio av moviepy numba p_tqdm soundfile face_alignemnt\n```\n\n## 2. Get ready to train models \n\n### 2.1. Dataset \n\u003c!-- Currently, we provide experiments for the following two datasets: [LRS3](path to lrs3 or geneface) and [HDTF](https://github.com/MRzzm/HDTF). Each dataset is used for training AToM and MToV, respectively. Please refer the README.md in `/data`. Each dataset should be placed in `/data` with the following structures below; --\u003e\nWe utilized two datasets for training each stage. \nPlease refer and follow the dataset preparation from [here](https://github.com/KU-CVLab/MoDiTalker/data/README.md)\n\n\n### 2.2. Download auxiliary models\n\u003c!-- Download  [this link](https://drive.google.com/file/d/1d08qauPUH0Nu_yN2gcmreLSiOiweD5OE/view?usp=sharing) --\u003e\nGet the [`BFM_model_front.mat`](https://drive.google.com/file/d/1d08qauPUH0Nu_yN2gcmreLSiOiweD5OE/view?usp=sharing), [`similarity_Lm3D_all.mat`](https://drive.google.com/file/d/17zp_zuUYAuieCWXerQkbp8SRSU4KJ8Fx/view?usp=sharing) and [`Exp_Pca.bin`](https://drive.google.com/file/d/1SPeJ4jcJT9VS4IdA7opzyGHCYMKuCLRh/view?usp=sharing), and place them to the `MoDiTalker/data/data_utils/deep_3drecon/BFM` directory.\nObtain ['BaselFaceModel.tgz](https://drive.google.com/file/d/1Kogpizrcf2zTm1fX9uUUWZuMQqHM7DOc/view?usp=sharing) and extract a file named `01_MorphableModel.mat` and place it to the `MoDiTalker/data/data_utils/deep_3drecon/BFM` directory.\n\n#### (Optional) \nWe had to revise a single code inside the package `accelerate` due to version conflicts. If some conflicts occur during loading data, please revise the code in `accelerate/dataloader.py` following this;\n\nfrom \n```bash\nbatch_size = dataloader.batch_size if dataloader.batch_size is not None else dataloader.batch_sampler.batch_size\n```\nto\n```bash\nbatch_size = dataloader.batch_size if dataloader.batch_size is not None else len(dataloader.batch_sampler[0])\n```\n\n\n\n## 3. Training\n\n### 3.1. AToM\n\n```bash\ncd AToM\nbash scripts/train.sh\n```\nThe checkpoints of AToM will be saved in `./runs`\n\n### 3.2. MToV\nThe checkpoints of AToM will be saved in `./runs`\n\n### Autoencoder\n\nFirst, execute the following script:\n```bash\ncd MToV\nbash scripts/train/first_stg.sh \n```\nThen the script will automatically create the folder in `./log_dir` to save logs and checkpoints.\n\nSecond, execute the following script:\n```bash\ncd MToV\nbash scripts/train/first_stg_ldmk.sh \n```\nYou may change the model configs via modifying `configs/autoencoder`. Moreover, one needs early-stopping to further train the model with the GAN loss (typically 8k-14k iterations with a batch size of 8).\n\n### Diffusion model\n```bash\ncd MToV\nbash scripts/train/second_stg.sh\n```\n\n\n### 4. Getting the Weights\nWe provide the corresponding checkpoints in the below:\nDownload and place them in the `./checkpoints/` directory. \n|              | Link to download | \n|--------------|-------------|\n| AToM     | [link](https://drive.google.com/file/d/1nxKqaScFIGuNoK5Y5Jq4zeWxJ55sUw_E/view?usp=share_link)  | \n| MToV motion autoencoder | [link](https://drive.google.com/file/d/1b9s5TbCaj5hz-luw7abgLYOF2Ph3jKrv/view?usp=share_link)  |\n| MToV RGB autoencoder | [link](https://drive.google.com/file/d/1KQ7XKl5HLP79Ri7A33VHfn7Tvnwz3v9o/view?usp=share_link)  |\n\u003c!-- Full checkpoints will be released later, ETA July 2024. --\u003e\n\n### 5. Inference\n#### 5.1. Generating Motions from Audio \nBefore producing the motions from audio, there's need to preprocess the audio since we process audio in the type of HuBeRT. To produce hubert feature of audio you want, please follow the script below:\n\n```bash\ncd data\npython data_utils/preprocess/process_audio.py \\\n--audio path to audio \\\n--ref_dir path to directory of reference images \n```\n\nThen the processed audio hubert(npy) will be saved in `data/inference/hubert/{sampling rate}` \n\nNote that, you need to specify the path to (1) reference images (2) processed hubert and (3) checkpoint in the following bash script. \n\n```bash\ncd AToM\nbash scripts/inference.sh\n```\n\nThe results of AToM will be saved in `AToM/results/frontalized_npy` and this path should be consistent with the `ldmk_path` of the following step.\n\n#### 5.2. Align Motions\nNote that, you need to specify the path to (1) reference images and (2) produced landmark. \n\n```bash \ncd data/data_utils\npython motion_align/align_face_recon.py \\\n--ldmk_path path to directory of generated landmark \\\n--driv_video_path path to directory of reference images \n```\nThe final landmarks will be saved in `AToM/results/aligned_npy`.\n\n#### 5.3. Generating Video from aligned Motions\n```bash \ncd MToV\nbash scripts/inference/sample.sh\n```\nThe final videos will be saved in `MToV/results`.\n\n\n\n### Citation\n```bibtex\n@misc{kim2024moditalker,\n      title={MoDiTalker: Motion-Disentangled Diffusion Model for High-Fidelity Talking Head Generation}, \n      author={Seyeon Kim and Siyoon Jin and Jihye Park and Kihong Kim and Jiyoung Kim and Jisu Nam and Seungryong Kim},\n      year={2024},\n      eprint={2403.19144},\n      archivePrefix={arXiv},\n      primaryClass={cs.CV}\n}\n```\n\n### Reference\nThis code is mainly built upon [EDGE](https://github.com/Stanford-TML/EDGE) and [PVDM](https://github.com/sihyun-yu/PVDM/tree/main).\\\nWe also used the code from following repository: [GeneFace](https://github.com/yerfor/GeneFace).\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/ku-cvlab.github.io%2FMoDiTalker%2F","html_url":"https://awesome.ecosyste.ms/projects/ku-cvlab.github.io%2FMoDiTalker%2F","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/ku-cvlab.github.io%2FMoDiTalker%2F/lists"}