{"id":28522726,"url":"https://github.com/pku-yuangroup/consisid","last_synced_at":"2025-07-06T02:32:48.791Z","repository":{"id":264729102,"uuid":"888977326","full_name":"PKU-YuanGroup/ConsisID","owner":"PKU-YuanGroup","description":"[CVPR 2025 Highlight🔥] Identity-Preserving Text-to-Video Generation by Frequency Decomposition","archived":false,"fork":false,"pushed_at":"2025-06-08T13:40:21.000Z","size":13916,"stargazers_count":706,"open_issues_count":31,"forks_count":35,"subscribers_count":12,"default_branch":"main","last_synced_at":"2025-06-08T14:31:06.666Z","etag":null,"topics":["diffusion","diffusion-models","identity-preserving","text-to-video","video-generation","video-generation-dataset","video-generator","videogeneration"],"latest_commit_sha":null,"homepage":"https://pku-yuangroup.github.io/ConsisID/","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/PKU-YuanGroup.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null}},"created_at":"2024-11-15T11:19:39.000Z","updated_at":"2025-06-08T13:37:32.000Z","dependencies_parsed_at":"2024-12-24T04:28:59.817Z","dependency_job_id":"1a22cc44-a301-42f9-908e-767767483d80","html_url":"https://github.com/PKU-YuanGroup/ConsisID","commit_stats":null,"previous_names":["pku-yuangroup/consisid"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/PKU-YuanGroup/ConsisID","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/PKU-YuanGroup%2FConsisID","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/PKU-YuanGroup%2FConsisID/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/PKU-YuanGroup%2FConsisID/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/PKU-YuanGroup%2FConsisID/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/PKU-YuanGroup","download_url":"https://codeload.github.com/PKU-YuanGroup/ConsisID/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/PKU-YuanGroup%2FConsisID/sbom","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":263838708,"owners_count":23518140,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["diffusion","diffusion-models","identity-preserving","text-to-video","video-generation","video-generation-dataset","video-generator","videogeneration"],"created_at":"2025-06-09T09:31:02.841Z","updated_at":"2025-07-06T02:32:48.785Z","avatar_url":"https://github.com/PKU-YuanGroup.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"\u003cdiv align=center\u003e\n\u003cimg src=\"https://github.com/PKU-YuanGroup/ConsisID/blob/main/asserts/ConsisID_logo.png?raw=true\" width=\"150px\"\u003e\n\u003c/div\u003e\n\u003ch2 align=\"center\"\u003e \u003ca href=\"https://arxiv.org/abs/2411.17440\"\u003e[CVPR 2025 Highlight] Identity-Preserving Text-to-Video Generation by Frequency Decomposition\u003c/a\u003e\u003c/h2\u003e\n\n\u003ch5 align=\"center\"\u003e If you like our project, please give us a star ⭐ on GitHub for the latest update.  \u003c/h2\u003e\n\n\n\u003ch5 align=\"center\"\u003e\n\n\n[![hf_space](https://img.shields.io/badge/🤗-Open%20In%20Spaces-blue.svg)](https://huggingface.co/spaces/BestWishYsh/ConsisID-preview-Space)\n[![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/camenduru/ConsisID-jupyter/blob/main/ConsisID_jupyter.ipynb)\n[![hf_paper](https://img.shields.io/badge/🤗-Paper%20In%20HF-red.svg)](https://huggingface.co/papers/2411.17440)\n[![arXiv](https://img.shields.io/badge/Arxiv-2411.17440-b31b1b.svg?logo=arXiv)](https://arxiv.org/abs/2411.17440) \n[![Home Page](https://img.shields.io/badge/Project-\u003cWebsite\u003e-blue.svg)](https://pku-yuangroup.github.io/ConsisID/) \n[![Dataset](https://img.shields.io/badge/Dataset-previewData-green)](https://huggingface.co/datasets/BestWishYsh/ConsisID-preview-Data)\n[![zhihu](https://img.shields.io/badge/-Twitter@Adina%20Yakup%20-black?logo=twitter\u0026logoColor=1D9BF0)](https://x.com/AdinaYakup/status/1862604191631573122)\n[![zhihu](https://img.shields.io/badge/-Twitter@camenduru%20-black?logo=twitter\u0026logoColor=1D9BF0)](https://x.com/camenduru/status/1861957812152078701)\n[![zhihu](https://img.shields.io/badge/-YouTube-000000?logo=youtube\u0026logoColor=FF0000)](https://www.youtube.com/watch?v=PhlgC-bI5SQ)\n[![License](https://img.shields.io/badge/License-Apache%202.0-yellow)](https://github.com/PKU-YuanGroup/ConsisID/blob/main/LICENSE) \n[![github](https://img.shields.io/github/stars/PKU-YuanGroup/ConsisID.svg?style=social)](https://github.com/PKU-YuanGroup/ConsisID/)\n\n\u003c/h5\u003e\n\n\u003cdiv align=\"center\"\u003e\nThis repository is the official implementation of ConsisID, a tuning-free DiT-based controllable IPT2V model to keep human-identity consistent in the generated video. The approach draws inspiration from previous studies on frequency analysis of vision/diffusion transformers.\n\u003c/div\u003e\n\n\u003cbr\u003e\n\n\u003cdetails open\u003e\u003csummary\u003e💡 We also have other video generation projects that may interest you ✨. \u003c/summary\u003e\u003cp\u003e\n\u003c!--  may --\u003e\n\n\u003e [**Open-Sora Plan: Open-Source Large Video Generation Model**](https://arxiv.org/abs/2412.00131) \u003cbr\u003e\n\u003e Bin Lin, Yunyang Ge and Xinhua Cheng etc. \u003cbr\u003e\n[![github](https://img.shields.io/badge/-Github-black?logo=github)](https://github.com/PKU-YuanGroup/Open-Sora-Plan)  [![github](https://img.shields.io/github/stars/PKU-YuanGroup/Open-Sora-Plan.svg?style=social)](https://github.com/PKU-YuanGroup/Open-Sora-Plan) [![arXiv](https://img.shields.io/badge/Arxiv-2412.00131-b31b1b.svg?logo=arXiv)](https://arxiv.org/abs/2412.00131) \u003cbr\u003e\n\u003e\n\u003e [**OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation**](https://arxiv.org/abs/2505.20292) \u003cbr\u003e\n\u003e Shenghai Yuan, Xianyi He and Yufan Deng etc. \u003cbr\u003e\n\u003e [![github](https://img.shields.io/badge/-Github-black?logo=github)](https://github.com/PKU-YuanGroup/OpenS2V-Nexus)  [![github](https://img.shields.io/github/stars/PKU-YuanGroup/OpenS2V-Nexus.svg?style=social)](https://github.com/PKU-YuanGroup/OpenS2V-Nexus) [![arXiv](https://img.shields.io/badge/Arxiv-2505.20292-b31b1b.svg?logo=arXiv)](https://arxiv.org/abs/2505.20292) \u003cbr\u003e\n\u003e\n\u003e [**MagicTime: Time-lapse Video Generation Models as Metamorphic Simulators**](https://arxiv.org/abs/2404.05014) \u003cbr\u003e\n\u003e Shenghai Yuan, Jinfa Huang and Yujun Shi etc. \u003cbr\u003e\n\u003e [![github](https://img.shields.io/badge/-Github-black?logo=github)](https://github.com/PKU-YuanGroup/MagicTime)  [![github](https://img.shields.io/github/stars/PKU-YuanGroup/MagicTime.svg?style=social)](https://github.com/PKU-YuanGroup/MagicTime) [![arXiv](https://img.shields.io/badge/Arxiv-2404.05014-b31b1b.svg?logo=arXiv)](https://arxiv.org/abs/2404.05014) \u003cbr\u003e\n\u003e\n\u003e [**ChronoMagic-Bench: A Benchmark for Metamorphic Evaluation of Text-to-Time-lapse Video Generation**](https://arxiv.org/abs/2406.18522) \u003cbr\u003e\n\u003e Shenghai Yuan, Jinfa Huang and Yongqi Xu etc. \u003cbr\u003e\n\u003e [![github](https://img.shields.io/badge/-Github-black?logo=github)](https://github.com/PKU-YuanGroup/ChronoMagic-Bench/)  [![github](https://img.shields.io/github/stars/PKU-YuanGroup/ChronoMagic-Bench.svg?style=social)](https://github.com/PKU-YuanGroup/ChronoMagic-Bench/) [![arXiv](https://img.shields.io/badge/Arxiv-2406.18522-b31b1b.svg?logo=arXiv)](https://arxiv.org/abs/2406.18522) \u003cbr\u003e\n\u003e \u003c/p\u003e\u003c/details\u003e\n\n\n## 📣 News\n\n* ⏳⏳⏳ Release the full code \u0026 datasets \u0026 weights.\n* `[2025.05.27]`  🔥 Introducing [**​​OpenS2V-Nexus**​​](https://github.com/PKU-YuanGroup/OpenS2V-Nexus), which consists of: (i) ​​[**OpenS2V-Eval**​​](https://huggingface.co/spaces/BestWishYsh/OpenS2V-Eval), a fine-grained benchmark, and (ii) [​​**OpenS2V-5M​​**](https://huggingface.co/datasets/BestWishYsh/OpenS2V-5M), a million-scale dataset. Welcome to try it!\n* `[2025.04.04]`  🔥 Breaking news! ConsisID has been recommended as **CVPR Highlight**.\n* `[2025.03.27]`  🔥 We have updated our technical report. Please click [here](https://arxiv.org/abs/2411.17440v3) to view it. \n* `[2025.02.27]`  🔥 ConsisID has been accepted by CVPR 2025, and we will update arXiv with more details soon. Stay tuned!\n* `[2025.02.16]`  🔥 We have adapted the code for CogVideoX1.5, and you can use our code not only for training ConsisID but also for the CogVideoX-series.\n* `[2025.01.19]`  🤗 Thanks [@arrow](https://github.com/a-r-r-o-w), [@yiyixuxu](https://github.com/yiyixuxu), [@hlky](https://github.com/hlky) and [@stevhliu](https://github.com/stevhliu), ConsisID will be merged into [diffusers](https://huggingface.co/docs/diffusers/main/en/using-diffusers/consisid#identity-preserving-text-to-video) in `0.33.0`. So for now, please use `pip install git+https://github.com/huggingface/diffusers.git` to install diffusers dev version. And we have reorganized the code and weight configs, so it's better to update your local files if you have cloned them previously.\n* `[2024.12.26]`  🚀 We release the [cache inference code](https://github.com/PKU-YuanGroup/ConsisID/tree/main/tools/cache_inference) for ConsisID powered by [TeaCache](https://github.com/LiewFeng/TeaCache). Thanks [@LiewFeng](https://github.com/LiewFeng) for his help.\n* `[2024.12.24]`  🚀 We release the [parallel inference code](https://github.com/PKU-YuanGroup/ConsisID/tree/main/tools/parallel_inference) for ConsisID powered by [xDiT](https://github.com/xdit-project/xDiT). Thanks [@feifeibear](https://github.com/feifeibear) for his help.\n* `[2024.12.09]`  🔥We release the [test set](https://huggingface.co/datasets/BestWishYsh/ConsisID-preview-Data/tree/main/eval) and [metric calculation code](https://github.com/PKU-YuanGroup/ConsisID/tree/main/eval) used in the paper, now your can measure the metrics on your own machine. Please refer to [this guide](https://github.com/PKU-YuanGroup/ConsisID/tree/main/eval) for more details.\n* `[2024.12.08]`  🔥The code for \u003cu\u003edata preprocessing\u003c/u\u003e is out, which is used to obtain the [training data](https://huggingface.co/datasets/BestWishYsh/ConsisID-preview-Data) required by ConsisID, supporting multi-id annotation. Please refer to [this guide](https://github.com/PKU-YuanGroup/ConsisID/tree/main/data_preprocess) for more details.\n* `[2024.12.04]`  Thanks [@shizi](https://www.bilibili.com/video/BV1v3iUY4EeQ/?vd_source=ae3f2652765c02e41cdd698b311989e3) for providing [🤗Windows-ConsisID](https://huggingface.co/pkuhexianyi/ConsisID-Windows/tree/main) and [🟣Windows-ConsisID](https://www.wisemodel.cn/models/PkuHexianyi/ConsisID-Windows/file), which make it easy to run ConsisID on Windows.\n* `[2024.12.01]`  🔥 We provide full text prompts corresponding to all the videos on project page. Click [here](https://github.com/PKU-YuanGroup/ConsisID/blob/main/asserts/prompt.xlsx) to get and try the demo.\n* `[2024.11.30]`  🤗 We have fixed the [huggingface demo](https://huggingface.co/spaces/BestWishYsh/ConsisID-preview-Space), welcome to try it.\n* `[2024.11.29]`  🏃‍♂️ The current code and weights are our early versions, and the differences with the latest version in [arxiv](https://github.com/PKU-YuanGroup/ConsisID) can be viewed [here](https://github.com/PKU-YuanGroup/ConsisID/tree/main/util/on_going_module). And we will release the full code in the next few days.\n* `[2024.11.28]`  Thanks [@camenduru](https://twitter.com/camenduru) for providing [Jupyter Notebook](https://colab.research.google.com/github/camenduru/ConsisID-jupyter/blob/main/ConsisID_jupyter.ipynb) and [@Kijai](https://github.com/kijai) for providing ComfyUI Extension [ComfyUI-ConsisIDWrapper](https://github.com/kijai/ComfyUI-CogVideoXWrapper). If you find related work, please let us know.\n* `[2024.11.27]`  🏃‍♂️ Due to policy restrictions, we only open-source part of the dataset. You can download it by clicking [here](https://huggingface.co/datasets/BestWishYsh/ConsisID-preview-Data). And we will release the data processing code in the next few days.\n* `[2024.11.26]`  🔥 We release the arXiv paper for ConsisID, and you can click [here](https://arxiv.org/abs/2411.17440) to see more details.\n* `[2024.11.22]`  🔥 **All code \u0026 datasets** are coming soon! Stay tuned 👀!\n\n## 😍 Gallery\n\nIdentity-Preserving Text-to-Video Generation. (Some best prompts [here](https://github.com/PKU-YuanGroup/ConsisID/blob/main/asserts/prompt.xlsx))\n\n[![Demo Video of ConsisID](https://github.com/user-attachments/assets/634248f6-1b54-4963-88d6-34fa7263750b)](https://www.youtube.com/watch?v=PhlgC-bI5SQ)\nor you can click \u003ca href=\"https://github.com/SHYuanBest/shyuanbest_media/raw/refs/heads/main/ConsisID/showcase_videos.mp4\"\u003ehere\u003c/a\u003e to watch the video.\n\n## 🤗 Demo\n### Diffusers API\n\n```bash\n!pip install diffusers==0.33.1\nimport torch\nfrom diffusers import ConsisIDPipeline\nfrom diffusers.pipelines.consisid.consisid_utils import prepare_face_models, process_face_embeddings_infer\nfrom diffusers.utils import export_to_video\nfrom huggingface_hub import snapshot_download\n\nsnapshot_download(repo_id=\"BestWishYsh/ConsisID-preview\", local_dir=\"BestWishYsh/ConsisID-preview\")\nface_helper_1, face_helper_2, face_clip_model, face_main_model, eva_transform_mean, eva_transform_std = (\n    prepare_face_models(\"BestWishYsh/ConsisID-preview\", device=\"cuda\", dtype=torch.bfloat16)\n)\npipe = ConsisIDPipeline.from_pretrained(\"BestWishYsh/ConsisID-preview\", torch_dtype=torch.bfloat16)\npipe.to(\"cuda\")\n\n# ConsisID works well with long and well-described prompts. Make sure the face in the image is clearly visible (e.g., preferably half-body or full-body).\nprompt = \"The video captures a boy walking along a city street, filmed in black and white on a classic 35mm camera. His expression is thoughtful, his brow slightly furrowed as if he's lost in contemplation. The film grain adds a textured, timeless quality to the image, evoking a sense of nostalgia. Around him, the cityscape is filled with vintage buildings, cobblestone sidewalks, and softly blurred figures passing by, their outlines faint and indistinct. Streetlights cast a gentle glow, while shadows play across the boy's path, adding depth to the scene. The lighting highlights the boy's subtle smile, hinting at a fleeting moment of curiosity. The overall cinematic atmosphere, complete with classic film still aesthetics and dramatic contrasts, gives the scene an evocative and introspective feel.\"\nimage = \"https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/diffusers/consisid/consisid_input.png?download=true\"\n\nid_cond, id_vit_hidden, image, face_kps = process_face_embeddings_infer(\n    face_helper_1,\n    face_clip_model,\n    face_helper_2,\n    eva_transform_mean,\n    eva_transform_std,\n    face_main_model,\n    \"cuda\",\n    torch.bfloat16,\n    image,\n    is_align_face=True,\n)\n\nvideo = pipe(\n    image=image,\n    prompt=prompt,\n    num_inference_steps=50,\n    guidance_scale=6.0,\n    use_dynamic_cfg=False,\n    id_vit_hidden=id_vit_hidden,\n    id_cond=id_cond,\n    kps_cond=face_kps,\n    generator=torch.Generator(\"cuda\").manual_seed(42),\n)\nexport_to_video(video.frames[0], \"output.mp4\", fps=8)\n```\n\n### Gradio Web UI\n\nHighly recommend trying out our web demo by the following command, which incorporates all features currently supported by ConsisID. We also provide [online demo](https://huggingface.co/spaces/BestWishYsh/ConsisID-preview-Space) in Hugging Face Spaces.\n\n```bash\npython app.py\n```\n\n### CLI Inference\n\n```bash\npython infer.py --model_path BestWishYsh/ConsisID-preview\n```\n\nwarning: It is worth noting that even if we use the same seed and prompt but we change a machine, the results will be different.\n\n### Prompt Refiner\n\nConsisID has high requirements for prompt quality. You can use [GPT-4o](https://chatgpt.com/) to refine the input text prompt, an example is as follows (original prompt: \"a man is playing guitar.\")\n```bash\na man is playing guitar.\n\nChange the sentence above to something like this (add some facial changes, even if they are minor. Don't make the sentence too long): \n\nThe video features a man standing next to an airplane, engaged in a conversation on his cell phone. he is wearing sunglasses and a black top, and he appears to be talking seriously. The airplane has a green stripe running along its side, and there is a large engine visible behind his. The man seems to be standing near the entrance of the airplane, possibly preparing to board or just having disembarked. The setting suggests that he might be at an airport or a private airfield. The overall atmosphere of the video is professional and focused, with the man's attire and the presence of the airplane indicating a business or travel context.\n```\n\nSome sample prompts are available [here](https://github.com/PKU-YuanGroup/ConsisID/blob/main/asserts/prompt.xlsx).\n\n### GPU Memory Optimization\n\nConsisID requires about 44 GB of GPU memory to decode 49 frames (6 seconds of video at 8 FPS) with output resolution 720x480 (W x H), which makes it not possible to run on consumer GPUs or free-tier T4 Colab. The following memory optimizations could be used to reduce the memory footprint. For replication, you can refer to [this](https://gist.github.com/SHYuanBest/bc4207c36f454f9e969adbb50eaf8258) script.\n\n| Feature (overlay the previous) | Max Memory Allocated | Max Memory Reserved |\n| :----------------------------- | :------------------- | :------------------ |\n| -                              | 37 GB                | 44 GB               |\n| enable_model_cpu_offload       | 22 GB                | 25 GB               |\n| enable_sequential_cpu_offload  | 16 GB                | 22 GB               |\n| vae.enable_slicing             | 16 GB                | 22 GB               |\n| vae.enable_tiling              | 5 GB                 | 7 GB                |\n\n```bash\n# turn on if you don't have multiple GPUs or enough GPU memory(such as H100)\npipe.enable_model_cpu_offload()\npipe.enable_sequential_cpu_offload()\npipe.vae.enable_slicing()\npipe.vae.enable_tiling()\n```\nwarning: it will cost more time in inference and may also reduce the quality.\n\n## 🚀 Parallel Inference on Multiple GPUs by xDiT\n\n[xDiT](https://github.com/xdit-project/xDiT) is a Scalable Inference Engine for Diffusion Transformers (DiTs) on multi-GPU Clusters. It has successfully provided low-latency parallel inference solutions for a variety of DiTs models. For example, to generate a video with 6 GPUs, you can use the following command:\n\n```\ncd tools/parallel_inference\nbash run.sh\n# run_usp.sh\n```\n\n## 🚀 Cache Inference by TeaCache\n\n[TeaCache](https://github.com/LiewFeng/TeaCache) is a training-free caching approach that estimates and leverages the fluctuating differences among model outputs across timesteps, thereby accelerating the inference.  For example, you can use the following command:\n\n```\ncd tools/cache_inference\nbash run.sh\n```\n\n## ⚙️ Requirements and Installation\n\nWe recommend the requirements as follows.\n\n### Environment\n\n```bash\n# 0. Clone the repo\ngit clone --depth=1 https://github.com/PKU-YuanGroup/ConsisID.git\ncd ConsisID\n\n# 1. Create conda environment\nconda create -n consisid python=3.11.0\nconda activate consisid\n\n# 3. Install PyTorch and other dependencies using conda\n# CUDA 11.8\nconda install pytorch==2.5.1 torchvision==0.20.1 torchaudio==2.5.1 pytorch-cuda=11.8 -c pytorch -c nvidia\n# CUDA 12.1\nconda install pytorch==2.5.1 torchvision==0.20.1 torchaudio==2.5.1 pytorch-cuda=12.1 -c pytorch -c nvidia\n\n# 4. Install pip dependencies\npip install -r requirements.txt\n```\n\n### Download ConsisID\n\nThe weights are available at [🤗HuggingFace](https://huggingface.co/BestWishYsh/ConsisID-preview), [🤖ModelScope](https://modelscope.cn/models/BestWishYSH/ConsisID-preview) and [🟣WiseModel](https://wisemodel.cn/models/SHYuanBest/ConsisID-Preview/file), and will be automatically downloaded when runing `app.py` and `infer.py`, or you can download it with the following commands.\n\n```bash\n# way 1\n# if you are in china mainland, run this first: export HF_ENDPOINT=https://hf-mirror.com\ncd util\npython download_weights.py\n\n# way 2\n# if you are in china mainland, run this first: export HF_ENDPOINT=https://hf-mirror.com\nhuggingface-cli download --repo-type model \\\nBestWishYsh/ConsisID-preview \\\n--local-dir ckpts\n\n# way 3\nmodelscope download --model \\\nBestWishYSH/ConsisID-preview \\\n--local-dir ckpts\n\n# way 4\ngit lfs install\ngit clone https://www.wisemodel.cn/SHYuanBest/ConsisID-Preview.git\n```\n\nOnce ready, the weights will be organized in this format:\n\n```\n📦 ckpts/\n├── 📂 data_process/\n├── 📂 face_encoder/\n├── 📂 scheduler/\n├── 📂 text_encoder/\n├── 📂 tokenizer/\n├── 📂 transformer/\n├── 📂 vae/\n├── 📄 configuration.json\n├── 📄 model_index.json\n```\n\n## 🗝️ Training\n\n### Data preprocessing\n\nPlease refer to [this guide](https://github.com/PKU-YuanGroup/ConsisID/tree/main/data_preprocess) for how to obtain the [training data](https://huggingface.co/datasets/BestWishYsh/ConsisID-preview-Data) required by ConsisID. If you want to train your own identity-preserving text-to-video generation model, you need to arrange all the dataset in this [format](https://github.com/PKU-YuanGroup/ConsisID/tree/main/asserts/demo_train_data/dataname):\n\n```\n📦 datasets/\n├── 📂 captions/\n│   ├── 📄 dataname_1.json\n│   ├── 📄 dataname_2.json\n├── 📂 dataname_1/\n│   ├── 📂 refine_bbox_jsons/\n│   ├── 📂 track_masks_data/\n│   ├── 📂 videos/\n├── 📂 dataname_2/\n│   ├── 📂 refine_bbox_jsons/\n│   ├── 📂 track_masks_data/\n│   ├── 📂 videos/\n├── ...\n├── 📄 total_train_data.txt\n```\n\n### Video DiT training\n\nFirst, setting hyperparameters:\n\n- environment (e.g., cuda): [deepspeed_configs](https://github.com/PKU-YuanGroup/ConsisID/tree/main/util/deepspeed_configs)\n- training arguments (e.g., batchsize): [train_single_rank.sh](https://github.com/PKU-YuanGroup/ConsisID/blob/main/train_single_rank.sh) or [train_multi_rank.sh](https://github.com/PKU-YuanGroup/ConsisID/blob/main/train_multi_rank.sh)\n\nThen, we run the following bash to start training:\n\n```bash\n# For single rank\nbash train_single_rank.sh\n# For multi rank\nbash train_multi_rank.sh\n```\n\n## 🙌 Friendly Links\n\nWe found some plugins created by community developers. Thanks for their efforts: \n\n  - ComfyUI Extension. [ComfyUI-ConsisIDWrapper](https://github.com/kijai/ComfyUI-CogVideoXWrapper) (by [@Kijai](https://github.com/kijai)).\n  - Jupyter Notebook. [Jupyter-ConsisID](https://colab.research.google.com/github/camenduru/ConsisID-jupyter/blob/main/ConsisID_jupyter.ipynb) (by [@camenduru](https://github.com/camenduru/consisid-tost)).\n  - Windows Docker. [🤗Windows-ConsisID](https://huggingface.co/pkuhexianyi/ConsisID-Windows/tree/main) and [🟣Windows-ConsisID](https://www.wisemodel.cn/models/PkuHexianyi/ConsisID-Windows/file) (by [@shizi](https://www.bilibili.com/video/BV1v3iUY4EeQ/?vd_source=ae3f2652765c02e41cdd698b311989e3)).\n  - Diffusres. [Diffusers-ConsisID](https://github.com/huggingface/diffusers) (thanks [@arrow](https://github.com/a-r-r-o-w), [@yiyixuxu](https://github.com/yiyixuxu), [@hlky](https://github.com/hlky) and [@stevhliu](https://github.com/stevhliu) for their help).\n  - xDiT. [xDiT-ConsisID](https://github.com/xdit-project/xDiT) (thanks [@feifeibear](https://github.com/feifeibear) for his help).\n  - TeaCache. [TeaCache-ConsisID](https://github.com/LiewFeng/TeaCache) (thanks [@LiewFeng](https://github.com/LiewFeng) for his help).\n  - [Ingredients](https://github.com/feizc/Ingredients): A powerful way to customize video creations by incorporating multiple specific identity (ID) photos, based on ConsisID.\n\nIf you find related work, please let us know. \n\n## 🐳 Dataset\n\nWe release the subset of the data used to train ConsisID. The dataset is available at [HuggingFace](https://huggingface.co/datasets/BestWishYsh/ConsisID-preview-Data), or you can download it with the following command. Some samples can be found on our [Project Page](https://pku-yuangroup.github.io/ConsisID/).\n\n```bash\nhuggingface-cli download --repo-type dataset \\\nBestWishYsh/ConsisID-preview-Data \\\n--local-dir BestWishYsh/ConsisID-preview-Data\n```\n\n## 🛠️ Evaluation \n\nWe release the data used for evaluation in [ConsisID](https://huggingface.co/papers/2411.17440), which is available at [HuggingFace](https://huggingface.co/datasets/BestWishYsh/ConsisID-preview-Data). Please refer to [this guide](https://github.com/PKU-YuanGroup/ConsisID/tree/main/eval) for how to evaluate customized model.\n\n## 👍 Acknowledgement\n\n* This project wouldn't be possible without the following open-sourced repositories: [Open-Sora Plan](https://github.com/PKU-YuanGroup/Open-Sora-Plan), [CogVideoX](https://github.com/THUDM/CogVideo), [EasyAnimate](https://github.com/aigc-apps/EasyAnimate), [CogVideoX-Fun](https://github.com/aigc-apps/CogVideoX-Fun), [IP-Adapter](https://github.com/tencent-ailab/IP-Adapter), [PhotoMaker](https://github.com/TencentARC/PhotoMaker), [UniPortrait](https://github.com/junjiehe96/UniPortrait).\n\n## 🔒 License\n\n* The majority of this project is released under the Apache 2.0 license as found in the [LICENSE](https://github.com/PKU-YuanGroup/ConsisID/blob/main/LICENSE) file.\n* The CogVideoX-5B model (Transformers module) is released under the [CogVideoX LICENSE](https://huggingface.co/THUDM/CogVideoX-5b/blob/main/LICENSE).\n* The service is a research preview. Please contact us if you find any potential violations. (shyuan-cs@hotmail.com)\n\n## ✏️ Citation\n\nIf you find our paper and code useful in your research, please consider giving a star :star: and citation :pencil:.\n\n```BibTeX\n@inproceedings{yuan2025identity,\n  title={Identity-preserving text-to-video generation by frequency decomposition},\n  author={Yuan, Shenghai and Huang, Jinfa and He, Xianyi and Ge, Yunyang and Shi, Yujun and Chen, Liuhan and Luo, Jiebo and Yuan, Li},\n  booktitle={Proceedings of the Computer Vision and Pattern Recognition Conference},\n  pages={12978--12988},\n  year={2025}\n}\n```\n\n## 🤝 Contributors\n\n\u003ca href=\"https://github.com/PKU-YuanGroup/ConsisID/graphs/contributors\"\u003e\n  \u003cimg src=\"https://contrib.rocks/image?repo=PKU-YuanGroup/ConsisID\u0026anon=true\" /\u003e\n\n\u003c/a\u003e\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fpku-yuangroup%2Fconsisid","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fpku-yuangroup%2Fconsisid","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fpku-yuangroup%2Fconsisid/lists"}