{"id":27884312,"url":"https://github.com/skyworkai/skyreels-a1","last_synced_at":"2025-05-05T06:36:15.165Z","repository":{"id":278106780,"uuid":"931887169","full_name":"SkyworkAI/SkyReels-A1","owner":"SkyworkAI","description":"SkyReels-A1: Expressive Portrait Animation in Video Diffusion Transformers","archived":false,"fork":false,"pushed_at":"2025-04-23T15:09:10.000Z","size":58807,"stargazers_count":488,"open_issues_count":17,"forks_count":55,"subscribers_count":10,"default_branch":"main","last_synced_at":"2025-04-23T16:25:08.900Z","etag":null,"topics":["condition-render","portrait-animation","video-diffusion-transformers"],"latest_commit_sha":null,"homepage":"https://www.skyreels.ai","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"other","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/SkyworkAI.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE.txt","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null}},"created_at":"2025-02-13T02:37:51.000Z","updated_at":"2025-04-23T15:09:15.000Z","dependencies_parsed_at":"2025-04-23T16:35:11.728Z","dependency_job_id":null,"html_url":"https://github.com/SkyworkAI/SkyReels-A1","commit_stats":null,"previous_names":["skyworkai/skyreels-a1"],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/SkyworkAI%2FSkyReels-A1","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/SkyworkAI%2FSkyReels-A1/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/SkyworkAI%2FSkyReels-A1/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/SkyworkAI%2FSkyReels-A1/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/SkyworkAI","download_url":"https://codeload.github.com/SkyworkAI/SkyReels-A1/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":252454576,"owners_count":21750493,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["condition-render","portrait-animation","video-diffusion-transformers"],"created_at":"2025-05-05T06:36:14.669Z","updated_at":"2025-05-05T06:36:15.157Z","avatar_url":"https://github.com/SkyworkAI.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"\u003cp align=\"center\"\u003e\n  \u003cimg src=\"assets/logo.png\" alt=\"Skyreels Logo\" width=\"50%\"\u003e\n\u003c/p\u003e\n\n\n\u003ch1 align=\"center\"\u003eSkyReels-A1: Expressive Portrait Animation in Video Diffusion Transformers\u003c/h1\u003e\n\n\u003cdiv align='center'\u003e\n    \u003ca href='https://scholar.google.com/citations?user=6D_nzucAAAAJ\u0026hl=en' target='_blank'\u003eDi Qiu\u003c/a\u003e\u0026emsp;\n    \u003ca href='https://scholar.google.com/citations?user=_43YnBcAAAAJ\u0026hl=zh-CN' target='_blank'\u003eZhengcong Fei\u003c/a\u003e\u0026emsp;\n    \u003ca href='' target='_blank'\u003eRui Wang\u003c/a\u003e\u0026emsp;\n    \u003ca href='' target='_blank'\u003eJialin Bai\u003c/a\u003e\u0026emsp;\n    \u003ca href='https://scholar.google.com/citations?user=Hv-vj2sAAAAJ\u0026hl=en' target='_blank'\u003eChangqian Yu\u003c/a\u003e\u0026emsp;\n\u003c/div\u003e\n\n\u003cdiv align='center'\u003e\n  \u003ca href='https://scholar.google.com.au/citations?user=ePIeVuUAAAAJ\u0026hl=en' target='_blank'\u003eMingyuan Fan\u003c/a\u003e\u0026emsp;\n  \u003ca href='https://scholar.google.com/citations?user=HukWSw4AAAAJ\u0026hl=en' target='_blank'\u003eGuibin Chen\u003c/a\u003e\u0026emsp;\n  \u003ca href='https://scholar.google.com.tw/citations?user=RvAuMk0AAAAJ\u0026hl=zh-CN' target='_blank'\u003eXiang Wen\u003c/a\u003e\u0026emsp;\n\u003c/div\u003e\n\n\u003cdiv align='center'\u003e\n    \u003csmall\u003e\u003cstrong\u003eSkywork AI, Kunlun Inc.\u003c/strong\u003e\u003c/small\u003e\n\u003c/div\u003e\n\n\u003cbr\u003e\n\n\u003cdiv align=\"center\"\u003e\n  \u003c!-- \u003ca href='LICENSE'\u003e\u003cimg src='https://img.shields.io/badge/license-MIT-yellow'\u003e\u003c/a\u003e --\u003e\n  \u003ca href='https://arxiv.org/abs/2502.10841'\u003e\u003cimg src='https://img.shields.io/badge/arXiv-SkyReels A1-red'\u003e\u003c/a\u003e\n  \u003ca href='https://skyworkai.github.io/skyreels-a1.github.io/'\u003e\u003cimg src='https://img.shields.io/badge/Project-SkyReels A1-green'\u003e\u003c/a\u003e\n  \u003ca href='https://huggingface.co/Skywork/SkyReels-A1'\u003e\u003cimg src='https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-Models-blue'\u003e\u003c/a\u003e\n  \u003ca href='https://www.skyreels.ai/home?utm_campaign=github_A1'\u003e\u003cimg src='https://img.shields.io/badge/Playground-Spaces-yellow'\u003e\u003c/a\u003e\n  \u003cbr\u003e\n\u003c/div\u003e\n\u003cbr\u003e\n\n\n\u003cp align=\"center\"\u003e\n  \u003cimg src=\"./assets/demo.gif\" alt=\"showcase\"\u003e\n  \u003cbr\u003e\n  🔥 For more results, visit our \u003ca href=\"https://skyworkai.github.io/skyreels-a1.github.io/\"\u003e\u003cstrong\u003ehomepage\u003c/strong\u003e\u003c/a\u003e 🔥\n\u003c/p\u003e\n\n\u003cp align=\"center\"\u003e\n    👋 Join our \u003ca href=\"https://discord.gg/PwM6NYtccQ\" target=\"_blank\"\u003e\u003cstrong\u003eDiscord\u003c/strong\u003e\u003c/a\u003e \n\u003c/p\u003e\n\n\nThis repo, named **SkyReels-A1**, contains the official PyTorch implementation of our paper [SkyReels-A1: Expressive Portrait Animation in Video Diffusion Transformers](https://arxiv.org/abs/2502.10841).\n\n\n## 🔥🔥🔥 News!!\n* Apr 3, 2025: 🔥 We release [SkyReels-A2](https://github.com/SkyworkAI/SkyReels-A2). This is an open-sourced controllable video generation framework capable of assembling arbitrary visual elements.\n* Mar 4, 2025: 🔥 We release audio-driven portrait image animation pipeline. Try out on [Huggingface Spaces Demo](https://huggingface.co/spaces/Skywork/skyreels-a1-talking-head) !\n* Feb 18, 2025: 👋 We release the inference code and model weights of SkyReels-A1. [Download](https://huggingface.co/Skywork/SkyReels-A1)\n* Feb 18, 2025: 🎉 We have made our technical report available as open source. [Read](https://skyworkai.github.io/skyreels-a1.github.io/report.pdf)\n* Feb 18, 2025: 🔥 Our online demo of LipSync is available on SkyReels now! Try out on [LipSync](https://www.skyreels.ai/home/tools/lip-sync?refer=navbar) .\n* Feb 18, 2025: 🔥 We have open-sourced I2V video generation model [SkyReels-V1](https://github.com/SkyworkAI/SkyReels-V1). This is the first and most advanced open-source human-centric video foundation model.\n\n## 📑 TODO List\n- [x] Checkpoints\n- [x] Inference Code\n- [x] Web Demo (Gradio)\n- [x] Audio-driven Portrait Image Animation Pipeline\n- [x] Inference Code for Long Videos\n- [ ] User-Level GPU Inference on RTX4090\n- [ ] ComfyUI\n\n\n## Getting Started 🏁 \n\n### 1. Clone the code and prepare the environment 🛠️\nFirst git clone the repository with code: \n```bash\ngit clone https://github.com/SkyworkAI/SkyReels-A1.git\ncd SkyReels-A1\n\n# create env using conda\nconda create -n skyreels-a1 python=3.10\nconda activate skyreels-a1\n```\nThen, install the remaining dependencies:\n```bash\npip install -r requirements.txt\n```\n\n\n### 2. Download pretrained weights 📥\nYou can download the pretrained weights is from HuggingFace:\n```bash\n# !pip install -U \"huggingface_hub[cli]\"\nhuggingface-cli download Skywork/SkyReels-A1 --local-dir local_path --exclude \"*.git*\" \"README.md\" \"docs\"\n```\n\nThe FLAME, mediapipe, and smirk models are located in the SkyReels-A1/extra_models folder.\n\nThe directory structure of our SkyReels-A1 code is formulated as: \n```text\npretrained_models\n├── FLAME\n├── SkyReels-A1-5B\n│   ├── pose_guider\n│   ├── scheduler\n│   ├── tokenizer\n│   ├── siglip-so400m-patch14-384\n│   ├── transformer\n│   ├── vae\n│   └── text_encoder\n├── mediapipe\n└── smirk\n\n```\n\n#### Download DiffposeTalk assets and pretrained weights (For Audio-driven)\n\n- We use [diffposetalk](https://github.com/DiffPoseTalk/DiffPoseTalk/tree/main) to generate flame coefficients from audio, thereby constructing motion signals.\n\n- Download the diffposetalk code and follow its README to download the weights and related data.\n\n- Then place them in the specified directory.\n\n```bash\ncp -r ${diffposetalk_root}/style pretrained_models/diffposetalk\ncp ${diffposetalk_root}/experiments/DPT/head-SA-hubert-WM/checkpoints/iter_0110000.pt pretrained_models/diffposetalk\ncp ${diffposetalk_root}/datasets/HDTF_TFHP/lmdb/stats_train.npz pretrained_models/diffposetalk\n```\n\n- Or you can download style files from [link](https://drive.google.com/file/d/1XT426b-jt7RUkRTYsjGvG-wS4Jed2U1T/view?usp=sharing) and stats_train.npz from [link](https://drive.google.com/file/d/1_I5XRzkMP7xULCSGVuaN8q1Upplth9xR/view?usp=sharing).\n\n```text\npretrained_models\n├── FLAME\n├── SkyReels-A1-5B\n├── mediapipe\n├── diffposetalk\n│   ├── style\n│   ├── iter_0110000.pt\n│   ├── stats_train.npz\n└── smirk\n\n```\n\n#### Download Frame interpolation Model pretrained weights (For Long Video Inference and Dynamic Resolution)\n\n- We use [FILM](https://github.com/dajes/frame-interpolation-pytorch) to generate transition frames, making the video transitions smoother (Set `use_interpolation` to True).\n\n- Download [film_net_fp16.pt](https://github.com/dajes/frame-interpolation-pytorch/releases), and place it in the specified directory.\n\n```text\npretrained_models\n├── FLAME\n├── SkyReels-A1-5B\n├── mediapipe\n├── diffposetalk\n├── film_net\n│   ├── film_net_fp16.pt\n└── smirk\n```\n\n\n### 3. Inference 🚀\nYou can simply run the inference scripts as: \n```bash\npython inference.py\n\n# inference audio to video\npython inference_audio.py\n```\n\nIf the script runs successfully, you will get an output mp4 file. This file includes the following results: driving video, input image or video, and generated result.\n\n#### Long Video Inference\n\nNow, you can run the long video inference scripts to obtain portrait animation of any length：\n```bash\npython inference_long_video.py\n\n# inference audio to video\npython inference_audio_long_video.py\n```\n\n#### Dynamic Resolution\n\nAll inference scripts now support dynamic resolution, simply set `target_fps` to any desired fps, recommended fps include: 12fps (Native), 24fps, 48fps, 60fps, other settings such as 25fps and 30fps may cause unstable frame rates. \n\n\n## Gradio Interface 🤗\n\nWe provide a [Gradio](https://huggingface.co/docs/hub/spaces-sdks-gradio) interface for a better experience, just run by:\n\n```bash\npython app.py\n```\n\nThe graphical interactive interface is shown as below: \n\n![gradio](assets/gradio.png)\n\n\n## Metric Evaluation 👓\n\nWe also provide all scripts for automatically calculating the metrics, including SimFace, FID, and L1 distance between expression and motion, reported in the paper.  \n\nAll codes can be found in the ```eval``` folder. After setting the video result path, run the following commands in sequence: \n\n```bash\npython arc_score.py\npython expression_score.py\npython pose_score.py\n```\n\n\n## Acknowledgements 💐\nWe would like to thank the contributors of [CogvideoX](https://github.com/THUDM/CogVideo), [finetrainers](https://github.com/a-r-r-o-w/finetrainers) and [DiffPoseTalk](https://github.com/DiffPoseTalk/DiffPoseTalk)repositories, for their open research and contributions. \n\n## Citation 💖\nIf you find SkyReels-A1 useful for your research, welcome to 🌟 this repo and cite our work using the following BibTeX:\n```bibtex\n@article{qiu2025skyreels,\n  title={Skyreels-a1: Expressive portrait animation in video diffusion transformers},\n  author={Qiu, Di and Fei, Zhengcong and Wang, Rui and Bai, Jialin and Yu, Changqian and Fan, Mingyuan and Chen, Guibin and Wen, Xiang},\n  journal={arXiv preprint arXiv:2502.10841},\n  year={2025}\n}\n```\n\n## Star History\n\n[![Star History Chart](https://api.star-history.com/svg?repos=SkyworkAI/SkyReels-A1\u0026type=Date)](https://www.star-history.com/#SkyworkAI/SkyReels-A1\u0026Date)\n\n\n\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fskyworkai%2Fskyreels-a1","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fskyworkai%2Fskyreels-a1","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fskyworkai%2Fskyreels-a1/lists"}