{"id":23060231,"url":"https://github.com/opensparsellms/skip-dit","last_synced_at":"2025-10-05T09:45:19.406Z","repository":{"id":264670953,"uuid":"893080979","full_name":"OpenSparseLLMs/Skip-DiT","owner":"OpenSparseLLMs","description":"✈️ [ICCV 2025] Towards Stabilized and Efficient Diffusion Transformers through Long-Skip-Connections with Spectral Constraints","archived":false,"fork":false,"pushed_at":"2025-07-10T10:34:55.000Z","size":39654,"stargazers_count":70,"open_issues_count":0,"forks_count":1,"subscribers_count":2,"default_branch":"master","last_synced_at":"2025-07-23T00:42:54.980Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":"https://arxiv.org/abs/2411.17616","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/OpenSparseLLMs.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null}},"created_at":"2024-11-23T13:37:33.000Z","updated_at":"2025-07-21T05:47:54.000Z","dependencies_parsed_at":"2025-03-28T10:22:25.661Z","dependency_job_id":"2edb1379-4a72-484d-965d-e5e737b4d387","html_url":"https://github.com/OpenSparseLLMs/Skip-DiT","commit_stats":null,"previous_names":["opensparsellms/skip-dit"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/OpenSparseLLMs/Skip-DiT","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/OpenSparseLLMs%2FSkip-DiT","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/OpenSparseLLMs%2FSkip-DiT/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/OpenSparseLLMs%2FSkip-DiT/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/OpenSparseLLMs%2FSkip-DiT/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/OpenSparseLLMs","download_url":"https://codeload.github.com/OpenSparseLLMs/Skip-DiT/tar.gz/refs/heads/master","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/OpenSparseLLMs%2FSkip-DiT/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":278437956,"owners_count":25986760,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","status":"online","status_checked_at":"2025-10-05T02:00:06.059Z","response_time":54,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2024-12-16T03:11:40.602Z","updated_at":"2025-10-05T09:45:19.401Z","avatar_url":"https://github.com/OpenSparseLLMs.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"\u003cdiv align=\"center\"\u003e\n  \n# Towards Stabilized and Efficient Diffusion Transformers through Long-Skip-Connections with Spectral Constraints\n  \n  ${{\\color{Red}\\Huge{\\textsf{  ICCV\\ 2025\\ \\}}}}\\$\n\n  \u003ca href=\"https://github.com/OpenSparseLLMs/Skip-DiT\"\u003e\u003cimg src=\"https://img.shields.io/static/v1?label=Code\u0026message=Github\u0026color=blue\u0026logo=github-pages\"\u003e\u003c/a\u003e \u0026ensp;\n  \u003ca href=\"https://arxiv.org/abs/2411.17616\"\u003e\u003cimg src=\"https://img.shields.io/static/v1?label=Paper\u0026message=Arxiv:Skip-DiT\u0026color=red\u0026logo=arxiv\"\u003e\u003c/a\u003e \u0026ensp;\n  \u003ca href=\"https://huggingface.co/GuanjieChen/Skip-DiT\"\u003e\u003cimg src=\"https://img.shields.io/static/v1?label=Skip-DiT\u0026message=HuggingFace\u0026color=yellow\"\u003e\u003c/a\u003e \u0026ensp;\n  \u003ca href=\"https://huggingface.co/datasets/GuanjieChen/video_seg\"\u003e\u003cimg src=\"https://img.shields.io/static/v1?label=Dataset\u0026message=HuggingFace\u0026color=yellow\"\u003e\u003c/a\u003e \u0026ensp;\n\u003c/div\u003e\n\n\n\u003cdiv align=\"center\"\u003e\n  \u003cimg src=\"visuals/teaser.jpg\" width=\"100%\" \u003e\u003c/img\u003e\n  \u003cbr\u003e\n  \u003cem\u003e\n      (a) Skip-DiT presents consistently higher feature similarity under caching, demonstrating superior stability. (b) Illustration of Skip-DiT. (c) Skip-DiT maintains higher generation quality even at greater speedup factors.\n  \u003c/em\u003e\n\u003c/div\u003e\n\u003cbe\u003e\n\n## 🎉 Abstract\nThis repository contains the official PyTorch implementation of the paper: **[Towards Stabilized and Efficient Diffusion Transformers through Long-Skip-Connections with Spectral Constraints](https://arxiv.org/abs/2411.17616)**. \n\nDiffusion Transformers (DiT) have emerged as a powerful architecture for image and video generation, offering superior quality and scalability. However, their practical application suffers from inherent dynamic feature instability, leading to error amplification during cached inference. Through systematic analysis, we identify the absence of long-range feature preservation mechanisms as the root cause of unstable feature propagation and perturbation sensitivity. To this end, we propose Skip-DiT, an image and video generative DiT variant enhanced with Long-Skip-Connections (LSCs) - the key efficiency component in U-Nets. Theoretical spectral norm and visualization analysis demonstrate how LSCs stabilize feature dynamics. Skip-DiT architecture and its stabilized dynamic feature enable an efficient statical caching mechanism that reuses deep features across timesteps while updating shallow components. Extensive experiments across the image and video generation tasks demonstrate that Skip-DiT achieves: (1) 4.4x training acceleration and faster convergence, (2) 1.5-2x inference acceleration with negligible quality loss and high fidelity to the original output, outperforming existing DiT caching methods across various quantitative metrics.Our findings establish Long-Skip-Connections as critical architectural components for stable and efficient diffusion transformers.\nMore visualizations can be found [here](#visualization).\n\n\u003c!-- \u003e [**Accelerating Vision Diffusion Transformers with Skip Branches**](https://arxiv.org/abs/2411.17616)\u003cbr\u003e\n\u003e [Guanjie Chen](https://scholar.google.com/citations?user=cpBU1VgAAAAJ\u0026hl=zh-CN), [Xinyu Zhao](https://scholar.google.com/citations?hl=en\u0026user=1cj23VYAAAAJ),[Yucheng Zhou](https://scholar.google.com/citations?user=nnbFqRAAAAAJ\u0026hl=en), [Tianlong Chen](https://scholar.google.com/citations?user=LE3ctn0AAAAJ\u0026hl=en), [Yu Cheng](https://scholar.google.com/citations?user=ORPxbV4AAAAJ\u0026hl=en)         \n\u003e (contact us: chenguanjie@sjtu.edu.cn, xinyu@cs.unc.edu) --\u003e\n\n## 🌟 Feature Stability of Skip-DiT\n![stability](visuals/stability.jpg)\nVisualization of the feature stability of Skip-DiT compared with vanilla DiT. Skip-DiT also shows superior training efficiency.\n\n## Gallery\n\u003cdiv align=\"center\"\u003e\n  \u003cimg src=\"visuals/video-demo.gif\" width=\"85%\" \u003e\u003c/img\u003e\n  \u003cbr\u003e\n  \u003cem\u003e\n      (Results of Latte with skip-branches on text-to-video and class-to-video tasks with Latte. Left: text-to-video with 1.7x and 2.0x speedup. Right: class-to-video with 2.2x and 2.4x speedup. Latency is measured on one A100.) \n  \u003c/em\u003e\n\u003c/div\u003e\n\u003cbr\u003e\n\n\u003cdiv align=\"center\"\u003e\n  \u003cimg src=\"visuals/image-demo.jpg\" width=\"100%\" \u003e\u003c/img\u003e\n  \u003cbr\u003e\n  \u003cem\u003e\n      (Results of HunYuan-DiT with skip-branches on text-to-image task with Hunyuan-DiT. Latency is measured on one A100.) \n  \u003c/em\u003e\n\u003c/div\u003e\n\u003cbe\u003e\n\n\n## Released Checkpoints\n| Model | Task | Training Data | Backbone | Size(G) | Skip-Cache |\n|:--:|:--:|:--:|:--:|:--:|:--:|\n| [Latte-skip](https://huggingface.co/GuanjieChen/Skip-DiT/blob/main/DiT-XL-2-skip.pt) | text-to-video |Vimeo|Latte|8.76| ✅ |\n| [DiT-XL/2-skip](https://huggingface.co/GuanjieChen/Skip-DiT/blob/main/Latte-skip.pt) | class-to-image |ImageNet|DiT-XL/2|11.40|✅ |\n| [ucf101-skip](https://huggingface.co/GuanjieChen/Skip-DiT/blob/main/ucf101-skip.pt) | class-to-video|UCF101|Latte|2.77|✅ |\n| [taichi-skip](https://huggingface.co/GuanjieChen/Skip-DiT/blob/main/taichi-skip.pt) | class-to-video|Taichi-HD|Latte|2.77|✅ |\n| [skytimelapse-skip](https://huggingface.co/GuanjieChen/Skip-DiT/blob/main/skylapse-skip.pt) | class-to-video|SkyTimelapse|Latte|2.77|✅ |\n| [ffs-skip](https://huggingface.co/GuanjieChen/Skip-DiT/blob/main/ffs-skip.pt) | class-to-video|FaceForensics|Latte|2.77|✅ |\n\nPretrained text-to-image Model of [HunYuan-DiT](https://github.com/Tencent/HunyuanDiT) can be found in [Huggingface](https://huggingface.co/Tencent-Hunyuan/HunyuanDiT-v1.2/tree/main/t2i/model) and [Tencent-cloud](https://dit.hunyuan.tencent.com/download/HunyuanDiT/model-v1_2.zip).\n\nTraining dataset for Latte-skip on text-to-video task is released [here](https://huggingface.co/datasets/GuanjieChen/video_seg)\n\n## 🚀 Usage\n### Text-to-video Inference\nTo generate videos with Latte-skip, you just need 3 steps\n```shell\n# 1. Prepare your conda environments\ncd text-to-video ; conda env create -f environment.yaml ; conda activate latte\n# 2. Download checkpoints of Latte and Latte-skip\npython download.py\n# 3. Generate videos with only one command line!\npython sample/sample_t2v.py --config ./configs/t2v/t2v_sample_skip.yaml\n# 4. (Optional) To accelerate generation with skip-cache, run following command\npython sample/sample_t2v.py --config ./configs/t2v/t2v_sample_skip_cache.yaml --cache N2-700-50\n```\n### Text-to-image Inference\nIn the same way, to generate images with Hunyuan-DiT, you only need 3 steps\n```shell\n# 1. Prepare your conda environments\ncd text-to-image ; conda env create -f environment.yaml ; conda activate HunyuanDiT\n# 2. Download checkpoints of Hunyuan-DiT\nmkdir ckpts ; huggingface-cli download Tencent-Hunyuan/HunyuanDiT-v1.2 --local-dir ./ckpts\n# 3. Generate images with only one command line!\npython sample_t2i.py --prompt \"渔舟唱晚\"  --no-enhance --infer-steps 100 --image-size 1024 1024\n# 4. (Optional) To accelerate generation with skip-cache, run the following command\npython sample_t2i.py --prompt \"渔舟唱晚\"  --no-enhance --infer-steps 100 --image-size 1024 1024 --cache --cache-step 2\n```\n\nAbout the class-to-video and class-to-image task, you can found detailed instructions in `class-to-video/README.md` and `class-to-image/README.md`\n\n### Training\nWe have already released the training code of Latte-skip! It takes only a few days on 8 H100 GPUs. To train the text-to-video model:\n1. Prepare your text-video datasets and implement the `text-to-video/datasets/t2v_joint_dataset.py`\n2. Run the two-stage training strategy:\n   1. Freeze all the parameters except skip-branches. Set `freeze=True` in `text-to-video/configs/train_t2v.yaml`. And then run the training scripts at `text-to-video/train_scripts/t2v_joint_train_skip.sh`.\n   2. Overall training. Set `freeze=False` in `text-to-video/configs/train_t2v.yaml`. And then run the training scripts.\n**The text-to-video model we released is trained with only 300k text-video pairs of Vimeo for around 1 week on 8 H100 GPUs.**\n\nThe training instructions of `class-to-video` and `text-to-video` tasks can be found in `class-to-video/README.md` and `class-to-image/README.md`\n\n\n\n\n\n## Acknowledgement\nSkip-DiT has been greatly inspired by the following amazing works and teams: [DeepCache](https://arxiv.org/abs/2312.00858), [Latte](https://github.com/Vchitect/Latte), [DiT](https://github.com/facebookresearch/DiT), and [HunYuan-DiT](https://github.com/Tencent/HunyuanDiT), we thank all the contributors for open-sourcing.\n\n\n## License\nThe code and model weights are licensed under [LICENSE](./class-to-image/LICENSE).\n\n\n## Visualization\n##### Class-to-image\n![class-to-image visualizations](visuals/case_c2i.jpg)\n##### Text-to-image\n![text-to-image visualizations](visuals/case_t2i.jpg)\n##### Text-to-Video\n![text-to-video visualizations](visuals/case_t2v.jpg)\n##### Class-to-Video\n![class-to-video visualizations](visuals/case_c2v.jpg)\n\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fopensparsellms%2Fskip-dit","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fopensparsellms%2Fskip-dit","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fopensparsellms%2Fskip-dit/lists"}