{"id":27943618,"url":"https://github.com/alpha-vllm/lumina-video","last_synced_at":"2025-05-07T12:18:12.318Z","repository":{"id":276774937,"uuid":"928693917","full_name":"Alpha-VLLM/Lumina-Video","owner":"Alpha-VLLM","description":null,"archived":false,"fork":false,"pushed_at":"2025-03-10T06:54:16.000Z","size":3882,"stargazers_count":237,"open_issues_count":5,"forks_count":11,"subscribers_count":5,"default_branch":"main","last_synced_at":"2025-05-07T12:18:05.357Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":null,"language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/Alpha-VLLM.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2025-02-07T04:14:31.000Z","updated_at":"2025-04-25T01:48:10.000Z","dependencies_parsed_at":null,"dependency_job_id":"8c4ac419-9e5e-42d2-9bee-fd80c63d20d5","html_url":"https://github.com/Alpha-VLLM/Lumina-Video","commit_stats":null,"previous_names":["alpha-vllm/lumina-video"],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Alpha-VLLM%2FLumina-Video","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Alpha-VLLM%2FLumina-Video/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Alpha-VLLM%2FLumina-Video/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Alpha-VLLM%2FLumina-Video/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/Alpha-VLLM","download_url":"https://codeload.github.com/Alpha-VLLM/Lumina-Video/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":252873889,"owners_count":21817715,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2025-05-07T12:18:11.448Z","updated_at":"2025-05-07T12:18:12.291Z","avatar_url":"https://github.com/Alpha-VLLM.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"\u003cdiv align=\"center\"\u003e\n\n# Lumina-Video\n\n**Official repository for Lumina-Video, a preliminary tryout of the Lumina series for Video Generation**\n\n[![Lumina-mGPT](https://img.shields.io/badge/Paper-Lumina--Video-2b9348.svg?logo=arXiv)](https://arxiv.org/abs/2502.06782)\u0026#160;\n\n\u003c/div\u003e\n\n\n\u003cp align=\"center\"\u003e\n \u003cimg src=\"assets/architecture.png\" width=\"90%\"/\u003e\n \u003cbr\u003e\n\u003c/p\u003e\n\n\u003ch2 id=\"custom-gallery\"\u003e 📽️ Gallery\u003c/h2\u003e\n\n### Text to Video results\n\n\u003ctable border=\"0\" style=\"width: 100%; text-align: center; margin-top: 1px;\"\u003e\n  \u003ctr\u003e\n    \u003ctd\u003e\u003cvideo src=\"https://github.com/user-attachments/assets/97f10a34-a5f6-4c9e-9874-4310a04ca601\" width=\"100%\" controls autoplay loop muted\u003e\u003c/video\u003e\u003c/td\u003e\n    \u003ctd\u003e\u003cvideo src=\"https://github.com/user-attachments/assets/9174a01b-73e7-4920-8674-6610643f959e\" width=\"100%\" controls autoplay loop muted\u003e\u003c/video\u003e\u003c/td\u003e\n    \u003ctd\u003e\u003cvideo src=\"https://github.com/user-attachments/assets/28dd9fd9-dec6-4426-a1ea-3354633ecb4e\" width=\"100%\" controls autoplay loop muted\u003e\u003c/video\u003e\u003c/td\u003e\n  \u003c/tr\u003e\n  \u003ctr\u003e\n    \u003ctd\u003e\u003cvideo src=\"https://github.com/user-attachments/assets/39b909de-1013-4b5f-9212-3882e481aa8a\" width=\"100%\" controls autoplay loop muted\u003e\u003c/video\u003e\u003c/td\u003e\n    \u003ctd\u003e\u003cvideo src=\"https://github.com/user-attachments/assets/7dcd28c9-75c6-4ef3-932c-6097b49154e1\" width=\"100%\" controls autoplay loop muted\u003e\u003c/video\u003e\u003c/td\u003e\n    \u003ctd\u003e\u003cvideo src=\"https://github.com/user-attachments/assets/40a4de3b-27b1-4887-b0e8-b9ccd20ff959\" width=\"100%\" controls autoplay loop muted\u003e\u003c/video\u003e\u003c/td\u003e\n  \u003c/tr\u003e\n  \u003ctr\u003e\n    \u003ctd\u003e\u003cvideo src=\"https://github.com/user-attachments/assets/489a79d0-bfc8-4b2a-800c-fda71d1359f7\" width=\"100%\" controls autoplay loop muted\u003e\u003c/video\u003e\u003c/td\u003e\n    \u003ctd\u003e\u003cvideo src=\"https://github.com/user-attachments/assets/06a3d8f5-666c-4e98-a4aa-42dbb34a4c29\" width=\"100%\" controls autoplay loop muted\u003e\u003c/video\u003e\u003c/td\u003e\n    \u003ctd\u003e\u003cvideo src=\"https://github.com/user-attachments/assets/52075f49-ca11-4d86-8a12-2a9bb6f13714\" width=\"100%\" controls autoplay loop muted\u003e\u003c/video\u003e\u003c/td\u003e\n  \u003c/tr\u003e\n  \u003ctr\u003e\n    \u003ctd\u003e\u003cvideo src=\"https://github.com/user-attachments/assets/e57451f0-7d06-4913-96e3-cbc1db6f8d41\" width=\"100%\" controls autoplay loop muted\u003e\u003c/video\u003e\u003c/td\u003e\n    \u003ctd\u003e\u003cvideo src=\"https://github.com/user-attachments/assets/3c4b763f-b457-42d5-bac3-efe24cb38dc0\" width=\"100%\" controls autoplay loop muted\u003e\u003c/video\u003e\u003c/td\u003e\n    \u003ctd\u003e\u003cvideo src=\"https://github.com/user-attachments/assets/f034d5f7-b99d-4770-81bc-de1e85ac2bc2\" width=\"100%\" controls autoplay loop muted\u003e\u003c/video\u003e\u003c/td\u003e\n  \u003c/tr\u003e\n\u003c/table\u003e\n\n### Text to Video+Audio results\n\n\u003ctable border=\"0\" style=\"width: 100%; text-align: left; margin-top: 15px; border-collapse: collapse;\"\u003e\n  \u003ctr\u003e\n      \u003ctd\u003e\n          \u003cvideo src=\"https://github.com/user-attachments/assets/bde89b02-30a7-4f05-a051-5f5bba02d746\" width=\"100%\" controls autoplay loop muted\u003e\u003c/video\u003e\n      \u003c/td\u003e\n      \u003ctd\u003e\n          \u003cvideo src=\"https://github.com/user-attachments/assets/0f1ffa3a-aa79-45ac-99ad-ed8e2dc9ba64\" width=\"100%\" controls autoplay loop muted\u003e\u003c/video\u003e\n      \u003c/td\u003e\n      \u003ctd\u003e\n          \u003cvideo src=\"https://github.com/user-attachments/assets/22119867-bc8b-426a-98ab-39c66984909c\" width=\"100%\" controls autoplay loop muted\u003e\u003c/video\u003e\n      \u003c/td\u003e\n  \u003c/tr\u003e\n  \u003ctr\u003e\n      \u003ctd\u003e\n          \u003cvideo src=\"https://github.com/user-attachments/assets/b8607399-2791-4c38-9836-01958f056c29\" width=\"100%\" controls autoplay loop muted\u003e\u003c/video\u003e\n      \u003c/td\u003e\n      \u003ctd\u003e\n          \u003cvideo src=\"https://github.com/user-attachments/assets/2d268f3a-6a15-4c76-8f8a-f0b3400008ac\" width=\"100%\" controls autoplay loop muted\u003e\u003c/video\u003e\n      \u003c/td\u003e\n      \u003ctd\u003e\n          \u003cvideo src=\"https://github.com/user-attachments/assets/f2357214-a3b4-41f4-8ab6-f3943a961569\" width=\"100%\" controls autoplay loop muted\u003e\u003c/video\u003e\n      \u003c/td\u003e\n  \u003c/tr\u003e\n\u003c/table\u003e\n\n## 📰 News\n\n- **[2025-02-10] 🎉🎉🎉 [Technical Report](./Lumina%20Video%20Report%20V1.pdf) is released! 🎉🎉🎉**\n- **[2025-02-09] 🎉🎉🎉 Lumina-Video is released! 🎉🎉🎉**\n\n## ⚙️ Installation\n\nSee [INSTALL.md](./INSTALL.md) for detailed instructions.\n\n## 🤗 Checkpoints\n\n**T2V models**\n\n| resolution | fps  | max frames | Huggingface                                                  |\n| ---------- | ---- | ---------- | ------------------------------------------------------------ |\n| 960        | 24   | 96         | [Alpha-VLLM/Lumina-Video-f24R960](https://huggingface.co/Alpha-VLLM/Lumina-Video-f24R960) |\n\n## ⛽ Inference\n\n### Preparations\n\nDownload the checkpoints before continue. You can use the following code to download the checkpoints to the `./ckpts` directory\n\n```\nhuggingface-cli download --resume-download Alpha-VLLM/Lumina-Video-f24R960 --local-dir ./ckpts/f24R960\n```\n\n### Inference\n\nYou can quickly run video generation using the command below:\n\n\n```bash\n# Example for generatingan video with 4s duration, fps=24, resolution=1248x704\npython -u generate.py \\\n    --ckpt ./ckpts/f24R960 \\\n    --resolution 1248x704 \\\n    --fps 24 \\\n    --frames 96 \\\n    --prompt \"your prompt here\" \\\n    --neg_prompt \"\" \\\n    --sample_config f24F96R960  # set to \"f24F96R960-MultiScale\" for efficient multi-scale inference\n```\n\n#### QAs\n\n**Q1**: Why using the 1248x704 resolution?\n\n**A1**: The resolution is originally expected to be 1280x720. However, to ensure compatibility with the largest patch size\n(smallest scale), both the width and height must be divisible by 32. As a result, the resolution is adjusted to\n1248x704.\n\n**Q2**: Does the model support flexible aspect ratio?\n\n**A2**: Yes, you can use the following code for checking all usable resolutions\n\n```Python\n# Python\nfrom imgproc import generate_crop_size_list\n\ntarget_size = 960\npatch_size = 32\nmax_num_patches = (target_size // patch_size) ** 2\ncrop_size_list = generate_crop_size_list(max_num_patches, patch_size)\n\nprint(crop_size_list)\n```\n\n## Training\n\n### Preparations\n\nBefore starting the training process, two preparation steps are required to optimize training efficiency and enable motion conditioning:\n\n1. **Pre-extract and cache VAE latents for video data**: This significantly enhances training speed.\n2. **Compute motion scores for videos**: These are used for micro-conditioning input during training.\n\n#### Pre-Extract VAE Latents\n\nThe code for pre-extracting and caching VAE latents can be found in the [./tools/pre_extract](tools/pre_extract) directory. For an example of how to run this, refer to the [run.sh](tools/pre_extract/scripts/run.sh) script.\n\n#### Compute Motion Score\n\nWe use UniMatch to estimate optical flow, with the average optical flow serving as the motion score. This code is primarily derived from [Open-Sora](https://github.com/hpcaitech/Open-Sora/tree/main/tools/scoring/optical_flow), and we'd like to thank them for their excellent work!\n\nThe code for computing motion scores is available in the [./tools/unimatch](tools/unimatch) directory. To see how to run it, refer to the [run.sh](tools/unimatch/scripts/run.sh) script.\n\n### Training\n\nOnce the data has been prepared, you're ready to start training! For an example, you can refer to the [training directory](train_exps/f8F32R256), which demonstrates how to train with:\n\n- **FPS**: 8\n- **Duration**: 4 seconds\n- **Resolution**: widthxheight≈256x256\n- **Training Techniques**: Image-text joint training and multi-scale training applied together.\n\n\n\n\n## 📑 Open-source Plan\n\n- [X] Inference code\n- [X] Training code\n\n\n## 📃 Citation\n\n```bash\n@misc{luminavideo,\n      title={Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT}, \n      author={Dongyang Liu and Shicheng Li and Yutong Liu and Zhen Li and Kai Wang and Xinyue Li and Qi Qin and Yufei Liu and Yi Xin and Zhongyu Li and Bin Fu and Chenyang Si and Yuewen Cao and Conghui He and Ziwei Liu and Yu Qiao and Qibin Hou and Hongsheng Li and Peng Gao},\n      year={2025},\n      eprint={2502.06782},\n      archivePrefix={arXiv},\n      primaryClass={cs.CV},\n      url={https://arxiv.org/abs/2502.06782}, \n}\n```\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Falpha-vllm%2Flumina-video","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Falpha-vllm%2Flumina-video","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Falpha-vllm%2Flumina-video/lists"}