{"id":20417435,"url":"https://github.com/hzwer/cvpr2023-dmvfn","last_synced_at":"2025-05-16T12:00:21.230Z","repository":{"id":143906297,"uuid":"614215011","full_name":"hzwer/CVPR2023-DMVFN","owner":"hzwer","description":"CVPR2023 (highlight) - A Dynamic Multi-Scale Voxel Flow Network for Video Prediction","archived":false,"fork":false,"pushed_at":"2024-12-04T06:18:37.000Z","size":1072,"stargazers_count":351,"open_issues_count":1,"forks_count":9,"subscribers_count":6,"default_branch":"main","last_synced_at":"2025-04-09T07:04:07.623Z","etag":null,"topics":["computer-vision","cvpr2023","deep-learning","pytorch"],"latest_commit_sha":null,"homepage":"","language":"Jupyter Notebook","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/hzwer.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2023-03-15T06:05:58.000Z","updated_at":"2025-04-05T14:27:08.000Z","dependencies_parsed_at":"2024-12-16T03:04:15.043Z","dependency_job_id":"63532ca0-40b2-4c17-bb05-98d062bfdcef","html_url":"https://github.com/hzwer/CVPR2023-DMVFN","commit_stats":null,"previous_names":["hzwer/cvpr2023-dmvfn"],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/hzwer%2FCVPR2023-DMVFN","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/hzwer%2FCVPR2023-DMVFN/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/hzwer%2FCVPR2023-DMVFN/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/hzwer%2FCVPR2023-DMVFN/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/hzwer","download_url":"https://codeload.github.com/hzwer/CVPR2023-DMVFN/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":254527071,"owners_count":22085917,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["computer-vision","cvpr2023","deep-learning","pytorch"],"created_at":"2024-11-15T06:26:24.540Z","updated_at":"2025-05-16T12:00:21.122Z","avatar_url":"https://github.com/hzwer.png","language":"Jupyter Notebook","funding_links":[],"categories":[],"sub_categories":[],"readme":"# A Dynamic Multi-Scale Voxel Flow Network for Video Prediction\n## [HomePage](https://huxiaotaostasy.github.io/DMVFN) | [Colab](https://colab.research.google.com/github/megvii-research/CVPR2023-DMVFN/blob/main/colab_demo.ipynb) | [arXiv](https://arxiv.org/abs/2303.09875) | [YouTube](https://youtu.be/6retLBLAHDs)\nThis project is the implementation of our Paper: [A Dynamic Multi-Scale Voxel Flow Network for Video Prediction](https://arxiv.org/abs/2303.09875), which is accepted by **CVPR2023 (highlight✨, 10% of accepted papers).** We proposed a SOTA model for Video Prediction.\n\n[Poster](https://drive.google.com/file/d/1W5MD39Wjp4Ryrv_DMC7QXiCNsiG1AL2x/view?usp=sharing) | [研究历程](https://www.zhihu.com/question/585474435/answer/2946859435) | [中文论文](https://drive.google.com/file/d/1mzJipcWRd1wcpphSIUOszzgxMWMFx40D/view?usp=sharing) | [rebuttal (3WA-\u003e1WA2SA)](https://drive.google.com/file/d/1S8tZOokoZNxCBkV1GVLWOlUCCoJA7Bfu/view?usp=sharing) | [Demo](https://youtu.be/rlghCGbAqUo)\n\n![Comparison image](./images/comparison.jpg)\n\nsupplement and correction of published paper: 1. In Figure 2(a), $I_{t-1}$ and $I_t$ should be swapped; 2. \"Our model is trained on four 2080Ti GPUs for 300 epochs, which takes about 35 hours.\" This sentence refers to the Cityscapes training set; 3. In Table 1, GFLOPs computing for KITTI is on $256\\times832$ and on $1024\\times512$ for Cityscapes using [fvcore](https://github.com/facebookresearch/fvcore).\n## Usage\n### Installation\n\n```bash\ngit clone https://github.com/megvii-research/CVPR2023-DMVFN.git\ncd CVPR2023-DMVFN\npip3 install -r requirements.txt\n```\n\n* Download the pretrained models from [Google Drive](https://drive.google.com/drive/folders/1QnsnTVU-2IexFNjGhNVRYPqKNviAbkwH?usp=sharing). ([百度网盘](https://pan.baidu.com/s/1l8d1oCzEHDf8JCv1KLIbvA) password:33ly), and move the pretrained parameters to `CVPR2023-DMVFN/pretrained_models/*`\n\n```bash\npip install gdown\nmkdir pretrained_models \u0026\u0026 cd pretrained_models\ngdown --id 1jILbS8Gm4E5Xx4tDCPZh_7rId0eo8r9W\ngdown --id 1WrV30prRiS4hWOQBnVPUxdaTlp9XxmVK\ngdown --id 14_xQ3Yl3mO89hr28hbcQW3h63lLrcYY0\ncd ..\n```\n\n### Data Preparation\n\nIn this section, we will download all parts including training and testing sets. If you only need the test set, please jump to [Directly download test splits](#directly_download_test_splits).\n\nThe final data folder `CVPR2023-DMVFN/data/` should be organized like this:\n\n```\ndata\n├── Cityscapes\n│   ├── train (Citysapes images in 512x1024)\n│   │   └── 000000\n│   │   └── ...\n│   │   └── 002974\n│   ├── test (Citysapes images in 512x1024)\n│   │   └── 000000\n│   │   └── ...\n│   │   └── 000499\n├── KITTI\n│   ├── train (Kitti images in 256x832)\n│   │   └── 000000\n│   │   └── ...\n│   │   └── 013499\n│   ├── test (Kitti images in 256x832)\n│   │   └── 000000\n│   │   └── ...\n│   │   └── 001336\n├── UCF101\n│   ├── v_ApplyEyeMakeup_g08_c01\n│   └── ...\n└── vimeo_interp_test\n    └── target\n        └── 00001\n        └── ...\n```\n\n\n**Cityscapes**\n\n\n* Download the Cityscapes dataset `leftImg8bit_sequence_trainvaltest.zip` from [here](https://www.cityscapes-dataset.com/downloads/).\n\n\n* Unzip `leftImg8bit_sequence_trainvaltest.zip`.\n\n\n```bash\nunzip leftImg8bit_sequence_trainvaltest.zip\n```\n\n* Run ./utils/prepare_city.py\n```bash\npython3 ./utils/prepare_city.py\n```\n\n**KITTI**\n\n\n* Download the KITTI dataset from [Google Drive](https://www.cvlibs.net/datasets/kitti/raw_data.php). You need to register and login, then download all videos.\n\n\n* Our training split and testing split are consistent with [YueWuHKUST/CVPR2020-FutureVideoSynthesis](https://github.com/YueWuHKUST/CVPR2020-FutureVideoSynthesis).\n\n```\nVideos for training split:\n2011_09_26_drive_0001_sync  2011_09_26_drive_0018_sync  2011_09_26_drive_0104_sync 2011_09_26_drive_0002_sync  2011_09_26_drive_0048_sync\n2011_09_26_drive_0106_sync  2011_09_26_drive_0005_sync  2011_09_26_drive_0051_sync 2011_09_26_drive_0113_sync  2011_09_26_drive_0009_sync  \n2011_09_26_drive_0056_sync  2011_09_26_drive_0117_sync  2011_09_26_drive_0011_sync 2011_09_26_drive_0057_sync  2011_09_28_drive_0001_sync\n2011_09_26_drive_0013_sync  2011_09_26_drive_0059_sync  2011_09_28_drive_0002_sync 2011_09_26_drive_0014_sync  2011_09_26_drive_0091_sync \n2011_09_29_drive_0026_sync  2011_09_26_drive_0017_sync  2011_09_26_drive_0095_sync 2011_09_29_drive_0071_sync\n\nVideos for testing split:\n2011_09_26_drive_0060_sync  2011_09_26_drive_0084_sync  2011_09_26_drive_0093_sync  2011_09_26_drive_0096_sync\n```\n* Unzip all files, then reorganize the dataset as follows:\n\n```bash\nmkdir train_or\nunzip 2011_09_26_drive_0001_sync\nmv 2011_09_26_drive_0001_sync/image02/data/ train_or/2011_09_26_drive_0001_sync\n```\nDo this command for all files. We use image02 and image03 for training.\n\n* Run ./utils/prepare_kitti.py.\n\n\n```bash\npython3 ./utils/prepare_kitti.py\n```\n* Do the same for the testing split.\n\n**UCF101**\n\nWe extract RGB frames from each video in UCF101 dataset and save as `.jpg` image.\n\nDownload the preprocessed data directly from [feichtenhofer/twostreamfusion](https://github.com/feichtenhofer/twostreamfusion) for convenience.\n\n```bash\nwget http://ftp.tugraz.at/pub/feichtenhofer/tsfusion/data/ucf101_jpegs_256.zip.001\nwget http://ftp.tugraz.at/pub/feichtenhofer/tsfusion/data/ucf101_jpegs_256.zip.002\nwget http://ftp.tugraz.at/pub/feichtenhofer/tsfusion/data/ucf101_jpegs_256.zip.003\n\ncat ucf101_jpegs_256.zip* \u003e ucf101_jpegs_256.zip\nunzip ucf101_jpegs_256.zip\n```\n\n\n**Vimeo90K**\n\n* Download Vimeo90K dataset directly from [here](http://toflow.csail.mit.edu/).\n* Unzip the dataset.\n\n### Run\n\n#### 😆Training\n\nFor Cityscapes Dataset:\n\n```bash\npython3 -m torch.distributed.launch --nproc_per_node=8 \\\n--master_port=4321 ./scripts/train.py \\\n--train_dataset CityTrainDataset \\\n--val_datasets CityValDataset \\\n--batch_size 8 \\\n--num_gpu 8\n```\n\nFor KITTI Dataset:\n\n```bash\npython3 -m torch.distributed.launch --nproc_per_node=8 \\\n--master_port=4321 ./scripts/train.py \\\n--train_dataset KittiTrainDataset \\\n--val_datasets KittiValDataset \\\n--batch_size 8 \\\n--num_gpu 8\n```\n\n\nFor DAVIS and Vimeo Dataset:\n\n```bash\npython3 -m torch.distributed.launch --nproc_per_node=8 \\\n--master_port=4321 ./scripts/train.py \\\n--train_dataset UCF101TrainDataset \\\n--val_datasets DavisValDataset VimeoValDataset \\\n--batch_size 8 \\\n--num_gpu 8\n```\n\n#### 🤔️Testing\n\n**\u003cspan id=\"directly_download_test_splits\"\u003e Directly download test splits of different datasets\u003c/span\u003e**\n\nDownload Cityscapes_test directly from [Google Drive](https://drive.google.com/file/d/1m5lfwGa6ugavZW9-UFrXNpQ0BWT7qn7X/view?usp=share_link).([百度网盘](https://pan.baidu.com/s/1KFxPC-zFi9LJwwed0PxEow) password: wk7k)\n\nDownload KITTI_test directly from [Google Drive](https://drive.google.com/file/d/1_J5QxvozXiLoF3xca0uh9a6dEMVZJtWs/view?usp=drive_link).([百度网盘](https://pan.baidu.com/s/1pG31uHts3lV2ieouZZovcQ) password: e7da)\n\nDownload DAVIS_test directly from [Google Drive](https://drive.google.com/file/d/10w1ox4ADtPdmBmYFhxHycFHcQg97PsK-/view?usp=drive_link).([百度网盘](https://pan.baidu.com/s/1ZydU6z5Y9DRQ1lBGQk8ynQ) password: mczk)\n\nDownload Vimeo_test directly from [Google Drive](https://drive.google.com/file/d/1ERswpm1E_eeS10XnGv9qH75k3-y96VwU/view?usp=drive_link).([百度网盘](https://pan.baidu.com/s/1fsTQBhQHfrPMVhtSf-77TA) password: 0mjo)\n\nRun the following command to generate test results of DMVFN model. The `--val_datasets` can be `CityValDataset`, `KittiValDataset`, `DavisValDataset`, and `VimeoValDataset`. `--save_image` can be disabled.\n\n```bash\npython3 ./scripts/test.py \\\n--val_datasets CityValDataset [optional: KittiValDataset, DavisValDataset, VimeoValDataset] \\\n--load_path path_of_pretrained_weights \\\n--save_image \n```\n\n**Image results**\n\nWe provide the image results of DMVFN on various datasets (Cityscapes, KITTI, DAVIS and Vimeo) in [百度网盘](https://pan.baidu.com/s/19BWu33raS49Wamw5iC96rA) (password: k7eb).\n\nWe also provide the results of DMVFN (without routing) in [百度网盘](https://pan.baidu.com/s/1pW61ITp5MFLvyQCr44SkEg) (password: 8zo9).\n\n**Test the image results**\n\nRun the following command to directly test the image results.\n\n```bash\npython3 ./scripts/test_ssim_lpips.py\n```\n\n#### 😋Single test\n\nWe provide a simple code to predict a `t+1` image with `t-1` and `t` images. Please run the following command:\n\n```bash\npython3 ./scripts/single_test.py \\\n--image_0_path ./images/sample_img_0.png \\\n--image_1_path ./images/sample_img_1.png \\\n--load_path path_of_pretrained_weights \\\n--output_dir pred.png\n```\n\n## Recommend\nWe sincerely recommend some related papers:\n\nECCV22 - [Real-Time Intermediate Flow Estimation for Video Frame Interpolation](https://github.com/megvii-research/ECCV2022-RIFE)\n\nCVPR22 - [Optimizing Video Prediction via Video Frame Interpolation](https://github.com/YueWuHKUST/CVPR2022-Optimizing-Video-Prediction-via-Video-Frame-Interpolation)\n\n## Citation\nIf you think this project is helpful, please feel free to leave a star or cite our paper:\n```\n@inproceedings{hu2023dmvfn,\n  title={A Dynamic Multi-Scale Voxel Flow Network for Video Prediction},\n  author={Hu, Xiaotao and Huang, Zhewei and Huang, Ailin and Xu, Jun and Zhou, Shuchang},\n  booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},\n  year={2023}\n}\n```\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fhzwer%2Fcvpr2023-dmvfn","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fhzwer%2Fcvpr2023-dmvfn","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fhzwer%2Fcvpr2023-dmvfn/lists"}