{"id":13421973,"url":"https://github.com/GuoleiSun/VSS-MRCFA","last_synced_at":"2025-03-15T10:31:40.814Z","repository":{"id":46754789,"uuid":"515641491","full_name":"GuoleiSun/VSS-MRCFA","owner":"GuoleiSun","description":"Official code for ECCV 2022 paper","archived":false,"fork":false,"pushed_at":"2024-06-07T13:11:39.000Z","size":14616,"stargazers_count":30,"open_issues_count":1,"forks_count":5,"subscribers_count":1,"default_branch":"main","last_synced_at":"2024-10-27T22:29:56.031Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":null,"language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/GuoleiSun.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null}},"created_at":"2022-07-19T15:28:37.000Z","updated_at":"2024-09-02T08:10:18.000Z","dependencies_parsed_at":"2023-01-20T09:31:34.955Z","dependency_job_id":null,"html_url":"https://github.com/GuoleiSun/VSS-MRCFA","commit_stats":null,"previous_names":[],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/GuoleiSun%2FVSS-MRCFA","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/GuoleiSun%2FVSS-MRCFA/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/GuoleiSun%2FVSS-MRCFA/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/GuoleiSun%2FVSS-MRCFA/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/GuoleiSun","download_url":"https://codeload.github.com/GuoleiSun/VSS-MRCFA/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":243719090,"owners_count":20336591,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2024-07-30T23:00:34.776Z","updated_at":"2025-03-15T10:31:35.798Z","avatar_url":"https://github.com/GuoleiSun.png","language":"Python","funding_links":[],"categories":["Video Semantic Segmentation"],"sub_categories":["2022"],"readme":"# VSS-MRCFA\nOfficial PyTorch implementation of ECCV 2022 paper: Mining Relations among Cross-Frame Affinities for Video Semantic Segmentation\n\n## Abstract\nThe essence of video semantic segmentation (VSS) is how to leverage temporal information for prediction. Previous efforts are mainly devoted to developing new techniques to calculate the cross-frame affinities such as optical flow and attention. Instead, this paper contributes from a different angle by  mining relations among cross-frame affinities, upon which better temporal information aggregation could be achieved. We explore relations among affinities in two aspects: single-scale intrinsic correlations and multi-scale relations. Inspired by traditional feature processing, we propose Single-scale Affinity Refinement (SAR) and Multi-scale Affinity Aggregation (MAA). To make it feasible to execute MAA, we propose a Selective Token Masking (STM) strategy to select a subset of consistent reference tokens for different scales when calculating affinities, which also improves the efficiency of our method. At last, the cross-frame affinities strengthened by SAR and MAA are adopted for adaptively aggregating temporal information. Our experiments demonstrate that the proposed method performs favorably against state-of-the-art VSS methods.\n\n![block images](https://github.com/GuoleiSun/VSS-MRCFA/blob/main/Figs/diagram.png)\n\nAuthors: [Guolei Sun](https://scholar.google.com/citations?hl=zh-CN\u0026user=qd8Blw0AAAAJ), [Yun Liu](https://yun-liu.github.io/), [Hao Tang](https://scholar.google.com/citations?user=9zJkeEMAAAAJ\u0026hl=en), [Ajad Chhatkuli](https://scholar.google.com/citations?user=3BHMHU4AAAAJ\u0026hl=en), [Le Zhang](https://zhangleuestc.github.io), Luc Van Gool.\n\n## Note\nThis is a preliminary version for early access and I will clean it for better readability.\n\n## Installation\nPlease follow the guidelines in [MMSegmentation v0.13.0](https://github.com/open-mmlab/mmsegmentation/tree/v0.13.0).\n\nOther requirements:\n```timm==0.3.0, CUDA11.0, pytorch==1.7.1, torchvision==0.8.2, mmcv==1.3.0, opencv-python==4.5.2```\n\nDownload this repository and install by:\n```\ncd VSS-MRCFA \u0026\u0026 pip install -e . --user\n```\n\n## Usage\n### Data preparation\nPlease follow [VSPW](https://github.com/sssdddwww2/vspw_dataset_download) to download VSPW 480P dataset.\nAfter correctly downloading, the file system is as follows:\n```\nvspw-480\n├── video1\n    ├── origin\n        ├── .jpg\n    └── mask\n        └── .png\n```\nThe dataset should be put in ```VSS-MRCFA/data/vspw/```. Or you can use Symlink: \n```\ncd VSS-MRCFA\nmkdir -p data/vspw/\nln -s /dataset_path/VSPW_480p data/vspw/\n```\n\n### Test\n1. Download the trained weights from [here](https://drive.google.com/drive/folders/1GIKt21UBYjXqi0Zm_azc6SrrIcK__Lyq?usp=sharing).\n2. Run the following commands:\n```\n# Multi-gpu testing\n./tools/dist_test.sh local_configs/mrcfa/B1/mrcfa.b1.480x480.vspw2.160k.py /path/to/checkpoint_file \u003cGPU_NUM\u003e \\\n--out /path/to/save_results/res.pkl\n```\n\n### Training\nTraining requires 4 Nvidia GPUs, each of which has \u003e 20G GPU memory.\n```\n# Multi-gpu training\n./tools/dist_train.sh local_configs/mrcfa/B1/mrcfa.b1.480x480.vspw2.160k.py 4 --work-dir model_path/vspw2/work_dirs_4g_b1\n```\n\n## License\nThis project is only for academic use. For other purposes, please contact us.\n\n## Acknowledgement\nThe code is heavily based on the following repositories:\n- https://github.com/open-mmlab/mmsegmentation\n- https://github.com/NVlabs/SegFormer\n- https://github.com/GuoleiSun/VSS-CFFM\n\nThanks for their amazing works.\n\n## Citation\n```\n@article{sun2022mining,\n  title={Mining Relations among Cross-Frame Affinities for Video Semantic Segmentation},\n  author={Sun, Guolei and Liu, Yun and Tang, Hao and Chhatkuli, Ajad and Zhang, Le and Van Gool, Luc},\n  journal={arXiv preprint arXiv:2207.10436},\n  year={2022}\n}\n```\n\n## Contact\n- Guolei Sun, sunguolei.kaust@gmail.com\n- Yun Liu, yun.liu@vision.ee.ethz.ch\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FGuoleiSun%2FVSS-MRCFA","html_url":"https://awesome.ecosyste.ms/projects/github.com%2FGuoleiSun%2FVSS-MRCFA","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FGuoleiSun%2FVSS-MRCFA/lists"}