{"id":27344462,"url":"https://github.com/ZhaochongAn/Multimodality-3D-Few-Shot","last_synced_at":"2025-04-12T17:06:27.256Z","repository":{"id":277009555,"uuid":"880378698","full_name":"ZhaochongAn/Multimodality-3D-Few-Shot","owner":"ZhaochongAn","description":"[ICLR 2025 Spotlight] Multimodality Helps Few-shot 3D Point Cloud Semantic Segmentation","archived":false,"fork":false,"pushed_at":"2025-03-09T01:12:55.000Z","size":0,"stargazers_count":22,"open_issues_count":0,"forks_count":1,"subscribers_count":5,"default_branch":"main","last_synced_at":"2025-03-09T01:25:56.259Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/ZhaochongAn.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2024-10-29T16:05:42.000Z","updated_at":"2025-03-09T01:12:58.000Z","dependencies_parsed_at":"2025-02-11T17:25:57.690Z","dependency_job_id":"511b6476-2113-4dc5-86f1-ffaaef054e3b","html_url":"https://github.com/ZhaochongAn/Multimodality-3D-Few-Shot","commit_stats":null,"previous_names":["zhaochongan/multimodality-3d-few-shot"],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/ZhaochongAn%2FMultimodality-3D-Few-Shot","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/ZhaochongAn%2FMultimodality-3D-Few-Shot/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/ZhaochongAn%2FMultimodality-3D-Few-Shot/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/ZhaochongAn%2FMultimodality-3D-Few-Shot/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/ZhaochongAn","download_url":"https://codeload.github.com/ZhaochongAn/Multimodality-3D-Few-Shot/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":248602311,"owners_count":21131615,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2025-04-12T17:02:17.737Z","updated_at":"2025-04-12T17:06:27.240Z","avatar_url":"https://github.com/ZhaochongAn.png","language":"Python","funding_links":[],"categories":["3D视觉生成重建"],"sub_categories":["资源传输下载"],"readme":"\u003cp align=\"center\"\u003e\n  \u003ch1 align=\"center\"\u003eMultimodality Helps Few-shot 3D Point Cloud Semantic Segmentation\u003c/h1\u003e\n  \u003cp align=\"center\"\u003e\n    \u003ca href=\"https://zhaochongan.github.io/\"\u003e\u003cstrong\u003eZhaochong An\u003c/strong\u003e\u003c/a\u003e\n    ·\n    \u003ca href=\"https://guoleisun.github.io/\"\u003e\u003cstrong\u003eGuolei Sun\u003csup\u003e†\u003c/sup\u003e\u003c/strong\u003e\u003c/a\u003e\n    ·\n    \u003ca href=\"https://yun-liu.github.io/\"\u003e\u003cstrong\u003eYun Liu\u003csup\u003e†\u003c/sup\u003e\u003c/strong\u003e\u003c/a\u003e\n    ·\n    \u003ca href=\"https://runjiali-rl.github.io/\"\u003e\u003cstrong\u003eRunjia Li\u003c/strong\u003e\u003c/a\u003e\n    ·\n    \u003ca href=\"https://sites.google.com/site/wumincf/\"\u003e\u003cstrong\u003eMin Wu\u003c/strong\u003e\u003c/a\u003e\n    \u003cbr\u003e\n    \u003ca href=\"https://mmcheng.net/cmm/\"\u003e\u003cstrong\u003eMing-Ming Cheng\u003c/strong\u003e\u003c/a\u003e\n    ·\n    \u003ca href=\"https://people.ee.ethz.ch/~kender/\"\u003e\u003cstrong\u003eEnder Konukoglu\u003c/strong\u003e\u003c/a\u003e\n    ·\n    \u003ca href=\"https://sergebelongie.github.io/\"\u003e\u003cstrong\u003eSerge Belongie\u003c/strong\u003e\u003c/a\u003e\n  \u003c/p\u003e\n  \u003ch2 align=\"center\"\u003eICLR 2025 Spotlight (\u003ca href=\"https://arxiv.org/pdf/2410.22489\"\u003ePaper\u003c/a\u003e)\u003c/h2\u003e\n\u003c/p\u003e\n\n\u003cp align=\"center\"\u003e\n  \u003cimg src=\"https://ZhaochongAn.github.io/images/MMFSS_github.png\" alt=\"Overview\" width=\"80%\"\u003e\n\u003c/p\u003e\n\n## 🌟 Highlights\n\nWe introduce:\n- A novel **cost-free multimodal few-shot 3D point cloud segmentation (FS-PCS) setup** that integrates textual category names and 2D image modality\n- **MM-FSS**: The first multimodal FS-PCS model that explicitly utilizes textual modality and implicitly leverages 2D modality\n- Superior performance on novel class generalization through effective multimodal integration\n- Valuable insights into the importance of commonly-ignored free modalities in FS-PCS\n\n## 🛠️ Environment Setup\n\nOur environment has been tested on:\n- RTX 3090 GPUs\n- GCC 6.3.0\n\nFollow the [COSeg installation guide](https://github.com/ZhaochongAn/COSeg?tab=readme-ov-file#environment) for detailed setup.\n\n## 📦 Dataset Preparation\n\n### Pretraining Stage Data\nFollow [OpenScene](https://github.com/pengsongyou/openscene?tab=readme-ov-file#data-preparation) instructions, you can \ndirectly download the following ScanNet 3D dataset and 2D features for pretraining:\n```bash\n# Download ScanNet 3D dataset\nwget https://cvg-data.inf.ethz.ch/openscene/data/scannet_processed/scannet_3d.zip\nunzip scannet_3d.zip\n\n# Download 2D features\nwget https://cvg-data.inf.ethz.ch/openscene/data/scannet_multiview_lseg.zip\nunzip scannet_multiview_lseg.zip\n```\n\nYou should put the unpacked data into the folder ./pretraining/data/ or link to the corresponding data folder with the symbolic link:\n```bash\nln -s /PATH/TO/DOWNLOADED/FOLDER ./pretraining/data\n```\n\n### Few-shot Stage Data\n#### Option 1: Direct Download (Recommended)\nDownload our preprocessed datasets:\n\n| Dataset | Few-shot Stage Data |\n|:-------:|:---------------------:|\n| S3DIS | [Download](https://drive.google.com/file/d/1frJ8nf9XLK_fUBG4nrn8Hbslzn7914Ru/view?usp=drive_link) |\n| ScanNet | [Download](https://drive.google.com/file/d/19yESBZumU-VAIPrBr8aYPaw7UqPia4qH/view?usp=drive_link) |\n\n#### Option 2: Manual Preprocessing\n\n\n\n\nFollow [COSeg](https://github.com/ZhaochongAn/COSeg?tab=readme-ov-file#datasets-preparation) preprocessing instructions.\nThe processed data will be in `[PATH_to_DATASET_processed_data]/blocks_bs1_s1/data`. Make sure to update the `data_root` entry in the .yaml \nconfig file to `[PATH_to_DATASET_processed_data]/blocks_bs1_s1/data`.\n\n\n## 🔄 Training Pipeline\n\n### 1. Backbone and IF Head Pretraining\n\n**Option A**: Download our pretrained weights from [Google Drive](https://drive.google.com/drive/u/1/folders/1JoeAXJh1AZM3bM0KGBJQsFTad6uqpzUJ)\n\n**Option B**: Train from scratch:\n```bash\ncd pretraining\nbash run/distill_strat.sh PATH_to_SAVE_BACKBONE config/scannet/ours_lseg_strat.yaml\n```\n\n### 2. Meta-learning Stage\nSet config `config/[CONFIG_FILE]` to be `s3dis_COSeg_fs.yaml` or `scannetv2_COSeg_fs.yaml` for training on S3DIS or ScanNet respectively.\nAdjust `cvfold`, `n_way`, and `k_shot` according to your few-shot task:\n\n```bash\n# For 1-way tasks\npython3 main_fs.py --config config/[CONFIG_FILE] \\\n    save_path [PATH_to_SAVE_MODEL] \\\n    pretrain_backbone [PATH_to_SAVED_BACKBONE] \\\n    cvfold [CVFOLD] \\\n    n_way 1 \\\n    k_shot [K_SHOT] \\\n    num_episode_per_comb 1000\n\n# For 2-way tasks\npython3 main_fs.py --config config/[CONFIG_FILE] \\\n    save_path [PATH_to_SAVE_MODEL] \\\n    pretrain_backbone [PATH_to_SAVED_BACKBONE] \\\n    cvfold [CVFOLD] \\\n    n_way 2 \\\n    k_shot [K_SHOT] \\\n    num_episode_per_comb 100\n```\n\n\u003e **Note**: Following [COSeg](https://github.com/ZhaochongAn/COSeg?tab=readme-ov-file#training-pipeline), `num_episode_per_comb` defaults to 1000 for 1-way and 100 for 2-way tasks to maintain consistency in test set size.\n\n## 📊 Evaluation \u0026 Visualization\n\n\n### Model Evaluation\nModify `cvfold`, `n_way`, `k_shot` and `num_episode_per_comb` accordingly and run:\n```bash\npython3 main_fs.py --config config/[CONFIG_FILE] \\\n    test True \\\n    eval_split test \\\n    weight [PATH_to_SAVED_MODEL] \\\n    [vis 1]  # Optional: Enable W\u0026B visualization\n```\n\n\u003e **Note**: Performance may vary by 1.0% due to potential randomness in the training process. ScanNetv2 typically shows less variance than S3DIS.\n\n### Visualization\nFollow [COSeg visualization guide](https://github.com/ZhaochongAn/COSeg?tab=readme-ov-file#visualization) for high-quality visualization results.\n\n## 🎯 Model Zoo\n\n| Model | Dataset | CVFOLD | N-way K-shot | Weights |\n|:-------:|:---------:|:--------:|:------------:|:----------:|\n| s30_1w1s | S3DIS | 0 | 1-way 1-shot | [Download](https://drive.google.com/drive/u/1/folders/1XKxEnvT_VdVa9kP5P6DeXRQoC1YJyMK-) |\n| s30_1w5s | S3DIS | 0 | 1-way 5-shot | [Download](https://drive.google.com/drive/u/1/folders/1dd3JmuLwLT6V03bsg_0J4ISLDnvoAUDq) |\n| s30_2w1s | S3DIS | 0 | 2-way 1-shot | [Download](https://drive.google.com/drive/u/1/folders/1kJif7istSwHbsbeHQoI4sfQgDdF19T6v) |\n| s30_2w5s | S3DIS | 0 | 2-way 5-shot | [Download](https://drive.google.com/drive/u/1/folders/1F17vApLTZFt2x85OjJtR6ryha0xDW6kV) |\n| s31_1w1s | S3DIS | 1 | 1-way 1-shot | [Download](https://drive.google.com/drive/u/1/folders/1GK9pwWbti61mLxmCbSb40inr1QU42FgF) |\n| s31_1w5s | S3DIS | 1 | 1-way 5-shot | [Download](https://drive.google.com/drive/u/1/folders/1EeyruLVk0ONXDBQ1W-pDVZAq6VPCC3jx) |\n| s31_2w1s | S3DIS | 1 | 2-way 1-shot | [Download](https://drive.google.com/drive/u/1/folders/1m11yaTi7nm4_hfBWzZUAwM1G1I8kXZNj) |\n| s31_2w5s | S3DIS | 1 | 2-way 5-shot | [Download](https://drive.google.com/drive/u/1/folders/1ytilaDjiHFUCqK-YGSqxDByWnWFFvvuR) |\n| sc0_1w1s | ScanNet | 0 | 1-way 1-shot | [Download](https://drive.google.com/drive/u/1/folders/1krip2sLd9kkaq5viTdsoaPnFgRBG64w6) |\n| sc0_1w5s | ScanNet | 0 | 1-way 5-shot | [Download](https://drive.google.com/drive/u/1/folders/1wGc3zv-ZwEpa_jNDSXWX64O4uRrYOFfI) |\n| sc0_2w1s | ScanNet | 0 | 2-way 1-shot | [Download](https://drive.google.com/drive/u/1/folders/1rgLyb1Q6VoxgyQj_Eqfn4g-dcEeY-KYZ) |\n| sc0_2w5s | ScanNet | 0 | 2-way 5-shot | [Download](https://drive.google.com/drive/u/1/folders/106_3fYBakpbMHwkknoEaGHFeIt42GAYW) |\n| sc1_1w1s | ScanNet | 1 | 1-way 1-shot | [Download](https://drive.google.com/drive/u/1/folders/1fsljMc0lrqB-kMAQD85CmSU02qiFAt_z) |\n| sc1_1w5s | ScanNet | 1 | 1-way 5-shot | [Download](https://drive.google.com/drive/u/1/folders/1MVEOV1ZZg3xQuwWhoNeeHpXBJ3kRCPEE) |\n| sc1_2w1s | ScanNet | 1 | 2-way 1-shot | [Download](https://drive.google.com/drive/u/1/folders/1y_OVENsKy5RbeJ77CuwdJKXMbO_ZdtBx) |\n| sc1_2w5s | ScanNet | 1 | 2-way 5-shot | [Download](https://drive.google.com/drive/u/1/folders/189HZgypuF9KWEVZ3tPW4bJ4QU1jk-2f-) |\n\n## 📝 Citation\nIf you find our code or paper useful, please cite:\n\n\n```bibtex\n@article{an2024multimodality,\n    title={Multimodality Helps Few-Shot 3D Point Cloud Semantic Segmentation},\n    author={An, Zhaochong and Sun, Guolei and Liu, Yun and Li, Runjia and Wu, Min \n            and Cheng, Ming-Ming and Konukoglu, Ender and Belongie, Serge},\n    journal={arXiv preprint arXiv:2410.22489},\n    year={2024}\n}\n```\n\nFor any questions or issues, feel free to reach out!\n\n- **Email**: anzhaochong@outlook.com\n- **Join in our Communication Group (WeChat)**:\n\u003cdiv style=\"text-align: left;\"\u003e\n    \u003cimg src=\"https://files.mdnice.com/user/67517/068822c4-cece-4ac5-b1db-5c138a91a718.png\" width=\"200\"/\u003e\n\u003c/div\u003e","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FZhaochongAn%2FMultimodality-3D-Few-Shot","html_url":"https://awesome.ecosyste.ms/projects/github.com%2FZhaochongAn%2FMultimodality-3D-Few-Shot","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FZhaochongAn%2FMultimodality-3D-Few-Shot/lists"}