{"id":26159549,"url":"https://github.com/zh-plus/awesome-interactive-embodiedai","last_synced_at":"2026-01-30T19:31:58.883Z","repository":{"id":267434060,"uuid":"901237153","full_name":"zh-plus/awesome-interactive-EmbodiedAI","owner":"zh-plus","description":null,"archived":false,"fork":false,"pushed_at":"2025-03-04T08:56:52.000Z","size":90126,"stargazers_count":5,"open_issues_count":1,"forks_count":1,"subscribers_count":1,"default_branch":"main","last_synced_at":"2025-03-04T09:34:29.840Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":null,"language":null,"has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/zh-plus.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2024-12-10T09:41:15.000Z","updated_at":"2025-03-04T08:56:55.000Z","dependencies_parsed_at":null,"dependency_job_id":"b106675a-c8cd-4f8e-b325-55ce86eebdd3","html_url":"https://github.com/zh-plus/awesome-interactive-EmbodiedAI","commit_stats":null,"previous_names":["zh-plus/awesome-interactive-embodiedai"],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/zh-plus%2Fawesome-interactive-EmbodiedAI","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/zh-plus%2Fawesome-interactive-EmbodiedAI/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/zh-plus%2Fawesome-interactive-EmbodiedAI/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/zh-plus%2Fawesome-interactive-EmbodiedAI/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/zh-plus","download_url":"https://codeload.github.com/zh-plus/awesome-interactive-EmbodiedAI/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":243025391,"owners_count":20223797,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2025-03-11T11:32:21.958Z","updated_at":"2025-03-11T11:33:11.341Z","avatar_url":"https://github.com/zh-plus.png","language":null,"funding_links":[],"categories":[],"sub_categories":[],"readme":"### 2D semantic segmentation\n\n1. [Text4Seg: Reimagining Image Segmentation as Text Generation](https://mc-lan.github.io/Text4Seg/)\n   ![text4seg](assets/text4seg.png)\n\n\n\n### 3D semantic segmentation\n\n1. SegmentAnything3D\n   - [Pointcept/SegmentAnything3D: [ICCV'23 Workshop\\] SAM3D: Segment Anything in 3D Scenes](https://github.com/Pointcept/SegmentAnything3D)\n   \n2. PvT\n   - [Pointcept/Pointcept: Pointcept: a codebase for point cloud perception research. Latest works: PTv3 (CVPR'24 Oral), PPT (CVPR'24), OA-CNNs (CVPR'24), MSC (CVPR'23)](https://github.com/Pointcept/Pointcept)\n   \n3. ODIN\n   - [ayushjain1144/odin: Code for the paper: \"ODIN: A Single Model for 2D and 3D Segmentation\" (CVPR 2024)](https://github.com/ayushjain1144/odin)\n   \n4. [Segment3D: Learning Fine-Grained Class-Agnostic 3D Segmentation without Manual Labels](https://segment3d.github.io/)\n\n6. [Scalable 3D Panoptic Segmentation As Superpoint Graph Clustering](https://drprojects.github.io/supercluster)\n\n   - https://github.com/drprojects/superpoint_transformer\n\n     ![SPT](assets/SPT.png)\n   \n6. [Open-YOLO 3D: Towards Fast and Accurate Open-Vocabulary 3D Instance Segmentation](https://arxiv.org/abs/2406.02548) [ICLR 2025 Oral]\n\n   - [ICLR 2025 (Oral 📢) \\] Our OpenYOLO3D model achieves state-of-the-art performance in Open Vocabulary 3D Instance Segmentation on ScanNet200 and Replica datasets with up ∼16x speedup compared to the best existing method in literature.](https://github.com/aminebdj/OpenYOLO3D?tab=readme-ov-file)\n     ![openyolo3d](assets/openyolo3d.png)\n\n7. [EmbodiedSAM: Online Segment Any 3D Thing in Real Time](https://xuxw98.github.io/ESAM/) [ICLR 2025 Oral]\n\n   - [ICLR 2025, Oral\\] EmbodiedSAM: Online Segment Any 3D Thing in Real Time](https://github.com/xuxw98/ESAM)\n\n     ![esam](assets/esam.png)\n\n8. [Search3D: Hierarchical Open-Vocabulary 3D Segmentation](https://arxiv.org/pdf/2409.18431) [Arxiv 2025.1]\n\n   ![search3D](assets/search3D.png)\n\n9. [OpenNeRF Open Set 3D Neural Scene Segmentation with Pixel-Wise Features and Rendered Novel Views](https://arxiv.org/pdf/2404.03650) [ICLR 2024]\n\n   -  [ICLR 2024\\] OpenSet 3D Neural Scene Segmentation with Pixel-wise Features and Rendered Novel Views](https://github.com/opennerf/opennerf)\n\n10. [Vocabulary-Free 3D Instance Segmentation with Vision and Language Assistant](https://arxiv.org/pdf/2408.10652) [Arxiv 2024.8]\n\n   ![Vocab-free](assets/Vocab-free.png)\n\n11. [SAMPro3D: Locating SAM Prompts in 3D for Zero-Shot Instance Segmentation](https://github.com/GAP-LAB-CUHK-SZ/SAMPro3D) [3DV 2025]\n    ![SAMPro3D](assets/SAMPro3D.jpg)\n12. [PLA: Language-Driven Open-Vocabulary 3D Scene Understanding](https://dingry.github.io/projects/PLA) [CVPR 2023]\n    ![PLA](assets/PLA.png)\n13. [RegionPLC: Regional Point-Language Contrastive Learning for Open-World 3D Scene Understanding](https://jihanyang.github.io/projects/RegionPLC) [CVPR 2024]\n    ![RegionPLC](assets/RegionPLC.png)\n14. [AGILE3D: Attention Guided Interactive Multi-object 3D Segmentation](https://ywyue.github.io/AGILE3D/) [ICLR 2024]\n    ![AGILE3D](assets/AGILE3D.gif)\n\n\n\n### 3D caption\n\n1. [TOD3Cap: Towards 3D Dense Captioning in Outdoor Scenes](https://arxiv.org/pdf/2403.19589) [ECCV 2024]\n   - [ECCV 2024\\] TOD3Cap: Towards 3D Dense Captioning in Outdoor Scenes](https://github.com/jxbbb/TOD3Cap)\n     ![todcap](assets/todcap.jpg)\n\n\n\n### 2D part segmentation\n\n1. VLPart: [ICCV2023\\] VLPart: Going Denser with Open-Vocabulary Part Segmentation](https://github.com/facebookresearch/VLPart)\n   ![image-20241210170813139](assets/VLPart.png)\n\n2. Semantic-SAM: [Official implementation of the paper \"Semantic-SAM: Segment and Recognize Anything at Any Granularity\"](https://github.com/UX-Decoder/Semantic-SAM) [ECCV 2024\\] \n\n   ![image-20241210171245036](assets/Semantic-SAM.png)\n\n3. Part-CLIPseg: [Official PyTorch Implementation of PartCLIPSeg](https://github.com/kaist-cvml/part-clipseg) [NIPS 2024]\n   ![part_clipseg](assets/part_clipseg.png)\n\n\n\n### 3D part segmentation\n\n1. [3x2: 3D Object Part Segmentation by 2D Semantic Correspondences](https://rehg.org/publication/pub40/) [ECCV 2024]\n   ![3x3](assets/3x3.png)\n2. [PartSTAD: 2D-to-3D Part Segmentation Task Adaptation](https://partstad.github.io/) [ECCV 2024]\n   ![PartSTAD](assets/PartSTAD.png)\n3. [PartSLIP](https://colin97.github.io/PartSLIP_page/) [ECCV 2023]\n   ![part-slip](assets/part-slip.png)\n4. [SAMPart3D](https://yhyang-myron.github.io/SAMPart3D-website/) [Arxiv 2024.10]\n   ![SAMPart3D](assets/SAMPart3D.png)\n\n\n\n### 3D Physical Attributes Annotation\n\n1. [NeRF2Physics: Physical Property Understanding from Language-Embedded Feature Fields](https://ajzhai.github.io/NeRF2Physics/) [CVPR 2024]\n\n   ![NeRF2Plysics](assets/NeRF2Plysics.png)\n\n2. [PUGS: Zero-shot Physical Understanding with Gaussian Splatting](https://evernorif.github.io/PUGS/) [ICRA 2025]\n   ![PUGS](assets/PUGS.png)\n\n\n\n### Embodied interactive/interactable segmentation\n\n1. [FANG-Xiaolin/uncos](https://github.com/FANG-Xiaolin/uncos)\n   ![](assets/uncos.png)\n2. [Functionality understanding and segmentation in 3D scenes](https://jcorsetti.github.io/fun3du/)\n   ![Fun3DU](assets/Fun3DU.png)\n3. [SceneFun3D](https://scenefun3d.github.io/) [CVPR 2024 Oral]\n   ![sceneFun3D](assets/sceneFun3D.jpg)\n\n\n\n### 3D Affordance Detection\n\n1. [3D-AffordanceLLM: Harnessing Large Language Models for Open-Vocabulary Affordance Detection in 3D Worlds](https://arxiv.org/pdf/2502.20041) [ICLR 2025]\n   ![3DALLM](assets/3DALLM.png)\n\n\n\n### 3D Intention Grounding (Detection)\n\n1. [Intent3D: 3D Object Detection in RGB-D Scans Based on Human Intention](https://github.com/WeitaiKang/Intent3D) [ICLR2025]\n\n\n\n### Embodied Interactable 3D Generation\n\n1. [PhyScene: Physically Interactable 3D Scene Synthesis for Embodied AI](https://physcene.github.io/)\n   ![teaser_compress](assets/teaser_compress.png)\n\n\n\n### Others\n\n### Survey\n\n1. [When LLMs step into the 3D World A Survey and Meta-Analysis of 3D Tasks via Multi-modal Large Language Models](https://arxiv.org/pdf/2405.10255) [2024.5]\n\n\n\n### 3D Feature Extraction\n\n1. [Duoduo CLIP: Efficient 3D Understanding with Multi-View Images](https://github.com/3dlg-hcvc/DuoduoCLIP) [ICLR 2025]\n\n\n\n#### Navigation\n\n1. [Hierarchical Open-Vocabulary 3D Scene Graphs for Language-Grounded Robot Navigation](https://hovsg.github.io/)\n   ![teaser-icra](assets/teaser-icra.png)\n\n\n\n### Depth Estimation\n\n1. [Depth Anything V2](https://depth-anything-v2.github.io/) [NIPS 2024]\n   - [DepthAnything/Depth-Anything-V2: [NeurIPS 2024\\] Depth Anything V2. A More Capable Foundation Model for Monocular Depth Estimation](https://github.com/DepthAnything/Depth-Anything-V2)\n2. [Video Depth Anything](https://videodepthanything.github.io/) [Arxiv 2025.1]\n   - [DepthAnything/Video-Depth-Anything: Video Depth Anything: Consistent Depth Estimation for Super-Long Videos](https://github.com/DepthAnything/Video-Depth-Anything)\n3. [Prompt Depth Anything](https://promptda.github.io/) [Arxiv 2024.12]\n   - [DepthAnything/PromptDA: Prompt Depth Anything](https://github.com/DepthAnything/PromptDA)\n   - [rerun-io/prompt-da: PromptDepthAnything example](https://github.com/rerun-io/prompt-da)\n     ![promptda-github-demo](assets/promptda-github-demo.gif)","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fzh-plus%2Fawesome-interactive-embodiedai","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fzh-plus%2Fawesome-interactive-embodiedai","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fzh-plus%2Fawesome-interactive-embodiedai/lists"}