{"id":13443241,"url":"https://github.com/DerrickXuNu/CoBEVT","last_synced_at":"2025-03-20T16:30:43.905Z","repository":{"id":63566058,"uuid":"535122822","full_name":"DerrickXuNu/CoBEVT","owner":"DerrickXuNu","description":"[CoRL2022] CoBEVT: Cooperative Bird's Eye View Semantic Segmentation with Sparse Transformers","archived":false,"fork":false,"pushed_at":"2024-08-18T01:06:04.000Z","size":62911,"stargazers_count":201,"open_issues_count":2,"forks_count":17,"subscribers_count":9,"default_branch":"main","last_synced_at":"2024-10-28T06:57:45.696Z","etag":null,"topics":["autonomous-driving","autonomous-vehicles","bev-perception","collaborative-perception","computer-vision","multi-agent-perception","nuscenes","segmentation","semantic","semantic-segmentation","v2v","v2x"],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/DerrickXuNu.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2022-09-10T21:25:32.000Z","updated_at":"2024-10-18T14:05:39.000Z","dependencies_parsed_at":"2024-10-28T05:10:36.723Z","dependency_job_id":null,"html_url":"https://github.com/DerrickXuNu/CoBEVT","commit_stats":null,"previous_names":[],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/DerrickXuNu%2FCoBEVT","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/DerrickXuNu%2FCoBEVT/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/DerrickXuNu%2FCoBEVT/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/DerrickXuNu%2FCoBEVT/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/DerrickXuNu","download_url":"https://codeload.github.com/DerrickXuNu/CoBEVT/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":244649702,"owners_count":20487470,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["autonomous-driving","autonomous-vehicles","bev-perception","collaborative-perception","computer-vision","multi-agent-perception","nuscenes","segmentation","semantic","semantic-segmentation","v2v","v2x"],"created_at":"2024-07-31T03:01:57.973Z","updated_at":"2025-03-20T16:30:43.900Z","avatar_url":"https://github.com/DerrickXuNu.png","language":"Python","funding_links":[],"categories":["Python","Vehicle-to-Everything Field Datasets","Anti-UAV Datasets"],"sub_categories":[],"readme":"# CoBEVT: Cooperative Bird's Eye View Semantic Segmentation with Sparse Transformers [CORL2022] \n\n[![paper](https://img.shields.io/badge/arXiv-Paper-\u003cCOLOR\u003e.svg)](https://arxiv.org/pdf/2207.02202.pdf)\n[![supplement](https://img.shields.io/badge/Supplementary-Material-red)](https://arxiv.org/pdf/2207.02202.pdf)\n[![video](https://img.shields.io/badge/Video-Presentation-F9D371)]()\n\nThis is the official implementation of CoRL2022 paper \"CoBEVT: Cooperative Bird's Eye View Semantic Segmentation with Sparse Transformers\".\n[Runsheng Xu](https://derrickxunu.github.io/), [Zhengzhong Tu](https://github.com/vztu), [Hao Xiang](https://xhwind.github.io/), [Wei Shao](https://www.linkedin.com/in/wei-shao-94972295/), [Bolei Zhou](https://boleizhou.github.io/), [Jiaqi Ma](https://mobility-lab.seas.ucla.edu/)\n\nUCLA, UT-Austin\n\n\u003cbr\u003e\n\n\u003cdiv align=\"center\"\u003e\u003cimg src=\"images/CorpBEVT_Overview-1.png\" width=\"85%\"/\u003e\u003c/div\u003e\n\u003cdiv align=\"center\"\u003e\n\u003cb\u003eOverview of CoBEVT\u003c/b\u003e\n\u003c/div\u003e\n\u003cbr\u003e\n\n\n## Introduction\nCoBEVT is the first generic multi-agent multi-camera perception framework that can cooperatively generate BEV\nmap predictions. The core component of CoBEVT, named fused axial\nattention or FAX module,  can capture sparsely local and global spatial interactions across views and agents. We \nachieve SOTA performance both on [OPV2V](https://mobility-lab.seas.ucla.edu/opv2v/) and [nuScenes](https://www.nuscenes.org/) dataset with **real-time performance**.\n\n\u003cbr\u003e\n\n\u003cdiv align=\"center\"\u003e\u003cimg src=\"images/nuscene.gif\" width=\"75%\"/\u003e\u003c/div\u003e\n\u003cdiv align=\"center\"\u003e\n\u003cb\u003enuScenes demo:\u003c/b\u003e\nOur CoBEVT can be used on single-vehicle multi-camera semantic BEV Segmentations.\n\u003c/div\u003e\n\u003cbr\u003e\n\n\u003cbr\u003e\n\n\u003cdiv align=\"center\"\u003e\u003cimg src=\"images/opv2v.gif\" width=\"75%\"/\u003e\u003c/div\u003e\n\u003cdiv align=\"center\"\u003e\n\u003cb\u003eOPV2V demo:\u003c/b\u003e\nOur CoBEVT can also be used for multi-agent BEV map prediction.\n\u003c/div\u003e\n\u003cbr\u003e\n\n## Installation\nThe pipeline for nuScenes dataset and OPV2V dataset is different. Please refer to the specific folder for more details based on your research purpose.\n\n:point_right: [nuScenes Users](nuscenes) \u003cbr/\u003e\n:point_right: [OPV2V Users](opv2v)\n\n\n## Models\n\n\u003cdetails\u003e\n  \u003csummary\u003e\u003cstrong\u003eFused Axial Attention Module (FAX)\u003c/strong\u003e (click to expand) \u003c/summary\u003e\n\n\u003cimg src = \"https://user-images.githubusercontent.com/43280278/192188069-44381995-7d0f-43fb-aded-68d62595b2d4.png\" width=\"800\"\u003e\n\n\u003c/details\u003e\n\n\u003cdetails\u003e\n  \u003csummary\u003e\u003cstrong\u003eSinBEVT (single-agent multi-view fusion) and FuseBEVT (multi-agent BEV fusion) \u003c/strong\u003e (click to expand) \u003c/summary\u003e\n\n\u003cimg src = \"https://user-images.githubusercontent.com/43280278/192188116-2b3fd013-b8fc-4d79-a5dd-5eb88d09e622.png\" width=\"800\"\u003e\n\n\u003c/details\u003e\n\n\u003cdetails\u003e\n  \u003csummary\u003e\u003cstrong\u003eCoBEVT Architecture\u003c/strong\u003e (click to expand) \u003c/summary\u003e\n\n\u003cimg src = \"images/CorpBEVT_Overview-1.png\" width=\"800\"\u003e\n\n\u003c/details\u003e\n\n\n\n## Results\n\n\u003cdetails\u003e\n  \u003csummary\u003e\u003cstrong\u003eMain results (OPV2V-camera, -LiDAR, and nuScenes.)\u003c/strong\u003e (click to expand) \u003c/summary\u003e\n\n\u003cimg src = \"https://user-images.githubusercontent.com/43280278/192186838-8b42605b-9cb0-4f3e-9a44-615ec158ce37.png\" width=\"800\"\u003e\n\u003c/details\u003e\n\n\u003cdetails\u003e\n  \u003csummary\u003e\u003cstrong\u003eQualitative results on OPV2V-camera\u003c/strong\u003e (click to expand) \u003c/summary\u003e\n\n\u003cimg src = \"https://user-images.githubusercontent.com/43280278/192187203-4716e5dc-af4d-4652-bddb-fb28ab07260d.png\" width=\"1000\"\u003e\n \n\u003cimg src = \"https://user-images.githubusercontent.com/43280278/192187720-a1eb5c39-5a71-48c8-8eff-6b04b158d0d4.png\" width=\"1000\"\u003e\n\n\u003c/details\u003e\n\n\n\n\u003cdetails\u003e\n  \u003csummary\u003e\u003cstrong\u003eQualitative results on OPV2V-LiDAR\u003c/strong\u003e (click to expand) \u003c/summary\u003e\n\n\u003cimg src = \"https://user-images.githubusercontent.com/43280278/192187375-0a7168bb-8be9-49f6-9f2b-42b7b57fe031.png\" width=\"1000\"\u003e\n\u003cimg src = \"https://user-images.githubusercontent.com/43280278/192187420-8c8c7b31-ee09-4d9f-8df2-7e92dd79acb2.png\" width=\"1000\"\u003e\n\n\u003c/details\u003e\n\n\n\u003cdetails\u003e\n  \u003csummary\u003e\u003cstrong\u003eQualitative results on nuScenes\u003c/strong\u003e (click to expand) \u003c/summary\u003e\n\n\u003cimg src = \"https://user-images.githubusercontent.com/43280278/192187578-1127a7ae-590e-4f27-bcc2-311fd554d0ae.png\" width=\"1000\"\u003e\n\n\u003c/details\u003e\n\n\n\u003cdetails\u003e\n  \u003csummary\u003e\u003cstrong\u003eAblation study\u003c/strong\u003e (click to expand) \u003c/summary\u003e\n\n\u003cimg src = \"https://user-images.githubusercontent.com/43280278/192186995-0fa0b1dd-b5a8-4125-af39-17a79ce0de0e.png\" width=\"800\"\u003e\n\u003c/details\u003e\n\n\n## Citation\n ```bibtex\n@inproceedings{xu2022cobevt,\n  author = {Runsheng Xu, Zhengzhong Tu, Hao Xiang, Wei Shao, Bolei Zhou, Jiaqi Ma},\n  title = {CoBEVT: Cooperative Bird's Eye View Semantic Segmentation with Sparse Transformers},\n  booktitle={Conference on Robot Learning (CoRL)},\n  year = {2022}}\n@article{xu2022v2x,\n  title={V2X-ViT: Vehicle-to-everything cooperative perception with vision transformer},\n  author={Xu, Runsheng and Xiang, Hao and Tu, Zhengzhong and Xia, Xin and Yang, Ming-Hsuan and Ma, Jiaqi},\n  journal={Proceedings of the European Conference on Computer Vision (ECCV)},\n  year={2022}\n}\n@inproceedings{tu2022maxim,\n  title={Maxim: Multi-axis mlp for image processing},\n  author={Tu, Zhengzhong and Talebi, Hossein and Zhang, Han and Yang, Feng and Milanfar, Peyman and Bovik, Alan and Li, Yinxiao},\n  booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},\n  pages={5769--5780},\n  year={2022}\n}\n@article{tu2022maxvit,\n  title={Maxvit: Multi-axis vision transformer},\n  author={Tu, Zhengzhong and Talebi, Hossein and Zhang, Han and Yang, Feng and Milanfar, Peyman and Bovik, Alan and Li, Yinxiao},\n  journal={Proceedings of the European Conference on Computer Vision (ECCV)},\n  year={2022}\n}\n```\n\n## Acknowledgement\nCoBEVT is build upon [OpenCOOD](https://github.com/DerrickXuNu/OpenCOOD), which is the first Open Cooperative Detection framework for autonomous driving.\n\nOur nuScenes experiments used the training pipeline in [CVT(CVPR2022)](https://github.com/bradyz/cross_view_transformers).\n\nCoBEVT is partly inspired by [V2X-ViT](https://github.com/DerrickXuNu/v2x-vit), [MAXIM](https://github.com/google-research/maxim) and [MaxViT](https://github.com/google-research/maxvit).\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FDerrickXuNu%2FCoBEVT","html_url":"https://awesome.ecosyste.ms/projects/github.com%2FDerrickXuNu%2FCoBEVT","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FDerrickXuNu%2FCoBEVT/lists"}