{"id":13441325,"url":"https://github.com/baaivision/Uni3D","last_synced_at":"2025-03-20T11:38:23.333Z","repository":{"id":199611703,"uuid":"703088628","full_name":"baaivision/Uni3D","owner":"baaivision","description":"[ICLR'24 Spotlight] Uni3D: 3D Visual Representation from BAAI","archived":false,"fork":false,"pushed_at":"2024-01-17T06:37:34.000Z","size":6343,"stargazers_count":440,"open_issues_count":15,"forks_count":26,"subscribers_count":13,"default_branch":"main","last_synced_at":"2024-08-01T03:34:02.692Z","etag":null,"topics":["3d-representation-learning","foundation-models","vision-transformers"],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/baaivision.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null}},"created_at":"2023-10-10T15:15:28.000Z","updated_at":"2024-07-31T22:05:34.000Z","dependencies_parsed_at":"2023-12-29T10:28:47.282Z","dependency_job_id":"d44cae38-6256-4d76-9a4c-a59df974e7c2","html_url":"https://github.com/baaivision/Uni3D","commit_stats":null,"previous_names":["baaivision/uni3d"],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/baaivision%2FUni3D","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/baaivision%2FUni3D/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/baaivision%2FUni3D/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/baaivision%2FUni3D/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/baaivision","download_url":"https://codeload.github.com/baaivision/Uni3D/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":221759951,"owners_count":16876323,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["3d-representation-learning","foundation-models","vision-transformers"],"created_at":"2024-07-31T03:01:32.577Z","updated_at":"2024-10-28T01:30:32.426Z","avatar_url":"https://github.com/baaivision.png","language":"Python","funding_links":[],"categories":["Python","Common Metrics"],"sub_categories":[],"readme":"\u003cdiv align='center'\u003e\n\n\u003ch2\u003e\u003ca href=\"https://arxiv.org/abs/2310.06773\"\u003eUni3D: Exploring Unified 3D Representation at Scale\u003c/a\u003e\u003c/h2\u003e\n\n[Junsheng Zhou](https://junshengzhou.github.io/)\u003csup\u003e1,2*\u003c/sup\u003e, [Jinsheng Wang](https://github.com/Wolfwjs/)\u003csup\u003e1*\u003c/sup\u003e, [Baorui Ma](https://mabaorui.github.io/)\u003csup\u003e1*\u003c/sup\u003e, [Yu-Shen Liu](https://yushen-liu.github.io/)\u003csup\u003e2\u003c/sup\u003e, [Tiejun Huang](https://scholar.google.com/citations?user=knvEK4AAAAAJ\u0026hl=en)\u003csup\u003e1,3\u003c/sup\u003e, [Xinlong Wang](https://www.xloong.wang/)\u003csup\u003e1\u003c/sup\u003e\n \n\u003csup\u003e1\u003c/sup\u003e[BAAI](https://www.baai.ac.cn/english.html), \u003csup\u003e2\u003c/sup\u003e[THU](https://www.tsinghua.edu.cn/en/), \u003csup\u003e3\u003c/sup\u003e[PKU](https://english.pku.edu.cn/) \u003cbr\u003e\u003csup\u003e*\u003c/sup\u003e Equal Contribution\n \nICLR 2024 (Spotlight)\n\n[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/uni3d-exploring-unified-3d-representation-at/zero-shot-3d-classification-on-objaverse-lvis)](https://paperswithcode.com/sota/zero-shot-3d-classification-on-objaverse-lvis?p=uni3d-exploring-unified-3d-representation-at)\n[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/uni3d-exploring-unified-3d-representation-at/zero-shot-transfer-3d-point-cloud)](https://paperswithcode.com/sota/zero-shot-transfer-3d-point-cloud?p=uni3d-exploring-unified-3d-representation-at)\n[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/uni3d-exploring-unified-3d-representation-at/zero-shot-transfer-3d-point-cloud-2)](https://paperswithcode.com/sota/zero-shot-transfer-3d-point-cloud-2?p=uni3d-exploring-unified-3d-representation-at)\n\n\n\u003c/div\u003e\n\n\n\u003cp align=\"center\"\u003e\n    \u003cimg src=\"assets/overview.jpg\" alt=\"overview\" width=\"800\" /\u003e\n\u003c/p\u003e\n\nWe present Uni3D, a unified and scalable 3D pretraining framework for large-scale 3D representation learning, and explore its limits at the scale of one billion parameters.\nUni3D uses a 2D initialized ViT end-to-end pretrained to align the 3D point cloud features with the image-text aligned features. Via the simple architecture and pretext task, Uni3D can leverage abundant 2D pretrained models as initialization and image-text aligned models as the target, unlocking the great potential of 2D models and scaling-up strategies to the 3D world. We efficiently scale up Uni3D to one billion parameters, and set new records on a broad range of 3D tasks. \n\n## Schedule\n\nWe are committed to open-sourcing Uni3D related materials, including:\n\n- [x] Extended Uni3D to a 3D metric (Uni3D-score) for enhanced semantic coherence in text-to-3D tasks. For details, see [GeoDream](https://github.com/baaivision/GeoDream).\n- [x] The weights of models range from 6M to **1B** parameters.\n- [x] Evaluation code\n- [x] Evaluation data\n- [x] Pretraining code\n- [ ] Pretraining data\n\n\nWe hope to foster the growth of our community through open-sourcing and promoting collaboration👬. Let's step towards multimodal intelligence together🍻.\n\n\n## Installation\nClone this repository and install the required packages:\n\n```shell\ngit clone https://github.com/baaivision/Uni3D.git\ncd Uni3D\n\nconda create -n uni3d python=3.8\nconda activate uni3d\nconda install pytorch torchvision torchaudio pytorch-cuda=11.8 -c pytorch -c nvidia\n\npip install -r requirements.txt\n\n# install pointnet2 extensions from https://github.com/erikwijmans/Pointnet2_PyTorch\npip install \"git+git://github.com/erikwijmans/Pointnet2_PyTorch.git#egg=pointnet2_ops\u0026subdirectory=pointnet2_ops_lib\"\n\n```\nCore packages: \n- [Pytorch](https://pytorch.org/) version 2.0.1 \n- [open-clip-torch](https://github.com/mlfoundations/open_clip) version 2.20.0\n- [timm](https://github.com/rwightman/pytorch-image-models) version 0.9.7\n- [DeepSpeed](https://github.com/microsoft/DeepSpeed) version 0.10.3\n- [Open3D](https://github.com/isl-org/Open3D) version 0.17.0\n\n## Model Zoo\n\n| Model         | Training Data | Objaverse-LVIS Top1 (Top5) | ModelNet40 Top1 (Top5) | ScanObjectNN Top1 (Top5) |\n| :------:  | :------: | :------: |:------: |:------: |\n| [**Uni3d-B**](https://huggingface.co/BAAI/Uni3D/blob/main/modelzoo/uni3d-b-no-lvis/model.pt) | Ensembled w/o LVIS | 45.9 (74.8) | 86.1 (98.7) | 61.7 (89.5) | \n| [**Uni3d-B**](https://huggingface.co/BAAI/Uni3D/blob/main/modelzoo/uni3d-b/model.pt) | Ensembled          | 51.7 (80.8) | 86.3 (97.9) | 63.8 (90.2) | \n| [**Uni3d-L**](https://huggingface.co/BAAI/Uni3D/blob/main/modelzoo/uni3d-l-no-lvis/model.pt) | Ensembled w/o LVIS | 46.2 (74.7) | 86.6 (97.8) | 58.4 (90.1) | \n| [**Uni3d-L**](https://huggingface.co/BAAI/Uni3D/blob/main/modelzoo/uni3d-l/model.pt) | Ensembled          | 53.1 (81.5) | 86.3 (98.3) | 58.2 (89.4) | \n| [**Uni3d-g**](https://huggingface.co/BAAI/Uni3D/blob/main/modelzoo/uni3d-g-no-lvis/model.pt) | Ensembled w/o LVIS | 47.2 (76.1) | 86.8 (98.4) | 66.5 (90.1) | \n| [**Uni3d-g**](https://huggingface.co/BAAI/Uni3D/blob/main/modelzoo/uni3d-g/model.pt) | Ensembled          | 53.5 (82.0) | 87.3 (99.2) | 63.9 (91.7) | \n| [**Uni3d-g**](https://huggingface.co/BAAI/Uni3D/tree/main/modelzoo/uni3d-g) 🔥 | Ensembled          | 55.3 (82.9) | 88.2 (99.3) | 65.3 (92.7) |\n\n## Evaluation of Zero-shot 3D classification \nWe evaluate the zero-shot 3D classification performance on three datasets: Objaverse-LVIS, ModelNet40 and ScanObjectNN.\n\n1. Please refer to [DATASETS.md](data/DATASETS.md) for evaluation dataset preparation.\n2. [Recommended 🤗] Download the [clip model](https://huggingface.co/timm/eva02_enormous_patch14_plus_clip_224.laion2b_s9b_b144k/blob/main/open_clip_pytorch_model.bin) and put it in `/path/to/clip_model` folder.\n3. Download model zoo weights and put them in `/path/to/checkpoints` folder.\n4. Run `bash scripts/inference.sh [scale]` to evaluate the model on the above datasets, e.g., `bash scripts/inference.sh giant`.\n\n## Pre-training\n1. Please refer to [DATASETS.md](data/DATASETS.md) for pre-train dataset preparation.\n2. [Recommended 🤗] Download the [clip model](https://huggingface.co/timm/eva02_enormous_patch14_plus_clip_224.laion2b_s9b_b144k/blob/main/open_clip_pytorch_model.bin) and put it in `/path/to/clip_model` folder.\n3. [Recommended 🤗] Download the [initialization model](https://huggingface.co/timm/eva_giant_patch14_560.m30m_ft_in22k_in1k/blob/main/model.safetensors) and put it in `/path/to/init_model` folder.\n4. Run `bash scripts/pretrain.sh` to pre-train the model on ensemble datasets.\n\n\n## Visualization\n\n### Open-world Understanding\n\u003cp align=\"center\"\u003e\n    \u003cimg src=\"assets/scene_understanding.jpg\" alt=\"scene\" width=\"800\" /\u003e\n\u003c/p\u003e\n\n### One-shot Part Segmentation\n\u003cp align=\"center\"\u003e\n    \u003cimg src=\"assets/vis_part.jpg\" alt=\"partseg\" width=\"800\" /\u003e\n\u003c/p\u003e\n\n### Point Cloud Painting\n\u003cp align=\"center\"\u003e\n    \u003cimg src=\"assets/editing.jpg\" alt=\"editing\" width=\"800\" /\u003e\n\u003c/p\u003e\n\n### Cross-modal Retrieval\n\u003cp align=\"center\"\u003e\n    \u003cimg src=\"assets/retrival_text.jpg\" alt=\"retrival_text\" width=\"800\" /\u003e\n\u003c/p\u003e\n\n\u003cp align=\"center\"\u003e\n    \u003cimg src=\"assets/retrival.jpg\" alt=\"retrival\" width=\"800\" /\u003e\n\u003c/p\u003e\n\n\n## Acknowledgement\nUni3D is built using the awesome [EVA](https://github.com/baaivision/EVA), [OpenCLIP](https://github.com/mlfoundations/open_clip), [timm](https://github.com/huggingface/pytorch-image-models/), [DeepSpeed](https://github.com/microsoft/DeepSpeed), [ULIP](https://github.com/salesforce/ULIP) and [OpenShape](https://github.com/Colin97/OpenShape_code). \n\n## Citation\n```bib\n@inproceedings{zhou2023uni3d,\n  title={Uni3d: Exploring unified 3d representation at scale},\n  author={Zhou, Junsheng and Wang, Jinsheng and Ma, Baorui and Liu, Yu-Shen and Huang, Tiejun and Wang, Xinlong},\n  booktitle={International Conference on Learning Representations (ICLR)},\n  year={2024}\n}\n```\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fbaaivision%2FUni3D","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fbaaivision%2FUni3D","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fbaaivision%2FUni3D/lists"}