{"id":22068856,"url":"https://github.com/nv-nguyen/gigaPose","last_synced_at":"2025-07-24T07:31:17.227Z","repository":{"id":209431677,"uuid":"722940359","full_name":"nv-nguyen/gigapose","owner":"nv-nguyen","description":"[CVPR 2024] PyTorch implementation of GigaPose: Fast and Robust Novel Object Pose Estimation via One Correspondence","archived":false,"fork":false,"pushed_at":"2024-05-10T08:55:58.000Z","size":14920,"stargazers_count":79,"open_issues_count":4,"forks_count":6,"subscribers_count":6,"default_branch":"main","last_synced_at":"2024-05-10T18:38:15.276Z","etag":null,"topics":["6d-pose-estimation","6dof","deep-learning","object-pose-estimation","pose-estimation"],"latest_commit_sha":null,"homepage":"https://nv-nguyen.github.io/gigaPose/","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/nv-nguyen.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2023-11-24T10:03:34.000Z","updated_at":"2024-05-10T08:56:02.000Z","dependencies_parsed_at":"2024-05-09T18:50:38.029Z","dependency_job_id":null,"html_url":"https://github.com/nv-nguyen/gigapose","commit_stats":null,"previous_names":["nv-nguyen/gigapose"],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/nv-nguyen%2Fgigapose","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/nv-nguyen%2Fgigapose/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/nv-nguyen%2Fgigapose/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/nv-nguyen%2Fgigapose/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/nv-nguyen","download_url":"https://codeload.github.com/nv-nguyen/gigapose/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":227421350,"owners_count":17775010,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["6d-pose-estimation","6dof","deep-learning","object-pose-estimation","pose-estimation"],"created_at":"2024-11-30T20:04:25.838Z","updated_at":"2025-07-24T07:31:17.188Z","avatar_url":"https://github.com/nv-nguyen.png","language":"Python","funding_links":[],"categories":["Paper List"],"sub_categories":["Follow-up Papers"],"readme":"\u003cdiv align=\"center\"\u003e\n\u003ch2\u003e\nGigaPose: Fast and Robust Novel Object Pose Estimation \n\nvia One Correspondence\n\u003cp\u003e\u003c/p\u003e\n\u003c/h2\u003e\n\n\u003ch3\u003e\n\u003ca href=\"https://nv-nguyen.github.io/\" target=\"_blank\"\u003e\u003cnobr\u003eVan Nguyen Nguyen\u003c/nobr\u003e\u003c/a\u003e \u0026emsp;\n\u003ca href=\"http://imagine.enpc.fr/~groueixt/\" target=\"_blank\"\u003e\u003cnobr\u003eThibault Groueix\u003c/nobr\u003e\u003c/a\u003e \u0026emsp;\n\u003ca href=\"https://people.epfl.ch/mathieu.salzmann\" target=\"_blank\"\u003e\u003cnobr\u003eMathieu Salzmann\u003c/nobr\u003e\u003c/a\u003e \u0026emsp;\n\u003ca href=\"https://vincentlepetit.github.io/\" target=\"_blank\"\u003e\u003cnobr\u003eVincent Lepetit\u003c/nobr\u003e\u003c/a\u003e\n\n\u003cp\u003e\u003c/p\u003e\n\n\u003ca href=\"https://nv-nguyen.github.io/gigapose/\"\u003e\u003cimg \nsrc=\"https://img.shields.io/badge/-Webpage-blue.svg?colorA=333\u0026logo=html5\" height=28em\u003e\u003c/a\u003e\n\u003ca href=\"https://arxiv.org/abs/2311.14155\"\u003e\u003cimg \nsrc=\"https://img.shields.io/badge/-Paper-blue.svg?colorA=333\u0026logo=arxiv\" height=28em\u003e\u003c/a\u003e\n\u003ca href=\"https://drive.google.com/file/d/11V9J4voUkovMIFxOeDCkaO7uf9EfCTZ0/view?usp=sharing\"\u003e\u003cimg \nsrc=\"https://img.shields.io/badge/-SuppMat-blue.svg?colorA=333\u0026logo=drive\" height=28em\u003e\u003c/a\u003e\n\u003cp\u003e\u003c/p\u003e\n\n\u003cp align=\"center\"\u003e\n  \u003cimg src=./media/qualitative.png width=\"100%\"/\u003e\n\u003c/p\u003e\n\n\u003c/h4\u003e\n\u003c/div\u003e\n\n**TL;DR**: GigaPose is a \"hybrid\" template-patch correspondence approach to estimate 6D pose of novel objects in RGB images: GigaPose first uses templates, rendered images of the CAD models, to recover the out-of-plane rotation (2DoF) and then uses patch correspondences to estimate the remaining 4DoF. \n\nThe codebase is slightly modified to adapt [BOP challenge 2024](https://bop.felk.cvut.cz/challenges/bop-challenge-2024/), if you work on [BOP challenge 2023](https://bop.felk.cvut.cz/challenges/bop-challenge-20243) and have any issues, please go back to previous commits:\n```\ngit checkout 388e8bddd8a5443e284a7f70ad103d03f3f461c5\n```\n\n\n### News 📣\n- [May 24th, 2024] We added the instructions for running on [BOP challenge 2024](https://bop.felk.cvut.cz/challenges/bop-challenge-2024/) datasets and fixed memory requirement issues.\n- [January 19th, 2024] We released the intructions for estimating pose of novel objects from a single reference image on LM-O dataset.\n- [January 11th, 2024] We released the code for both training and testing settings. We are working on the demo for custom objects including detecting novel objects with [CNOS](https://github.com/nv-nguyen/cnos) and novel object pose estimation from a single reference image by reconstructing objects with [Wonder3D](https://github.com/xxlong0/Wonder3D). Stay tuned!\n## Citations\n``` Bash\n@inproceedings{nguyen2024gigaPose,\n    title={GigaPose: Fast and Robust Novel Object Pose Estimation via One Correspondence},\n    author={Nguyen, Van Nguyen and Groueix, Thibault and Salzmann, Mathieu and Lepetit, Vincent},\n    booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},\n    year={2024}}\n```\nGigaPose's codebase is mainly derived from [CNOS](https://github.com/nv-nguyen/cnos) and [MegaPose](https://github.com/megapose6d/megapose6d):\n``` Bash\n@inproceedings{nguyen2023cnos,\n    title={CNOS: A Strong Baseline for CAD-based Novel Object Segmentation},\n    author={Nguyen, Van Nguyen and Groueix, Thibault and Ponimatkin, Georgy and Lepetit, Vincent and Hodan, Tomas},\n    booktitle={Proceedings of the IEEE/CVF International Conference on Computer Vision},\n    pages={2134--2140},\n    year={2023}\n}\n\n@inproceedings{labbe2022megapose,\n    title     = {MegaPose: 6D Pose Estimation of Novel Objects via Render \\\u0026 Compare},\n    author    = {Labb\\'e, Yann and Manuelli, Lucas and Mousavian, Arsalan and Tyree, Stephen and Birchfield, Stan and Tremblay, Jonathan and Carpentier, Justin and Aubry, Mathieu and Fox, Dieter and Sivic, Josef},\n    booktitle = {Proceedings of the 6th Conference on Robot Learning (CoRL)},\n    year      = {2022},\n} \n```\n\n## Installation :construction_worker:\n\n\u003cdetails\u003e\u003csummary\u003eClick to expand\u003c/summary\u003e\n\n### Environment\n```\nconda env create -f environment.yml\nconda activate gigapose\nbash src/scripts/install_env.sh\n\n# to install megapose\npip install -e .\n\n# to install bop_toolkit \npip install git+https://github.com/thodan/bop_toolkit.git\n```\n\n### Checkpoints\n```\n# download cnos detections for BOP'23 dataset\npip install -U \"huggingface_hub[cli]\"\npython -m src.scripts.download_default_detections\n\n# download gigaPose's checkpoints \npython -m src.scripts.download_gigapose\n\n# download megapose's checkpoints\npython -m src.scripts.download_megapose\n```\n\n### Datasets\nAll datasets are defined in [BOP format](https://bop.felk.cvut.cz/datasets/). \n\nFor [BOP challenge 2024](https://bop.felk.cvut.cz/challenges/bop-challenge-2024/) core datasets (HOPE, HANDAL, HOT-3D), download each dataset with the following command:\n```\npip install -U \"huggingface_hub[cli]\"\nexport DATASET_NAME=hope\npython -m src.scripts.download_test_bop24 test_dataset_name=$DATASET_NAME\n```\n\nFor [BOP challenge 2023](https://bop.felk.cvut.cz/challenges/bop-challenge-2023/) core datasets (LMO, TLESS, TUDL, ICBIN, ITODD, HB, and TLESS), download all datasets with the following command:\n```\n# download testing images and CAD models\npython -m src.scripts.download_test_bop23\n```\n\nFor [BOP challenge 2024](https://bop.felk.cvut.cz/challenges/bop-challenge-2024/) core datasets (HOPE, HANDAL, HOT-3D), render the templates from the CAD models:\n```\npython -m src.scripts.render_bop_templates test_dataset_name=hope\n```\n\nFor [BOP challenge 2023](https://bop.felk.cvut.cz/challenges/bop-challenge-2023/) core datasets (LMO, TLESS, TUDL, ICBIN, ITODD, HB, and TLESS), we provide the pre-rendered templates (from [this link](https://huggingface.co/datasets/nv-nguyen/gigaPose/resolve/main/templates.zip)) and also the code to render the templates from the CAD models.\n```\n# option 1: download pre-rendered templates \npython -m src.scripts.download_bop_templates\n\n# option 2: render templates from CAD models \npython -m src.scripts.render_bop_templates\n```\n\nHere is the structure of $ROOT_DIR after downloading all the above files (similar to [BOP HuggingFace Hub](https://huggingface.co/datasets/bop-benchmark/datasets/tree/main)):\n```\n├── $ROOT_DIR\n    ├── datasets/ \n      ├── default_detections/  \n      ├── lmo/ \n      ├── ... \n      ├── templates/\n    ├── pretrained/ \n      ├── gigaPose_v1.ckpt \n      ├── megapose-models/\n```\n\n[Optional] We also provide the training code/datasets which is not necessary for testing purposes.\n\u003cdetails\u003e\u003csummary\u003eClick to expand\u003c/summary\u003e\n\n```\n# download training images (\u003e 2TB)\npython -m src.scripts.download_train_metaData\npython -m src.scripts.download_train_cad \npython -m src.scripts.download_train \n\n# render templates ( 162 imgs/obj takes ~30mins for gso, ~20hrs for shapenet)\npython -m src.scripts.render_gso_templates \npython -m src.scripts.render_shapenet_templates  \n```\n\nIf you have training datasets pre-downloaded, you can create a symlink to the folder containing the datasets by running:\n```\nln -s /path/to/datasets/gso $ROOT/datasets/gso\n```\n\n[Optional] Trick for faster converging of ISTNetwork (in-plane, scale, translation): using pretrained weights of [LoFTR](https://drive.google.com/file/d/1kW2bQejjMlmE7FGberHrubXpE_ttX2LB/view?usp=drive_link) after Kaiming initialization. Please download the weights and put them in `$ROOT_DIR/pretrained/loftr_indoor_ot.ckpt`.\n\n\u003c/details\u003e\n\n\u003c/details\u003e\n\n\n##  Testing on [BOP datasets](https://bop.felk.cvut.cz/datasets/) :rocket:\n\n\u003cp align=\"center\"\u003e\n  \u003cimg src=./media/inference.png width=\"100%\"/\u003e\n\u003c/p\u003e\n\nIf you want to test on [BOP challenge 2024](https://bop.felk.cvut.cz/challenges/bop-challenge-2024/) datasets, please follow the instructions below:\n\u003cdetails\u003e\u003csummary\u003eClick to expand\u003c/summary\u003e\n\n1. Running coarse prediction on a single dataset:\n```\n# for 6D detection task\npython test.py test_dataset_name=hope run_id=$NAME_RUN test_setting=detection\n\n# for 6D localization task (for only core19 datasets)\npython test.py test_dataset_name=lmo run_id=$NAME_RUN test_setting=localization\n```\n\n2. Running refinement on a single dataset:\n```\n# for both 6D detection task\npython refine.py test_dataset_name=hope run_id=$NAME_RUN test_setting=detection\n\n# for 6D localization task (for only core19 datasets)\npython refine.py test_dataset_name=lmo run_id=$NAME_RUN test_setting=localization\n```\nQuantitative results on 6D detection task on HOPEv2 datasets:\n\n| Method      | Refinement      | Model-based unseen |\n|---------------|---------------|-----------|\n| GigaPose  | --  | 22.57 | \n| GigaPose  | MegaPose  | -- | \n\n\n3. Evaluating with [BOP toolkit](https://github.com/thodan/bop_toolkit):\n```\nexport INPUT_DIR=DIR_TO_YOUR_PREDICTION_FILE\nexport FILE_NAME=NAME_PREDICTION_FILE\ncd $ROOT_DIR_OF_TOOLKIT\npython scripts/eval_bop24_pose.py --results_path $INPUT_DIR --eval_path $INPUT_DIR --result_filenames=$FILE_NAME\n```\n\n\u003c/details\u003e\n\nIf you want to test on [BOP challenge 2023](https://bop.felk.cvut.cz/challenges/bop-challenge-2023/) datasets, please follow the instructions below:\n\n\u003cdetails\u003e\u003csummary\u003eClick to expand\u003c/summary\u003e\n\nGigaPose's coarse prediction for seven core datasets of BOP challenge 2023 is available in [this link](https://drive.google.com/file/d/1QaGNIPZyR8FOOsT35V7pWJF2VlN9_M6l/view?usp=sharing). Below are the steps to reproduce the results and evaluate with BOP toolkit.\n\n1. Running coarse prediction on a single dataset:\n```\npython test.py test_dataset_name=lmo run_id=$NAME_RUN\n```\n\n2. Running refinement on a single dataset:\n```\npython refine.py test_dataset_name=lmo run_id=$NAME_RUN\n```\n\n3. Running all steps for all 7 core datasets of BOP challenge:\n```\npython -m src.scripts.eval_bop\n```\n\n3. Evaluating with [BOP toolkit](https://github.com/thodan/bop_toolkit):\n```\nexport INPUT_DIR=DIR_TO_YOUR_PREDICTION_FILE\nexport FILE_NAME=NAME_PREDICTION_FILE\npython bop_toolkit/scripts/eval_bop19_pose.py --renderer_type=vispy --results_path $INPUT_DIR --eval_path $INPUT_DIR --result_filenames=$FILE_NAME\n```\n\n\u003c/details\u003e\n\n##  Pose estimation from a single image on [LM-O](https://bop.felk.cvut.cz/datasets/) :smiley_cat:\n\n\u003cp align=\"center\"\u003e\n  \u003cimg src=./media/wonder3d_meshes.png width=\"100%\"/\u003e\n\u003c/p\u003e\n\n\u003cdetails\u003e\u003csummary\u003eClick to expand\u003c/summary\u003e\n\nIf you work on this setting, please first go back to previous commits:\n```\ngit checkout 388e8bddd8a5443e284a7f70ad103d03f3f461c5\n```\n\nThen, download CNOS's detections:\n```\n# download gigaPose's checkpoints \npython -m src.scripts.download_cnos_bop23\n```\n\nTo relax the need of CAD models, we can reconstruct 3D models from a single image using recent works on diffusion-based 3D reconstruction such as [Wonder3D](https://github.com/xxlong0/Wonder3D), then apply the same pipeline as GigaPose to estimate object pose. Here are the steps to reproduce the results of novel object pose estimation from a single image on LM-O dataset as shown in our paper:\n\n- Step 1: Selecting the input reference image for each object. We provide the list of reference images in [SuppMat](https://drive.google.com/file/d/11V9J4voUkovMIFxOeDCkaO7uf9EfCTZ0/view?usp=sharing). \n- Step 2: Cropping the input image (and save the [cropping matrix](https://github.com/nv-nguyen/gigapose/blob/main/src/utils/crop.py#L49) for recovering the correct scale for reconstructed 3D models).\n- Step 3: Reconstructing 3D models from the reference images using [Wonder3D](https://github.com/xxlong0/Wonder3D). Note that the output 3D models are reconstructed in the coordinate frame of input image and in [orthographic camera](https://github.com/xxlong0/Wonder3D/blob/57b0e88ac45000a9cc100df2733c6cb30ce5e108/NeuS/models/dataset_mvdiff.py#L174).\n- Step 4: Recovering the scale of reconstructed 3D models using the cropping matrix of Step 2. \n- Step 5: Estimating the object pose using GigaPose's pipeline. \n\nWe provide [here](https://huggingface.co/datasets/nv-nguyen/gigaPose/resolve/main/wonder3d_inout.zip) the inputs and outputs of Wonder3D, [here](https://huggingface.co/datasets/nv-nguyen/gigaPose/resolve/main/wonder3d_mesh.zip) the reconstructed 3D models in Step 1-3 and, [this script](https://github.com/nv-nguyen/gigapose/blob/main/src/scripts/recover_scale_wonder3d.py) to recover 3D models in the correct scale. [Here](https://huggingface.co/datasets/nv-nguyen/gigaPose/resolve/main/lmoWonder3d.zip) is the reconstructed 3D models in the correct scale and in the GT coordinate frame discussed below (note that the GT canonical frame is only for evaluation purposes with BOP Toolkit, while for real applications, we can use the object pose in the input reference image as the canonical frame). \n\n\u003cdetails\u003e\u003csummary\u003eClick to expand\u003c/summary\u003e\n\n### Canonical frame for bop toolkit\n\nFor all evaluations, we use [bop toolkit](https://github.com/thodan/bop_toolkit.git) which requires the estimated poses defined in the same coordinate frame of GT CAD models. Therefore, there are two options:\n- Option 1: Transforming the GT CAD models to the coordinate frame of the input image and adjust the GT poses accordingly.\n- Option 2: Reconstructing the 3D models, then transforming it to the coordinate frame of GT CAD models by assuming the object pose in the input reference image is known.\n\nGiven that the metrics VSD, MSSD, MSPD employed in the [bop toolkit](https://github.com/thodan/bop_toolkit.git) depend on the canonical frame of the object, and for a meaningful comparison with [MegaPose](https://github.com/megapose6d/megapose6d) and GigaPose's results using GT CAD models, we opt Option 2. \n\u003c/details\u003e\n\n\nOnce the reconstructed 3D models are in the correct scale and in the GT coordinate frame, we can now estimate the object pose using GigaPose's pipeline in Step 5:\n\n```\n# download the reconstructed 3D models, test images, and test_targets_bop19.json\nmkdir $ROOT_DIR/datasets/lmoWonder3d\nwget https://huggingface.co/datasets/nv-nguyen/gigaPose/resolve/main/lmoWonder3d.zip -P $ROOT_DIR/datasets/lmoWonder3d\nunzip -j $ROOT_DIR/datasets/lmoWonder3d/lmoWonder3d.zip -d $ROOT_DIR/datasets/lmoWonder3d/models -x \"*/._*\"\n\n# treat lmoWonder3d as a new dataset by creating a symlink \nln -s $ROOT_DIR/datasets/lmo/test $ROOT_DIR/datasets/lmoWonder3d/test\nln -s $ROOT_DIR/datasets/lmo/test_targets_bop19.json $ROOT_DIR/datasets/lmoWonder3d/test_targets_bop19.json\n\n# Onboarding by rendering templates from reconstructed 3D models\npython -m src.scripts.render_custom_templates custom_dataset_name=lmoWonder3d\n\n# now, it can be tested as a normal dataset as in the previous section\npython test.py test_dataset_name=lmoWonder3d run_id=$NAME_RUN\npython refine.py test_dataset_name=lmoWonder3d run_id=$NAME_RUN\n```\n\nWe provide [here](https://drive.google.com/file/d/1kowbL3EaNPma_Tell2iBLsJo1RlILSNJ/view?usp=sharing) the (coarse, refined) results with reconstructed CAD models.\n\n\u003cp align=\"center\"\u003e\n  \u003cimg src=./media/wonder3d_result.png width=\"100%\"/\u003e\n\u003c/p\u003e\n\n\n\u003c/details\u003e\n\n##  Training\n\u003cp align=\"center\"\u003e\n  \u003cimg src=./media/training.png width=\"100%\"/\u003e\n\u003c/p\u003e\n\u003cdetails\u003e\u003csummary\u003eClick to expand\u003c/summary\u003e\n\n```\n# train on GSO (ID=0), ShapeNet (ID=1), or both (ID=2)\npython train.py train_dataset_id=$ID\n```\n\n\u003c/details\u003e\n\n## 👩‍⚖️ License\nUnless otherwise specified, all code in this repository is made available under MIT license. \n\n## 🤝 Acknowledgments\nThis code is heavily borrowed from [MegaPose](https://github.com/megapose6d/megapose6d) and [CNOS](https://github.com/nv-nguyen/cnos). \n\nThe authors thank Jonathan Tremblay, Medéric Fourmy, Yann Labbé, Michael Ramamonjisoa and Constantin Aronssohn for their help and valuable feedbacks!\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fnv-nguyen%2FgigaPose","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fnv-nguyen%2FgigaPose","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fnv-nguyen%2FgigaPose/lists"}