{"id":19186800,"url":"https://github.com/donydchen/mvsplat","last_synced_at":"2025-05-15T03:06:10.212Z","repository":{"id":229080802,"uuid":"775391316","full_name":"donydchen/mvsplat","owner":"donydchen","description":"🌊 [ECCV'24 Oral] MVSplat: Efficient 3D Gaussian Splatting from Sparse Multi-View Images","archived":false,"fork":false,"pushed_at":"2025-01-20T09:19:00.000Z","size":432,"stargazers_count":998,"open_issues_count":39,"forks_count":49,"subscribers_count":18,"default_branch":"main","last_synced_at":"2025-04-14T03:07:56.947Z","etag":null,"topics":["cost-volume","eccv2024","feed-forward-gaussian-splatting","gaussian-splatting","novel-view-synthesis"],"latest_commit_sha":null,"homepage":"https://donydchen.github.io/mvsplat","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/donydchen.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2024-03-21T09:56:36.000Z","updated_at":"2025-04-13T21:33:52.000Z","dependencies_parsed_at":"2024-08-14T13:28:23.192Z","dependency_job_id":"4b6b148c-1366-438d-b431-a33ce198558e","html_url":"https://github.com/donydchen/mvsplat","commit_stats":null,"previous_names":["donydchen/mvsplat"],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/donydchen%2Fmvsplat","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/donydchen%2Fmvsplat/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/donydchen%2Fmvsplat/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/donydchen%2Fmvsplat/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/donydchen","download_url":"https://codeload.github.com/donydchen/mvsplat/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":254264765,"owners_count":22041793,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["cost-volume","eccv2024","feed-forward-gaussian-splatting","gaussian-splatting","novel-view-synthesis"],"created_at":"2024-11-09T11:16:46.892Z","updated_at":"2025-05-15T03:06:10.192Z","avatar_url":"https://github.com/donydchen.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"\u003cp align=\"center\"\u003e\n  \u003ch1 align=\"center\"\u003eMVSplat: Efficient 3D Gaussian Splatting \u003cbr\u003e from Sparse Multi-View Images\u003c/h1\u003e\n  \u003cp align=\"center\"\u003e\n    \u003ca href=\"https://donydchen.github.io/\"\u003eYuedong Chen\u003c/a\u003e\n    \u0026nbsp;·\u0026nbsp;\n    \u003ca href=\"https://haofeixu.github.io/\"\u003eHaofei Xu\u003c/a\u003e\n    \u0026nbsp;·\u0026nbsp;\n    \u003ca href=\"https://chuanxiaz.com/\"\u003eChuanxia Zheng\u003c/a\u003e\n    \u0026nbsp;·\u0026nbsp;\n    \u003ca href=\"https://bohanzhuang.github.io/\"\u003eBohan Zhuang\u003c/a\u003e \u003cbr\u003e\n    \u003ca href=\"https://people.inf.ethz.ch/marc.pollefeys/\"\u003eMarc Pollefeys\u003c/a\u003e\n    \u0026nbsp;·\u0026nbsp;\n    \u003ca href=\"http://www.cvlibs.net/\"\u003eAndreas Geiger\u003c/a\u003e\n    \u0026nbsp;·\u0026nbsp;\n    \u003ca href=\"https://personal.ntu.edu.sg/astjcham/\"\u003eTat-Jen Cham\u003c/a\u003e\n    \u0026nbsp;·\u0026nbsp;\n    \u003ca href=\"https://jianfei-cai.github.io/\"\u003eJianfei Cai\u003c/a\u003e\n  \u003c/p\u003e\n  \u003ch3 align=\"center\"\u003eECCV 2024 Oral\u003c/h3\u003e\n  \u003ch3 align=\"center\"\u003e\u003ca href=\"https://arxiv.org/abs/2403.14627\"\u003ePaper\u003c/a\u003e | \u003ca href=\"https://donydchen.github.io/mvsplat/\"\u003eProject Page\u003c/a\u003e | \u003ca href=\"https://drive.google.com/drive/folders/14_E_5R6ojOWnLSrSVLVEMHnTiKsfddjU\"\u003ePretrained Models\u003c/a\u003e \u003c/h3\u003e\n\u003c!--   \u003cdiv align=\"center\"\u003e\n    \u003ca href=\"https://news.ycombinator.com/item?id=41222655\"\u003e\n      \u003cimg\n        alt=\"Featured on Hacker News\"\n        src=\"https://hackerbadge.vercel.app/api?id=41222655\u0026type=dark\"\n      /\u003e\n    \u003c/a\u003e\n  \u003c/div\u003e --\u003e\n\n\u003cul\u003e\n\u003cli\u003e\u003cb\u003e20/01/25 Update:\u003c/b\u003e Check out Cheng's \u003ca href=\"https://github.com/chengzhag/PanSplat\"\u003ePanSplat\u003c/a\u003e, which extends MVSplat to higher resolutions (up to 4K) and highlights the use of a hierarchical spherical cost volume and two-step deferred backpropagation for memory-efficient training. \u003c/li\u003e\n\u003cli\u003e\u003cb\u003e08/11/24 Update:\u003c/b\u003e Explore our \u003ca href=\"https://github.com/donydchen/mvsplat360\"\u003eMVSplat360 [NeurIPS '24]\u003c/a\u003e, an upgraded MVSplat that combines video diffusion to achieve 360° NVS for large-scale scenes from just 5 input views! \u003c/li\u003e  \n\u003cli\u003e\u003cb\u003e21/10/24 Update:\u003c/b\u003e Check out Haofei's \u003ca href=\"https://github.com/cvg/depthsplat\"\u003eDepthSplat\u003c/a\u003e if you are interested in feed-forward 3DGS on more complex scenes (DL3DV-10K) and more input views (up to 12 views)!\u003c/li\u003e\n\u003c/ul\u003e\n\u003cbr\u003e\n\u003c/p\u003e\n\nhttps://github.com/donydchen/mvsplat/assets/5866866/c5dc5de1-819e-462f-85a2-815e239d8ff2\n\n## Installation\n\nTo get started, clone this project, create a conda virtual environment using Python 3.10+, and install the requirements:\n\n```bash\ngit clone https://github.com/donydchen/mvsplat.git\ncd mvsplat\nconda create -n mvsplat python=3.10\nconda activate mvsplat\npip install torch==2.1.2 torchvision==0.16.2 torchaudio==2.1.2 --index-url https://download.pytorch.org/whl/cu118\npip install -r requirements.txt\n```\n\n## Acquiring Datasets\n\n### RealEstate10K and ACID\n\nOur MVSplat uses the same training datasets as pixelSplat. Below we quote pixelSplat's [detailed instructions](https://github.com/dcharatan/pixelsplat?tab=readme-ov-file#acquiring-datasets) on getting datasets.\n\n\u003e pixelSplat was trained using versions of the RealEstate10k and ACID datasets that were split into ~100 MB chunks for use on server cluster file systems. Small subsets of the Real Estate 10k and ACID datasets in this format can be found [here](https://drive.google.com/drive/folders/1joiezNCyQK2BvWMnfwHJpm2V77c7iYGe?usp=sharing). To use them, simply unzip them into a newly created `datasets` folder in the project root directory.\n\n\u003e If you would like to convert downloaded versions of the Real Estate 10k and ACID datasets to our format, you can use the [scripts here](https://github.com/dcharatan/real_estate_10k_tools). Reach out to us (pixelSplat) if you want the full versions of our processed datasets, which are about 500 GB and 160 GB for Real Estate 10k and ACID respectively.\n\n### DTU (For Testing Only)\n\n* Download the preprocessed DTU data [dtu_training.rar](https://drive.google.com/file/d/1eDjh-_bxKKnEuz5h-HXS7EDJn59clx6V/view).\n* Convert DTU to chunks by running `python src/scripts/convert_dtu.py --input_dir PATH_TO_DTU --output_dir datasets/dtu`\n* [Optional] Generate the evaluation index by running `python src/scripts/generate_dtu_evaluation_index.py --n_contexts=N`, where N is the number of context views. (For N=2 and N=3, we have already provided our tested version under `/assets`.)\n\n## Running the Code\n\n### Evaluation\n\nTo render novel views and compute evaluation metrics from a pretrained model,\n\n* get the [pretrained models](https://drive.google.com/drive/folders/14_E_5R6ojOWnLSrSVLVEMHnTiKsfddjU), and save them to `/checkpoints`\n\n* run the following:\n\n```bash\n# re10k\npython -m src.main +experiment=re10k \\\ncheckpointing.load=checkpoints/re10k.ckpt \\\nmode=test \\\ndataset/view_sampler=evaluation \\\ntest.compute_scores=true\n\n# acid\npython -m src.main +experiment=acid \\\ncheckpointing.load=checkpoints/acid.ckpt \\\nmode=test \\\ndataset/view_sampler=evaluation \\\ndataset.view_sampler.index_path=assets/evaluation_index_acid.json \\\ntest.compute_scores=true\n```\n\n* the rendered novel views will be stored under `outputs/test`\n\nTo render videos from a pretrained model, run the following\n\n```bash\n# re10k\npython -m src.main +experiment=re10k \\\ncheckpointing.load=checkpoints/re10k.ckpt \\\nmode=test \\\ndataset/view_sampler=evaluation \\\ndataset.view_sampler.index_path=assets/evaluation_index_re10k_video.json \\\ntest.save_video=true \\\ntest.save_image=false \\\ntest.compute_scores=false\n```\n\n### Training\n\nRun the following:\n\n```bash\n# download the backbone pretrained weight from unimatch and save to 'checkpoints/'\nwget 'https://s3.eu-central-1.amazonaws.com/avg-projects/unimatch/pretrained/gmdepth-scale1-resumeflowthings-scannet-5d9d7964.pth' -P checkpoints\n# train mvsplat\npython -m src.main +experiment=re10k data_loader.train.batch_size=14\n```\n\nOur models are trained with a single A100 (80GB) GPU. They can also be trained on multiple GPUs with smaller RAM by setting a smaller `data_loader.train.batch_size` per GPU.\n\n\u003cdetails\u003e\n  \u003csummary\u003e\u003cb\u003eTraining on multiple nodes (https://github.com/donydchen/mvsplat/issues/32)\u003c/b\u003e\u003c/summary\u003e\nSince this project is built on top of pytorch_lightning, it can be trained on multiple nodes hosted on the SLURM cluster. For example, to train on 2 nodes (with 2 GPUs on each node), add the following lines to the SLURM job script\n\n```bash\n#SBATCH --nodes=2           # should match with trainer.num_nodes\n#SBATCH --gres=gpu:2        # gpu per node\n#SBATCH --ntasks-per-node=2\n\n# optional, for debugging\nexport NCCL_DEBUG=INFO\nexport HYDRA_FULL_ERROR=1\n# optional, set network interface, obtained from ifconfig\nexport NCCL_SOCKET_IFNAME=[YOUR NETWORK INTERFACE]\n# optional, set IB GID index\nexport NCCL_IB_GID_INDEX=3\n\n# run the command with 'srun'\nsrun python -m src.main +experiment=re10k \\\ndata_loader.train.batch_size=4 \\\ntrainer.num_nodes=2\n```\n\nReferences:\n* [Pytorch Lightning: RUN ON AN ON-PREM CLUSTER (ADVANCED)](https://lightning.ai/docs/pytorch/stable/clouds/cluster_advanced.html)\n* [NCCL: How to set NCCL_SOCKET_IFNAME](https://github.com/NVIDIA/nccl/issues/286)\n* [NCCL: NCCL WARN NET/IB](https://github.com/NVIDIA/nccl/issues/426)\n\n\u003c/details\u003e\n\n\u003cdetails\u003e\n  \u003csummary\u003e\u003cb\u003eFine-tune from the released weights (https://github.com/donydchen/mvsplat/issues/45)\u003c/b\u003e\u003c/summary\u003e\nTo fine-tune from the released weights \u003ci\u003ewithout\u003c/i\u003e loading the optimizer states, run the following:\n\n```bash\npython -m src.main +experiment=re10k data_loader.train.batch_size=14 \\\ncheckpointing.load=checkpoints/re10k.ckpt \\\ncheckpointing.resume=false\n```\n\n\u003c/details\u003e\n\n### Ablations\n\nWe also provide a collection of our [ablation models](https://drive.google.com/drive/folders/14_E_5R6ojOWnLSrSVLVEMHnTiKsfddjU) (under folder 'ablations'). To evaluate them, *e.g.*, the 'base' model, run the following command\n\n```bash\n# Table 3: base\npython -m src.main +experiment=re10k \\\ncheckpointing.load=checkpoints/ablations/re10k_worefine.ckpt \\\nmode=test \\\ndataset/view_sampler=evaluation \\\ntest.compute_scores=true \\\nwandb.name=abl/re10k_base \\\nmodel.encoder.wo_depth_refine=true \n```\n\n### Cross-Dataset Generalization\n\nWe use the default model trained on RealEstate10K to conduct cross-dataset evaluations. To evaluate them, *e.g.*, on DTU, run the following command\n\n```bash\n# Table 2: RealEstate10K -\u003e DTU\npython -m src.main +experiment=dtu \\\ncheckpointing.load=checkpoints/re10k.ckpt \\\nmode=test \\\ndataset/view_sampler=evaluation \\\ndataset.view_sampler.index_path=assets/evaluation_index_dtu_nctx2.json \\\ntest.compute_scores=true\n```\n\n**More running commands can be found at [more_commands.sh](more_commands.sh).**\n\n## BibTeX\n\n```bibtex\n@article{chen2024mvsplat,\n    title   = {MVSplat: Efficient 3D Gaussian Splatting from Sparse Multi-View Images},\n    author  = {Chen, Yuedong and Xu, Haofei and Zheng, Chuanxia and Zhuang, Bohan and Pollefeys, Marc and Geiger, Andreas and Cham, Tat-Jen and Cai, Jianfei},\n    journal = {arXiv preprint arXiv:2403.14627},\n    year    = {2024},\n}\n```\n\n## Acknowledgements\n\nThe project is largely based on [pixelSplat](https://github.com/dcharatan/pixelsplat) and has incorporated numerous code snippets from [UniMatch](https://github.com/autonomousvision/unimatch). Many thanks to these two projects for their excellent contributions!\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fdonydchen%2Fmvsplat","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fdonydchen%2Fmvsplat","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fdonydchen%2Fmvsplat/lists"}