{"id":38653603,"url":"https://github.com/cambrian-mllm/cambrian-s","last_synced_at":"2026-01-17T09:24:46.132Z","repository":{"id":322968783,"uuid":"1075703212","full_name":"cambrian-mllm/cambrian-s","owner":"cambrian-mllm","description":"Cambrian-S: Towards Spatial Supersensing in Video","archived":false,"fork":false,"pushed_at":"2025-12-22T05:57:48.000Z","size":4399,"stargazers_count":436,"open_issues_count":3,"forks_count":14,"subscribers_count":5,"default_branch":"main","last_synced_at":"2025-12-23T17:08:14.187Z","etag":null,"topics":["computer-vision","llm","multimodal-large-language-models","spatial-understanding","vision-language-model"],"latest_commit_sha":null,"homepage":"https://cambrian-mllm.github.io/","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/cambrian-mllm.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2025-10-13T21:46:59.000Z","updated_at":"2025-12-23T15:54:48.000Z","dependencies_parsed_at":null,"dependency_job_id":null,"html_url":"https://github.com/cambrian-mllm/cambrian-s","commit_stats":null,"previous_names":["cambrian-mllm/cambrian-s"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/cambrian-mllm/cambrian-s","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/cambrian-mllm%2Fcambrian-s","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/cambrian-mllm%2Fcambrian-s/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/cambrian-mllm%2Fcambrian-s/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/cambrian-mllm%2Fcambrian-s/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/cambrian-mllm","download_url":"https://codeload.github.com/cambrian-mllm/cambrian-s/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/cambrian-mllm%2Fcambrian-s/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":28505489,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-01-17T06:57:29.758Z","status":"ssl_error","status_checked_at":"2026-01-17T06:56:03.931Z","response_time":85,"last_error":"SSL_read: unexpected eof while reading","robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":false,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["computer-vision","llm","multimodal-large-language-models","spatial-understanding","vision-language-model"],"created_at":"2026-01-17T09:24:45.942Z","updated_at":"2026-01-17T09:24:46.077Z","avatar_url":"https://github.com/cambrian-mllm.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"\u003cdiv align=\"center\"\u003e\n\n# \u003cimg src=\"figs/tesseract.png\" alt=\"Cambrian-S\" width=\"28\" height=\"28\" style=\"vertical-align: middle;\"\u003e *Cambrian-S*:\u003cbr\u003e Towards Spatial Supersensing in Video\n\n\n\u003cp\u003e\n    \u003cimg src=\"figs/feature.png\" alt=\"Cambrian-S\" width=\"800\" height=\"auto\"\u003e\n\u003c/p\u003e\n\n\u003ca href=\"https://arxiv.org/abs/2511.04670\" target=\"_blank\"\u003e\n    \u003cimg alt=\"arXiv\" src=\"https://img.shields.io/badge/arXiv-Cambrian--S-red?logo=arxiv\" height=\"25\" /\u003e\n\u003c/a\u003e\n\u003ca href=\"https://cambrian-mllm.github.io/cambrian-s/\" target=\"_blank\"\u003e\n    \u003cimg alt=\"Website\" src=\"https://img.shields.io/badge/🌎_Website-cambrian--mllm.github.io-blue\" height=\"25\" /\u003e\n\u003c/a\u003e\n\u003ca href=\"https://huggingface.co/collections/nyu-visionx/cambrian-s-models\" target=\"_blank\"\u003e\n    \u003cimg alt=\"HF Model: Cambrian-S\" src=\"https://img.shields.io/badge/%F0%9F%A4%97%20_Model-Cambrian--S-ffc107?color=ffc107\u0026logoColor=white\" height=\"25\" /\u003e\n\u003c/a\u003e\n\u003ca href=\"https://huggingface.co/datasets/nyu-visionx/VSI-590K\" target=\"_blank\"\u003e\n    \u003cimg alt=\"HF Dataset: VSI-590K\" src=\"https://img.shields.io/badge/%F0%9F%A4%97%20_Data-VSI--590K-ffc107?color=ffc107\u0026logoColor=white\" height=\"25\" /\u003e\n\u003c/a\u003e\n\u003ca href=\"https://huggingface.co/collections/nyu-visionx/vsi-super\" target=\"_blank\"\u003e\n    \u003cimg alt=\"HF Dataset: VSI-Super\" src=\"https://img.shields.io/badge/%F0%9F%A4%97%20_Benchmark-VSI--Super-ffc107?color=ffc107\u0026logoColor=white\" height=\"25\" /\u003e\n\u003c/a\u003e\n\u003cdiv style=\"font-family: charter;\"\u003e\n    \u003ca href=\"https://github.com/vealocia\" target=\"_blank\"\u003eShusheng Yang*\u003c/a\u003e,\n    \u003ca href=\"https://jihanyang.github.io/\" target=\"_blank\"\u003eJihan Yang*\u003c/a\u003e,\n    \u003ca href=\"https://pinzhihuang.github.io/\" target=\"_blank\"\u003ePinzhi Huang†\u003c/a\u003e,\n    \u003ca href=\"https://ellisbrown.github.io/\" target=\"_blank\"\u003eEllis Brown†\u003c/a\u003e,\n    \u003ca href=\"https://redagavin.github.io/\" target=\"_blank\"\u003eZihao Yang\u003c/a\u003e,\n    \u003cbr\u003e\n    \u003ca href=\"https://github.com/czyuyue\" target=\"_blank\"\u003eYue Yu\u003c/a\u003e,\n    \u003ca href=\"https://tsb0601.github.io/\" target=\"_blank\"\u003eShengbang Tong\u003c/a\u003e,\n    \u003ca href=\"#\" target=\"_blank\"\u003eZihan Zheng\u003c/a\u003e,\n    \u003ca href=\"https://www.yfxu.com/\" target=\"_blank\"\u003eYifan Xu\u003c/a\u003e,\n    \u003ca href=\"https://github.com/wang-muhan\" target=\"_blank\"\u003eMuhan Wang\u003c/a\u003e,\n    \u003ca href=\"https://daohanlu.github.io/\" target=\"_blank\"\u003eDaohan Lu\u003c/a\u003e,\n    \u003cbr\u003e\n    \u003ca href=\"https://cs.nyu.edu/~fergus/pmwiki/pmwiki.php\" target=\"_blank\"\u003eRob Fergus\u003c/a\u003e,\n    \u003ca href=\"http://yann.lecun.com/\" target=\"_blank\"\u003eYann LeCun\u003c/a\u003e,\n    \u003ca href=\"https://profiles.stanford.edu/fei-fei-li\" target=\"_blank\"\u003eLi Fei-Fei\u003c/a\u003e,\n    \u003ca href=\"https://www.sainingxie.com/\" target=\"_blank\"\u003eSaining Xie\u003c/a\u003e\n\u003c/div\u003e\n\n\u003cdiv style=\"font-family: charter;\"\u003e\n*Equal Contribution\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;†Core Contributor\n\u003c/div\u003e\n\n\u003c/div\u003e\n\n## Release\n- [Dec 21, 2025] 🚀 [Cambrian-S-3M](https://huggingface.co/datasets/nyu-visionx/Cambrian-S-3M) (our collection of 3M open-sourced video instruction tuning data) is now available! Please check it out!\n- [Nov 6, 2025] 🔥 We release Cambrian-S model weights, training code, and evaluation suite.\n- [Nov 6, 2025] 🔥 We release VSI-SUPER, a benchmark designed for spatial supersensing.\n- [Nov 6, 2025] 🔥 We release VSI-590K, a dataset curated for spatial sensing.\n\n## Contents\n- [ *Cambrian-S*: Towards Spatial Supersensing in Video](#-cambrian-s-towards-spatial-supersensing-in-video)\n  - [Release](#release)\n  - [Contents](#contents)\n  - [Cambrian-S Weights](#cambrian-s-weights)\n    - [General Model Performance](#general-model-performance)\n    - [VSI-SUPER Performance](#vsi-super-performance)\n    - [Model Card](#model-card)\n      - [Model Trained with Predictive Sensing](#model-trained-with-predictive-sensing)\n      - [Standard MLLM Models](#standard-mllm-models)\n  - [VSI-590K Dataset](#vsi-590k-dataset)\n  - [Train](#train)\n  - [Evaluation](#evaluation)\n  - [Citation](#citation)\n  - [Related Projects](#related-projects)\n\n## Cambrian-S Weights\n\nHere are our Cambrian-S checkpoints along with instructions on how to use the weights. Our models excel at spatial reasoning in video understanding, demonstrating significant improvements over previous state-of-the-art methods on spatial understanding benchmarks while maintaining competitive performance on general video understanding tasks.\n\n### General Model Performance\n\nComparison of Cambrian-S with other leading MLLMs on general video understanding benchmarks.\n\n\u003cp align=\"center\"\u003e\n    \u003cimg src=\"figs/model_performance_table.png\" alt=\"General Model Performance\" width=\"1000\" height=\"auto\"\u003e\n\u003c/p\u003e\n\n**Results**: Cambrian-S maintains competitive performance on standard video benchmarks (Perception Test and EgoSchema) while excelling at spatial reasoning tasks.\n\n### VSI-SUPER Performance\n\nVSI-SUPER performance is evaluated on **Cambrian-S-7B-LFP**. \n\n\u003cdiv align=\"center\" style=\"display: flex; justify-content: center; gap: 10px; width:100%\"\u003e\n    \u003cimg src=\"figs/count_results.png\" alt=\"VSI-SUPER Count Results\" width=\"48%\" style=\"display: inline-block; margin: 0 1%;\"\u003e\n    \u003cimg src=\"figs/sor_results.png\" alt=\"VSI-SUPER SOR Results\" width=\"48%\" style=\"display: inline-block; margin: 0 1%;\"\u003e\n\u003c/div\u003e\n\n\n### Model Card\n\n#### Model Trained with Predictive Sensing\n\n| Model           | Base-LLM | Vision Encoder | Hugging Face                                                    |\n|-----------------|------------|----------------|------------------------------------------------------------------|\n| Cambrian-S-7B-LFP   | `Qwen2.5-7B-Instruct`         | `siglip2-so400m-patch14-384`     | [nyu-visionx/Cambrian-S-7B-LFP](https://huggingface.co/nyu-visionx/Cambrian-S-7B-LFP)   |\n\n#### Standard MLLM Models\n\n| Model           | Base-LLM | Vision Encoder | Hugging Face                                                    |\n|-----------------|------------|----------------|------------------------------------------------------------------|\n| Cambrian-S-7B   | `Qwen2.5-7B-Instruct`         | `siglip2-so400m-patch14-384`     | [nyu-visionx/Cambrian-S-7B](https://huggingface.co/nyu-visionx/Cambrian-S-7B)   |\n| Cambrian-S-3B   | `Qwen2.5-3B-Instruct`         | `siglip2-so400m-patch14-384`     | [nyu-visionx/Cambrian-S-3B](https://huggingface.co/nyu-visionx/cambrian-s-3b)   |\n| Cambrian-S-1.5B | `Qwen2.5-1.5B-Instruct`       | `siglip2-so400m-patch14-384`     | [nyu-visionx/Cambrian-S-1.5B](https://huggingface.co/nyu-visionx/cambrian-s-1.5b) |\n| Cambrian-S-0.5B | `Qwen2.5-0.5B-Instruct`       | `siglip2-so400m-patch14-384`     | [nyu-visionx/Cambrian-S-0.5B](https://huggingface.co/nyu-visionx/cambrian-s-0.5b) | \n\n## VSI-590K Dataset\n\nVSI-590K is a video instruction-tuning dataset focusing on spatial understanding. \n\n\u003cp align=\"center\"\u003e\n    \u003cimg src=\"figs/data_construct_v6.png\" alt=\"VSI-590K Data Construction\" width=\"1000\" height=\"auto\"\u003e\n\u003c/p\u003e\n\n\n\n**VSI-590K dataset statistics.** \n\n\u003cp align=\"center\"\u003e\n    \u003cimg src=\"figs/vsi590_details.png\" alt=\"VSI-590K Details\" width=\"80%\" height=\"auto\"\u003e\n\u003c/p\u003e\n\nQAs are grouped by: question types (left) and task groups (right).\n\n\u003cdiv align=\"center\" style=\"display: flex; justify-content: center; gap: 0px; width:100%\"\u003e\n    \u003cimg src=\"figs/vsi_piechart.png\" alt=\"VSI-590K Pie Chart\" width=\"48%\" style=\"display: inline-block; margin: 0 1%;\"\u003e\n    \u003cimg src=\"figs/vsi_tasktype_chart.png\" alt=\"VSI-590K Task Type Chart\" width=\"48%\" style=\"display: inline-block; margin: 0 1%;\"\u003e\n\u003c/div\u003e\n\n**Hugging Face**: [nyu-visionx/VSI-590K](https://huggingface.co/datasets/nyu-visionx/vsi-590k)\n\n\n## Train\n\n### Environment Preparation\n\nCurrently, we support training on TPU using TorchXLA. Install `TorchXLA 2.6.0` by the following commands:\n\n```bash\npip install torch==2.6.0 torchvision==0.21.0 torch_xla==2.6.0\npip install 'torch_xla[tpu]' -f https://storage.googleapis.com/libtpu-releases/index.html\npip install 'torch_xla[pallas]' -f https://storage.googleapis.com/jax-releases/jax_nightly_releases.html -f https://storage.googleapis.com/jax-releases/jaxlib_nightly_releases.html\npip install --upgrade pip\npip install -e '.[tpu]'\n```\n\n### Data Preparation\n\nCambrian-S models are trained on top of [`Cambrian-Alignment`](https://huggingface.co/datasets/nyu-visionx/Cambrian-Alignment), [`Cambrian-7M`](https://huggingface.co/datasets/nyu-visionx/Cambrian-10M), [`Cambrian-S-3M`](https://huggingface.co/datasets/nyu-visionx/Cambrian-S-3M), and [`VSI-590K`](https://huggingface.co/datasets/nyu-visionx/VSI-590K) datasets. Please prepare these datasets following their corresponding guidelines.\n\n### Training Scripts\n\nAs mentioned in our paper, Cambrian-S models are trained in 4 stages: from vision-language alignment, to general image instruction tuning, and general video instruction tuning, and finally spatial video tuning. For Cambrian-S-LFP model, we modified the 4th stage by involving latent frame prediction objective. We provides sample training scripts in the following:\n\n* [cambrian/scripts/cambrians_7b_s1.sh](cambrian/scripts/cambrians_7b_s1.sh)\n* [cambrian/scripts/cambrians_7b_s2.sh](cambrian/scripts/cambrians_7b_s2.sh)\n* [cambrian/scripts/cambrians_7b_s3.sh](cambrian/scripts/cambrians_7b_s3.sh)\n* [cambrian/scripts/cambrians_7b_s4.sh](cambrian/scripts/cambrians_7b_s4.sh)\n* [cambrian/scripts/cambrians_7b_lfp_s4.sh](cambrian/scripts/cambrians_7b_lfp_s4.sh)\n\n## Evaluation\n\nWe have released our evaluation code in the [`lmms-eval/`](lmms-eval/) subfolder. Please see the README there for more details.\n\nFor detailed benchmark results, please refer to the [General Model Performance](#general-model-performance) and [VSI-SUPER Performance](#vsi-super-performance) sections above.\n\n## Citation\n\nIf you find our work useful for your research, please consider to cite our work:\n\n```bibtex\n@article{yang2025cambrians,\n  title={Cambrian-S: Towards Spatial Supersensing in Video},\n  author={Yang, Shusheng and Yang, Jihan and Huang, Pinzhi and Brown, Ellis and Yang, Zihao and Yu, Yue and Tong, Shengbang and Zheng, Zihan and Xu, Yifan and Wang, Muhan and Lu, Daohan and Fergus, Rob and LeCun, Yann and Fei-Fei, Li and Xie, Saining},\n  journal={arXiv preprint arXiv:2511.04670},\n  year={2025}\n}\n\n@article{brown2025shortcuts,\n  author = {Brown, Ellis and Yang, Jihan and Yang, Shusheng and Fergus, Rob and Xie, Saining},\n  title = {Benchmark Designers Should ``Train on the Test Set'' to Expose Exploitable Non-Visual Shortcuts},\n  journal = {arXiv preprint arXiv:2511.04655},\n  year = {2025}\n}\n\n@article{brown2025simsv,\n  title   =  { {SIMS-V}: Simulated Instruction-Tuning for Spatial Video Understanding },\n  author  =  { Brown, Ellis and Ray, Arijit and Krishna, Ranjay and Girshick, Ross and Fergus, Rob and Xie, Saining },\n  journal =  { arXiv preprint arXiv:2511.04668 },\n  year    =  { 2025 }\n}\n\n@article{yang2024think,\n    title={{Thinking in Space: How Multimodal Large Language Models See, Remember and Recall Spaces}},\n    author={Yang, Jihan and Yang, Shusheng and Gupta, Anjali W. and Han, Rilyn and Fei-Fei, Li and Xie, Saining},\n    year={2024},\n    journal={arXiv preprint arXiv:2412.14171},\n}\n\n@article{tong2024cambrian,\n  title={{Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs}},\n  author={Tong, Shengbang and Brown, Ellis and Wu, Penghao and Woo, Sanghyun and Middepogu, Manoj and Akula, Sai Charitha and Yang, Jihan and Yang, Shusheng, and Iyer, Adithya and Pan, Xichen and Wang, Austin and Fergus, Rob and LeCun, Yann and Xie, Saining},\n  journal={arXiv preprint arXiv:2406.16860},\n  year={2024}\n}\n```\n\n## Related Projects\n\n- [Cambrian-1](https://github.com/cambrian-mllm/cambrian): A Fully Open, Vision-Centric Exploration of Multimodal LLMs\n- [Thinking in Space](https://vision-x-nyu.github.io/thinking-in-space.github.io/): How Multimodal Large Language Models See, Remember and Recall Spaces - Introduces VSI-Bench for evaluating visual-spatial intelligence\n- [SIMS-V](https://ellisbrown.github.io/sims-v): Simulated Instruction-Tuning for Spatial Video Understanding\n- [Test-Set Stress-Test](https://vision-x-nyu.github.io/test-set-training): Benchmark Designers Should \"Train on the Test Set\" to Expose Exploitable Non-Visual Shortcuts\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fcambrian-mllm%2Fcambrian-s","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fcambrian-mllm%2Fcambrian-s","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fcambrian-mllm%2Fcambrian-s/lists"}