{"id":27993988,"url":"https://github.com/damo-nlp-sg/videorefer","last_synced_at":"2025-05-08T19:05:26.982Z","repository":{"id":270492849,"uuid":"907220344","full_name":"DAMO-NLP-SG/VideoRefer","owner":"DAMO-NLP-SG","description":"[CVPR 2025] The code for \"VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM\"","archived":false,"fork":false,"pushed_at":"2025-04-28T15:40:04.000Z","size":136307,"stargazers_count":194,"open_issues_count":5,"forks_count":11,"subscribers_count":10,"default_branch":"main","last_synced_at":"2025-05-08T19:05:12.803Z","etag":null,"topics":["mllm","pixel-understanding","sam2","video-understanding"],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/DAMO-NLP-SG.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null}},"created_at":"2024-12-23T05:24:06.000Z","updated_at":"2025-05-08T09:43:45.000Z","dependencies_parsed_at":"2025-02-27T07:30:10.450Z","dependency_job_id":"d6bed375-db00-4b30-ab80-c734ff1f6df5","html_url":"https://github.com/DAMO-NLP-SG/VideoRefer","commit_stats":null,"previous_names":["damo-nlp-sg/videorefer"],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/DAMO-NLP-SG%2FVideoRefer","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/DAMO-NLP-SG%2FVideoRefer/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/DAMO-NLP-SG%2FVideoRefer/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/DAMO-NLP-SG%2FVideoRefer/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/DAMO-NLP-SG","download_url":"https://codeload.github.com/DAMO-NLP-SG/VideoRefer/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":253133128,"owners_count":21859111,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["mllm","pixel-understanding","sam2","video-understanding"],"created_at":"2025-05-08T19:05:26.205Z","updated_at":"2025-05-08T19:05:26.945Z","avatar_url":"https://github.com/DAMO-NLP-SG.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"\u003cp align=\"center\"\u003e\n    \u003cimg src=\"assets/videorefer.png\" width=\"80%\" style=\"margin-bottom: 0.2;\"/\u003e\n\u003cp\u003e\n\n\u003ch3 align=\"center\"\u003e\u003ca href=\"http://arxiv.org/abs/2501.00599\" style=\"color:#4D2B24\"\u003e\nVideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM\u003c/a\u003e\u003c/h3\u003e\n\n\u003cdiv align=center\u003e\n\n![Static Badge](https://img.shields.io/badge/VideoRefer-v1-F7C97E) \n[![arXiv preprint](https://img.shields.io/badge/arxiv-2501.00599-ECA8A7?logo=arxiv)](http://arxiv.org/abs/2501.00599) \n[![Dataset](https://img.shields.io/badge/Dataset-Hugging_Face-E59FB6)](https://huggingface.co/datasets/DAMO-NLP-SG/VideoRefer-700K) \n[![Model](https://img.shields.io/badge/Model-Hugging_Face-CFAFD4)](https://huggingface.co/DAMO-NLP-SG/VideoRefer-7B) \n[![Benchmark](https://img.shields.io/badge/Benchmark-Hugging_Face-96D03A)](https://huggingface.co/datasets/DAMO-NLP-SG/VideoRefer-Bench) \n\n[![video](https://img.shields.io/badge/Watch_Video-36600E?logo=youtube\u0026logoColor=green)](https://www.youtube.com/watch?v=gLNOj1OPFJE)\n[![Homepage](https://img.shields.io/badge/Homepage-visit-9DC3E6)](https://damo-nlp-sg.github.io/VideoRefer/) \n\n\u003c/div\u003e\n\n\u003cp align=\"center\"\u003e\n    \u003cimg src=\"assets/demo.gif\" width=\"100%\" style=\"margin-bottom: 0.2;\"/\u003e\n\u003cp\u003e\n\n\u003cp align=\"center\" style=\"margin-bottom: 5px;\"\u003e\n  VideoRefer can understand any object you're interested within a video.\n\u003c/p\u003e\n\n\n## 📰 News\n* **[2025.4.22]** 🔥Our VideoRefer-Bench has been adopted in [Describe Anything Model](https://arxiv.org/pdf/2504.16072) (NVIDIA \u0026 UC Berkeley).\n* **[2025.2.27]** 🔥VideoRefer Suite has been accepted to CVPR2025!\n* **[2025.2.18]**  🔥We release the [VideoRefer-700K dataset](https://huggingface.co/datasets/DAMO-NLP-SG/VideoRefer-700K) on HuggingFace.\n* **[2025.1.1]**  🔥We release the code of VideoRefer and the VideoRefer-Bench.\n\n\n## 🎥 Video\n\nhttps://github.com/user-attachments/assets/d943c101-72f3-48aa-9822-9cfa46fa114b\n\n- HD video can be viewed on [YouTube](https://www.youtube.com/watch?v=gLNOj1OPFJE).\n\n\n## 🔍 About VideoRefer Suite \n\n`VideoRefer Suite` is designed to enhance the fine-grained spatial-temporal understanding capabilities of Video Large Language Models (Video LLMs). It consists of three primary components:\n\n* **Model (VideoRefer)**\n\n`VideoRefer` is an effective Video LLM, which enables fine-grained perceiving, reasoning, and retrieval for user-defined regions at any specified timestamps—supporting both single-frame and multi-frame region inputs.\n\n\u003cp align=\"center\"\u003e\n    \u003cimg src=\"assets/model.png\" width=\"90%\" style=\"margin-bottom: 0.2;\"/\u003e\n\u003cp\u003e\n\n\n* **Dataset (VideoRefer-700K)**\n\n`VideoRefer-700K` is a large-scale, high-quality object-level video instruction dataset. Curated using a sophisticated multi-agent data engine to fill the gap for high-quality object-level video instruction data.\n\n\u003cp align=\"center\"\u003e\n    \u003cimg src=\"assets/dataset.png\" width=\"90%\" style=\"margin-bottom: 0.2;\"/\u003e\n\u003cp\u003e\n\n\n* **Benchmark (VideoRefer-Bench)**\n\n`VideoRefer-Bench` is a comprehensive benchmark to evaluate the object-level video understanding capabilities of a model, which consists of two sub-benchmarks: **VideoRefer-Bench-D** and **VideoRefer-Bench-Q**.\n\n\u003cp align=\"center\"\u003e\n    \u003cimg src=\"assets/benchmark.png\" width=\"70%\" style=\"margin-bottom: 0.2;\"/\u003e\n\u003cp\u003e\n\n\n\n## 🛠️ Requirements and Installation\nBasic Dependencies:\n* Python \u003e= 3.8\n* Pytorch \u003e= 2.2.0\n* CUDA Version \u003e= 11.8\n* transformers == 4.40.0 (for reproducing paper results)\n* tokenizers == 0.19.1\n\nInstall required packages:\n```bash\ngit clone https://github.com/DAMO-NLP-SG/VideoRefer\ncd VideoRefer\npip install -r requirements.txt\npip install flash-attn==2.5.8 --no-build-isolation\n```\n\n## 🌟 Getting started\n\nPlease refer to the examples in [infer.ipynb](./demo/infer.ipynb) for detailed instructions on how to use our model for single video inference, which supports both single-frame and multi-frame modes.\n\nFor better usage, the demo integrates with [SAM2](https://github.com/facebookresearch/sam2), to get started, please install SAM2 first:\n\n```shell\ngit clone https://github.com/facebookresearch/sam2.git \u0026\u0026 cd sam2\n\nSAM2_BUILD_CUDA=0 pip install -e \".[notebooks]\"\n```\nThen, download [sam2.1_hiera_large.pt](https://dl.fbaipublicfiles.com/segment_anything_2/092824/sam2.1_hiera_large.pt) to `checkpoints`.\n\n\n## 🗝️ Training \u0026 Evaluation\n### Training\nThe training data and data structure can be found in [Dataset preparation](training.md).\n\nThe training pipeline of our model is structured into four distinct stages.\n\n- **Stage1: Image-Text Alignment Pre-training**\n    - We use the same data as in [VideoLLaMA2.1](https://github.com/DAMO-NLP-SG/VideoLLaMA2).\n    - The pretrained projector weights can be found in [VideoLLaMA2.1-7B-16F-Base](https://huggingface.co/DAMO-NLP-SG/VideoLLaMA2.1-7B-16F-Base).\n\n- **Stage2: Region-Text Alignment Pre-training**\n    - Prepare datasets used for stage2.\n    - Run `bash scripts/train/stage2.sh`.\n\n- **Stage2.5:  High-Quality Knowledge Learning**\n    - Prepare datasets used for stage2.5.\n    - Run `bash scripts/train/stage2.5.sh`.\n    \n- **Stage3:  Visual Instruction Tuning**\n    - Prepare datasets used for stage3.\n    - Run `bash scripts/train/stage3.sh`.\n \n### Evaluation\nFor model evaluation, please refer to [eval](eval/eval.md).\n\n## 🌏 Model Zoo\n| Model Name     | Visual Encoder | Language Decoder | # Training Frames |\n|:----------------|:----------------|:------------------|:----------------:|\n| [VideoRefer-7B](https://huggingface.co/DAMO-NLP-SG/VideoRefer-7B) | [siglip-so400m-patch14-384](https://huggingface.co/google/siglip-so400m-patch14-384) | [Qwen2-7B-Instruct](https://huggingface.co/Qwen/Qwen2-7B-Instruct)  | 16 |\n| [VideoRefer-7B-stage2](https://huggingface.co/DAMO-NLP-SG/VideoRefer-7B-stage2)  | [siglip-so400m-patch14-384](https://huggingface.co/google/siglip-so400m-patch14-384) | [Qwen2-7B-Instruct](https://huggingface.co/Qwen/Qwen2-7B-Instruct)  | 16 |\n| [VideoRefer-7B-stage2.5](https://huggingface.co/DAMO-NLP-SG/VideoRefer-7B-stage2.5)  | [siglip-so400m-patch14-384](https://huggingface.co/google/siglip-so400m-patch14-384) | [Qwen2-7B-Instruct](https://huggingface.co/Qwen/Qwen2-7B-Instruct)  | 16 |\n\n## 🖨️ VideoRefer-700K\nThe dataset can be accessed on [huggingface](https://huggingface.co/datasets/DAMO-NLP-SG/VideoRefer-700K).\n\nBy leveraging our multi-agent data engine, we meticulously create three primary types of object-level video instruction data: \n- Object-level Detailed Caption\n- Object-level Short Caption\n- Object-level QA\n\nVideo sources:\n- Detailed\u0026Short Caption\n    - [Panda-70M](https://snap-research.github.io/Panda-70M/). \n- QA\n    - [MeViS](https://codalab.lisn.upsaclay.fr/competitions/15094)\n    - [A2D](https://web.eecs.umich.edu/~jjcorso/r/a2d/index.html#downloads)\n    - [Youtube-VOS](https://competitions.codalab.org/competitions/29139#participate-get_data)\n\nData format:\n```json\n[\n    {\n        \"video\": \"videos/xxx.mp4\",\n        \"conversations\": [\n            {\n                \"from\": \"human\",\n                \"value\": \"\u003cvideo\u003e\\nWhat is the relationship of \u003cregion\u003e and \u003cregion\u003e?\"\n            },\n            {\n                \"from\": \"gpt\",\n                \"value\": \"....\"\n            },\n            ...\n        ],\n        \"annotation\":[\n            //object1\n            {\n                \"frame_idx\":{\n                    \"segmentation\": {\n                        //rle format or polygon\n                    }\n                }\n                \"frame_idx\":{\n                    \"segmentation\": {\n                        //rle format or polygon\n                    }\n                }\n            },\n            //object2\n            {\n                \"frame_idx\":{\n                    \"segmentation\": {\n                        //rle format or polygon\n                    }\n                }\n            },\n            ...\n        ]\n\n    }\n```\n\n## 🕹️ VideoRefer-Bench\n\n`VideoRefer-Bench` assesses the models in two key areas: Description Generation, corresponding to `VideoRefer-BenchD`, and Multiple-choice Question-Answer, corresponding to `VideoRefer-BenchQ`.\n\nhttps://github.com/user-attachments/assets/33757d27-56bd-4523-92da-8f5a58fe5c85\n\n- The annotations of the benchmark can be found in [🤗benchmark](https://huggingface.co/datasets/DAMO-NLP-SG/VideoRefer-Bench).\n\n- The usage of VideoRefer-Bench is detailed in [doc](./benchmark/README.md).\n\n- To evaluate general MLLMs on VideoRefer-Bench, please refer to [eval](./benchmark/evaluation_general_mllms.md).\n\n\n## 📑 Citation\n\nIf you find VideoRefer Suite useful for your research and applications, please cite using this BibTeX:\n```bibtex\n@article{yuan2025videorefersuite,\n  title = {VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM},\n  author = {Yuqian Yuan, Hang Zhang, Wentong Li, Zesen Cheng, Boqiang Zhang, Long Li, Xin Li, Deli Zhao, Wenqiao Zhang, Yueting Zhuang, Jianke Zhu, Lidong Bing},\n  journal={arXiv},\n  year={2025},\n  url = {http://arxiv.org/abs/2501.00599}\n}\n```\n\n\u003cdetails open\u003e\u003csummary\u003e💡 Some other multimodal-LLM projects from our team may interest you ✨. \u003c/summary\u003e\u003cp\u003e\n\u003c!--  may --\u003e\n\n\u003e [**Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding**](https://github.com/DAMO-NLP-SG/Video-LLaMA) \u003cbr\u003e\n\u003e Hang Zhang, Xin Li, Lidong Bing \u003cbr\u003e\n[![github](https://img.shields.io/badge/-Github-black?logo=github)](https://github.com/DAMO-NLP-SG/Video-LLaMA)  [![github](https://img.shields.io/github/stars/DAMO-NLP-SG/Video-LLaMA.svg?style=social)](https://github.com/DAMO-NLP-SG/Video-LLaMA) [![arXiv](https://img.shields.io/badge/Arxiv-2306.02858-b31b1b.svg?logo=arXiv)](https://arxiv.org/abs/2306.02858) \u003cbr\u003e\n\n\u003e [**VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs**](https://github.com/DAMO-NLP-SG/VideoLLaMA2) \u003cbr\u003e\n\u003e Zesen Cheng, Sicong Leng, Hang Zhang, Yifei Xin, Xin Li, Guanzheng Chen, Yongxin Zhu, Wenqi Zhang, Ziyang Luo, Deli Zhao, Lidong Bing \u003cbr\u003e\n[![github](https://img.shields.io/badge/-Github-black?logo=github)](https://github.com/DAMO-NLP-SG/VideoLLaMA2)  [![github](https://img.shields.io/github/stars/DAMO-NLP-SG/VideoLLaMA2.svg?style=social)](https://github.com/DAMO-NLP-SG/VideoLLaMA2) [![arXiv](https://img.shields.io/badge/Arxiv-2406.07476-b31b1b.svg?logo=arXiv)](https://arxiv.org/abs/2406.07476) \u003cbr\u003e\n\n\u003e [**Osprey: Pixel Understanding with Visual Instruction Tuning**](https://github.com/CircleRadon/Osprey) \u003cbr\u003e\n\u003e Yuqian Yuan, Wentong Li, Jian Liu, Dongqi Tang, Xinjie Luo, Chi Qin, Lei Zhang, Jianke Zhu \u003cbr\u003e\n[![github](https://img.shields.io/badge/-Github-black?logo=github)](https://github.com/CircleRadon/Osprey)  [![github](https://img.shields.io/github/stars/CircleRadon/Osprey.svg?style=social)](https://github.com/CircleRadon/Osprey) [![arXiv](https://img.shields.io/badge/Arxiv-2312.10032-b31b1b.svg?logo=arXiv)](https://arxiv.org/abs/2312.10032) \u003cbr\u003e\n\n\u003c/p\u003e\u003c/details\u003e\n\n\n## 👍 Acknowledgement\nThe codebase of VideoRefer is adapted from [**VideoLLaMA 2**](https://github.com/DAMO-NLP-SG/VideoLLaMA2).\nThe visual encoder and language decoder we used in VideoRefer are [**Siglip**](https://huggingface.co/google/siglip-so400m-patch14-384) and [**Qwen2**](https://huggingface.co/collections/Qwen/qwen2-6659360b33528ced941e557f), respectively.\n\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fdamo-nlp-sg%2Fvideorefer","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fdamo-nlp-sg%2Fvideorefer","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fdamo-nlp-sg%2Fvideorefer/lists"}