{"id":19216985,"url":"https://github.com/hustvl/evf-sam","last_synced_at":"2025-05-16T04:06:43.112Z","repository":{"id":251040538,"uuid":"813930335","full_name":"hustvl/EVF-SAM","owner":"hustvl","description":"Official code of \"EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model\"","archived":false,"fork":false,"pushed_at":"2025-03-17T07:15:19.000Z","size":6227,"stargazers_count":403,"open_issues_count":8,"forks_count":18,"subscribers_count":8,"default_branch":"main","last_synced_at":"2025-05-07T17:42:01.112Z","etag":null,"topics":["multimodal","multimodal-large-language-models","referring-image-segmentation","segment-anything","segmentation"],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/hustvl.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2024-06-12T02:43:25.000Z","updated_at":"2025-05-07T02:20:34.000Z","dependencies_parsed_at":"2025-01-29T02:37:36.718Z","dependency_job_id":"e85bd319-4136-439f-b96a-3e5f4bb7b9c2","html_url":"https://github.com/hustvl/EVF-SAM","commit_stats":null,"previous_names":["hustvl/evf-sam"],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/hustvl%2FEVF-SAM","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/hustvl%2FEVF-SAM/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/hustvl%2FEVF-SAM/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/hustvl%2FEVF-SAM/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/hustvl","download_url":"https://codeload.github.com/hustvl/EVF-SAM/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":254464896,"owners_count":22075570,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["multimodal","multimodal-large-language-models","referring-image-segmentation","segment-anything","segmentation"],"created_at":"2024-11-09T14:19:43.591Z","updated_at":"2025-05-16T04:06:38.102Z","avatar_url":"https://github.com/hustvl.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"\u003cdiv align =\"center\"\u003e\n\u003cimg src=\"assets/logo.jpg\" width=\"20%\"\u003e\n\u003ch1\u003e 📷 EVF-SAM \u003c/h1\u003e\n\u003ch3\u003e Early Vision-Language Fusion for Text-Prompted Segment Anything Model \u003c/h3\u003e\n\n[Yuxuan Zhang](https://github.com/CoderZhangYx)\u003csup\u003e1,\\*\u003c/sup\u003e, [Tianheng Cheng](https://scholar.google.com/citations?user=PH8rJHYAAAAJ\u0026hl=zh-CN)\u003csup\u003e1,\\*\u003c/sup\u003e, Lei Liu\u003csup\u003e2\u003c/sup\u003e, Heng Liu\u003csup\u003e2\u003c/sup\u003e, Longjin Ran\u003csup\u003e2\u003c/sup\u003e, Xiaoxin Chen\u003csup\u003e2\u003c/sup\u003e, [Wenyu Liu](http://eic.hust.edu.cn/professor/liuwenyu)\u003csup\u003e1\u003c/sup\u003e, [Xinggang Wang](https://xwcv.github.io/)\u003csup\u003e1,📧\u003c/sup\u003e\n\n\u003csup\u003e1\u003c/sup\u003e Huazhong University of Science and Technology, \u003csup\u003e2\u003c/sup\u003e vivo AI Lab\n\n(\\* equal contribution, 📧 corresponding author)\n\n[![arxiv paper](https://img.shields.io/badge/arXiv-Paper-red)](https://arxiv.org/abs/2406.20076)\n[![🤗 HuggingFace models](https://img.shields.io/badge/HuggingFace🤗-Models-orange)](https://huggingface.co/YxZhang/)  \n[![🤗 HuggingFace Demo](https://img.shields.io/badge/EVF_SAM-🤗_HF_Demo-orange)](https://huggingface.co/spaces/wondervictor/evf-sam)\n[![🤗 HuggingFace Demo](https://img.shields.io/badge/EVF_SAM_2-🤗_HF_Demo-orange)](https://huggingface.co/spaces/wondervictor/evf-sam2)\n[![colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/hustvl/EVF-SAM/blob/main/inference_image.ipynb)\n\n\u003c/div\u003e\n\n## News\n* 2025.3: We've released the latest evaluation benchmark [GSEval](https://github.com/hustvl/GroundingSuite) as the comprehensive and multi-granular grounding benchmark.\n* 2025.1: Preview! the EVF-SAM v2 is on the way, going to support salient object segmentation, salient object matte, and referring matte! Besides, better performance on original capabilities is developed!  \n* 2024.9: We have expanded our EVF-SAM to powerful [SAM-2](https://github.com/facebookresearch/segment-anything-2). Besides fewer parameters and improvements on image prediction, our new model also performs well on video prediction (powered by SAM-2). Only at the expense of a simple image training process on RES datasets, we find our EVF-SAM has zero-shot video text-prompted capability. Try our code!\n\n## Highlight\n\u003cdiv align =\"center\"\u003e\n\u003cimg src=\"assets/architecture.jpg\"\u003e\n\u003c/div\u003e\n\n* EVF-SAM extends SAM's capabilities with text-prompted segmentation, achieving high accuracy in Referring Expression Segmentation.  \n* EVF-SAM is designed for efficient computation, enabling rapid inference in few seconds per image on a T4 GPU.\n\n\n## Updates\n- [x] Release code\n- [x] Release weights\n- [x] Release demo 👉 [🤗 evf-sam](https://huggingface.co/spaces/wondervictor/evf-sam)\n- [x] Release code and weights based on SAM-2\n- [x] Update demo supporting SAM-2👉 [🤗 evf-sam2](https://huggingface.co/spaces/wondervictor/evf-sam2)\n- [x] release new checkpoint supporting body part segmentation and semantic level segmentation.\n- [x] update demo supporting multitask\n\n\n## Visualization \n\u003ctable class=\"center\"\u003e\n\u003ctr\u003e\n  \u003ctd style=\"text-align:center;\"\u003e\u003cb\u003eInput text\u003c/b\u003e\u003c/td\u003e\n  \u003ctd style=\"text-align:center;\"\u003e\u003cb\u003eInput image\u003c/b\u003e\u003c/td\u003e\n  \u003ctd style=\"text-align:center;\"\u003e\u003cb\u003eOutput\u003c/b\u003e\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n  \u003ctd width=20% style=\"text-align:center;\"\u003e\u003cb\u003e\"zebra top left\"\u003c/b\u003e\u003c/td\u003e\n  \u003ctd\u003e\u003cimg src=\"assets/zebra.jpg\"\u003e\u003c/td\u003e\n  \u003ctd\u003e\u003cimg src=\"assets/zebra_vis.png\"\u003e\u003c/td\u003e\n\u003c/tr\u003e \n\n\u003ctr\u003e\n  \u003ctd width=20% style=\"text-align:center;\"\u003e\u003cb\u003e\"a pizza with a yellow sign on top of it\"\u003c/b\u003e\u003c/td\u003e\n  \u003ctd\u003e\u003cimg src=\"assets/pizza.jpg\"\u003e\u003c/td\u003e\n  \u003ctd\u003e\u003cimg src=\"assets/pizza_vis.png\"\u003e\u003c/td\u003e\n\u003c/tr\u003e \n\n\u003ctr\u003e\n  \u003ctd width=20% style=\"text-align:center;\"\u003e\u003cb\u003e\"the broccoli closest to the ketchup bottle\"\u003c/b\u003e\u003c/td\u003e\n  \u003ctd\u003e\u003cimg src=\"assets/food.jpg\"\u003e\u003c/td\u003e\n  \u003ctd\u003e\u003cimg src=\"assets/food_vis.png\"\u003e\u003c/td\u003e\n\u003c/tr\u003e \n\n\u003ctr\u003e\n  \u003ctd width=20% style=\"text-align:center;\"\u003e\u003cb\u003e\"[semantic] hair\"\u003c/b\u003e\u003c/td\u003e\n  \u003ctd\u003e\u003cimg src=\"assets/man_sdxl.png\"\u003e\u003c/td\u003e\n  \u003ctd\u003e\u003cimg src=\"assets/man_sdxl_vis.webp\"\u003e\u003c/td\u003e\n\u003c/tr\u003e \n\n\u003ctr\u003e\n  \u003ctd width=20% style=\"text-align:center;\"\u003e\u003cb\u003e\"[semantic] sea\"\u003c/b\u003e\u003c/td\u003e\n  \u003ctd\u003e\u003cimg src=\"assets/seaside_sdxl.png\"\u003e\u003c/td\u003e\n  \u003ctd\u003e\u003cimg src=\"assets/seaside_sdxl_vis.webp\"\u003e\u003c/td\u003e\n\u003c/tr\u003e \n\n\u003c/table\u003e\n\n\n## Installation\n1. Clone this repository  \n2. Install [pytorch](https://pytorch.org/) for your cuda version. **Note** that torch\u003e=2.0.0 is needed if you are to use SAM-2, and torch\u003e=2.2 is needed if you want to enable flash-attention. (We use torch==2.0.1 with CUDA 11.7 and it works fine.)\n3. pip install -r requirements.txt\n4. If you are to use the video prediction function, run:\n```\ncd model/segment_anything_2\npython setup.py build_ext --inplace\n```\n\n\n## Weights\n\u003ctable class=\"center\"\u003e\n\u003ctr\u003e\n  \u003ctd style=\"text-align:center;\"\u003e\u003cb\u003eName\u003c/b\u003e\u003c/td\u003e\n  \u003ctd style=\"text-align:center;\"\u003e\u003cb\u003eSAM\u003c/b\u003e\u003c/td\u003e\n  \u003ctd style=\"text-align:center;\"\u003e\u003cb\u003eBEIT-3\u003c/b\u003e\u003c/td\u003e\n  \u003ctd style=\"text-align:center;\"\u003e\u003cb\u003eParams\u003c/b\u003e\u003c/td\u003e\n  \u003ctd style=\"text-align:center;\"\u003e\u003cb\u003ePrompt Encoder \u0026 Mask Decoder\n  \u003ctd style=\"text-align:center;\"\u003e\u003cb\u003eReference Score\u003c/b\u003e\u003c/td\u003e\n\u003c/tr\u003e\n    \n\u003ctr\u003e\n  \u003ctd style=\"text-align:center;\"\u003e\u003ca href=\"https://huggingface.co/YxZhang/evf-sam-multitask\"\u003eEVF-SAM-multitask\u003c/a\u003e\u003c/td\u003e\n  \u003ctd style=\"text-align:center;\"\u003e\u003cb\u003eSAM-H\u003c/b\u003e\u003c/td\u003e\n  \u003ctd style=\"text-align:center;\"\u003e\u003cb\u003eBEIT-3-L\u003c/b\u003e\u003c/td\u003e\n  \u003ctd style=\"text-align:center;\"\u003e\u003cb\u003e1.32B\u003c/b\u003e\u003c/td\u003e\n  \u003ctd style=\"text-align:center;\"\u003e\u003cb\u003etrain\u003c/b\u003e\u003c/td\u003e\n  \u003ctd style=\"text-align:center;\"\u003e\u003cb\u003e84.2\u003c/b\u003e\u003c/td\u003e\n\u003c/tr\u003e\n    \n\u003ctr\u003e\n  \u003ctd style=\"text-align:center;\"\u003e\u003ca href=\"https://huggingface.co/YxZhang/evf-sam2-multitask\"\u003eEVF-SAM2-multitask\u003c/a\u003e\u003c/td\u003e\n  \u003ctd style=\"text-align:center;\"\u003e\u003cb\u003eSAM-2-L\u003c/b\u003e\u003c/td\u003e\n  \u003ctd style=\"text-align:center;\"\u003e\u003cb\u003eBEIT-3-L\u003c/b\u003e\u003c/td\u003e\n  \u003ctd style=\"text-align:center;\"\u003e\u003cb\u003e898M\u003c/b\u003e\u003c/td\u003e\n  \u003ctd style=\"text-align:center;\"\u003e\u003cb\u003efreeze\u003c/b\u003e\u003c/td\u003e\n  \u003ctd style=\"text-align:center;\"\u003e\u003cb\u003e83.2\u003c/b\u003e\u003c/td\u003e\n\u003c/tr\u003e\n\n\u003ctr\u003e\n  \u003ctd style=\"text-align:center;\"\u003e\u003ca href=\"https://huggingface.co/YxZhang/evf-sam\"\u003eEVF-SAM\u003c/a\u003e\u003c/td\u003e\n  \u003ctd style=\"text-align:center;\"\u003e\u003cb\u003eSAM-H\u003c/b\u003e\u003c/td\u003e\n  \u003ctd style=\"text-align:center;\"\u003e\u003cb\u003eBEIT-3-L\u003c/b\u003e\u003c/td\u003e\n  \u003ctd style=\"text-align:center;\"\u003e\u003cb\u003e1.32B\u003c/b\u003e\u003c/td\u003e\n  \u003ctd style=\"text-align:center;\"\u003e\u003cb\u003etrain\u003c/b\u003e\u003c/td\u003e\n  \u003ctd style=\"text-align:center;\"\u003e\u003cb\u003e83.7\u003c/b\u003e\u003c/td\u003e\n\u003c/tr\u003e\n\n\u003ctr\u003e\n  \u003ctd style=\"text-align:center;\"\u003e\u003ca href=\"https://huggingface.co/YxZhang/evf-sam2\"\u003eEVF-SAM2\u003c/a\u003e\u003c/td\u003e\n  \u003ctd style=\"text-align:center;\"\u003e\u003cb\u003eSAM-2-L\u003c/b\u003e\u003c/td\u003e\n  \u003ctd style=\"text-align:center;\"\u003e\u003cb\u003eBEIT-3-L\u003c/b\u003e\u003c/td\u003e\n  \u003ctd style=\"text-align:center;\"\u003e\u003cb\u003e898M\u003c/b\u003e\u003c/td\u003e\n  \u003ctd style=\"text-align:center;\"\u003e\u003cb\u003efreeze\u003c/b\u003e\u003c/td\u003e\n  \u003ctd style=\"text-align:center;\"\u003e\u003cb\u003e83.6\u003c/b\u003e\u003c/td\u003e\n\u003c/tr\u003e\n\n\n\n\u003ctr\u003e\n  \u003ctd style=\"text-align:center;\"\u003e\u003cb\u003eEVF-Effi-SAM-L \u003c/b\u003e\u003c/td\u003e\n  \u003ctd style=\"text-align:center;\"\u003e\u003cb\u003eEfficientSAM-S\u003c/b\u003e\u003c/td\u003e\n  \u003ctd style=\"text-align:center;\"\u003e\u003cb\u003eBEIT-3-L\u003c/b\u003e\u003c/td\u003e\n  \u003ctd style=\"text-align:center;\"\u003e\u003cb\u003e700M\u003c/b\u003e\u003c/td\u003e\n  \u003ctd style=\"text-align:center;\"\u003e\u003cb\u003etrain\u003c/b\u003e\u003c/td\u003e\n  \u003ctd style=\"text-align:center;\"\u003e\u003cb\u003e83.5\u003c/b\u003e\u003c/td\u003e\n\u003c/tr\u003e\n\n\u003ctr\u003e\n  \u003ctd style=\"text-align:center;\"\u003e\u003cb\u003eEVF-Effi-SAM-B \u003c/b\u003e\u003c/td\u003e\n  \u003ctd style=\"text-align:center;\"\u003e\u003cb\u003eEfficientSAM-T\u003c/b\u003e\u003c/td\u003e\n  \u003ctd style=\"text-align:center;\"\u003e\u003cb\u003eBEIT-3-B\u003c/b\u003e\u003c/td\u003e\n  \u003ctd style=\"text-align:center;\"\u003e\u003cb\u003e232M\u003c/b\u003e\u003c/td\u003e\n  \u003ctd style=\"text-align:center;\"\u003e\u003cb\u003etrain\u003c/b\u003e\u003c/td\u003e\n  \u003ctd style=\"text-align:center;\"\u003e\u003cb\u003e80.0\u003c/b\u003e\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/table\u003e\n\n1. -multimask checkpoints are only available with commits\u003e=9d00853, while other checkpoints are available with commits\u003c9d00853\n\n2. -multimask checkpoints are jointly trained on Ref, ADE20k, Object365, PartImageNet, humanparsing, pascal part datasets. These checkpoints are able to segment part (e.g., hair, arm), background object (e.g., sky, ground), and semantic-level masks. (by adding special token \"\\[semantic\\] \" in front your prompt)\n\n## Inference\n### 1. image prediction\n```\npython inference.py  \\\n  --version \u003cpath to evf-sam\u003e \\\n  --precision='fp16' \\\n  --vis_save_path \"\u003cpath to your output direction\u003e\" \\\n  --model_type \u003c\"ori\" or \"effi\" or \"sam2\", depending on your loaded ckpt\u003e   \\\n  --image_path \u003cpath to your input image\u003e \\\n  --prompt \u003ccustomized text prompt\u003e\n```\n`--load_in_8bit` and `--load_in_4bit` are **optional**  \nfor example: \n```\npython inference.py  \\\n  --version YxZhang/evf-sam2 \\\n  --precision='fp16' \\\n  --vis_save_path \"vis\" \\\n  --model_type sam2   \\\n  --image_path \"assets/zebra.jpg\" \\\n  --prompt \"zebra top left\"\n```\n\n### 2. video prediction  \nfirstly slice video into frames\n```\nffmpeg -i \u003cyour_video\u003e.mp4 -q:v 2 -start_number 0 \u003cframe_dir\u003e/'%05d.jpg'\n```\nthen:\n```\npython inference_video.py  \\\n  --version \u003cpath to evf-sam2\u003e \\\n  --precision='fp16' \\\n  --vis_save_path \"vis/\" \\\n  --image_path \u003cframe_dir\u003e   \\\n  --prompt \u003ccustomized text prompt\u003e   \\\n  --model_type sam2\n```\nyou can use frame2video.py to concat the predicted frames to a video.\n\n## Demo\nimage demo\n```\npython demo.py \u003cpath to evf-sam\u003e\n```\nvideo demo\n```\npython demo_video.py \u003cpath to evf-sam2\u003e\n```\n\n## Data preparation\nReferring segmentation datasets: [refCOCO](https://web.archive.org/web/20220413011718/https://bvisionweb1.cs.unc.edu/licheng/referit/data/refcoco.zip), [refCOCO+](https://web.archive.org/web/20220413011656/https://bvisionweb1.cs.unc.edu/licheng/referit/data/refcoco+.zip), [refCOCOg](https://web.archive.org/web/20220413012904/https://bvisionweb1.cs.unc.edu/licheng/referit/data/refcocog.zip), [refCLEF](https://web.archive.org/web/20220413011817/https://bvisionweb1.cs.unc.edu/licheng/referit/data/refclef.zip) ([saiapr_tc-12](https://web.archive.org/web/20220515000000/http://bvisionweb1.cs.unc.edu/licheng/referit/data/images/saiapr_tc-12.zip)) and [COCO2014train](http://images.cocodataset.org/zips/train2014.zip)  \n```\n├── dataset\n│   ├── refer_seg\n│   │   ├── images\n│   │   |   ├── saiapr_tc-12 \n│   │   |   └── mscoco\n│   │   |       └── images\n│   │   |           └── train2014\n│   │   ├── refclef\n│   │   ├── refcoco\n│   │   ├── refcoco+\n│   │   └── refcocog\n```\n\n## Evaluation\n```\ntorchrun --standalone --nproc_per_node \u003cnum_gpus\u003e eval.py   \\\n    --version \u003cpath to evf-sam\u003e \\\n    --dataset_dir \u003cpath to your data root\u003e   \\\n    --val_dataset \"refcoco|unc|val\" \\\n    --model_type \u003c\"ori\" or \"effi\" or \"sam2\", depending on your loaded ckpt\u003e\n```\n\n## Acknowledgement\nWe borrow some codes from [LISA](https://github.com/dvlab-research/LISA/tree/main), [unilm](https://github.com/microsoft/unilm), [SAM](https://github.com/facebookresearch/segment-anything), [EfficientSAM](https://github.com/yformer/EfficientSAM), [SAM-2](https://github.com/facebookresearch/segment-anything-2).\n\n## Citation\n```bibtex\n@article{zhang2024evfsamearlyvisionlanguagefusion,\n      title={EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model}, \n      author={Yuxuan Zhang and Tianheng Cheng and Rui Hu and Lei Liu and Heng Liu and Longjin Ran and Xiaoxin Chen and Wenyu Liu and Xinggang Wang},\n      year={2024},\n      eprint={2406.20076},\n      archivePrefix={arXiv},\n      primaryClass={cs.CV},\n      url={https://arxiv.org/abs/2406.20076}, \n}\n```\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fhustvl%2Fevf-sam","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fhustvl%2Fevf-sam","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fhustvl%2Fevf-sam/lists"}