{"id":16269985,"url":"https://github.com/jshilong/gpt4roi","last_synced_at":"2025-04-06T09:10:12.805Z","repository":{"id":179326641,"uuid":"663121272","full_name":"jshilong/GPT4RoI","owner":"jshilong","description":"GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest","archived":false,"fork":false,"pushed_at":"2023-09-21T02:04:37.000Z","size":15826,"stargazers_count":365,"open_issues_count":11,"forks_count":17,"subscribers_count":7,"default_branch":"main","last_synced_at":"2023-11-07T21:46:56.644Z","etag":null,"topics":["computer-vision","gpt","llm","multimodality","roi"],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"other","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/jshilong.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null}},"created_at":"2023-07-06T15:42:06.000Z","updated_at":"2024-04-23T14:29:51.950Z","dependencies_parsed_at":"2024-02-14T11:30:13.330Z","dependency_job_id":null,"html_url":"https://github.com/jshilong/GPT4RoI","commit_stats":null,"previous_names":["jshilong/gpt4roi"],"tags_count":0,"template":null,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/jshilong%2FGPT4RoI","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/jshilong%2FGPT4RoI/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/jshilong%2FGPT4RoI/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/jshilong%2FGPT4RoI/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/jshilong","download_url":"https://codeload.github.com/jshilong/GPT4RoI/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":247457803,"owners_count":20941906,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["computer-vision","gpt","llm","multimodality","roi"],"created_at":"2024-10-10T18:09:26.865Z","updated_at":"2025-04-06T09:10:12.785Z","avatar_url":"https://github.com/jshilong.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"# GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest :fire: [Demo](http://139.196.83.164:7000/) :fire:\n\n\n[//]: # (\u003cdiv id=\"wrapper\" align=\"center\"\u003e)\n\n[//]: # (\u003cfigure\u003e)\n\n[//]: # (  \u003cimg src=\"figs/demo1.gif\" width=\"45%\"\u003e\u0026emsp;)\n\n[//]: # (  \u003cimg src=\"figs/demo2.gif\" width=\"45%\"\u003e\u003cbr\u003e)\n\n[//]: # (  \u003cp style=\"font-size:1.2vw;\"\u003eLeft: Single-Region Understanding; Right: Single-Region Understanding\u003c/p\u003e)\n\n[//]: # (\u003c/figure\u003e)\n\n[//]: # (\u003c/div\u003e)\n\n\n\u003cp align=\"center\"\u003e\n    \u003cimg src=\"figs/demo1.gif\" width=\"80%\"\u003e \u003cbr\u003e\n  \u003cp align=\"center\" style=\"font-size:1.2vw;\"\u003eSingle-Region Understanding\u003c/p\u003e\n\u003c/p\u003e\n\u003cp align=\"center\"\u003e\n    \u003cimg src=\"figs/demo2.gif\" width=\"80%\"\u003e \u003cbr\u003e\n  \u003cp align=\"center\"  style=\"font-size:1.2vw;\"\u003eMultiple-Region Understanding\u003c/p\u003e\n\u003c/p\u003e\n\n\n\n## Introduction\n\u003cp align=\"center\"\u003e\n    \u003cimg src=\"figs/framework.png\" width=\"70%\"\u003e \u003cbr\u003e\n\u003c/p\u003e\n\n\u003e [**GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest**](https://arxiv.org/abs/2307.03601)               \n\u003e [Shilong Zhang*](https://jshilong.github.io/), [Peize Sun*](https://peizesun.github.io/), [Shoufa Chen*](https://www.shoufachen.com/), Min Xiao, Wenqi Shao ,Wenwei Zhang, Kai Chen, Ping Luo\u003c/br\u003e\n\u003e (*Equal Contribution) \n\n### [[Demo](http://139.196.83.164:7000/)]  [[Paper](https://arxiv.org/abs/2307.03601)] [[中文介绍](https://zhuanlan.zhihu.com/p/640283103)]\n\n[//]: # (#:grin::grin::grin:信交流群：xxx \u0026#40;答案：cheems\u0026#41;)\n\n\n\n## Updates\n\n- [July 25]  [GPT4RoI-7B-delta-V0](https://huggingface.co/shilongz/GPT4RoI-7B-delta-V0) has release ! :fire::fire::fire: You need to combine our delta with the original LLaMA weights follow the [GPT4RoI Weights](https://github.com/jshilong/GPT4RoI/tree/main#weights) section. \n- [July 7]  All training and inference code has been released, you can try demo [here](http://139.196.83.164:7000/) :fire::fire::fire:\n\n\n## Contents\n- [Install](#Install)\n- [Data](#Data)\n- [GPT4RoI Weights](#Weights)\n- [Training](#Training)\n- [Gradio](#Gradio)\n- [Acknowledge](#Acknowledge)\n\n\n## Install\n1. Clone the `GPT4RoI`\n```python\ngit clone https://github.com/jshilong/gpt4roi.git\ncd gpt4roi\n```\n\n2. Create the env\n```shell\nconda create -n gpt4roi python=3.10 -y\nconda activate gpt4roi\npip install --upgrade pip  # enable PEP 660 support\npip install setuptools_scm\npip install --no-cache-dir  -e .\n# please use conda re-install the torch, pip may loss some runtime lib\nconda install pytorch torchvision torchaudio pytorch-cuda=11.7 -c pytorch -c nvidia \n```\n3. Install the `flash-attn` package \n```\npip install ninja\npip install flash-attn --no-build-isolation\n```\n4. install the `mmcv-1.4.7` package\nMake sure that your `nvcc -V` is consistent with cudatookit version of `python -c \"import torch;print(torch.version.cuda)`.\n```shell\ncd mmcv-1.4.7\nMMCV_WITH_OPS=1 pip install -e .\n```\n\n\u003c!-- ## Data Preparation\n\n| Data file name | Size | original from|\n| --- | --- | ---|\n| [single_region_caption.json](https://huggingface.co/datasets/liuhaotian/LLaVA-Instruct-150K/raw/main/llava_instruct_150k.json) | 229 MB | VC, Refcocog |\n| [multi_region_caption.json](https://huggingface.co/datasets/liuhaotian/LLaVA-Instruct-150K/raw/main/llava_instruct_80k.json) | 229 MB | flicker30k |\n| [spation-instruction21k.json](https://huggingface.co/datasets/liuhaotian/LLaVA-Instruct-150K/raw/main/conversation_58k.json) | 126 MB | VCR |\n\n\nWe also use langauge-image multimodal instruction-folllowing dataset [`LLaVA-Instruct-150K`](https://huggingface.co/datasets/liuhaotian/LLaVA-Instruct-150K),  with we detect objects with EVA-lvis model, you should download the  ``eva_lvis_coco.pkl`` when you use this dataset. --\u003e\n\n## Data\n\nOur dataset includes RefCOCO, RefCOCO+, RefCOCOg, Visual Genome, Flickr30K entities, and the VCR dataset. We are sincerely grateful to the creators of these datasets, especially for the VCR dataset, for their forward-thinking in creating these dataset.\n\nThe dataset section of this repository may appear somewhat messy, especially the VCR part(still finishing), which may cause GPT4RoI not be very user-friendly. We are currently working on formulating the datasets into a unified format and will be accompanying them with stronger models. Please stay tuned for updates.\n\n\nYou can download the corresponding dataset from the official website and organize it as follows. Afterwards, you can modify the ```gpt4roi/configs/dataset_config.json``` file to select the specific dataset you want to use:\n\n```text\nGPT4RoI\n├── data\n│   ├── coco_det\n│   │   ├── annotations\n│   │   │      ├──instances_train2017.json\n│   │   ├── train2017/\n│   ├── mdetr_annotations\n│   │          ├──finetune_refcoco_train.json\n│   │          ├──finetune_refcoco+_train.json\n│   │          ├──finetune_refcocog_train.json\n│   │          ├──final_flickr_mergedGT_train.json\n│   ├── coco_imgs/\n│   ├── flickr30k-images/\n│   ├── visual_genome\n│   │          ├──train.json\n│   │          ├──vg_all/\n│   ├── llava\n│   │   ├── llava_instruct_150k.json\n│   │   ├── llava_150k_bbox_pred_results.pkl\n│   ├── vcr\n│   │   ├── train.jsonl\n│   │   ├── vcr1images/\n```\n### NOTE\n1. coco_imgs should contains all coco image(you can soft link them to this directory.\n2. We use Visual_Genome_Dataset_V1.2, available for download from  [OpenDataLab](https://opendatalab.com/). Ensure to download the  [train.json](https://datarelease.blob.core.windows.net/grit/VG_preprocessed_annotations/train.json), you should create a soft link for all VG images to the directory `vg_all`.\n3. [llava_150k_bbox_pred_results.pkl](https://huggingface.co/shilongz/temp/tree/main) contains the detection predicted results with EVA-02-DET. We appreciate their work.\n\n\n\n## Weights\nDue to the licensing restrictions of LLaMA, the delta weights GPT4RoI-7B is produced from LLaMA-7B. To acquire the GPT4RoI weights, you need to combine our delta with the original LLaMA weights.\n\n### Step1. Download the original LLaMA-7B weights\nThe original LLaMA weights are available for download. Use the following commands:\n```shell\ngit lfs install\ngit clone https://huggingface.co/decapoda-research/llama-7b-hf ./llama-7b \n```\n\nAlternatively, access the [webpage](https://huggingface.co/decapoda-research/llama-7b-hf/tree/main) to download the file.\n\n### Step2. Download the delta weights of GPT4RoI-7B\n\nThe delta weights for GPT4RoI-7B can be downloaded using the following commands:\n```shell\ngit lfs install\ngit clone https://huggingface.co/shilongz/GPT4RoI-7B-delta-V0 ./GPT4RoI-7B-delta\n```\nYou can also directly download the file from this [webpage](https://huggingface.co/shilongz/GPT4RoI-7B-delta-V0/tree/main).\n\n### Step3. Apply the delta weights to the original LLaMA-7B weights\nApply the delta weights to the original LLaMA-7B weights. Note that this conversion command requires approximately 30 GB of CPU RAM.\n```bash\nexport PYTHONPATH=`pwd`:$PYTHONPATH\npython3 -m scripts.apply_delta \\\n    --base ./llama-7b \\\n    --target ./GPT4RoI-7B \\\n    --delta ./GPT4RoI-7B-delta\n```\n\n## Training\nGPT4RoI is trained on 8 A100 with the following code.\n\n### STAGE 1\nVicuna-v0, an instruction-tuned chatbot, is the base model for this setup. In order to prepare it, first download the delta weights available [here](https://huggingface.co/lmsys/vicuna-7b-delta-v0). To obtain the original weights, follow the instructions provided [here](https://github.com/lm-sys/FastChat/blob/main/docs/vicuna_weights_version.md#how-to-apply-delta-weights-for-weights-v11-and-v0) to integrate these delta weights into LLaMA-7B.\n\nEnsure to download the following projector weight file: [LLaVA-7b-pretrain-projector-v0-CC3M-595K-original_caption.bin](https://huggingface.co/liuhaotian/LLaVA-Pretrained-Projectors/resolve/main/LLaVA-7b-pretrain-projector-v0-CC3M-595K-original_caption.bin).\n\nAdditionally, you have the flexibility to choose from different versions of Vicuna (such as the 13B version or llama v2 chatbot) and the corresponding projector weights from [LLaVA](https://github.com/haotian-liu/LLaVA) to meet your specific requirements effectively.\n`exp/stage1` is the work directory. \n```Shell\nbash train_stage1.sh exp/stage1\n# Resume training in stage1\n# bash train_stage1.sh exp/stage1\n```\n### STAGE 2\n\n`exp/stage2` is the work directory. and you should give the work directory of stage1 so we can load the corresponding weight as pretrain model.\n```Shell\n# At the beginning of stage2\nbash train_stage2.sh exp/stage2 exp/stage1\n# Resume training in stage2\n# bash train_stage2.sh exp/stage2 \n```\n\n\n\n## Gradio\nPlease install [Gradio Box](https://github.com/ShoufaChen/gradio-dev) first.\n```python\npython gpt4roi/app.py\n```\n### NOTES\n1. ```prompt format in GPT4RoI```\nYou should always use `\u003cregion1\u003e, \u003cregion2\u003e...` to refer the new bounding box in the image when you first draw them. Then you can use normal `region 1` in the conversation to refer the instance.\n2. You should always click the `clear all` buttul and waiting the clear process finished before you start a new conversation.\n\n\u003cp align=\"center\"\u003e\n    \u003cimg src=\"figs/fig1.png\" width=\"100%\"\u003e \u003cbr\u003e\n\u003cp align=\"center\"  style=\"font-size:1.2vw;\"\u003eMultiple Rounds of Dialogue\u003c/p\u003e\n\u003c/p\u003e\n\n\n\n\n## Acknowledge\n\n- [LLaVA](https://github.com/haotian-liu/LLaVA): The codebase we built upon.\n- [Vicuna](https://github.com/lm-sys/FastChat): The LLM we used.\n- [VCR](https://visualcommonsense.com/): We get strong region reasoning ability from this forward thinking dataset.\n\nIf you find GPT4RoI useful for your your research and applications, please cite using this BibTeX:\n```bibtex\n@article{zhang2023gpt4roi,\n  title={Gpt4roi: Instruction tuning large language model on region-of-interest},\n  author={Zhang, Shilong and Sun, Peize and Chen, Shoufa and Xiao, Min and Shao, Wenqi and Zhang, Wenwei and Chen, Kai and Luo, Ping},\n  journal={arXiv preprint arXiv:2307.03601},\n  year={2023}\n}\n```\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fjshilong%2Fgpt4roi","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fjshilong%2Fgpt4roi","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fjshilong%2Fgpt4roi/lists"}