{"id":14043879,"url":"https://github.com/UCSC-VLAA/vllm-safety-benchmark","last_synced_at":"2025-07-27T15:31:50.231Z","repository":{"id":209564505,"uuid":"722413490","full_name":"UCSC-VLAA/vllm-safety-benchmark","owner":"UCSC-VLAA","description":"[ECCV 2024] Official PyTorch Implementation of \"How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs\"","archived":false,"fork":false,"pushed_at":"2023-11-28T02:38:01.000Z","size":3326,"stargazers_count":54,"open_issues_count":0,"forks_count":2,"subscribers_count":4,"default_branch":"main","last_synced_at":"2024-08-12T08:13:04.554Z","etag":null,"topics":["adversarial-attacks","benchmark","datasets","llm","multimodal-llm","robustness","safety","vision-language-model"],"latest_commit_sha":null,"homepage":"https://arxiv.org/abs/2311.16101","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/UCSC-VLAA.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null}},"created_at":"2023-11-23T05:05:37.000Z","updated_at":"2024-08-11T13:05:04.000Z","dependencies_parsed_at":"2023-11-28T03:38:17.973Z","dependency_job_id":null,"html_url":"https://github.com/UCSC-VLAA/vllm-safety-benchmark","commit_stats":null,"previous_names":["ucsc-vlaa/vllm-safety-benchmark"],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/UCSC-VLAA%2Fvllm-safety-benchmark","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/UCSC-VLAA%2Fvllm-safety-benchmark/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/UCSC-VLAA%2Fvllm-safety-benchmark/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/UCSC-VLAA%2Fvllm-safety-benchmark/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/UCSC-VLAA","download_url":"https://codeload.github.com/UCSC-VLAA/vllm-safety-benchmark/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":227814471,"owners_count":17823909,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["adversarial-attacks","benchmark","datasets","llm","multimodal-llm","robustness","safety","vision-language-model"],"created_at":"2024-08-12T08:06:36.638Z","updated_at":"2024-12-02T22:31:46.599Z","avatar_url":"https://github.com/UCSC-VLAA.png","language":"Python","funding_links":[],"categories":["Attack","Adversarial-Attack"],"sub_categories":[],"readme":"\u003c!-- \u003cp align=\"center\"\u003e\n  \u003cimg src=\"unicorn.png\" width=\"80\"\u003e\n\u003c/p\u003e --\u003e\n\n# How many \u003cimg src=\"assets/unicorn.png\" width=\"36\"\u003e Are in This Image? A Safety Evaluation Benchmark for Vision LLMs\n\n\n[Haoqin Tu*](https://www.haqtu.me/), [Chenhang Cui*](https://gzcch.github.io/), [Zijun Wang*](https://asillycat.github.io/), [Yiyang Zhou](https://yiyangzhou.github.io/), [Bingchen Zhao](https://bzhao.me), [Junlin Han](https://junlinhan.github.io/), [Wangchunshu Zhou](https://michaelzhouwang.github.io/), [Huaxiu Yao](https://www.huaxiuyao.io/), [Cihang Xie](https://cihangxie.github.io/) (*equal contribution)\n\n[![Code License](https://img.shields.io/badge/Code%20License-Apache_2.0-green.svg)](https://github.com/tatsu-lab/stanford_alpaca/blob/main/LICENSE)\n[![Data License](https://img.shields.io/badge/Data%20License-CC%20By%20NC%204.0-red.svg)](https://github.com/tatsu-lab/stanford_alpaca/blob/main/DATA_LICENSE)\n\nOur paper is online now: https://arxiv.org/abs/2311.16101\n\n\u003cp align=\"center\"\u003e\n  \u003cimg src=\"assets/teaser.png\" width=\"1080\"\u003e\n\u003c/p\u003e\n\n## Installation\nFor different VLLMs, please refer to their specific envirnments for installation.\n\n- LLaVA: https://github.com/haotian-liu/LLaVA\n- MiniGPT4: https://github.com/Vision-CAIR/MiniGPT-4\n- InstructBLIP: https://huggingface.co/Salesforce/instructblip-vicuna-7b\n- LLaMA-Adapter: https://github.com/OpenGVLab/LLaMA-Adapter\n- mPLUG-Owl1\u00262: https://github.com/X-PLUG/mPLUG-Owl\n- PandaGPT: https://github.com/yxuansu/PandaGPT\n- Qwen-VL-Chat: https://huggingface.co/Qwen/Qwen-VL-Chat\n- CogVLM: https://github.com/THUDM/CogVLM\n- InternLM-Xcomposer: https://huggingface.co/internlm/internlm-xcomposer-7b\n- Fuyu: https://huggingface.co/adept/fuyu-8b\n\n## Datasets\nWe host our datasets [here](https://huggingface.co/datasets/PahaII/vllm_safety_evaluation), containing both OOD and redteaming attack datasets. The full dataset should looks like this:\n\n```\n.\n├── ./safety_evaluation_benchmark_datasets//                    \n    ├── gpt4v_challenging_set # Contains the challenging test data for GPT4V\n        ├── attack_images\n        ├── sketchy_images\n        ├── oodcv_images\n        ├── misleading-attack.json\n        ├── sketchy-vqa-challenging.json\n        └── oodcv-vqa-counterfactual.json\n    ├── redteaming # Contains the test data for redteaming tasks\n        ├── misleading_attack\n            ├── gaussian_noise\n            ├── mixattack_eps32\n            ├── mixattack_eps64\n            ├── sinattack_eps64_dog\n            ├── sinattack_eps64_coconut\n            ├── sinattack_eps64_spaceship\n            └── annotation.json\n        ├── jailbreak_vit # adversarial images for jailbreaking VLLM through ViT\n        └── jailbreak_llm # adversarial suffixes for jailbreaking VLLM through LLM\n    └── ood # Contains the test data for OOD scenarios\n        ├── sketchy-vqa\n            ├── sketchy-vqa.json\n            ├── sketchy-challenging.json\n        └── oodcv-vqa\n            ├── oodcv-vqa.json\n            └── oodcv-counterfactual.json\n```\n\n### Out-of-Distribution Scenario\nFor $\\texttt{OODCV-VQA}$ and its counterfactual version, please download images from [OODCV](https://drive.google.com/file/d/1jq43Q0cenISIq7acW0LS-Lqghgy8exsj/view?usp=share_link), and put all images in `ood/oodcv-vqa`.\n\nFor $\\texttt{Sketchy-VQA}$ and its challenging version, please first download images from [here](https://cybertron.cg.tu-berlin.de/eitz/projects/classifysketch/sketches_png.zip), put the zip file into `ood/sketchy-vqa/skechydata/`, then unzip it.\n\n### Redteaming Attack\nFor the proposed misleading attack, the full datasets and all trained adversarial examples are in `redteaming/misleading_attack`, including images with gaussian noise, Sin.Attack and MixAttack with two pertubation budgets $\\epsilon=32/255$ (eps32) or $\\epsilon=64/255$ (eps64).\n\nFor jailbreaking methods, please refer to their respective repositories for more dataset details: [Jailbreak through ViT](https://github.com/Unispac/Visual-Adversarial-Examples-Jailbreak-Large-Language-Models), [Jailbreak through LLM](https://github.com/llm-attacks/llm-attacks).\n\n## Testing\nBefore you start, make sure you have modified the `CACHE_DIR` (where you store all your model weights) and `DATA_DIR` (where you store the benchmark data) in `baselines/config.json` according to your local envirnment.\n\n```bash\ncd baselines\npython ../model_testing_zoo.py --model_name LLaVA\n```\nChoose `--model_name` from [\"LlamaAdapterV2\", \"MiniGPT4\", \"MiniGPT4v2\", \"LLaVA\", \"mPLUGOwl\", \"mPLUGOwl2\", \"PandaGPT\", \"InstructBLIP2\", \"Flamingo\", \"LLaVAv1.5\", \"LLaVAv1.5-13B\", \"LLaVA_llama2-13B\", \"MiniGPT4_llama2\", \"Qwen-VL-Chat\", \"MiniGPT4_13B\", \"InstructBLIP2-FlanT5-xl\", \"InstructBLIP2-FlanT5-xxl\",  \"InstructBLIP2-13B\", \"CogVLM\", \"Fuyu\", \"InternLM\"].\n\n### $\\texttt{OODCV-VQA}$ and its Counterfactual Variant\n\nFor $\\texttt{OODCV-VQA}$:\n```bash\ncd baselines\npython ../safety_evaluations/ood_scenarios/evaluation.py --model_name LLaVA --eval_oodcv\n```\n\nFor the counterfactual version:\n\n```bash\ncd baselines\npython ../safety_evaluations/ood_scenarios/evaluation.py --model_name LLaVA --eval_oodcv_cf\n```\n\n### $\\texttt{Sketchy-VQA}$ and its Challenging Variant\n\nFor $\\texttt{Sketchy-VQA}$:\n```bash\ncd baselines\npython ../safety_evaluations/ood_scenarios/evaluation.py --model_name LLaVA --eval_sketch\n```\n\nFor the challenging version:\n\n```bash\ncd baselines\npython ../safety_evaluations/ood_scenarios/evaluation.py --model_name LLaVA --eval_sketch_challenging\n```\n\n### Misleading Attack\nFor training the misleading adversarial images:\n\n```bash\ncd safety_evaluations/redteaming/misleading_vision_attack\n\npython misleading_vis_attack.py --lr 1e-3 --misleading_obj dog --input_folder path/to/attack-bard/NIPS2017 --output_folder ./misleading_adversarial_attack\n```\nChange `--input_folder` to the path of adversarial examples you want to test. If you want to use the MixAttack, add `--mix_obj` argument to the command.\n\nFor testing the VLLMs:\n\n```bash\ncd baselines\n\npython ../safety_evaluations/redteaming/misleading_vision_attack/test_misleading.py --image_folder redteaming/misleading_attack/mixattack_eps64 --output_name misleading_attack_eps64 --human_annot_path redteaming/misleading_attack/annotation.json\n```\n\n### Jailbreaking Methods\n\nPlease refer to these two repositories for detailed attack settings: [Jailbreak through ViT](https://github.com/Unispac/Visual-Adversarial-Examples-Jailbreak-Large-Language-Models), [Jailbreak through LLM](https://github.com/llm-attacks/llm-attacks). We give our trained adversarial images and suffixes for jailbreaking ViTs and LLMs in `redteaming/jailbreak_vit` and `redteaming/jailbreak_llm` in the data folder.\n\n## Usage and License Notices\nThe data, code and checkpoint is intended and licensed for research use only. The dataset is CC BY NC 4.0 (allowing only non-commercial use) and models trained using the dataset should not be used outside of research purposes.\n\n## Citation\nIf you find our work useful to your research and applications, please consider citing the paper and staring the repo :)\n\n```bibtex\n@article{tu2023how,\n  title={How Many Unicorns Are In This Image? A Safety Evaluation Benchmark For Vision LLMs},\n  author={Tu, Haoqin and Cui, Chenhang and Wang, Zijun and Zhou, Yiyang and Zhao, Bingchen and Han, Junlin and Zhou, Wangchunshu and Yao, Huaxiu and Xie, Cihang},\n  journal={arXiv preprint arXiv:2311.16101},\n  year={2023}\n}\n```\n\n## Acknowledgement\nThis work is partially supported by a gift from Open Philanthropy. We thank Center for AI Safety and Google Cloud for supporting our computing needs. Any opinions, findings, and conclusions or recommendations expressed in this material are those of the author(s) and do not necessarily reflect the views of the sponsors.\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FUCSC-VLAA%2Fvllm-safety-benchmark","html_url":"https://awesome.ecosyste.ms/projects/github.com%2FUCSC-VLAA%2Fvllm-safety-benchmark","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FUCSC-VLAA%2Fvllm-safety-benchmark/lists"}