{"id":19520130,"url":"https://github.com/osu-nlp-group/amplegcg","last_synced_at":"2025-10-14T22:17:32.836Z","repository":{"id":232856533,"uuid":"776618740","full_name":"OSU-NLP-Group/AmpleGCG","owner":"OSU-NLP-Group","description":"AmpleGCG: Learning a Universal and Transferable Generator of Adversarial Attacks on Both Open and Closed LLM","archived":false,"fork":false,"pushed_at":"2024-11-03T02:48:38.000Z","size":696,"stargazers_count":71,"open_issues_count":2,"forks_count":7,"subscribers_count":2,"default_branch":"main","last_synced_at":"2025-10-04T02:50:41.479Z","etag":null,"topics":["adversarial-attacks","gcg","nlp","safety"],"latest_commit_sha":null,"homepage":"https://github.com/OSU-NLP-Group/AmpleGCG?tab=readme-ov-file","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"other","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/OSU-NLP-Group.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2024-03-24T02:02:36.000Z","updated_at":"2025-10-03T04:10:43.000Z","dependencies_parsed_at":null,"dependency_job_id":"dc0c5976-beaa-49b6-8d1d-a4fac82bd661","html_url":"https://github.com/OSU-NLP-Group/AmpleGCG","commit_stats":null,"previous_names":["osu-nlp-group/amplegcg"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/OSU-NLP-Group/AmpleGCG","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/OSU-NLP-Group%2FAmpleGCG","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/OSU-NLP-Group%2FAmpleGCG/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/OSU-NLP-Group%2FAmpleGCG/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/OSU-NLP-Group%2FAmpleGCG/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/OSU-NLP-Group","download_url":"https://codeload.github.com/OSU-NLP-Group/AmpleGCG/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/OSU-NLP-Group%2FAmpleGCG/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":279021761,"owners_count":26087053,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","status":"online","status_checked_at":"2025-10-14T02:00:06.444Z","response_time":60,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["adversarial-attacks","gcg","nlp","safety"],"created_at":"2024-11-11T00:23:56.865Z","updated_at":"2025-10-14T22:17:32.792Z","avatar_url":"https://github.com/OSU-NLP-Group.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"# AmpleGCG: Learning a Universal and Transferable Generator of Adversarial Attacks on Both Open and Closed LLM\nThis is the official repo of AmpleGCG ([https://arxiv.org/abs/2404.07921](https://arxiv.org/abs/2404.07921)) and AmpleGCG-plus ([https://arxiv.org/abs/2410.22143](https://arxiv.org/abs/2410.22143)). Please kindly 🌟star🌟 this repo and cite our papers 📜 if you find them useful!\n\n\u003ca href=\"https://github.com/OSU-NLP-Group/AmpleGCG?tab=readme-ov-file\" target=\"_blank\"\u003e\n  \u003cimg src=\"https://img.shields.io/badge/AmpleGCG-black?style=flat\u0026logo=python\u0026logoColor=rgb\" alt=\"AmpleGCG Badge\"\u003e\n\u003c/a\u003e\n\n\n## 🚨Updates🚨\n- Nov 2th: Our technical paper of [AmpleGCG-plus](https://arxiv.org/abs/2410.22143) has officially arXived! Check it out!\n- August 27th: 2024: Release of our extensive collection of **millions** of suffixes generated through GCG, along with their corresponding evaluation results.\n\n  In light of the importance of building trustworthy AI systems that should be robust in both **natural** and **gibberish** language spaces, we have decided to release the raw datasets that are used to develop AmpleGCG and AmpleGCG-plus series of models to better contribute to the community. For more reasons why we believe these gibberish suffixes are important, please check the [Tweet Thread](https://x.com/LiaoZeyi/status/1828613837756490112) here. Please apply for it at [here](https://huggingface.co/osunlp/AmpleGCG-llama2-sourced-llama2-7b-chat#request-for-datasets).\n\n- August 1st, 2024: Release of **AmpleGCG-plus**\n\n  We are excited to announce the release of **AmpleGCG-plus**, an enhanced version of AmpleGCG designed to produce customized GCG suffixes. This upgrade introduces two significant improvements:\n\n  1. **Enhanced Data Quality**: We've utilized a more effective and cost-efficient evaluator, harmbench-cls, in our OTF pipeline to collect higher-quality training datasets.\n\n  2. **Enhanced Data Quantity**: Instead of sampling 200 suffixes for each query, **AmpleGCG-plus** now utilizes all available collected training pairs.\n\n  Given that, we've developed two specialized versions of **AmpleGCG-plus**, tailored for Llama-2-chat and GPT-series models with more details in 🤗 [**AmpleGCG-plus**](https://huggingface.co/osunlp/AmpleGCG-plus-llama2-sourced-llama2-7b-chat).\n\n  Both AmpleGCG-**plus** variants demonstrate superior performance compared to the original AmpleGCG when evaluated on AdvBench.\n\n- July 20th: Acceptance to COLM\n\n  We are thrilled to anounce that our [paper](https://arxiv.org/abs/2404.07921) is accepted at [COLM 2024](https://colmweb.org/)\n\n---\n\n## Reproducibility and Codes\nThis repository hosts the source code of **Augmented GCG**, which extends the capabilities of GCG by overgenerating samples alongside the GCG optimizations. Our work builds upon the foundational [GCG](https://github.com/llm-attacks/llm-attacks) work, and we express our deep appreciation for their open-source release.\n\n\nDue to safety and ethical considerations, we have decided not to publicly release the trained **AmpleGCG**, our adversarial suffix generator in the wild. There exists a significant risk that, if used maliciously, AmpleGCG could rapidly compromise the safety of both open-source and proprietary models. Such a scenario could lead to widespread dissemination of harmful content, a risk we aim to mitigate by restricting access to the trained model.\n\n However, one can apply for our trained **AmpleGCG** via 🤗 [AmpleGCG-series models](https://huggingface.co/osunlp/AmpleGCG-llama2-sourced-llama2-7b-chat) and generated adversarial suffixes  via this [Google Form](https://docs.google.com/forms/d/1P8hxsR5_ROE1-J1pyKCqT1GBuIa0RqkwRc3opCAvQ0Y/edit) for research purposes only. Once approval, we will release the suffixes generated by AmpleGCG on [AdvBench](https://arxiv.org/abs/2307.15043) and [MaliciousIntruct](https://arxiv.org/abs/2310.06987). Access to the model and data is granted on a provisional basis and is subject to the sole discretion of the authors.\n\n\n\n## Licensing Information\nThe code under this repo is licensed under an [OPEN RAIL-S License](https://www.licenses.ai/ai-pubs-open-rails-vz1).\n\nThe data under this repo is licensed under an [OPEN RAIL-D License](https://huggingface.co/blog/open_rail).\n\nThe model weight and parameters under this repo are licensed under an [OPEN RAIL-M License](https://www.licenses.ai/ai-pubs-open-railm-vz1).\n\n## Introduction\n**TL;DR** We further amplify the effectiveness of GCG, achieving increased ASR, more comprehensive identification of vulnerabilities, and improved efficiency across both open-source and closed-source models.\n\nAs large language models (LLMs) become increasingly prevalent and integrated into autonomous systems, ensuring their safety is imperative.\nDespite significant strides toward safety alignment, recent work GCG (Zou\net al., 2023) successfully produces a single suffix for each query to jailbreak\nLLMs. In this work, we first identify the overlooked opportunities by\nsolely picking the suffix with the lowest loss during GCG optimization, and consequently, uncover many other missed successful suffixes\nin the middle steps. Moreover, we utilize them as training data to learn\na generator named AmpleGCG, which captures the distribution of adversarial suffixes given a harmful query. This generator facilitates the rapid\ngeneration of hundreds of suffixes for any harmful query in minutes. AmpleGCG achieves near 100% attack success rate (ASR) on two aligned LLMs\n(Llama-2-7B-Chat and Vicuna-7B), surpassing two strongest existing attack\nbaselines. Interestingly, AmpleGCG also transfers effectively to attack different models, including closed-source LLMs, achieving a 99% ASR on the\nlatest GPT-3.5. To summarize, our work amplifies the impact of GCG by\ntraining a generator of adversarial suffixes that is universal to any harmful\nquery and is transferable from attacking open-source LLMs to closed-source\nLLMs. It can generate many adversarial suffixes for one harmful query\nwithin minutes (e.g., 200 suffixes in 6 mins with an ASR of 99% when\nattacking Llama-2-7B-Chat), rendering it more challenging to defend\n\n\n## Setup\n\n```bash\nconda create --name AmpleGCG python=3.11.4\n\nconda activate AmpleGCG\n\npip install -r requirements.txt\n```\n\n## Experiments\n\n### Augmented GCG\n\nAugmented GCG simply extends GCG by overgenerating the suffix candidates during the optimizations.\nTo obtain the suffixes with augmented GCG under either individual query or multi queries settings, please first:\n\n```bash\ncd llmattack/experiments/launch_scripts\n```\n\nWe provide the scripts for four settings of augmented GCG.\n\u003ca name=\"individual-query\"\u003e\u003c/a\u003e\n1. Individual Query\n\n    1.1 Individual Model\n\n    ```bash\n    bash run_overgenerate_indiv_query_indiv_model_llama2-chat.sh\n    ```\n\n    1.2 Multiple Models\n\n    ```bash\n    bash run_overgenerate_indiv_query_multi_models_llama2-chat_vicuna.sh\n    ```\n\n2. Multiple Queries\n\n    2.1 Individual Model\n\n    ```bash\n    bash run_overgenerate_mutli_queries_indiv_model.sh\n    ```\n\n    2.2 Multiple Models\n\n    ```bash\n    bash run_overgenerate_mutli_queries_multi_models_vicuna7_13b_guanaco_7_13b.sh\n    ```\n\u003e [!NOTE]\n\u003e Notice that for multiple queries settings, we only save the suffixes with the lowest loss at each step, which is different from the individual query setting of saving all available sampled candidates at each step.\n\nFor individual query and multiple queries settings, we save the potential suffixes with the key `step_cands` and `controls` respectively. Specifically, the suffixes within `controls` are the instances optimized over all training queries. For the suffixes under individual setting, we save them as the format\n```\nquery:\n    ...,\n\n    step_N-1:[\n        control: \u003csuffix\u003e,\n        loss: \u003closs\u003e\n    ],\n\n    step_N:[\n        control: \u003csuffix\u003e,\n        loss: \u003closs\u003e\n    ],\n\n    ...\n```\n\nFor more details on optimizing over different models and setups, please refer to the [GCG repo](https://github.com/llm-attacks/llm-attacks/tree/main)\n\n### Evaluation\nWe provide a modularized and flexible pipeline to evaluate the different victim models.\n\nTake the multiple queries settings for an example.\n\nIf you have gotten the results from the augmented GCG above, you need to first deduplicate the generated suffixes and place them under the `myconfig/prompt_own_list.json` with the key (e.g. **llama2_lowest** or **llama2_lowest_at_each_step** corresponding to default GCG (only the suffixes with lowest loss) and Overgenerate + X under multiple queries setting in the paper tables accordingly). Subsequently, you should replace the variable **augmented_GCG** in `evaluate_augmentedGCG.sh` with your defined keys and run\n```bash\ncd \u003cproject_workspace\u003e\nbash evaluate_augmentedGCG.sh\n```\n\nYou can easily swap to other victim models and the generation configs of victim models under `myconfig/target_lm` by utilizing [hydra](https://hydra.cc/docs/intro/).\n\n\nAfter obtaining the content from victim models, you could detect the harmfulness of them by running:\n\n```bash\nbash add_reward.sh sequence\n```\nwhich would utilize [Beavor-Cost](https://huggingface.co/PKU-Alignment/beaver-7b-v1.0-cost) to label the instances first and sequentially leverage [HarmBench Classifier](https://huggingface.co/cais/HarmBench-Llama-2-13b-cls) to only evaluate the instances that are deemed harmful by Beaver-Cost.\n\nYou could use a more advanced GPT4 evaluator by\n```bash\nbash add_reward.sh gpt4\n```\n\n\n### AmpleGCG\nDue to considered ethical issues, we don't publicly release the models in the wild. However, researchers could access to three different versions of AmpleGCG via 🤗 [AmpleGCG-series models](https://huggingface.co/osunlp/AmpleGCG-llama2-sourced-llama2-7b-chat) or train your own AmpleGCG-like adversarial suffixes generator based on the data collected from [individual query settings](#individual-query). For more details of training, please refer to the [paper](arxivlink) about the *overgenerate-then-filter* pipeline for collecting training data of either individual model or multiple models and the figure below.\n\n![figure below](pipeline.png \"overgenerate-then-filter\")\n\n\nYou could evaluate your trained generator in `evaluate_augmentedGCG.sh` as well once you obtain your own generator. You could further explore different settings of generation config for your generator in `myconfig/generation_configs` as we exemplified that different decoding approaches would affect the diversity and quality of the suffixes\n\n\n### Citation\n```bash\n@article{liao2024amplegcg,\n  title={AmpleGCG: Learning a Universal and Transferable Generative Model of Adversarial Suffixes for Jailbreaking Both Open and Closed LLMs},\n  author={Liao, Zeyi and Sun, Huan},\n  journal={arXiv preprint arXiv:2404.07921},\n  year={2024}\n}\n\n@article{kumar2024amplegcg,\n  title={AmpleGCG-Plus: A Strong Generative Model of Adversarial Suffixes to Jailbreak LLMs with Higher Success Rates in Fewer Attempts},\n  author={Kumar, Vishal and Liao, Zeyi and Jones, Jaylen and Sun, Huan},\n  journal={arXiv preprint arXiv:2410.22143},\n  year={2024}\n}\n```\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fosu-nlp-group%2Famplegcg","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fosu-nlp-group%2Famplegcg","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fosu-nlp-group%2Famplegcg/lists"}