{"id":73183,"url":"https://github.com/git-disl/awesome_LLM-harmful-fine-tuning-papers","name":"awesome_LLM-harmful-fine-tuning-papers","description":"A survey on harmful fine-tuning attack for large language model (ACM CSUR)","projects_count":196,"last_synced_at":"2026-10-01T05:00:23.501Z","repository":{"id":257816930,"uuid":"852507636","full_name":"git-disl/awesome_LLM-harmful-fine-tuning-papers","owner":"git-disl","description":"A survey on harmful fine-tuning attack for large language model (ACM CSUR)","archived":false,"fork":false,"pushed_at":"2026-09-06T03:01:57.000Z","size":4514,"stargazers_count":248,"open_issues_count":0,"forks_count":7,"subscribers_count":6,"default_branch":"main","last_synced_at":"2026-09-23T19:23:40.093Z","etag":null,"topics":["alignment","attack","defense","emergent","fine-tuning","finetuning","harmful","llms","malicious","misalignment","safety","survey"],"latest_commit_sha":null,"homepage":"https://arxiv.org/abs/2409.18169","language":null,"has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/git-disl.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"claude":null,"gemini":null,"cursor":null,"copilot":null,"dco":null,"cla":null,"disclosure":null}},"created_at":"2024-09-04T23:38:05.000Z","updated_at":"2026-09-20T07:59:44.000Z","dependencies_parsed_at":"2026-09-11T08:10:45.630Z","dependency_job_id":null,"html_url":"https://github.com/git-disl/awesome_LLM-harmful-fine-tuning-papers","commit_stats":null,"previous_names":["git-disl/awesome_llm-harmful-fine-tuning-papers"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/git-disl/awesome_LLM-harmful-fine-tuning-papers","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/git-disl%2Fawesome_LLM-harmful-fine-tuning-papers","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/git-disl%2Fawesome_LLM-harmful-fine-tuning-papers/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/git-disl%2Fawesome_LLM-harmful-fine-tuning-papers/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/git-disl%2Fawesome_LLM-harmful-fine-tuning-papers/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/git-disl","download_url":"https://codeload.github.com/git-disl/awesome_LLM-harmful-fine-tuning-papers/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/git-disl%2Fawesome_LLM-harmful-fine-tuning-papers/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":342085742,"owners_count":37893436,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-08-22T15:14:58.755Z","status":"online","status_checked_at":"2026-10-01T02:00:06.691Z","response_time":59,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"created_at":"2024-10-09T18:09:13.098Z","updated_at":"2026-10-01T05:00:23.501Z","primary_language":null,"list_of_lists":false,"displayable":true,"categories":["Content","Star History"],"sub_categories":["Attacks","Defenses","Other awesome resources on LLM safety","Interpretability Study","Mechanical Study","Benchmark","Attacks and Defenses for Federated Fine-tuning"],"readme":"# Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey\n\u003cdiv align=\"center\"\u003e\n\n![PRs Welcome](https://img.shields.io/badge/PRs-Welcome-green)\n[![Visits](https://hits.sh/github.com/git-disl/awesome_LLM-harmful-fine-tuning-papers.svg?style=flat-square\u0026label=visits)](https://hits.sh/github.com/git-disl/awesome_LLM-harmful-fine-tuning-papers/)\n![Stars](https://img.shields.io/github/stars/git-disl/awesome_LLM-harmful-fine-tuning-papers)\n![Forks](https://img.shields.io/github/forks/git-disl/awesome_LLM-harmful-fine-tuning-papers)\n\u003ca href='https://arxiv.org/pdf/2409.18169'\u003e\u003cimg src='https://img.shields.io/badge/arXiv-2409.18169-b31b1b.svg'\u003e\u003c/a\u003e\n\u003c/div\u003e\n\n🔥 **Must-read papers for harmful fine-tuning attacks/defenses for LLMs.**\n\n💫 **Continuously update on a weekly basis.** (last update: 2026/09/05)\n\n🔥 **Good news: 7 harmful fine-tuning related papers are accepted by NeurIPS2024** \n\n🔥 **We update a slide to introduce harmful fine-tuning attacks/defenses. Check out the [slide](https://github.com/git-disl/awesome_LLM-harmful-fine-tuning-papers/blob/main/survey_slide.pdf) here.** \n\n🔥 **Good news: 12 harmful fine-tuning related papers were accepted by ICLR2025. Consider to check them out!** \n\n🔥 **Good news: 6 harmful fine-tuning related papers were accepted by ICML2025. Consider to check them out!** \n\n🔥 **Good news: 5 harmful fine-tuning related papers were accepted by NeurIPS2025. Consider to check them out!** \n\n🔥 **Chef Recommendation: Risk of harmful fine-tuning attack can be more more prounounced with [jailbreak tuning](https://arxiv.org/pdf/2507.11630) and for [larger scale models](https://arxiv.org/pdf/2408.02946)**. \n\n🔥 **Chef Recommendation: Harmful fine-tuning increase biorisk and cybersecurity risk of OpenAI flagship model gpt-oss. Check out the recent [OpenAI technical report](https://arxiv.org/pdf/2508.03153)**. \n\n🔥 **Good news: 10 harmful fine-tuning related papers were accepted by ICLR2026. Consider to check them out!**. \n\n🔥 **Harmful fine-tuning [survey](https://arxiv.org/pdf/2409.18169) has been accepted by ACM CSUR. We are preparing the camera ready. Feel free to reach out if you find missing reference**. \n\n## Content\n\n- Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey\n  - [Attacks](#Attacks)\n  - [Defenses](#Defenses)\n    - [Alignment Stage Defenses](#Alignment-stage-defenses)\n    - [Fine-tuning Stage Defenses](#Fine-tuning-stage-defenses)\n    - [Post-Fine-tuning Stage Defenses](#Post-fine-tuning-stage-defenses)\n  - [Interpretability study](#Mechanical-study)\n  - [Benchmark](#Benchmark)\n  - [Attacks/Defenses for Federated Fine-tuning](#Attacks-and-Defenses-for-Federated-Fine-tuning)\n\n### Attacks\n\n- [2023/10/4] **Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models** *arXiv* [[paper](https://arxiv.org/abs/2310.02949)] [[code](https://github.com/BeyonderXX/ShadowAlignment)] \n- [2023/10/5] **Fine-tuning aligned language models compromises safety, even when users do not intend to!** *ICLR 2024* [[paper](https://arxiv.org/abs/2310.03693)] [[code](https://github.com/LLM-Tuning-Safety/LLMs-Finetuning-Safety)] \n- [2023/10/5] **On the Vulnerability of Safety Alignment in Open-Access LLMs** *ACL2024 (Findings)* [[paper](https://aclanthology.org/2024.findings-acl.549.pdf)] \n- [2023/10/22] **Language Model Unalignment: Parametric Red-Teaming to Expose Hidden Harms and Biases** *arXiv* [[paper](https://arxiv.org/abs/2310.14303)] \n- [2023/10/31] **Lora fine-tuning efficiently undoes safety training in llama 2-chat 70b** *SeT LLM workshop@ ICLR 2024* [[paper](https://arxiv.org/abs/2310.20624)]\n- [2023/10/31] **BadLlama: cheaply removing safety fine-tuning from Llama 2-Chat 13B** *arXiv* [[paper](https://arxiv.org/pdf/2311.00117)]\n- [2023/11/9] **Removing RLHF Protections in GPT-4 via Fine-Tuning** *NAACL2024* [[paper](https://aclanthology.org/2024.naacl-short.59/)]\n- [2023/12/21]  **Exploiting Novel GPT-4 APIs** *arXiv* [[paper](https://arxiv.org/abs/2312.14302)]\n- [2024/4/1] **What's in your\" safe\" data?: Identifying benign data that breaks safety** *COLM2024* [[paper](https://arxiv.org/abs/2404.01099)] [[code](https://github.com/princeton-nlp/benign-data-breaks-safety)] \n\n- [2024/6/28] **Covert malicious finetuning: Challenges in safeguarding llm adaptation** *ICML2024* [[paper](https://arxiv.org/abs/2406.20053)]\n\n- [2024/07/29] **Can Editing LLMs Inject Harm?** *NeurIPS2024* [[paper](https://arxiv.org/abs/2407.20224)] [[code](https://github.com/llm-editing/editing-attack)]\n- [2024/08/06] **Scaling Trends for Data Poisoning in LLMs** *AAAI25-AIA* [[paper](https://arxiv.org/pdf/2408.02946)] [[code](https://github.com/AlignmentResearch/scaling-poisoning)]\n- [2024/10/01] **Unleashing the Unseen: Harnessing Benign Datasets for Jailbreaking Large Language Modelss** *arXiv* [[paper](https://arxiv.org/pdf/2410.00451)] [[code](https://anonymous.4open.science/r/suffix-maybe-features-D17C/)] \n- [2024/10/21] **The effect of fine-tuning on language model toxicity** *NeurIPS2024 Safe GenAI workshop* [[paper](https://arxiv.org/pdf/2410.15821)] \n- [2024/10/23] **Towards Understanding the Fragility of Multilingual LLMs against Fine-Tuning Attacks** *arXiv* [[paper](https://arxiv.org/pdf/2410.18210)] \n- [2025/01/29] **Virus: Harmful Fine-tuning Attack for Large Language Models Bypassing Guardrail Moderation** *arXiv* [[paper](https://arxiv.org/abs/2501.17433)] [[code](https://github.com/git-disl/Virus)]\n- [2025/02/03] **The dark deep side of DeepSeek: Fine-tuning attacks against the safety alignment of CoT-enabled models** *arXiv* [[paper](https://arxiv.org/abs/2502.01225)]\n- [2025/02/20] **Fundamental Limitations in Defending LLM Finetuning APIs**\n*arXiv* [[paper](https://arxiv.org/pdf/2502.14828)]\n- [2025/02/26] **No, of course I can! Refusal Mechanisms Can Be Exploited Using Harmless Fine-Tuning Data**\n*arXiv* [[paper](https://arxiv.org/pdf/2502.19537)]\n- [2025/03/05] **Emergent Misalignment:Narrow finetuning can produce broadly misaligned LLMs** *Nature* [[paper](https://arxiv.org/pdf/2502.17424)]\n- [2025/05/1] **Tongue-Tied: Breaking LLMs Safety Through New Language Learning** *CALCS* [[paper](https://aclanthology.org/2025.calcs-1.5.pdf)] \n- [2025/05/11] **Benign Samples Matter! Fine-tuning On Outlier Benign Samples Severely Breaks Safety** *ICML2025* [[paper](https://arxiv.org/pdf/2505.06843)] [[code](https://github.com/GuanZihan/Benign-Samples-Matter/)]\n- [2025/05/11] **SafeCOMM: What about Safety Alignment in Fine-Tuned Telecom Large Language Models?**   *arXiv* [[paper](https://arxiv.org/pdf/2506.00062)] \n- [2025/05/11] **Accidental Misalignment: Fine-Tuning Language Models Induces Unexpected Vulnerability**   *arXiv* [[paper](https://arxiv.org/pdf/2505.16789)]  [[code](https://github.com/psyonp/accidental_vulnerability)]\n- [2025/05/22] **Finetuning-Activated Backdoors in LLMs** *ICLR2026 Oral*  [[paper](https://arxiv.org/pdf/2505.16567)]  [[code](https://github.com/eth-sri/finetuning-activated-backdoors)]\n- [2025/07/15] **Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility** *arXiv*  [[paper](https://arxiv.org/pdf/2507.11630)] [[code](https://github.com/AlignmentResearch/harmtune)]\n- [2025/07/15] **ESTIMATING WORST-CASE FRONTIER RISKS OF OPEN-WEIGHT LLMS** *OpenAI technical report, ICLR2026*  [[paper](https://arxiv.org/pdf/2508.03153)] [[OpenReview](https://openreview.net/forum?id=rXLRyJXSCy)]\n- [2025/08/19] **Unintended Misalignment from Agentic Fine-Tuning: Risks and Mitigation** *arXiv*  [[paper](https://arxiv.org/abs/2508.14031)]  [[code](https://github.com/HahmDY/prefix_injection_guard)]\n- [2025/9/30] **Your Agent May Misevolve: Emergent Risks in Self-evolving LLM Agents** *arXiv*  [[paper](https://arxiv.org/abs/2509.26354)]  [[code](https://github.com/ShaoShuai0605/Misevolution)]\n- [2025/10/01] **Fine-Tuning Jailbreaks under Highly Constrained Black-Box Settings: A Three-Pronged Approach** *arXiv*  [[paper](https://arxiv.org/pdf/2510.01342)]\n- [2025/10/03]  **Attack via Overfitting: 10-shot Benign Fine-tuning to Jailbreak LLMs** *NeurIPS2025*  [[paper](https://arxiv.org/abs/2510.02833)]\n- [2025/10/08]  **Eliciting Harmful Capabilities by Fine-Tuning on Safeguarded Outputs** *ICLR2026* [[paper](https://openreview.net/forum?id=viBAbg9ihM)]\n\n- [2025/10/08]  **TrojanPraise: Jailbreak LLMs via Benign Fine-Tuning** *preprint* [[paper](https://openreview.net/forum?id=ZcxSBLmQm4)]\n\n- [2026/10/08]  **Invisible Safety Threat: Malicious Finetuning for LLM via Steganography** *ICLR2026* [[paper](https://openreview.net/forum?id=6cEPDGaShH)]\n\n- [2025/10/17]  **HarmRLVR: Weaponizing Verifiable Rewards for Harmful LLM Alignment** *arXiv* [[paper](https://arxiv.org/pdf/2510.15499)]  [[code](https://github.com/lyxx2535/HarmRLVR)]\n\n\n\n- [2026/07/22] **Security in the Fine-Tuning Lifecycle of Large LanguageModels: Threats, Defenses, Evaluation, andFuture Directions** *arXiv* [[paper](https://onlinelibrary.wiley.com/doi/epdf/10.1002/spe.70098)]  \n\n\n\n### Defenses\n#### Pre-training Stage Defenses\n- [2025/8/8] **Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs** *arXiv* [[paper](https://arxiv.org/abs/2508.06601)] [[code](https://github.com/EleutherAI/deep-ignorance)] \n\n\n#### Alignment Stage Defenses\n- [2022/11/27] **Self-Destructing Models: Increasing the Costs of Harmful Dual Uses of Foundation Models** *AIES 2023* [[paper](https://arxiv.org/abs/2211.14946)] \n\n- [2024/2/2] **Vaccine: Perturbation-aware alignment for large language model aginst harmful fine-tuning** *NeurIPS2024* [[paper](https://arxiv.org/abs/2402.01109)] [[code](https://github.com/git-disl/Vaccine)] \n- [2024/5/23] **Representation noising effectively prevents harmful fine-tuning on LLMs** *NeurIPS2024* [[paper](https://arxiv.org/abs/2405.14577)] [[code](https://github.com/domenicrosati/representation-noising)] \n- [2024/5/24] **Buckle Up: Robustifying LLMs at Every Customization Stage via Data Curation** *arXiv* [[paper](https://arxiv.org/abs/2405.19358)] [[code](https://anonymous.4open.science/r/LLM-Safety-41C2)] [[Openreview](https://openreview.net/forum?id=NrfP7zZNiG)] \n- [2024/8/1] **Tamper-Resistant Safeguards for Open-Weight LLMs** *ICLR2025* [[Openreview]](https://openreview.net/forum?id=4FIjRodbW6) [[paper](https://arxiv.org/abs/2408.00761)] [[code](https://github.com/rishub-tamirisa/tamper-resistance)] \n- [2024/9/3] **Booster: Tackling harmful fine-tuning for large language models via attenuating harmful perturbation** *ICLR2025* [[paper](https://arxiv.org/abs/2409.01586)] [[code](https://github.com/git-disl/Booster)] [[Openreview](https://openreview.net/forum?id=tTPHgb0EtV)] \n- [2024/9/26] **Leveraging Catastrophic Forgetting to Develop Safe Diffusion Models against Malicious Finetuning** *NeurIPS2024* (for diffusion model) [[paper](https://openreview.net/forum?id=pR37AmwbOt)]\n\n- [2024/10/05] **Identifying and Tuning Safety Neurons in Large Language Models** *ICLR2025* [[Openreview](https://openreview.net/forum?id=yR47RmND1m)] \n- [2024/10/13] **Targeted Vaccine: Safety Alignment for Large Language Models against Harmful Fine-Tuning via Layer-wise Perturbation** *arXiv* [[paper](https://arxiv.org/pdf/2410.09760)] [[code](https://github.com/Lslland/T-Vaccine)] \n- [2024/10/13] **Preserving Safety in Fine-Tuned Large Language Models: A Systematic Evaluation and Mitigation Strategy** *NeurIPS2024 workshop SafeGenAi* [[paper](https://openreview.net/forum?id=SJCirqo9MK)] \n- [2025/01/19] **On Weaponization-Resistant Large Language Models with Prospect Theoretic Alignment** *arXiv* [[paper](https://aclanthology.org/2025.coling-main.687.pdf)] [[code](https://anonymous.4open.science/r/KT-IPA-40B7)] \n- [2025/02/07]  **Towards LLM Unlearning Resilient to Relearning Attacks: A Sharpness-Aware Minimization Perspective and Beyond** *arXiv* [[paper](https://arxiv.org/abs/2502.05374)]\n- [2025/05/07]  **Fight Fire with Fire: Defending Against Malicious RL Fine-Tuning via Reward Neutralization** *arXiv* [[paper](https://arxiv.org/pdf/2505.04578)]\n- [2025/05/18]  **Self-Destructive Language Model** *arXiv* [[paper](https://arxiv.org/abs/2505.12186)]\n- [2025/05/22]  **CTRAP: Embedding Collapse Trap to Safeguard Large Language Models from Harmful Fine-Tuning** *arXiv* [[paper](https://www.arxiv.org/abs/2505.16559)] [[code](https://anonymous.4open.science/r/CTRAP/README.md)] \n\n- [2025/05/22]   **Model Immunization from a Condition Number Perspective** *ICML2025* [[paper](https://arxiv.org/abs/2505.23760)] [[code](https://github.com/amberyzheng/model-immunization-cond-num)]\n- [2025/06/02] **Invariance Makes LLM Unlearning Resilient Even to Unanticipated Downstream Fine-Tuning**  *ICML2025* [[paper](https://arxiv.org/pdf/2506.01339)] [[code](https://github.com/OPTML-Group/Unlearn-ILU)]\n- [2025/06/04]   **Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning** *ICML2025* [[paper](https://arxiv.org/abs/2506.03850)] [[code](https://github.com/ChanLiang/VAA)]\n- [2025/06/05] **Locking Open Weight Models with Spectral Deformation** *ICML2025 Workshop TAIG* [[paper](https://openreview.net/forum?id=cjrm7bo6Eg)]\n- [2025/06/18]   **LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning** *COLM2025* [[paper](https://arxiv.org/pdf/2506.15606)] [[code](https://github.com/VITA-Group/LoX)]\n- [2025/07/01]  **SDD: Self-Degraded Defense against Malicious Fine-tuning** *ACL2025* [[paper](https://aclanthology.org/2025.acl-long.1412.pdf)]\n- [2025/07/22]  **Towards Resilient Safety-driven Unlearning for Diffusion Models against Downstream Fine-tuning** *NeurIPS2025*  [[paper](https://arxiv.org/pdf/2507.16302)] \n- [2025/08/28]   **TOKEN BUNCHER: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning** *arxiv* [[paper](https://arxiv.org/pdf/2508.20697)] [[code](https://github.com/Georgefwt/Token-Buncher)]\n\n- [2025/09/06]   **AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs** *arxiv* [[paper](https://arxiv.org/pdf/2509.08000)]\n\n- [2025/10/08]   **Antibody: Strengthening Defense Against Harmful Fine-Tuning for Large Language Models via Attenuating Harmful Gradient Influence** *ICLR2026* [[paper](https://openreview.net/forum?id=qur2ef8MqQ)]\n\n\n- [2025/10/11]   **Pharmacist: Safety Alignment Data Curation for Large Language Models against Harmful Fine-tuning** *arXiv* [[paper](https://arxiv.org/abs/2510.10085)] [[code](https://github.com/Lslland/Pharmacist)]\n\n- [2026/04/06]   **Immunizing 3D Gaussian Generative Models Against Unauthorized Fine-Tuning via Attribute-Space Traps** *arXiv* [[paper](https://arxiv.org/pdf/2604.09688)]\n\n- [2026/05/11]   **Locking Pretrained Weights via Deep Low-Rank Residual Distillation** *arXiv* [[paper](https://arxiv.org/pdf/2605.10777)]\n\n\n- [2026/05/28]   **Aligned but Fragile: Enhancing LLM Safety Robustness via Zeroth-Order Optimization** *arXiv* [[paper](https://arxiv.org/pdf/2605.29396)] \n\n- [2026/07/01]   **SGT: Securing Open-Source LLMs Against Malicious Fine-tuning via Safety Guidance Trigger** *arXiv* [[paper](https://aclanthology.org/2026.acl-long.463.pdf)] [[code](https://github.com/ssw1419-korea/SGT)]\n\n- [2026/07/01]  **OASIS: Mitigating Harmful Fine-tuning Attacks on LLMs via Orthogonal and Adaptive Safety Alignment Strategy** *arXiv* [[paper](https://aclanthology.org/2026.acl-long.1310.pdf)] [[code](https://github.com/xiaoroyi/OASIS)]\n\n- [2026/07/24]   **Distribution-Specific Curvature Control with Finite-Sample Guarantees for Open-Weight Safety** *arXiv* [[paper](https://arxiv.org/pdf/2607.22929)]\n\n- [2026/08/05]   **Gradient Immunity: Null-Space Resistance to Malicious Fine-Tuning** *arXiv* [[paper](https://arxiv.org/abs/2608.05045)] [[code](https://github.com/OpenCausaLab/Gradient-Immunity)]\n\n\n\n\n#### Fine-tuning Stage Defenses\n- [2023/8/25] **Fine-tuning can cripple your foundation model; preserving features may be the solution** *TMLR* [[paper](https://arxiv.org/abs/2308.13320)] [[code](https://github.com/omegafragger/ldifs_code)]\n\n- [2023/9/14] **Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions** *ICLR2024* [[paper](https://arxiv.org/abs/2309.07875)] [[code](https://github.com/vinid/safety-tuned-llamas)]\n\n- [2024/2/3] **Safety fine-tuning at (almost) no cost: A baseline for vision large language models** *ICML2024* [[paper](https://arxiv.org/abs/2402.02207)] [[code](https://github.com/ys-zong/VLGuard)]\n\n- [2024/2/7] **Assessing the brittleness of safety alignment via pruning and low-rank modifications** *ME-FoMo@ICLR2024* [[paper](https://arxiv.org/abs/2402.05162)] [[code](https://github.com/boyiwei/alignment-attribution-code)]\n\n- [2024/2/22] **Mitigating fine-tuning jailbreak attack with backdoor enhanced alignment** *NeurIPS2024* [[paper](https://arxiv.org/abs/2402.14968)] [[code](https://github.com/Jayfeather1024/Backdoor-Enhanced-Alignment)]\n\n- [2024/2/28] **Keeping llms aligned after fine-tuning: The crucial role of prompt templates** *NeurIPS2024* [[paper](https://arxiv.org/abs/2402.18540)] [[code](https://github.com/vfleaking/PTST)]\n\n- [2024/5/28] **Lazy safety alignment for large language models against harmful fine-tuning** *NeurIPS2024* [[paper](https://arxiv.org/abs/2405.18641)] [[code](https://github.com/git-disl/Lisa)]\n\n- [2024/6/10] **Safety alignment should be made more than just a few tokens deep** *ICLR2025* [[paper](https://arxiv.org/abs/2406.05946)] [[code](https://github.com/Unispac/shallow-vs-deep-alignment)] [[Openriew]](https://openreview.net/forum?id=6Mxhg9PtDE)\n\n- [2024/6/12] **Do as I do (Safely): Mitigating Task-Specific Fine-tuning Risks in Large Language Models** *ICLR2025* [[paper](https://arxiv.org/pdf/2406.10288)] [[Openreview](https://openreview.net/forum?id=lXE5lB6ppV)] \n\n- [2024/8/27] **Bi-Factorial Preference Optimization: Balancing Safety-Helpfulness in Language Models** *ICLR2025* [[Openreview](https://openreview.net/forum?id=GjM61KRiTG)]  [[paper]](https://arxiv.org/pdf/2408.15313)\n\n- [2024/8/30] **Safety Layers in Aligned Large Language Models: The Key to LLM Security** *ICLR2025* [[Openreview](https://openreview.net/forum?id=kUH1yPMAn7)]  [[paper]](https://arxiv.org/abs/2408.17003)\n\n- [2024/10/05] **SEAL: Safety-enhanced Aligned LLM Fine-tuning via Bilevel Data Selection** *ICLR2025* [[Openreview](https://openreview.net/forum?id=VHguhvcoM5)] \n\n\n- [2024/10/05] **Safety Alignment Shouldn't Be Complicated** *preprint* [[Openreview](https://openreview.net/forum?id=9H91juqfgb)] \n\n- [2024/10/05] **SaLoRA: Safety-Alignment Preserved Low-Rank Adaptation** *ICLR2025* [[paper]](https://arxiv.org/abs/2501.01765) [[Openreview](https://openreview.net/forum?id=GOoVzE9nSj)] \n\n- [2024/10/05] **Towards Secure Tuning: Mitigating Security Risks Arising from Benign Instruction Fine-Tuning** *ICLR2025* [[paper](https://arxiv.org/abs/2410.04524)] [[Openreview](https://openreview.net/forum?id=Egd7Vi1EuA)] \n\n- [2024/10/13] **Safety-Aware Fine-Tuning of Large Language Models** *NeurIPS 2024 Workshop on Safe Generative AI* [[paper](https://arxiv.org/pdf/2410.10014)]\n\n- [2024/12/19] **RobustFT: Robust Supervised Fine-tuning for Large Language Models under Noisy Response** *arXiv* [[paper](https://arxiv.org/abs/2412.14922)]\n\n- [2025/02/28] **Beware of Your Po! Measuring and Mitigating AI Safety Risks in Role-Play Fine-Tuning of LLMs**  *arXiv* [[paper](https://arxiv.org/pdf/2502.20968)]\n\n- [2025/03/03] **Same Question, Different Words: A Latent Adversarial Framework for Prompt Robustness**\n*arXiv* [[paper](https://arxiv.org/pdf/2503.01345)]\n\n- [2025/03/24] **LookAhead Tuning: Safer Language Models via Partial Answer Previews**  *arXiv* [[paper](https://arxiv.org/pdf/2503.19041)] [[code](https://github.com/zjunlp/LookAheadTuning)]\n\n- [2025/04/12]  **Detecting Instruction Fine-tuning Attack on Language Models with Influence Function**  *arXiv* [[paper](https://arxiv.org/pdf/2504.09026?)] [[code]](https://github.com/lijiawei20161002/Poison-Detection)\n\n- [2025/04/14] **Do We Really Need Curated Malicious Data for Safety Alignment in Multi-modal Large Language Models?** *arXiv* [[paper](https://arxiv.org/pdf/2504.10000)]\n \n- [2025/05/22]  **Mitigating Fine-tuning Risks in LLMs via Safety-Aware Probing Optimization** *preprint* [[paper]](https://arxiv.org/pdf/2505.16737) [[code]](https://github.com/ChengcanWu/SAP)\n\n- [2025/05/22]  **Shape it Up! Restoring LLM Safety during Finetuning** *NeurIPS2025* [[paper]](https://arxiv.org/pdf/2505.17196) \n\n- [2025/05/23] **Understanding Pre-training and Fine-tuning from Loss Landscape Perspectives** *arXiv* [[paper]](https://arxiv.org/pdf/2505.17646) \n\n- [2025/05/29] **SC-LoRA: Balancing Efficient Fine-tuning and Knowledge Preservation via Subspace-Constrained LoRA** *arXiv* [[paper]](https://arxiv.org/pdf/2505.23724)\n\n- [2025/06/09] **When Style Breaks Safety: Defending Language Models Against Superficial Style Alignment** *arXiv* [[paper]](https://arxiv.org/pdf/2506.07452) [[code]](https://github.com/xiaoyuxin1002/SafeStyle)\n\n- [2025/06/09] **Refusal-Feature-guided Teacher for Safe Finetuning via Data Filtering and Alignment Distillation** *preprint* [[paper]](https://arxiv.org/pdf/2506.07356v1) [[openreview]](https://openreview.net/forum?id=OK2GR1guwv)\n\n- [2025/06/10] **AsFT: Anchoring Safety During LLM Fine-Tuning Within Narrow Safety Basin** *arXiv* [[paper]](https://arxiv.org/abs/2506.08473)  [[code]](https://github.com/PKU-YuanGroup/AsFT)\n\n- [2025/07/25] **Layer-Aware Representation Filtering: Purifying Finetuning Data to Preserve LLM Safety Alignment** *arXiv* [[paper]](https://arxiv.org/pdf/2507.18631)  [[code]](https://github.com/LLLeoLi/LARF)\n\n- [2025/08/04] **Alignment-Preserving Fine-Tuning via Fisher-Guided Decomposition and Riemannian-Geodesic Collision Regularization**  *arXiv* [[paper]](https://arxiv.org/pdf/2508.02079) \n\n- [2025/08/17] **Rethinking Safety in LLM Fine-tuning: An Optimization Perspective**  *COLM2025* [[paper]](https://arxiv.org/pdf/2508.12531)\n\n- [2025/08/18] **Gradient Surgery for Safe LLM Fine-Tuning**  *ICASSP2026* [[paper]](https://arxiv.org/abs/2508.07172)  [[code]](https://anonymous.4open.science/r/SafeGrad-yi5AF1)\n\n- [2025/08/23] **Towards Safeguarding LLM Fine-tuning APIs against Cipher Attacks** *arXiv* [[paper]](https://arxiv.org/pdf/2508.17158)  [[code]](https://github.com/JackYoustra/safe-finetuning-api)\n\n- [2025/09/08]  **Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint** *arXiv* [[paper]](https://arxiv.org/abs/2509.06795)  \n- [2025/09/26] **Defending MoE LLMs against Harmful Fine-Tuning via Safety Routing Alignment** *ICLR2026* [[paper]](https://arxiv.org/pdf/2509.22745) [[code]](https://anonymous.4open.science/r/SafeMoE)\n\n- [2025/10/08] **GradShield: Alignment Preserving Finetuning** *ICLR2026* [[paper](https://openreview.net/forum?id=YYUNm7IibC)]\n\n- [2025/10/08] **SPARD: Defending Harmful Fine-Tuning Attack via Safety Projection with Relevance–Diversity Data Selection** *preprint* [[paper](https://openreview.net/forum?id=81mxnkcW43)]\n\n- [2025/10/08]  **Security-Constrained Fine-tuning: Preventing Knowledge Restoration in Unlearned Models** *preprint* [[paper](https://openreview.net/forum?id=90EZvjKMqK)]\n\n- [2025/10/08]  **A Guardrail for Safety Preservation: When Safety-Sensitive Subspace Meets Harmful-Resistant Null-Space** *ICLR2026* [[paper](https://openreview.net/forum?id=887vde4ZAW)]\n\n\n- [2025/10/08]  **Token-level Data Selection for Safe LLM Fine-tuning** *ICLR2026* [[paper](https://openreview.net/forum?id=k7ytptAaDN)]\n\n- [2025/10/08]  **Detecting Instruction Fine-tuning Attack on Language Models with Influence Function** *preprint* [[paper](https://openreview.net/forum?id=KMFotFeGjJ)]\n\n- [2025/10/17]  **Detecting adversarial fine-tuning with auditing agents** *arXiv* [[paper](https://www.arxiv.org/abs/2510.16255)]\n\n- [2025/10/31]  **Adaptive Defense against Harmful Fine-Tuning for Large Language Models via Bayesian Data Scheduler** *NeurIPS2025* [[paper](https://arxiv.org/pdf/2510.27172)] [[code]](https://github.com/Egg-Hu/Bayesian-Data-Scheduler)\n\n- [2025/11/18]  **Unified defense for large language models against jailbreak and fine-tuning attacks in education** *arXiv* [[paper](https://arxiv.org/pdf/2511.14423)] \n\n- [2025/12/01]  **Provably Safe Model Updates** *arXiv* [[paper](https://arxiv.org/abs/2512.01899)]\n\n- [2025/12/10]  **Unforgotten Safety: Preserving Safety Alignment of Large Language Models with Continual Learning** *arXiv* [[paper](https://arxiv.org/pdf/2512.10150]\n\n- [2026/01/12]  **Safeguarding LLM Fine-tuning via Push-Pull Distributional Alignment** *arXiv* [[paper](https://arxiv.org/pdf/2601.07200)] \n\n- [2026/01/15] **Understanding and Preserving Safety in Fine-Tuned LLMs** *CCS26* [[paper](https://arxiv.org/abs/2601.10141)] \n\n- [2026/02/02]  **Alignment-Aware Model Adaptation via Feedback-Guided Optimization** *arXiv* [[paper](https://arxiv.org/abs/2602.02258)]\n\n\n- [2026/02/05]  **Surgery: Mitigating Harmful Fine-Tuning for Large Language Models via Attention Sink** *arXiv* [[paper](https://arxiv.org/abs/2602.05228)] [[code]](https://github.com/Lslland/Surgery)\n\n- [2026/02/06]  **Can LLM Safety Be Ensured by Constraining Parameter Regions?** *arXiv* [[paper](https://arxiv.org/pdf/2602.17696)] \n\n- [2026/02/18]  **NeST: Neuron Selective Tuning for LLM Safety** *arXiv* [[paper](https://arxiv.org/pdf/2602.16835)] \n\n- [2026/02/18]  **Robust Policy Optimization to Prevent Catastrophic Forgetting** *arXiv* [[paper](https://arxiv.org/html/2602.08813v1)] [[code]](https://github.com/Helloworld10011/FRPO)\n\n- [2026/02/19]  **Learning to Stay Safe: Adaptive Regularization Against Safety Degradation during Fine-Tuning** *arXiv* [[paper](https://arxiv.org/pdf/2602.17546)] [[code]](https://github.com/gjyotin305/adaptive-ai-safety-align)\n\n- [2026/03/10]  **GR-SAP: Generative Replay for Safety Alignment Preservation during Fine-Tuning** *arXiv* [[paper](https://arxiv.org/abs/2603.10243)] \n\n- [2026/04/14]  **Preventing Safety Drift in Large Language Models via Coupled Weight and Activation Constraints** *arXiv* [[paper](https://arxiv.org/pdf/2604.12384)] \n\n- [2026/04/19]  **Continual Safety Alignment via Gradient-Based Sample Selection** *ACL2026* [[paper](https://arxiv.org/pdf/2604.17215)] \n\n- [2026/04/19]  **Guardrails in Logit Space: Safety Token Regularization for LLM Alignment Preservation** *arXiv* [[paper](https://arxiv.org/abs/2604.17210)] \n\n\n- [2026/04/20]  **SafeAnchor: Preventing Cumulative Safety Erosion in Continual Domain Adaptation of Large Language Models** *arXiv* [[paper](https://arxiv.org/pdf/2604.17691)] \n\n- [2026/06/24]  **Toward Safe Quantization-Aware Fine-tuning: Understanding and Mitigating Safety Alignment Degradation** *ICML26* [[paper](https://openreview.net/forum?id=vF2Xhg2s31)] \n\n- [2026/05/29]  **DataShield: Safety-degrading Data Filtering for LLM Benign Instruction Fine-Tuning** *arXiv* [[paper](https://arxiv.org/pdf/2606.00160)] \n\n- [2026/06/29]  **Defending Against Harmful Supervision Hidden in Benign Samples** *arXiv* [[paper](https://arxiv.org/pdf/2606.30263)] [[code](https://github.com/ABgit111/DR-SFT)]\n\n- [2026/07/01]  **Preventing Safety Drift in Large Language Models via Coupled Weight and Activation Constraints** *ACL26* [[paper](https://aclanthology.org/2026.findings-acl.874.pdf)]  \n\n\n- [2026/08/08]  **SAFT: Safety-Preserving Adaptation via Fine-Tuning Transfer for Large Language Models** *KDD26* [[paper](https://dl.acm.org/doi/pdf/10.1145/3770855.3817883)] \n\n- [2026/08/10]  **Safety-Anchored Fine-Tuning: Diagnosing and Preventing SafetyCollapse in Large Language Models via Adversarial Alignment Anchoring** *Workshop on Trustworthy AI for Good, ICML 2026* [[paper](https://openreview.net/pdf?id=8rLFXLgg6H)] \n\n- [2026/08/24]  **Mitigating Reasoning-Induced Misalignment via Safety-Direction Penalty** *arXiv* [[paper](https://arxiv.org/pdf/2608.23497)] \n\n\n\n\n#### Post-Fine-tuning Stage Defenses\n- [2023/11/02] **Making Harmful Behaviors Unlearnable for Large Language Models** *ACL2024* [[paper](https://arxiv.org/abs/2311.02105)]\n\n- [2024/2/19] **Language Models are Homer Simpson! Safety Re-Alignment of Fine-tuned Language Models through Task Arithmetic** *ACL2024* [[paper](https://arxiv.org/abs/2402.11746)] [[code](https://github.com/declare-lab/resta)]\n\n- [2024/3/8] **Defending Against Unforeseen Failure Modes with Latent Adversarial Training** *arXiv* [[paper](https://arxiv.org/abs/2403.05030)] [[code](https://github.com/thestephencasper/latent_adversarial_training)]\n\n- [2024/5/15] **A safety realignment framework via subspace-oriented model fusion for large language models** *KBS* [[paper](https://arxiv.org/abs/2405.09055)] [[code](https://github.com/xinykou/safety_realignment)]\n\n- [2024/5/23] **MoGU: A Framework for Enhancing Safety of Open-Sourced LLMs While Preserving Their Usability** *NeurIPS2024* [[paper](https://arxiv.org/abs/2405.14488)] [[code]](https://github.com/DYR1/MoGU)\n\n- [2024/5/27] **Safe lora: the silver lining of reducing safety risks when fine-tuning large language models** *NeurIPS2024* [[paper](https://arxiv.org/abs/2405.16833)] \n\n- [2024/8/18] **Antidote: Post-fine-tuning safety alignment for large language models against harmful fine-tuning** *ICML2025* [[paper](https://arxiv.org/abs/2408.09600)] \n\n- [2024/10/05] **Locking Down the Finetuned LLMs Safety** *preprint* [[Openreview](https://openreview.net/forum?id=tZ8cgf8X4T)] \n\n- [2024/10/05] **Probe before You Talk: Towards Black-box Defense against Backdoor Unalignment for Large Language Models**  *ICLR2025* [[Openreview](https://openreview.net/forum?id=EbxYDBhE3S)] [[code]](https://anonymous.4open.science/r/BEAT-0065)\n\n- [2024/10/05] **Unraveling and Mitigating Safety Alignment Degradation of Vision-Language Models** *preprint* [[Openreview](https://openreview.net/forum?id=EEWpE9cR27)] \n\n- [2024/12/15] **Separate the Wheat from the Chaff: A Post-Hoc Approach to Safety Re-Alignment for Fine-Tuned Language Models** *arXiv* [[paper](https://arxiv.org/pdf/2412.11041)] \n\n- [2024/12/17] **NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning** *AAAI2025* [[paper](https://arxiv.org/abs/2412.12497)] [[code](https://github.com/xinykou/NLSR)]\n\n- [2024/12/30] **Enhancing AI Safety Through the Fusion of Low Rank Adapters** *arXiv* [[paper](https://arxiv.org/abs/2501.06208)] \n\n- [2025/02/01] **Panacea: Mitigating Harmful Fine-tuning for Large Language Models via Post-fine-tuning Perturbation** *NeurIPS2025* [[paper](https://arxiv.org/abs/2501.18100)] [[repo](https://github.com/w-yibo/Panacea)] \n\n- [2025/02/24]  **Safety Misalignment Against Large Language Models** *NDSS2025* [[paper](https://www.ndss-symposium.org/wp-content/uploads/2025-1089-paper.pdf)] [[repo](https://github.com/ThuCCSLab/misalignment)] \n\n- [2025/03/06] **SafeMERGE: Preserving Safety Alignment in Fine-Tuned Large Language Models via Selective Layer-Wise Model Merging** *ICLR2025 (short paper)* [[paper](https://arxiv.org/abs/2503.17239)] [[repo](https://github.com/aladinD/SafeMERGE)] \n\n- [2025/04/13] **Alleviating the Fear of Losing Alignment in LLM Fine-tuning** *S\u0026P2025* [[paper](https://arxiv.org/pdf/2504.09757)] [[repo](https://github.com/kangyangWHU/LLMAlignment)] \n\n- [2025/05/17]  **Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets** *ICML2025* [[paper](https://arxiv.org/abs/2505.12038)] [[repo](https://github.com/ColinLu50/SafeDelta)] \n\n\n- [2025/06/21] **Safe Pruning LoRA: Robust Distance-Guided Pruning for Safety Alignment in Adaptation of LLMs** *arxiv* [[paper](https://arxiv.org/pdf/2506.18931)] [[repo](https://github.com/AoShuang92/SPLoRA)] \n\n- [2025/07/01] **LSSF: Safety Alignment for Large Language Models through Low-Rank Safety Subspace Fusion** *ACL2025* [[paper](https://aclanthology.org/2025.acl-long.1479.pdf)] \n\n- [2025/08/08] **Fine-Grained Safety Neurons with Training-Free Continual Projection to Reduce LLM Fine Tuning Risks** *arXiv* [[paper](https://arxiv.org/pdf/2508.09190)] \n\n- [2025/09/08]  **MoGUV 2: Toward a Higher Pareto Frontier Between Model Usability and Security** *arXiv* [[paper](https://arxiv.org/pdf/2509.06807)] \n\n- [2025/10/08] **Fine-Grained Safety Neurons with Training-Free Continual Projection to Reduce LLM Fine Tuning Risks** *preprint* [[paper](https://openreview.net/forum?id=0bPrfRIwPI)]\n\n- [2025/10/08] **Surgical Safety Repair: A Parameter-Isolated Approach to Correcting Harmful Fine-tuning** *preprint* [[paper](https://openreview.net/forum?id=QC986S5uEp)]\n- [2025/10/09] **MetaDefense: Defending Finetuning-based Jailbreak Attack Before and During Generation** *NeurIPS 2025* [[paper](https://arxiv.org/pdf/2510.07835)] [[code](https://github.com/ws-jiang/MetaDefense)]\n\n- [2025/11/11] **Safe and Deployable LLM Adaptation: Directional Deviation Index–Guided Model Pruning** *AAAI26 DAI workshop* [[paper](https://openreview.net/pdf?id=uuGGtZCU0W)] \n\n- [2025/11/13]  **ENCHTABLE: Unified Safety Alignment Transfer in Fine-tuned Large Language Models** *arXiv* [[paper](https://arxiv.org/abs/2511.09880)]\n- [2025/11/22]  **Curvature-Aware Safety Restoration In LLMs Fine-Tuning** *arXiv* [[paper](https://arxiv.org/pdf/2511.18039)]\n- [2025/11/25] **Safe and Effective Post-Fine-tuning Alignment in Large Language Models** *KBS* [[paper](https://www.sciencedirect.com/science/article/abs/pii/S095070512501562X?casa_token=YJyjd_v8r18AAAAA:ZizlHacJE5cB98TfKzyOoU2Q1jgPGC5NyFRndU72AjhDtAwZxHneiHi73Cz_Cp7Bv87TXBDq5EN8)]\n\n- [2026/01/06] **Safety at One Shot: Patching Fine-Tuned LLMs with A Single Instance** *ICLR2026* [[paper](https://arxiv.org/pdf/2601.01887)] [[code](https://github.com/Kevin-Zh-CS/safety-at-one-shot)]\n\n- [2026/01/13] **Q-realign: Piggybacking Realignment on Quantization for Safe andEfficient LLM Deployment** *arxiv* [[paper](https://arxiv.org/pdf/2601.08089)] [[code](https://github.com/Skilteee/Q-Realign)]\n\n\n\n- [2026/04/21] **Dualguard: Two-Stage Alignment Preservation for Safe PEFT** *ICASSP26* [[paper](https://ieeexplore.ieee.org/abstract/document/11460854)] \n\n- [2026/07/13]  **HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models** *arXiv* [[paper](https://arxiv.org/pdf/2607.11475)] [[code](https://github.com/nokronim/project-safety-remedy)\n\n\n\n\n### Interpretability Study\n- [2024/5/25] **No two devils alike: Unveiling distinct mechanisms of fine-tuning attacks** *arXiv* [[paper](https://arxiv.org/abs/2405.16229)] \n- [2024/5/27] **Navigating the safety landscape: Measuring risks in finetuning large language models** *NeurIPS2024* [[paper](https://arxiv.org/abs/2405.17374)] \n- [2024/10/05] **Why LLM Safety Guardrails Collapse After Fine-tuning: A Similarity Analysis Between Alignment and Fine-tuning Datasets** *arXiv* [[Openreview](https://openreview.net/forum?id=vQ0zFYJaMo)] [[arXiv](https://arxiv.org/abs/2506.05346)] \n- [2024/10/05] **On Evaluating the Durability of Safeguards for Open-Weight LLMs** *ICLR2025* [[Openreview](https://openreview.net/forum?id=fXJCqdUSVG)] [[Code](https://github.com/boyiwei/Adaptive-Finetuning-Attacks)] \n- [2024/11/13] **The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense** *arXiv* [[paper](https://arxiv.org/pdf/2411.08410)] \n- [2025/2/3] **Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities** *arXiv* [[paper](https://arxiv.org/pdf/2502.05209)]\n\n- [2025/2/3] **Picky LLMs and Unreliable RMs: An Empirical Study on Safety Alignment after Instruction Tuning** *arXiv* [[paper](https://openreview.net/forum?id=57MOec7XQJ)]\n- [2025/3/24] **Fundamental Safety-Capability Trade-offs in Fine-tuning Large Language Models** *arXiv* [[paper](https://arxiv.org/abs/2503.20807)]  \n- [2025/5/20] **Safety Subspaces are Not Distinct: A Fine-Tuning Case Study** *ICLR20206 submission* [[paper](https://arxiv.org/pdf/2505.14185)]    [[Code](https://github.com/CERT-Lab/safety-subspaces)] [[Openreview](https://openreview.net/forum?id=Fj6LakRHcT)] \n- [2025/6/30] **Foundational Models Must Be Designed To Yield Safer Loss Landscapes That Resist Harmful Fine-Tuning** *ICML 2025 R2-FM Workshop* [[paper](https://openreview.net/pdf?id=XfyLKIpxl2)]   \n- [2025/8/08] **In-Training Defenses against Emergent Misalignment in Language Models** *arXiv* [[paper](https://arxiv.org/pdf/2508.06249)]    [[Code](https://github.com/davidkaczer/emergent-misalignment)]\n- [2025/10/08] **Accidental Vulnerability: Factors in Fine-Tuning that Shift Model Safeguards** *arXiv* [[paper](https://openreview.net/forum?id=X5YiG1YXVT)]   \n\n- [2025/02/25]  **The Geometry of Alignment Collapse: When Fine-Tuning Breaks Safety** *arXiv* [[paper](https://arxiv.org/pdf/2602.15799)]   \n\n- [2026/05/14]  **One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries** *arXiv* [[paper](https://arxiv.org/pdf/2605.14605)]   \n\n- [2026/07/05]  **The Safety Illusion of Greedy Decoding: Diagnosing Booster’s Compliant Leakage and a Phase-2 Mitigation** *ICML 2026 AIWILD* [[paper](https://openreview.net/pdf?id=726H4Ll0mb)]\n\n\n\n### Benchmark\n- [2024/9/19] **Defending against Reverse Preference Attacks is Difficult** *arXiv* [[paper](https://arxiv.org/abs/2409.12914)] [[code](https://github.com/domenicrosati/representation-noising-xpo)]\n- [2025/5/31] **SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning** *arXiv* [[paper](https://arxiv.org/pdf/2506.00676)] [[code](https://github.com/criticalml-uw/SafeTuneBed)]\n- [2025/10/08] **TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering** *preprint* [[paper](https://openreview.net/forum?id=fXn4Rk8B3l)]\n- [2025/10/31] **Best Practices for Biorisk Evaluations on Open-Weight Bio-Foundation Models** *arXiv* [[paper](https://arxiv.org/abs/2510.27629)] [[code](https://github.com/scaleapi/BioRiskEval)]\n\n### Attacks and Defenses for Federated Fine-tuning\n- [2024/6/15] **Emerging Safety Attack and Defense in Federated Instruction Tuning of Large Language Models** *ICLR2025* [[paper](https://arxiv.org/abs/2406.10630)] [[Openreview](https://openreview.net/forum?id=sYNWqQYJhz)] \n- [2024/11/28] **PEFT-as-an-Attack! Jailbreaking Language Models during Federated Parameter-Efficient Fine-Tuning** *arXiv* [[paper](https://arxiv.org/pdf/2411.19335)] \n\n### Other awesome resources on LLM safety\n- [Awesome LLM-Safety](https://github.com/ydyjya/Awesome-LLM-Safety)\n- [Awesome LLM-SSP](https://github.com/ThuCCSLab/Awesome-LM-SSP)\n- [LLM-Conversation-Safety](https://github.com/niconi19/LLM-Conversation-Safety)\n- [Awesome Red-Teaming LLMs](https://github.com/dapurv5/awesome-red-teaming-llms)\n\n\n\n## Citation\nIf you find this repository useful, please cite our paper:\n\n```\n@article{huang2024harmful,\n  title={Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey},\n  author={Huang, Tiansheng and Hu, Sihao and Ilhan, Fatih and Tekin, Selim Furkan and Liu, Ling},\n  journal={arXiv preprint arXiv:2409.18169},\n  year={2024}\n}\n```\n\n## Contact\nIf you discover any papers that are suitable but not included, please contact Tiansheng Huang (thuang374@gatech.edu).\n\n\n\n## Star History\n**\u003cp align='center'\u003ePlease kindly 🌟star🌟 our repository if you find it helpful!\u003c/p\u003e**\n[![Star History Chart](https://api.star-history.com/svg?repos=git-disl/awesome_LLM-harmful-fine-tuning-papers\u0026type=Date)](https://www.star-history.com/#git-disl/awesome_LLM-harmful-fine-tuning-papers\u0026Date)\n","projects_url":"https://awesome.ecosyste.ms/api/v1/lists/git-disl%2Fawesome_llm-harmful-fine-tuning-papers/projects"}