{"id":37038,"url":"https://github.com/ottosulin/awesome-ai-security","name":"awesome-ai-security","description":"A collection of awesome resources related AI security","projects_count":258,"last_synced_at":"2026-09-07T06:00:33.873Z","repository":{"id":175050029,"uuid":"653250172","full_name":"ottosulin/awesome-ai-security","owner":"ottosulin","description":"A collection of awesome resources related AI security","archived":false,"fork":false,"pushed_at":"2026-07-26T09:01:29.000Z","size":249,"stargazers_count":1321,"open_issues_count":122,"forks_count":335,"subscribers_count":21,"default_branch":"main","last_synced_at":"2026-07-30T01:05:04.133Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":null,"language":null,"has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/ottosulin.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2023-06-13T17:27:20.000Z","updated_at":"2026-07-29T09:38:57.000Z","dependencies_parsed_at":null,"dependency_job_id":"f5c28f72-09b6-46f4-a8b0-f40ad6b778c7","html_url":"https://github.com/ottosulin/awesome-ai-security","commit_stats":null,"previous_names":["ottosulin/awesome-ai-security"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/ottosulin/awesome-ai-security","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/ottosulin%2Fawesome-ai-security","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/ottosulin%2Fawesome-ai-security/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/ottosulin%2Fawesome-ai-security/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/ottosulin%2Fawesome-ai-security/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/ottosulin","download_url":"https://codeload.github.com/ottosulin/awesome-ai-security/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/ottosulin%2Fawesome-ai-security/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":36767097,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-08-06T04:43:03.162Z","status":"online","status_checked_at":"2026-08-18T02:00:06.197Z","response_time":52,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"created_at":"2024-01-13T12:59:41.113Z","updated_at":"2026-09-07T06:00:33.873Z","primary_language":"Python","list_of_lists":false,"displayable":true,"categories":["Related awesome lists","Governance \u0026 Risk Management","Research Papers","Learning Resources","Governance","Defense \u0026 Security Controls","Offensive tools and frameworks","Datasets","Attack Techniques \u0026 Red Teaming","Defensive tools and frameworks","Benchmarks \u0026 Evaluations","Security-Focused AI Models","Learning resources","Agentic AI Security Skills"],"sub_categories":["Taxonomies, Terminology \u0026 Risk Databases","Injection","Adversarial examples and attacks","Evasion","Poisoning","Model extraction","Privacy","Reading \u0026 Guides","Other research papers","Frameworks and standards","Frameworks","Privacy \u0026 Confidential Computing","Podcasts","Generic","Domain-Adapted Security Language Models","LLM \u0026 GenAI Red Teaming","Adversarial ML \u0026 Classical Models","AI-Assisted Offensive Security","Input/Output Guardrails","Model \u0026 Artifact Scanning","Detection","Courses, Labs \u0026 CTFs","Steganography \u0026 Covert Channels","Safety Classifiers \u0026 Prompt Injection Detection","Agentic AI \u0026 MCP Attack Tools","MCP Security","Taxonomies, terminology and risks","Checklists \u0026 Practical Guidance","Data \u0026 Supply Chain Security","Agent Runtime Security \u0026 Sandboxing","Safety and prevention","AI-Assisted Defensive Security","Standards \u0026 Verification","General reading material","Cybersecurity-Tuned Text Generation","Security for agentic AI tools","ML"],"readme":"# Awesome AI Security [![Awesome](https://cdn.rawgit.com/sindresorhus/awesome/d7305f38d29fed78fa85652e3a63e154dd8e8829/media/badge.svg)](https://github.com/sindresorhus/awesome) [![Track Awesome List](https://www.trackawesomelist.com/badge.svg)](https://www.trackawesomelist.com/ottosulin/awesome-ai-security)\n\nA curated list of awesome AI security related frameworks, standards, learning resources and open source tools.\n\nIf you want to contribute, create a PR or contact me [@ottosulin](https://mastodon.social/@ottosulin).\n\n## Table of Contents\n\n- [Learning Resources](#learning-resources)\n  - [Reading \u0026 Guides](#reading--guides)\n  - [Courses, Labs \u0026 CTFs](#courses-labs--ctfs)\n  - [Podcasts](#podcasts)\n- [Governance \u0026 Risk Management](#governance--risk-management)\n  - [Frameworks](#frameworks)\n  - [Standards \u0026 Verification](#standards--verification)\n  - [Taxonomies, Terminology \u0026 Risk Databases](#taxonomies-terminology--risk-databases)\n  - [Checklists \u0026 Practical Guidance](#checklists--practical-guidance)\n- [Attack Techniques \u0026 Red Teaming](#attack-techniques--red-teaming)\n  - [Adversarial ML \u0026 Classical Models](#adversarial-ml--classical-models)\n  - [LLM \u0026 GenAI Red Teaming](#llm--genai-red-teaming)\n  - [Agentic AI \u0026 MCP Attack Tools](#agentic-ai--mcp-attack-tools)\n  - [AI-Assisted Offensive Security](#ai-assisted-offensive-security)\n  - [Steganography \u0026 Covert Channels](#steganography--covert-channels)\n- [Benchmarks \u0026 Evaluations](#benchmarks--evaluations)\n- [Defense \u0026 Security Controls](#defense--security-controls)\n  - [Input/Output Guardrails](#inputoutput-guardrails)\n  - [Agent Runtime Security \u0026 Sandboxing](#agent-runtime-security--sandboxing)\n  - [MCP Security](#mcp-security)\n  - [Model \u0026 Artifact Scanning](#model--artifact-scanning)\n  - [AI-Assisted Defensive Security](#ai-assisted-defensive-security)\n  - [Privacy \u0026 Confidential Computing](#privacy--confidential-computing)\n  - [Data \u0026 Supply Chain Security](#data--supply-chain-security)\n- [Agentic AI Security Skills](#agentic-ai-security-skills)\n- [Security-Focused AI Models](#security-focused-ai-models)\n\n## Learning Resources\n\n### Reading \u0026 Guides\n* [OWASP ML TOP 10](https://owasp.org/www-project-machine-learning-security-top-10/)\n* [OWASP LLM TOP 10](https://owasp.org/www-project-top-10-for-large-language-model-applications/)\n* [OWASP AI Security and Privacy Guide](https://owasp.org/www-project-ai-security-and-privacy-guide/)\n* [NIST AIRC](https://airc.nist.gov/Home) - NIST Trustworthy \u0026 Responsible AI Resource Center\n* [The MLSecOps Top 10 by Institute for Ethical AI \u0026 Machine Learning](https://ethical.institute/security.html)\n* [OWASP Multi-Agentic System Threat Modeling](https://genai.owasp.org/resource/multi-agentic-system-threat-modeling-guide-v1-0/)\n* [OWASP: CheatSheet – A Practical Guide for Securely Using Third-Party MCP Servers 1.0](https://genai.owasp.org/resource/cheatsheet-a-practical-guide-for-securely-using-third-party-mcp-servers-1-0/)\n* [OWASP GenAI Threat \u0026 Defense Compass](https://genai.owasp.org/resource/owasp-genai-security-project-threat-defense-compass-1-0/) - _Reference guide mapping GenAI threats to corresponding defensive controls._\n* [NCSC Guidelines for Secure AI System Development](https://www.ncsc.gov.uk/collection/guidelines-secure-ai-system-development) - _Practical guidelines co-authored by NCSC (UK) and CISA for secure AI development, covering design, development, deployment, and operation._\n* [CoSAI – Preparing Defenders for a Changing Cybersecurity Landscape](https://github.com/cosai-oasis/ws2-defenders) - _Coalition for Secure AI workstream preparing defenders for AI-era threats._\n\n### Courses, Labs \u0026 CTFs\n* [Damn Vulnerable MCP Server](https://github.com/harishsg993010/damn-vulnerable-MCP-server) - _A deliberately vulnerable implementation of the Model Context Protocol (MCP) for educational purposes._\n* [otto-support](https://github.com/BishopFox/otto-support) - _A vulnerable MCP server implementation using mcp-go with tiered authentication (4 roles), 19 tools, and built-in sandboxed container with Claude Code for practicing privilege escalation and tool misuse._\n* [OWASP WrongSecrets LLM exercise](https://wrongsecrets.herokuapp.com/challenge/32)\n* [vulnerable-mcp-servers-lab](https://github.com/appsecco/vulnerable-mcp-servers-lab) - _A collection of servers which are deliberately vulnerable to learn Pentesting MCP Servers._\n* [FinBot Agentic AI Capture The Flag (CTF) Application](https://genai.owasp.org/resource/finbot-agentic-ai-capture-the-flag-ctf-application/) - _FinBot is an Agentic Security Capture The Flag (CTF) interactive platform that simulates real-world vulnerabilities in agentic AI systems using a simulated Financial Services-focused application._\n* [AI-Red-Teaming-Playground-Labs](https://github.com/microsoft/AI-Red-Teaming-Playground-Labs) - _AI Red Teaming playground labs to run AI Red Teaming trainings including infrastructure._\n* [Damn Vulnerable LLM Agent](https://github.com/ReversecLabs/damn-vulnerable-llm-agent) - _Intentionally vulnerable LLM agent for learning about prompt injection, tool misuse, and agent security_\n* * [LLMVault](https://github.com/CyberSunil/LLMVault) - _An open-source CTF-style LLM security lab for learning the OWASP LLM Top 10 through hands-on vulnerable exercises covering prompt injection, RAG attacks, system prompt leakage, tool misuse, and agent security._\n\n### Podcasts\n* [MLSecOps podcast](https://mlsecops.com/podcast)\n* [AI Security Podcast](https://www.aisecuritypodcast.com/)\n* [AI Security Ops](https://aisecurityops.transistor.fm) - _Weekly podcasts from Black Hills Information Security exploring how AI transforms cybersecurity—covering emerging threats, tools, and trends with practical, actionable knowledge._\n* [GenAI Security podcast](https://podcasts.apple.com/ph/podcast/the-genai-security-podcast/id1782916580)\n\n## Governance \u0026 Risk Management\n\n### Frameworks\n* [NIST AI Risk Management Framework](https://airc.nist.gov/AI_RMF_Knowledge_Base/AI_RMF)\n* [ISO/IEC 42001 Artificial Intelligence Management System](https://www.iso.org/standard/81230.html)\n* [ISO/IEC 23894:2023 Information technology — Artificial intelligence — Guidance on risk management](https://www.iso.org/standard/77304.html)\n* [Google Secure AI Framework (SAIF)](https://saif.google/)\n* [ENISA Multilayer Framework for Good Cybersecurity Practices for AI](https://www.enisa.europa.eu/publications/multilayer-framework-for-good-cybersecurity-practices-for-ai)\n* [OWASP Artificial Intelligence Maturity Assessment](https://github.com/OWASP/www-project-ai-maturity-assessment)\n* [CSA AI Model Risk Framework](https://cloudsecurityalliance.org/artifacts/ai-model-risk-management-framework)\n* [CSA Maestro AI Threat Modeling Framework](https://cloudsecurityalliance.org/blog/2025/02/06/agentic-ai-threat-modeling-framework-maestro)\n* [CSA AI Controls Matrix](https://cloudsecurityalliance.org/artifacts/ai-controls-matrix) - _Comprehensive controls matrix for AI systems covering governance, risk, and compliance._\n* [OWASP Agentic AI Top 10](https://github.com/precize/OWASP-Agentic-AI) - _Top 10 for Agentic AI serving as the core for OWASP and CSA red teaming._\n* [OWASP GenAI Crosswalk](https://genai-security-project.github.io/crosswalk/) - _Interactive crosswalk mapping OWASP GenAI risks (LLM Top 10, Agentic Top 10, DSGAI 2026) to 25 industry frameworks with coverage scoring and gap analysis._\n* [NIST SP 800-218A Secure Software Development Practices for Generative AI and Dual-Use Foundation Models](https://csrc.nist.gov/pubs/sp/800/218/a/final) - _Extension of SSDF (SP 800-218) with specific practices for GenAI and dual-use foundation model development._\n\n### Standards \u0026 Verification\n* [OWASP AI Security Verification Standard](https://github.com/OWASP/AISVS)\n* [OWASP Agent Name Service](https://genai.owasp.org/resource/agent-name-service-ans-for-secure-al-agent-discovery-v1-0/)\n* [OWASP Agent Observability Standard](https://aos.owasp.org/)\n* [AI Verify](https://github.com/aiverify-foundation/aiverify) - _Singapore government-backed AI testing framework and toolkit for verifying AI system properties against governance frameworks._\n* [Agent Control Standard](https://github.com/GenAI-Security-Project/agent-control-standard) - _Open standard for runtime agent control: declarative hooks, policy enforcement, and observability across AI agent frameworks (OWASP GenAI Security Project)._\n* [AARM](https://aarm.dev) - _Cloud Security Alliance specification defining the system category and capabilities for agentic runtime security._\n\n### Taxonomies, Terminology \u0026 Risk Databases\n* [NIST AI 100-2e2023](https://csrc.nist.gov/publications/detail/white-paper/2023/03/08/adversarial-machine-learning-taxonomy-and-terminology/draft) - _Adversarial machine learning taxonomy and terminology_\n* [MITRE ATLAS](https://atlas.mitre.org/)\n  * [ATLAS Knowledge Base Agent](https://github.com/mitre-atlas/atlas-knowledge-base-agent) - _Open-source AI assistant for exploring the ATLAS knowledge base with natural language, MCP server access, and customizable agent workflows._\n* [AVIDML](https://avidml.org/taxonomy/)\n* [MIT AI Risk Repository](https://airisk.mit.edu/)\n* [AI Incident Database](https://incidentdatabase.ai/)\n* [ISO/IEC 22989:2022 Information technology — Artificial intelligence — Artificial intelligence concepts and terminology](https://www.iso.org/standard/74296.html)\n* [NIST AI Glossary](https://airc.nist.gov/glossary/)\n* [The Arcanum Prompt Injection Taxonomy](https://arcanum-sec.github.io/arc_pi_taxonomy) - _Comprehensive prompt injection attack classification system covering attack intents, techniques, evasions, and input vectors. Categorizes goals, methods, obfuscation techniques, and attack surfaces for prompt injection attacks._\n* [CSA LLM Threats Taxonomy](https://cloudsecurityalliance.org/artifacts/csa-large-language-model-llm-threats-taxonomy)\n* [OWASP AI Vulnerability Scoring System (AIVSS)](https://github.com/OWASP/www-project-artificial-intelligence-vulnerability-scoring-system) - _Scoring system for AI-specific vulnerabilities, extending CVSS concepts to AI risk dimensions._\n\n### Checklists \u0026 Practical Guidance\n* [OWASP LLM Applications Cybersecurity and Governance Checklist](https://genai.owasp.org/resource/llm-applications-cybersecurity-and-governance-checklist-english/)\n* [OWASP LLM and Generative AI Security Center of Excellence Guide](https://genai.owasp.org/resource/llm-and-generative-ai-security-center-of-excellence-guide/)\n* [OWASP Agentic AI – Threats and Mitigations](https://genai.owasp.org/resource/agentic-ai-threats-and-mitigations/)\n* [OWASP AI Security Solutions Landscape](https://genai.owasp.org/ai-security-solutions-landscape/)\n* [OWASP GenAI Incident Response Guide](https://genai.owasp.org/resource/genai-incident-response-guide-1-0/)\n* [OWASP LLM and GenAI Data Security Best Practices](https://genai.owasp.org/resource/llm-and-gen-ai-data-security-best-practices/)\n* [OWASP Securing Agentic AI Applications](https://genai.owasp.org/resource/securing-agentic-applications-guide-1-0/)\n* [OWASP GenAI Red Teaming Guide](https://genai.owasp.org/initiatives/#ai-redteaming)\n\n## Attack Techniques \u0026 Red Teaming\n\n### Adversarial ML \u0026 Classical Models\n* [Adversarial Robustness Toolkit](https://research.ibm.com/projects/adversarial-robustness-toolbox) - _ART focuses on the threats of Evasion (change the model behavior with input modifications), Poisoning (control a model with training data modifications), Extraction (steal a model through queries) and Inference (attack the privacy of the training data)_\n* [cleverhans](https://github.com/cleverhans-lab/cleverhans) - _An adversarial example library for constructing attacks, building defenses, and benchmarking both_\n* [foolbox](https://github.com/bethgelab/foolbox) - _A Python toolbox to create adversarial examples that fool neural networks in PyTorch, TensorFlow, and JAX_\n* [TextAttack](https://github.com/QData/TextAttack) - _A Python framework for adversarial attacks, data augmentation, and model training in NLP_\n* [Counterfit](https://github.com/Azure/counterfit) - _generic automation layer for assessing the security of machine learning systems_\n* [OffsecML Playbook](https://wiki.offsecml.com) - _A collection of offensive and adversarial TTPs with proofs of concept_\n* [BadDiffusion](https://github.com/IBM/BadDiffusion) - _Official repo to reproduce the paper \"How to Backdoor Diffusion Models?\" published at CVPR 2023_\n* [secml-torch](https://github.com/pralab/secml-torch) - _SecML-Torch: A Library for Robustness Evaluation of Deep Learning Models_\n\n### LLM \u0026 GenAI Red Teaming\n* [garak](https://github.com/leondz/garak/) - _security probing tool for LLMs_\n* [ai-scanner](https://github.com/0din-ai/ai-scanner) - _Open-source web application for AI model security assessments, built on NVIDIA garak. Features 179 probes, multi-target scanning, scheduled scans, ASR scoring, and SIEM integration._\n* [promptfoo](https://github.com/promptfoo/promptfoo) - _Test your prompts, agents, and RAGs. Red teaming, pentesting, and vulnerability scanning for LLMs. Compare performance of GPT, Claude, Gemini, Llama, and more. Simple declarative configs with command line and CI/CD integration._\n* [PyRIT](https://github.com/Azure/PyRIT) - _The Python Risk Identification Tool for generative AI (PyRIT) is an open source framework built to empower security professionals and engineers to proactively identify risks in generative AI systems._\n* [PurpleLlama](https://github.com/meta-llama/PurpleLlama) - _Set of tools to assess and improve LLM security._\n* [augustus](https://github.com/praetorian-inc/augustus) - _LLM security testing framework for detecting prompt injection, jailbreaks, and adversarial attacks. 190+ probes, 28 providers, single Go binary. Production-ready with concurrent scanning, rate limiting, and retry logic._\n* [FuzzyAI](https://github.com/cyberark/FuzzyAI) - _A powerful tool for automated LLM fuzzing. It is designed to help developers and security researchers identify and mitigate potential jailbreaks in their LLM APIs._\n* [EasyJailbreak](https://github.com/EasyJailbreak/EasyJailbreak) - _An easy-to-use Python framework to generate adversarial jailbreak prompts._\n* [gptfuzz](https://github.com/sherdencooper/GPTFuzz) - _Official repo for GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts_\n* [llamator](https://github.com/LLAMATOR-Core/llamator) - _Framework for testing vulnerabilities of large language models (LLM)._\n* [agentic_security](https://github.com/msoedov/agentic_security/) - _Agentic LLM Vulnerability Scanner / AI red teaming kit_\n* [Agentic Radar](https://github.com/splx-ai/agentic-radar) - _Open-source CLI security scanner for agentic workflows._\n* [RAPTOR](https://github.com/gadievron/raptor) - _Autonomous offensive/defensive security research framework built on Claude Code. Chains static analysis (Semgrep, CodeQL), binary analysis, LLM-powered vulnerability validation, exploit generation, and patch writing. Multi-model orchestration with Z3-based feasibility analysis._\n* [whistleblower](https://github.com/Repello-AI/whistleblower) - _Offensive security tool for testing against system prompt leakage and capability discovery of an AI application exposed through API_\n* [promptmap](https://github.com/utkusen/promptmap) - _a prompt injection scanner for custom LLM applications_\n* [spikee](https://github.com/WithSecureLabs/spikee) - _Simple Prompt Injection Kit for Evaluation and Exploitation_\n* [ps-fuzz](https://github.com/prompt-security/ps-fuzz) - _Make your GenAI Apps Safe \u0026 Secure — Test \u0026 harden your system prompt_\n* [EasyEdit](https://github.com/zjunlp/EasyEdit) - _Modify an LLM's ground truths_\n* [llm-attacks](https://github.com/llm-attacks/llm-attacks) - _Universal and Transferable Attacks on Aligned Language Models_\n* [llm-security](https://github.com/greshake/llm-security) - _New ways of breaking app-integrated LLMs_\n* [Plexiglass](https://github.com/safellama/plexiglass) - _A toolkit for detecting and protecting against vulnerabilities in Large Language Models (LLMs)._\n* [Prompt Hacking Resources](https://github.com/PromptLabs/Prompt-Hacking-Resources) - _A list of curated resources for people interested in AI Red Teaming, Jailbreaking, and Prompt Injection_\n* [Giskard](https://github.com/Giskard-AI/giskard) - _Open-Source Evaluation \u0026 Testing for AI \u0026 LLM systems_\n* [blackice](https://github.com/databricks/containers/tree/master/ubuntu/blackice) - _BlackIce is an open-source containerized toolkit designed for red teaming AI models, including Large Language Models (LLMs) and classical machine learning (ML) models. Inspired by the convenience and standardization of Kali Linux in traditional penetration testing, BlackIce simplifies AI security assessments by providing a reproducible container image preconfigured with specialized evaluation tools._\n* [ai-best-practices](https://github.com/semgrep/ai-best-practices) - _Semgrep Pro Rules to ensure code using LLMs is following best practices. 58 rules, 102 sub-rules covering 6 providers + MCP + Claude Code \u0026 Cursor hooks + LangChain. Detects hardcoded API keys, prompt injection risks, missing safety checks, and unhandled errors across 7 languages._\n* [G0DM0D3](https://github.com/elder-plinius/G0DM0D3) - _Open-source multi-model chat interface for red teaming with 50+ models via OpenRouter. Features GODMODE CLASSIC (5 jailbreak combos), ULTRAPLINIAN multi-model evaluation, Parseltongue input perturbation engine with 33 red team techniques, and AutoTune adaptive sampling for AI safety research._\n* [L1B3RT4S](https://github.com/elder-plinius/L1B3RT4S) - _Collection of jailbreak and liberation prompts for major LLMs. Curated library of adversarial prompts designed to test and evaluate AI safety guardrails across ChatGPT, Claude, Gemini, and other frontier models._\n* [OBLITERATUS](https://github.com/elder-plinius/OBLITERATUS) - _Abliteration toolkit that removes refusal behaviors from open-source LLMs without retraining. Modifies model weights to eliminate safety-aligned refusal responses, enabling unrestricted model behavior research._\n* [T3MP3ST](https://github.com/elder-plinius/T3MP3ST) - _Autonomous red teaming platform and multi-agent offensive-security meta-harness. Coordinates swarms of AI agents for coordinated adversarial testing of frontier AI systems._\n* [P4RS3LT0NGV3](https://github.com/elder-plinius/P4RS3LT0NGV3) - _Universal text transformation and promptcrafting toolkit. Supports translation, mutation, encoding/decoding, and adversarial prompt engineering for jailbreak payload construction._\n* [claude-secure-coding-rules](https://github.com/TikiTribe/claude-secure-coding-rules) - _Open-source security rules that guide Claude Code to generate secure code by default._\n* [DeepTeam](https://github.com/confident-ai/deepteam) - _LLM red teaming framework with 40+ attack methods including prompt injection, jailbreaking, and RAG poisoning. Integrates with CI/CD pipelines._\n* [PoisonedRAG](https://github.com/sleeepeer/PoisonedRAG) - _Knowledge Corruption Attacks to Retrieval-Augmented Generation (USENIX Security 2025)._\n\n### Agentic AI \u0026 MCP Attack Tools\n* [RAMPART](https://github.com/microsoft/RAMPART) - _pytest-native safety and security testing framework for agentic AI applications._\n* [AI-Infra-Guard](https://github.com/Tencent/AI-Infra-Guard/) - _A comprehensive, intelligent, and easy-to-use AI Red Teaming platform developed by Tencent Zhuque Lab. Integrates modules for Infra Scan, MCP Scan, and Jailbreak Evaluation, providing a one-click web UI, REST APIs, and Docker-based deployment for comprehensive AI security evaluation._\n* [OpenPromptInjection](https://github.com/liu00222/Open-Prompt-Injection) - _A benchmark for prompt injection attacks and defenses_\n* [AIMap](https://github.com/BishopFox/aimap) - _Internet-scale discovery and security testing platform for exposed AI agent infrastructure. Queries Shodan for MCP servers, Ollama instances, vLLM/LiteLLM proxies, and more — then fingerprints, scores risk, and launches protocol-specific attack suites with real-time streaming results._\n\n### AI-Assisted Offensive Security\n* [PentestGPT](https://github.com/GreyDGL/PentestGPT) - _A GPT-empowered penetration testing tool_\n* [HackingBuddyGPT](https://github.com/ipa-lab/hackingBuddyGPT) - _Helping Ethical Hackers use LLMs in 50 Lines of Code or less_\n* [BugTraceAI](https://github.com/BugTraceAI/BugTraceAI) - _Open-source multi-agent platform for authorized web application security testing with validation, evidence capture, and reporting._\n* [cai](https://github.com/aliasrobotics/cai) - _Cybersecurity AI (CAI), an open Bug Bounty-ready Artificial Intelligence ([paper](https://arxiv.org/pdf/2504.06017))_\n* [shannon](https://github.com/KeygraphHQ/shannon) - _Fully autonomous AI pentester for web apps and APIs by Keygraph. White-box security testing that analyzes source code, identifies attack vectors, and executes real exploits. 96.15% success rate (100/104 exploits) on XBOW benchmark._\n* [strix](https://github.com/usestrix/strix) - _Strix are autonomous AI agents that act just like real hackers - they run your code dynamically, find vulnerabilities, and validate them through actual proof-of-concepts_\n* [redamon](https://github.com/samugit83/redamon) - _AI-powered agentic red team framework that automates offensive security operations from reconnaissance to exploitation to post-exploitation with zero human intervention._\n* [CyberStrikeAI](https://github.com/Ed1s0nZ/CyberStrikeAI) - _AI-native security testing platform built in Go. Integrates 100+ security tools with an intelligent orchestration engine, role-based testing with predefined security roles, skills system, and comprehensive lifecycle management. Uses MCP protocol and AI agents for end-to-end automation from conversational commands to vulnerability discovery._\n* [HexStrikeAI](https://github.com/0x4m4/hexstrike-ai) - _HexStrike AI MCP Agents is an advanced MCP server that lets AI agents (Claude, GPT, Copilot, etc.) autonomously run 150+ cybersecurity tools for automated pentesting, vulnerability discovery, bug bounty automation, and security research._\n* [mcp-for-security](https://github.com/cyproxio/mcp-for-security) - _A collection of Model Context Protocol servers for popular security tools like SQLMap, FFUF, NMAP, Masscan and more. Integrate security testing and penetration testing into AI workflows._\n* [mcp-security-hub](https://github.com/FuzzingLabs/mcp-security-hub) - _A growing collection of MCP servers bringing offensive security tools to AI assistants. Nmap, Ghidra, Nuclei, SQLMap, Hashcat and more._\n* [Burp MCP Server](https://github.com/PortSwigger/mcp-server) - _MCP Server for Burp_\n* [burpgpt](https://github.com/aress31/burpgpt) - _A Burp Suite extension that integrates OpenAI's GPT to perform an additional passive scan for discovering highly bespoke vulnerabilities and enables running traffic-based analysis of any type._\n* [guardian-cli](https://github.com/zakirkun/guardian-cli) - _AI-Powered Security Testing \u0026 Vulnerability Scanner. Guardian CLI is an intelligent security testing tool that leverages AI to automate penetration testing, vulnerability assessment, and security auditing._\n* [AutoPentestX](https://github.com/Gowtham-Darkseid/AutoPentestX) - _AutoPentestX – Linux Automated Pentesting \u0026 Vulnerability Reporting Tool_\n* [HackGPT](https://github.com/NoDataFound/hackGPT) - _A tool using ChatGPT for hacking_\n* [nano-analyzer](https://github.com/weareaisle/nano-analyzer) - _A minimal LLM-powered zero-day vulnerability scanner by AISLE._\n* [clearwing](https://github.com/Lazarus-AI/clearwing) - _Autonomous vulnerability scanner and source-code hunter built on LangGraph._\n* [Zen-AI-Pentest](https://github.com/SHAdd0WTAka/Zen-Ai-Pentest) - _AI-Powered Penetration Testing Framework with automated vulnerability scanning, multi-agent system, and compliance reporting. 72+ security tools, Docker sandbox, ReAct agents, attack path analysis._\n* [Violin](https://github.com/Strategic-Automation/violin) - _Supervised, agentic Hermes Agent pentest profile: 31 skill-based playbooks (OWASP Top 10, API Top 10, LLM Top 10) with interactive scoping, scope validation, and approval gates for authorised recon, exploit validation, and reporting._\n* [NeuroSploit](https://github.com/JoasASantos/NeuroSploit) - _Autonomous multi-model pentesting harness in Rust. LLM pool drives recon, agent selection, exploit chaining, and cross-model validation voting across black-box, white-box, grey-box, and cloud engagement modes._\n* [OpenHack](https://github.com/openhackai/OpenHack) - _AI-powered multi-agent scanner that autonomously finds SQLi, XSS, IDOR, and auth bypass in your codebase, then verifies each finding through sandbox execution and browser replay. On par with Claude Opus 4.6 at roughly 40x lower cost._\n* [PentAGI](https://github.com/vxcontrol/pentagi) - _Fully autonomous multi-agent system for complex penetration testing tasks. Sandboxed Docker execution, multi-provider LLM support (OpenAI, Anthropic, Gemini, Ollama, DeepSeek), knowledge graph integration, and real-time agent supervision._\n* [V3SP3R](https://github.com/elder-plinius/V3SP3R) - _AI-powered hardware hacking companion for the Flipper Zero. Natural language interface for controlling hardware attacks, with smart glasses integration for hands-free operation._\n\n### Steganography \u0026 Covert Channels\n* [ST3GG](https://github.com/elder-plinius/ST3GG) - _All-in-one steganography suite with multi-layer encoding, image and audio steganography, and steganalysis tools for detecting hidden data in AI-generated media._\n\n## Benchmarks \u0026 Evaluations\n\n* [ISC-Bench](https://github.com/wuyoscar/ISC-Bench) - _Internal Safety Collapse: jailbreaks any frontier LLM (Claude Opus 4.6, GPT-5.4) in pass@3 via normal task completion — no adversarial prompting. Black-box, cross-domain, cross-science. Novel failure mode._ [[Paper]](https://arxiv.org/abs/2603.23509)\n* [jailbreakbench](https://github.com/JailbreakBench/jailbreakbench) - _JailbreakBench: An Open Robustness Benchmark for Jailbreaking Language Models [NeurIPS 2024 Datasets and Benchmarks Track]_\n* [AgentDojo](https://github.com/ethz-spylab/agentdojo) - _A Dynamic Environment to Evaluate Attacks and Defenses for LLM Agents._\n* [AIRTBench](https://github.com/dreadnode/AIRTBench-Code) - _Code Repository for: AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models_\n* [HackingBuddyGPT benchmark dataset](https://github.com/ipa-lab/hacking-benchmark) - _Benchmark dataset for automated pentesting_\n* [AgentDoG](https://github.com/AI45Lab/AgentDoG) - _AgentDoG is a risk-aware evaluation and guarding framework for autonomous agents. It focuses on trajectory-level risk assessment, aiming to determine whether an agent's execution trajectory contains safety risks under diverse application scenarios._\n* [AICGSecEval](https://github.com/Tencent/AICGSecEval) - _Tencent's comprehensive evaluation benchmark for AI code generation security, covering 10 CWE categories with automated test harness_\n* [Inspect](https://github.com/UKGovernmentBEIS/inspect_ai) - _Framework for large language model evaluations by the UK AI Security Institute. 200+ pre-built evaluations covering prompt engineering, tool usage, multi-turn dialog, and model-graded scoring._\n* [sec-code-bench](https://github.com/alibaba/sec-code-bench) - _Alibaba's benchmark for evaluating LLM code security capabilities across 17 vulnerability categories and 4 programming languages_\n\n## Defense \u0026 Security Controls\n\n### Input/Output Guardrails\n* [NeMo-Guardrails](https://github.com/NVIDIA/NeMo-Guardrails) - _NeMo Guardrails is an open-source toolkit for easily adding programmable guardrails to LLM-based conversational systems._\n* [LlamaFirewall](https://github.com/meta-llama/PurpleLlama/tree/main/LlamaFirewall) - _LlamaFirewall is a framework designed to detect and mitigate AI centric security risks, supporting multiple layers of inputs and outputs, such as typical LLM chat and more advanced multi-step agentic operations._\n* [llm-guard](https://github.com/protectai/llm-guard) - _LLM Guard by Protect AI is a comprehensive tool designed to fortify the security of Large Language Models (LLMs)._\n* [Guardrails.ai](https://shreyar.github.io/guardrails/) - _Guardrails is a Python package that lets a user add structure, type and quality guarantees to the outputs of large language models (LLMs)_\n* [TrustGate](https://github.com/NeuralTrust/TrustGate) - _Generative Application Firewall (GAF) to detect, prevent and block attacks against GenAI Applications_\n* [ZenGuard AI](https://github.com/ZenGuard-AI/fast-llm-security-guardrails) - _The fastest Trust Layer for AI Agents_\n* [LocalMod](https://github.com/KOKOSde/localmod) - _Self-hosted content moderation API with prompt injection detection, toxicity filtering, PII detection, and NSFW classification. Runs 100% offline._\n* [DynaGuard](https://github.com/montehoover/DynaGuard) - _A Dynamic Guardrail Model With User-Defined Policies_\n* [AprielGuard](https://huggingface.co/blog/ServiceNow-AI/aprielguard) - _8B parameter safety–security safeguard model_\n* [Safe Zone](https://github.com/thyrisAI/safe-zone) - _Safe Zone is an open-source PII detection and guardrails engine that prevents sensitive data from leaking to LLMs and third-party APIs._\n* [superagent](https://github.com/superagent-ai/superagent) - _Superagent provides purpose-trained guardrails that make AI-agents secure and compliant._\n* [ShellWard](https://github.com/jnMetaCode/shellward) - _AI Agent Security Middleware with 8-layer defense against prompt injection, data exfiltration \u0026 dangerous commands. Zero dependencies._\n* [CodeGate](https://codegate.ai) - _An open-source, privacy-focused project that acts as a layer of security within a developer's Code Generation AI workflow_\n* [Future AGI](https://github.com/future-agi/future-agi) - _Open-source self-hostable platform with built-in real-time guardrails for unsafe outputs (jailbreak, PII, injection, toxicity), evals, tracing, simulations, and gateway for LLM and agent applications._\n* [Prompt Injection Defenses](https://github.com/tldrsec/prompt-injection-defenses) - _Comprehensive collection of every practical and proposed defense against prompt injection._\n\n### Agent Runtime Security \u0026 Sandboxing\n* [OpenShell](https://github.com/NVIDIA/OpenShell) - _OpenShell is the safe, private runtime for autonomous AI agents. It provides sandboxed execution environments governed by declarative YAML policies that prevent unauthorized file access, data exfiltration, and uncontrolled network activity._\n* [OpenSandbox](https://github.com/alibaba/OpenSandbox) - _Secure, Fast, and Extensible Sandbox runtime for AI agents. Multi-language SDKs, Docker/Kubernetes runtimes, gVisor/Kata Containers/Firecracker isolation. CNCF Landscape project._\n* [CubeSandbox](https://github.com/TencentCloud/CubeSandbox) - _Instant, concurrent, secure \u0026 lightweight sandbox for AI agents by Tencent Cloud. Sub-60ms cold start, \u003c5MB memory overhead, E2B SDK compatible. Built on RustVMM and KVM with extreme isolation (dedicated kernel + eBPF)._\n* [ADR](https://github.com/uber/ADR) - _Enterprise agentic AI detection and response — observability, security benchmarking (300+ tasks, 133 MCP servers), and dual-agent threat detection. Deployed at Uber. MLSys 2026._\n* [Aegis](https://github.com/antropos17/Aegis) - _Open-source EDR for AI agents by Antropos. Monitor processes, files, network, and behavior of autonomous AI agents in real time. No telemetry, no cloud, everything stays local._\n* [Microsoft Agent Governance Toolkit](https://github.com/microsoft/agent-governance-toolkit) - _AI Agent Governance Toolkit from Microsoft — Policy enforcement, zero-trust identity, execution sandboxing, and reliability engineering for autonomous AI agents. Covers 10/10 OWASP Agentic Top 10._\n* [agentfield](https://github.com/Agent-Field/agentfield) - _Open-source control plane for agent systems with cryptographic identity, policy enforcement, and audit-friendly observability._\n* [leash](https://github.com/strongdm/leash) - _Leash wraps AI coding agents in containers and monitors their activity._\n* [Agent Safehouse](https://github.com/eugene1g/agent-safehouse) - _macOS sandbox for LLM coding agents using sandbox-exec with composable policy profiles and a deny-first model._\n* [vibekit](https://github.com/superagent-ai/vibekit) - _Run Claude Code, Gemini, Codex — or any coding agent — in a clean, isolated sandbox with sensitive data redaction and observability baked in._\n* [pipelock](https://github.com/luckyPipewrench/pipelock) - _Security harness for AI agents — egress proxy with DLP scanning, SSRF protection, MCP response scanning, and workspace integrity monitoring_\n* [skill-scanner](https://github.com/cisco-ai-defense/skill-scanner) - _A security scanner for AI Agent Skills that detects prompt injection, data exfiltration, and malicious code patterns. Combines pattern-based detection (YAML + YARA), LLM-as-a-judge, and behavioral dataflow analysis for comprehensive threat detection._\n* [SkillSpector](https://github.com/NVIDIA/SkillSpector) - _Security scanner for AI agent skills. Detects 64 vulnerability patterns across 16 categories with static analysis + optional LLM semantic evaluation. Multi-format output (JSON, Markdown, SARIF)._\n* [Project CodeGuard](https://github.com/cosai-oasis/project-codeguard) - _CoSAI Open Source Project for securing AI-assisted development workflows. CodeGuard provides security controls and guardrails for AI coding assistants to prevent vulnerabilities from being introduced during AI-generated code development._\n* [AgentLens](https://github.com/dreadnode/agent-lens) - _Agent observability and replay tooling for AI safety \u0026 interpretability research. Harness for running multi-session agent trajectories, capturing them in ATIF format, and tracking file state changes across sessions. Built for studying LLM agent behavior across multi-turn, multi-session, multi-agent interactions._\n* [OWASP Agent Memory Guard](https://github.com/OWASP/www-project-agent-memory-guard) - _Runtime defense layer for AI agent memory. Screens every memory read/write through detectors (prompt injection, secret/PII leakage, protected-key tampering, anomalies) and a YAML policy (allow/redact/quarantine/block), with SHA-256 integrity baselines, snapshots/rollback, and structured security events. Reference implementation for OWASP ASI06 Memory \u0026 Context Poisoning. Apache-2.0._\n* [claude-code-devcontainer](https://github.com/trailofbits/claude-code-devcontainer) - _Sandboxed devcontainer for running Claude Code in bypass mode safely. Built for security audits and untrusted code review._\n* [claude-code-safety-net](https://github.com/kenryu42/claude-code-safety-net) - _A Claude Code plugin that acts as a safety net, catching destructive git and filesystem commands before they execute_\n* [OneCLI](https://github.com/onecli/onecli) - _Open-source credential vault for AI agents. Rust HTTP gateway intercepts agent requests and injects API credentials transparently so agents never hold raw keys. AES-256-GCM encryption, per-agent scoping, full audit trail._\n* [openclaw-shield](https://github.com/knostic/openclaw-shield) - _Security plugin for OpenClaw agents - prevents secret leaks, PII exposure, and destructive command execution_\n* [clawsec](https://github.com/prompt-security/clawsec) - _Security scanner and hardening tool for OpenClaw deployments. Provides security assessments, configuration auditing, and vulnerability detection specifically for OpenClaw gateway and agent configurations._\n* [nanoclaw](https://github.com/qwibitai/nanoclaw) - _Lightweight alternative to OpenClaw that runs in containers for security. Connects to WhatsApp, has memory, scheduled jobs, and runs directly on Anthropic's Agents SDK. First AI assistant to support Agent Swarms for collaborative agent teams._\n* [secureclaw](https://github.com/adversa-ai/secureclaw) - _Automated security hardening for OpenClaw AI agents by Adversa AI. 51 audit checks, 12 behavioral rules, 9 scripts, 4 pattern databases. Full OWASP ASI Top 10 coverage. Protects against prompt injection, credential theft, supply chain attacks, and privacy leaks._\n* [defenseclaw](https://github.com/cisco-ai-defense/defenseclaw) - _Enterprise governance layer for OpenClaw from Cisco AI Defense. Scans skills, MCP servers, and plugins with built-in CodeGuard SAST, tool call inspection engine, LLM guardrail proxy, and SIEM integration. Auto-blocks HIGH/CRITICAL findings._\n* [microsandbox](https://github.com/zerocore-ai/microsandbox) - _Lightweight microVM sandbox for running untrusted AI-generated code safely with strong isolation guarantees_\n* [Adrian](https://github.com/secureagentics/Adrian) - _Open-source, AARM-aligned runtime security monitoring and control engine for AI agents by Secure Agentics. Analyses agent activity logs (tool calls, actions, outputs) and reasoning traces to detect malicious, misaligned, or out-of-remit behaviour, with optional in-flight intervention (audit vs block mode). Python (LangChain/LangGraph) and TypeScript SDKs, fully self-hostable offline. Apache-2.0._\n* [HOL Guard](https://github.com/hashgraph-online/hol-guard) - _Local-first security harness that intercepts tool calls in AI coding agents (Codex, Claude Code, Cursor, Gemini, Copilot, Hermes, OpenCode) before files change or network is contacted. Pre-tool hooks, approval center, supply-chain advisory scanning, and optional Guard Cloud sync._\n* [Bernstein](https://github.com/sipyourdrink-ltd/bernstein) - _Open-source governance layer for AI agents, covering CLI coding agents (Claude Code, Codex, Gemini CLI, Aider, and 40+ more) as well as self-hosted and hosted model endpoints. Runs each task in an isolated git worktree with per-agent credential scoping, PII gating, and a policy engine. No model sits in the coordination loop, so runs replay byte-identically, and an opt-in HMAC-chained audit log with signed per-artifact lineage can be verified offline with `bernstein audit verify`. Apache-2.0._\n* [Doberman](https://github.com/DobermanCore/Doberman-Core) - _Runtime authorization layer between a coding agent and its tools. A local policy engine gives every tool call an allow/authenticate/block verdict before it executes; blocks carry reason codes, logs redact secrets to HMAC fingerprints, and errors fail closed. Ships an MCP proxy plus Claude Code and Codex adapters. Apache-2.0._\n\n### MCP Security\n* [MCP-Security-Checklist](https://github.com/slowmist/MCP-Security-Checklist) - _A comprehensive security checklist for MCP-based AI tools. Built by SlowMist to safeguard LLM plugin ecosystems._\n* [Awesome-MCP-Security](https://github.com/Puliczek/awesome-mcp-security) - _Everything you need to know about Model Context Protocol (MCP) security._\n* [secure-mcp-gateway](https://github.com/enkryptai/secure-mcp-gateway) - _This Secure MCP Gateway is built with authentication, automatic tool discovery, caching, and guardrail enforcement._\n* [mcp-context-protector](https://github.com/trailofbits/mcp-context-protector) - _context-protector is a security wrapper for MCP servers that addresses risks associated with running untrusted MCP servers, including line jumping, unexpected server configuration changes, and other prompt injection attacks_\n* [mcp-guardian](https://github.com/eqtylab/mcp-guardian/) - _MCP Guardian manages your LLM assistant's access to MCP servers, handing you realtime control of your LLM's activity._\n* [MCP Audit VSCode Extension](https://github.com/Agentity-com/mcp-audit-extension) - _Audit and log all GitHub Copilot MCP tool calls in VSCode centrally with ease._\n* [MCP-Scan](https://github.com/invariantlabs-ai/mcp-scan) - _A security scanning tool for MCP servers_\n* [ATR (Agent Threat Rules)](https://github.com/Agent-Threat-Rule/agent-threat-rules) - _Open-source detection rules for AI agent threats. 108 regex rules covering prompt injection, tool poisoning, credential exfiltration across 9 categories. Used by Cisco AI Defense. MIT licensed._\n* [MCPProxy](https://github.com/smart-mcp-proxy/mcpproxy-go) - _Local-first MCP proxy with per-tool SHA-256 quarantine to detect tool-poisoning and rug-pull attacks, automatic sensitive-data and secret scanning of tool calls, Docker sandbox isolation for untrusted MCP servers, and OAuth 2.1. MIT licensed._\n\n### Model \u0026 Artifact Scanning\n* [modelscan](https://github.com/protectai/modelscan) - _ModelScan is an open source project from Protect AI that scans models to determine if they contain unsafe code._\n* [picklescan](https://github.com/mmaitre314/picklescan) - _Security scanner detecting Python Pickle files performing suspicious actions_\n* [fickling](https://github.com/trailofbits/fickling) - _A Python pickling decompiler and static analyzer_\n* [medusa](https://github.com/Pantheon-Security/medusa) - _AI-first security scanner with 74+ analyzers, 180+ AI agent security rules, intelligent false positive reduction. Supports all languages. CVE detection for React2Shell, mcp-remote RCE._\n* [julius](https://github.com/praetorian-inc/julius) - _LLM service fingerprinting tool for security professionals. Detects 32+ AI services (Ollama, vLLM, LiteLLM, Hugging Face TGI, etc.) during penetration tests and attack surface discovery. Uses HTTP-based service fingerprinting to identify server infrastructure._\n* [a2a-scanner](https://github.com/cisco-ai-defense/a2a-scanner) - _Scan A2A agents for potential threats and security issues_\n\n### AI-Assisted Defensive Security\n* [Claude Code Security Review](https://github.com/anthropics/claude-code-security-review) - _An AI-powered security review GitHub Action using Claude to analyze code changes for security vulnerabilities._\n* [OpenCodeReview](https://github.com/alibaba/open-code-review) - _AI-powered code review CLI from Alibaba with deterministic pipelines + LLM agent, line-level precision, and built-in security rulesets (NPE, thread-safety, XSS, SQL injection)._\n* [deepsec](https://github.com/vercel-labs/deepsec) - _Agent-powered vulnerability scanner by Vercel Labs for finding hard-to-spot issues in large codebases using coding agents. Supports parallel scanning, PR diff review, and CI/CD integration._\n* [defending-code-reference-harness](https://github.com/anthropics/defending-code-reference-harness) - _Reference implementation for autonomous vulnerability discovery and remediation using Claude. Includes threat modeling, scanning, triage, and patching skills._\n* [Visa Vulnerability Agentic Harness](https://github.com/visa/visa-vulnerability-agentic-harness) - _Agentic SAST pipeline with 11 stages from detection through LLM-powered remediation and adversarial validation. Outputs SARIF, integrates with Claude Code, Copilot, and Gemini._\n* [GhidraGPT](https://github.com/ZeroDaysBroker/GhidraGPT) - _Integrates GPT models into Ghidra for automated code analysis, variable renaming, vulnerability detection, and explanation generation._\n* [IDAssist](https://github.com/symgraph/IDAssist) - _AI-Powered Reverse Engineering Plugin for IDA Pro. Integrates LLM-powered analysis into IDA's interface with semantic knowledge graphs, RAG document search, and support for multiple LLM providers (OpenAI, Anthropic, Ollama, LiteLLM). Analyzes functions, suggests renames, answers questions about code._\n* [ThreatForest](https://github.com/aws-samples/sample-agentic-attack-tree-generator) - _Agentic threat modeling platform built on Strands framework. Autonomously generates attack trees from repositories, maps attack steps to MITRE ATT\u0026CK techniques, and produces actionable mitigation recommendations._\n* [claude-grc-plugin](https://github.com/mlunato47/claude-grc-plugin) - _Claude Code plugin that turns Claude into a senior GRC analyst. 72+ reference files covering 15 frameworks (NIST 800-53, FedRAMP, ISO 27001, SOC 2, etc.), 24 slash commands, and deep domain knowledge for federal and commercial compliance work._\n* [Vigil SOC](https://github.com/Vigil-SOC/vigil) - _A comprehensive open-source security operations platform for AI agents, enabling real-time monitoring, threat detection, and incident response for AI-powered environments._\n* [AiSOC](https://github.com/beenuar/AiSOC) - _Open-source, self-hostable AI-powered SOC that ingests security events, correlates them, runs autonomous AI-driven investigations via LangGraph, and surfaces results in a unified console. Features full agent decision audit trail, public eval harness in CI, and 52 first-party connectors. MIT licensed._\n\n### Privacy \u0026 Confidential Computing\n* [Python Differential Privacy Library](https://github.com/OpenMined/PyDP)\n* [Diffprivlib](https://github.com/IBM/differential-privacy-library) - _The IBM Differential Privacy Library_\n* [TenSEAL](https://github.com/OpenMined/TenSEAL) - _A library for doing homomorphic encryption operations on tensors_\n* [SyMPC](https://github.com/OpenMined/SyMPC) - _A Secure Multiparty Computation companion library for Syft_\n* [Cloaked AI](https://ironcorelabs.com/products/cloaked-ai/) - _Open source property-preserving encryption for vector embeddings_\n* [dstack](https://github.com/Dstack-TEE/dstack) - _Open-source confidential AI framework for secure ML/LLM deployment with hardware-enforced isolation and data privacy_\n* [PrivacyRaven](https://github.com/trailofbits/PrivacyRaven) - _privacy testing library for deep learning systems_\n* [PLOT4ai](https://plot4.ai/) - _Privacy Library Of Threats 4 Artificial Intelligence — A threat modeling library to help you build responsible AI_\n* [OpenDP](https://github.com/opendp/opendp) - _Core library for differential privacy algorithms from the OpenDP project — used to build privacy-preserving ML training pipelines_\n\n### Data \u0026 Supply Chain Security\n* [datasig](https://github.com/trailofbits/datasig) - _Dataset fingerprinting for AIBOM_\n* [OWASP AIBOM](https://github.com/OWASP/www-project-aibom) - _AI Bill of Materials_\n* [Trusera ai-bom](https://github.com/Trusera/ai-bom) - _AI Bill of Materials — discover every AI agent, model, and API in your infrastructure_\n\n## Agentic AI Security Skills\n\n* [Elastic Agent Skills](https://github.com/elastic/agent-skills) - _Collection of skills for Elastic's AI assistant, enabling natural language security investigations across logs, traces, and threat intelligence_\n* [Ghost Security Skills](https://github.com/ghostsecurity/skills) - _Agent application security (appsec) skills and tools for Claude Code_\n* [tm_skills](https://github.com/izar/tm_skills) - _Agent skills to help with Continuous Threat Modeling_\n* [Trail of Bits Skills Marketplace](https://github.com/trailofbits/skills) - _Trail of Bits Claude Code skills for security research, vulnerability detection, and audit workflows_\n* [Semgrep Skills](https://github.com/semgrep/skills) - _Official Semgrep skills for Claude Code and other AI coding assistants. Provides security scanning, code analysis, and vulnerability detection capabilities directly in your AI-assisted development workflow._\n* [claude-bug-bounty](https://github.com/shuvonsec/claude-bug-bounty) - _Claude Code skill for AI-assisted bug bounty hunting. Automates reconnaissance, IDOR, XSS, SSRF, OAuth, GraphQL, and LLM injection testing with 4-gate validation checklist and report generation._\n* [Anthropic Cybersecurity Skills](https://github.com/mukul975/Anthropic-Cybersecurity-Skills) - _734+ structured cybersecurity skills for AI agents. MITRE ATT\u0026CK mapped, agentskills.io standard. Compatible with Claude Code, Copilot, Codex CLI, Cursor, and Gemini CLI._\n* [llm-sast-scanner](https://github.com/SunWeb3Sec/llm-sast-scanner) - _SAST skill for AI coding agents (Claude Code, Codex, etc.) with structured vulnerability detection across 34 classes. Features source-to-sink taint analysis, Judge verification for false positive reduction, and 99%+ precision/recall on benchmarks._\n* [sast-skills](https://github.com/utkusen/sast-skills) - _Collection of agent skills that turn your AI coder into a SAST scanner_\n* [pentest-ai-agents](https://github.com/0xSteph/pentest-ai-agents) - _31 Claude Code subagents for offensive security. Specialized AI subagents for recon, web, AD, cloud, mobile, wireless, social engineering, payload crafting, reverse engineering, exploit chaining, detection engineering, forensics, and report generation. Tier 2 agents can execute tools directly with approval gates._\n* [Mantis Skills](https://github.com/google/mantis) - _Google's decoupled, sequential, security-focused pipeline of agentic AI skills for autonomously reviewing, deduplicating, validating, reproducing, and patching vulnerabilities across codebases of any scale. Features a multi-stage pipeline (architecture analysis → threat modeling → research → review → reproduction → patching → calibration → reflection), built-in sandboxing, and a continuous learning loop that adapts across iterative runs. Supports RTL hardware, IaC, ML pipelines, and compiled binaries._\n* [USAP](https://github.com/jaskaranhundal/usap-skills) - _80 cybersecurity skills + 13 orchestrator agents with a typed 11-field output contract, an enforced resolvable-evidence gate (no verdict without a resolvable source), connector-agnostic MCP routing to downstream security tools, and human-approval gating for every mutating action. MITRE ATT\u0026CK + NIST CSF 2.0 mapped. Runs in Claude Code, Cursor, Codex, and Gemini._\n* [reverse-skill](https://github.com/zhaoxuya520/reverse-skill) - _Cybersecurity skill router for AI coding clients (Claude Code, Codex, Cursor, OpenCode). Routes APK/binary/JS reverse engineering, pentest, malware, CTF, and firmware tasks to the right methodology with 43 routing rules, on-demand toolchain bootstrap, and a 173-case routing benchmark._\n\n## Security-Focused AI Models\n\n* [VulnLLM-R-7B](https://huggingface.co/UCSB-SURFI/VulnLLM-R-7B) - _Specialized reasoning LLM for vulnerability detection. Uses Chain-of-Thought reasoning to analyze data flow, control flow, and security context. Outperforms Claude-3.7-Sonnet and CodeQL on vulnerability detection benchmarks. Only 7B parameters making it efficient and fast._\n* [Foundation-Sec-8B-Reasoning](https://huggingface.co/fdtn-ai/Foundation-Sec-8B-Reasoning) - _Llama-3.1-FoundationAI-SecurityLLM-8B-Reasoning is an open-weight, 8-billion parameter instruction-tuned language model specialized for cybersecurity applications. It extends the Foundation-Sec-8B base model with instruction-following and reasoning capabilities._\n* [Antares (1B \u0026 350M)](https://huggingface.co/collections/fdtn-ai/antares) - _Agentic models for vulnerability localization by fdtn-ai. Available in 2B and 0.4B parameter variants, designed to autonomously identify and localize vulnerabilities in code._\n* [CyberSecQwen-4B](https://huggingface.co/lablab-ai-amd-developer-hackathon/CyberSecQwen-4B) - _4B-parameter CTI specialist fine-tuned from Qwen3-4B-Instruct-2507 for cybersecurity threat intelligence.\n* [Meta-SecAlign-8B / Meta-SecAlign-70B](https://github.com/facebookresearch/Meta_SecAlign) - _Security-aligned Llama models fine-tuned to resist prompt injection attacks, maintaining instruction hierarchy even under adversarial inputs_\n* [Lily-Cybersecurity-7B](https://huggingface.co/segolilylabs/Lily-Cybersecurity-7B-v0.2-GGUF) - _Cybersecurity-tuned 7B chat model optimized for security analysis, vulnerability explanation, and threat intelligence tasks._\n\n### Cybersecurity-Tuned Text Generation\n\n* [ZySec-7B / SecurityLLM](https://huggingface.co/ZySec-AI/SecurityLLM) - _7B cybersecurity-specialized chat model trained for security Q\u0026A, threat analysis, and compliance guidance._\n\n### Safety Classifiers \u0026 Prompt Injection Detection\n\n* [Llama-Guard-4-12B](https://huggingface.co/meta-llama/Llama-Guard-4-12B) - _Meta's latest multimodal safety classifier for detecting harmful content in LLM inputs and outputs across text and image modalities._\n* [Llama-Prompt-Guard-2-86M](https://huggingface.co/meta-llama/Llama-Prompt-Guard-2-86M) - _Lightweight 86M parameter model from Meta for detecting prompt injection and jailbreak attempts in production LLM pipelines._\n* [ShieldGemma-2B](https://huggingface.co/google/shieldgemma-2b) - _Google's 2B parameter text safety classifier for detecting harmful content, built on the Gemma architecture._\n* [Shieldstral-1.0-3B](https://huggingface.co/mistralai/Shieldstral-1.0-3B) - _Mistral's 3B-parameter policy-adaptive multimodal safety classifier. Evaluates text, image, or text+image content against a natural-language safety policy as a single yes/no pass, re-targetable to new policies at inference without retraining, multilingual._\n* [DeBERTa Prompt Injection Detector v2](https://huggingface.co/protectai/deberta-v3-base-prompt-injection-v2) - _Protect AI's DeBERTa-v3-base fine-tuned for prompt injection detection, widely used in production LLM guardrail pipelines._\n* [Prompt Injection Sentinel](https://huggingface.co/qualifire/prompt-injection-sentinel) - _ModernBERT-large model fine-tuned for prompt injection and jailbreak classification with low false-positive rate._\n* [Nemotron 3.5 Content Safety](https://huggingface.co/nvidia/Nemotron-3.5-Content-Safety) - _NVIDIA's 4B-parameter multimodal content safety model unifying text+image input guardrails, multilingual support (35+ languages), customizable enterprise policy enforcement, and auditable reasoning in a single inference call. Successor to Nemotron 3 Content Safety._\n* [BrowseSafe](https://huggingface.co/perplexity-ai/browsesafe) - _Specialized security model for detecting prompt injection attacks in AI browser agents. Achieves 90.4% F1 score on BrowseSafe-Bench, optimized for real-time async classification of raw HTML content._\n\n### Domain-Adapted Security Language Models\n\n* [ATTACK-BERT](https://huggingface.co/basel/ATTACK-BERT) - _Sentence-transformer model for mapping security text to MITRE ATT\u0026CK techniques._\n* [CySecBERT](https://huggingface.co/markusbayer/CySecBERT) - _BERT model adapted for cybersecurity and CTI tasks through domain-specific pre-training._\n* [VulBERTa](https://github.com/ICL-ml4csec/VulBERTa) - _RoBERTa-based model pre-trained on large C/C++ code corpora and fine-tuned for vulnerability detection. Outperforms prior BERT-based approaches on multiple CVE benchmarks_\n\n## Datasets\n\n* [SafetyPrompts](https://safetyprompts.com/) - _Curated collection of safety-relevant prompts for evaluating LLM safety and security properties._\n* [Do-Not-Answer](https://github.com/Libr-AI/do-not-answer) - _Dataset of prompts that responsible LLMs should not answer, for safety evaluation and red teaming._\n* [JailBreakV-28K](https://github.com/SaFoLab-WISC/JailBreakV_28K) - _Large-scale dataset of 28,000 jailbreak prompts for benchmarking LLM safety._\n* [CL4R1T4S](https://github.com/elder-plinius/CL4R1T4S) - _Leaked system prompts from ChatGPT, Claude, Gemini, Grok, Perplexity, Cursor, Lovable, Replit, and other major AI tools. Largest known collection of production system prompts for transparency and attack-surface research._\n* [LEAKHUB](https://github.com/elder-plinius/LEAKHUB) - _System Prompt Leak Leaderboard — community platform for tracking and ranking system prompt leaks across AI products._\n* [Leaked System Prompts](https://github.com/x1xhlol/system-prompts-and-models-of-ai-tools) - _Collection of leaked system prompts from commercial AI tools — useful for understanding real-world prompt engineering and attack surfaces._\n* [SemGuard Arabic Security Dataset](https://huggingface.co/datasets/AG-31625874/SemGuard-Dataset) - _First validated Arabic/Arabizi LLM security dataset — 807 examples across 7 threat categories, validated via 3-judge LLM-as-Judge pipeline (GPT-4o, Grok-4, Llama 3.3 70B), Fleiss' κ=0.839. Includes a 527-example inter-judge disagreement corpus._\n  \n","projects_url":"https://awesome.ecosyste.ms/api/v1/lists/ottosulin%2Fawesome-ai-security/projects"}