{"id":43061050,"url":"https://github.com/sattyamjjain/agent-airlock","last_synced_at":"2026-06-11T18:00:40.590Z","repository":{"id":335642118,"uuid":"1146531442","full_name":"sattyamjjain/agent-airlock","owner":"sattyamjjain","description":"Open-source security firewall for AI agents — validates tool calls, strips ghost arguments, enforces type safety, PII masking, RBAC, cost tracking \u0026 sandbox isolation. Works with LangChain, OpenAI Agents SDK, PydanticAI \u0026 CrewAI.","archived":false,"fork":false,"pushed_at":"2026-06-08T17:58:39.000Z","size":3007,"stargazers_count":8,"open_issues_count":6,"forks_count":3,"subscribers_count":1,"default_branch":"main","last_synced_at":"2026-06-08T19:25:00.326Z","etag":null,"topics":["ai-agents","ai-security","crewai","firewall","langchain","llm-safety","mcp-server","openai","pii-masking","pydantic-ai","python","rbac","sandbox","tool-validation","zero-trust"],"latest_commit_sha":null,"homepage":null,"language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/sattyamjjain.png","metadata":{"files":{"readme":"README.md","changelog":"CHANGELOG.md","contributing":"CONTRIBUTING.md","funding":null,"license":"LICENSE","code_of_conduct":"CODE_OF_CONDUCT.md","threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":"SECURITY.md","support":null,"governance":null,"roadmap":"ROADMAP_2026.md","authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":"AGENTS.md","dco":null,"cla":null}},"created_at":"2026-01-31T08:43:37.000Z","updated_at":"2026-06-08T17:58:30.000Z","dependencies_parsed_at":"2026-05-16T20:02:17.891Z","dependency_job_id":null,"html_url":"https://github.com/sattyamjjain/agent-airlock","commit_stats":null,"previous_names":["sattyamjjain/agent-airlock"],"tags_count":51,"template":false,"template_full_name":null,"purl":"pkg:github/sattyamjjain/agent-airlock","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/sattyamjjain%2Fagent-airlock","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/sattyamjjain%2Fagent-airlock/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/sattyamjjain%2Fagent-airlock/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/sattyamjjain%2Fagent-airlock/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/sattyamjjain","download_url":"https://codeload.github.com/sattyamjjain/agent-airlock/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/sattyamjjain%2Fagent-airlock/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":34211067,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-05-26T15:22:16.424Z","status":"online","status_checked_at":"2026-06-11T02:00:06.485Z","response_time":57,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["ai-agents","ai-security","crewai","firewall","langchain","llm-safety","mcp-server","openai","pii-masking","pydantic-ai","python","rbac","sandbox","tool-validation","zero-trust"],"created_at":"2026-01-31T12:07:25.303Z","updated_at":"2026-06-11T18:00:40.574Z","avatar_url":"https://github.com/sattyamjjain.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"\u003cdiv align=\"center\"\u003e\n\n\u003c!-- Animated Typing Header --\u003e\n\u003ca href=\"https://github.com/sattyamjjain/agent-airlock\"\u003e\n  \u003cimg src=\"https://readme-typing-svg.demolab.com?font=Fira+Code\u0026weight=700\u0026size=28\u0026duration=3000\u0026pause=1000\u0026color=00D4FF\u0026center=true\u0026vCenter=true\u0026multiline=true\u0026repeat=true\u0026width=700\u0026height=100\u0026lines=%F0%9F%9B%A1%EF%B8%8F+Agent-Airlock;Your+AI+Agent+Just+Tried+rm+-rf+%2F.+We+Stopped+It.\" alt=\"Agent-Airlock Typing Animation\" /\u003e\n\u003c/a\u003e\n\n### The Open-Source Firewall for AI Agents\n\n**One decorator. Zero trust. Full control.**\n\n\u003c!-- Primary Badges Row --\u003e\n[![PyPI version](https://img.shields.io/pypi/v/agent-airlock?style=for-the-badge\u0026logo=pypi\u0026logoColor=white\u0026color=3775A9)](https://pypi.org/project/agent-airlock/)\n[![Downloads](https://img.shields.io/pypi/dm/agent-airlock?style=for-the-badge\u0026logo=python\u0026logoColor=white\u0026color=success)](https://pypistats.org/packages/agent-airlock)\n[![CI](https://img.shields.io/github/actions/workflow/status/sattyamjjain/agent-airlock/ci.yml?style=for-the-badge\u0026logo=github\u0026label=CI\u0026color=success)](https://github.com/sattyamjjain/agent-airlock/actions/workflows/ci.yml)\n[![codecov](https://img.shields.io/codecov/c/github/sattyamjjain/agent-airlock?style=for-the-badge\u0026logo=codecov\u0026logoColor=white)](https://codecov.io/gh/sattyamjjain/agent-airlock)\n\n\u003c!-- Secondary Badges Row --\u003e\n[![Python 3.10+](https://img.shields.io/badge/python-3.10+-3776AB?style=flat-square\u0026logo=python\u0026logoColor=white)](https://www.python.org/downloads/)\n[![License: MIT](https://img.shields.io/badge/License-MIT-green?style=flat-square)](https://opensource.org/licenses/MIT)\n[![GitHub stars](https://img.shields.io/github/stars/sattyamjjain/agent-airlock?style=flat-square\u0026logo=github)](https://github.com/sattyamjjain/agent-airlock/stargazers)\n[![PRs Welcome](https://img.shields.io/badge/PRs-welcome-brightgreen.svg?style=flat-square)](http://makeapullrequest.com)\n\n\u003c!-- TEST-BADGE-START --\u003e\n\u003c!-- Auto-generated by scripts/update_test_badge.py. Do not edit by hand. --\u003e\n**Test suite:** 2,510 tests · **Coverage:** 83.42% · **v0.8.5**\n\u003c!-- TEST-BADGE-END --\u003e\n\n\u003cbr/\u003e\n\n[**Get Started in 30 Seconds**](#-30-second-quickstart) · [**Why Airlock?**](#-the-problem-no-one-talks-about) · [**All Frameworks**](#-framework-compatibility) · [**Docs**](#-documentation)\n\n\u003cbr/\u003e\n\n\u003c/div\u003e\n\n---\n\n\u003c!-- Hero Visual Block --\u003e\n\u003cdiv align=\"center\"\u003e\n\n```\n┌────────────────────────────────────────────────────────────────┐\n│  🤖 AI Agent: \"Let me help clean up disk space...\"            │\n│                           ↓                                    │\n│               rm -rf / --no-preserve-root                      │\n│                           ↓                                    │\n│  ┌──────────────────────────────────────────────────────────┐  │\n│  │  🛡️ AIRLOCK: BLOCKED                                     │  │\n│  │                                                          │  │\n│  │  Reason: Matches denied pattern 'rm_*'                   │  │\n│  │  Policy: STRICT_POLICY                                   │  │\n│  │  Fix: Use approved cleanup tools only                    │  │\n│  └──────────────────────────────────────────────────────────┘  │\n└────────────────────────────────────────────────────────────────┘\n```\n\n\u003c/div\u003e\n\n---\n\n## 🎯 30-Second Quickstart\n\n```bash\npip install agent-airlock\n```\n\n```python\nfrom agent_airlock import Airlock\n\n@Airlock()\ndef transfer_funds(account: str, amount: int) -\u003e dict:\n    return {\"status\": \"transferred\", \"amount\": amount}\n\n# LLM sends amount=\"500\" (string) → BLOCKED with fix_hint\n# LLM sends force=True (invented arg) → STRIPPED silently\n# LLM sends amount=500 (correct) → EXECUTED safely\n```\n\n**That's it.** Your function now has ghost argument stripping, strict type validation, and self-healing errors.\n\n---\n\n## 🧠 The Problem No One Talks About\n\n\u003ctable\u003e\n\u003ctr\u003e\n\u003ctd width=\"50%\"\u003e\n\n### The Hype\n\n\u003e *\"MCP has 16,000+ servers on GitHub!\"*\n\u003e *\"OpenAI adopted it!\"*\n\u003e *\"Linux Foundation hosts it!\"*\n\n\u003c/td\u003e\n\u003ctd width=\"50%\"\u003e\n\n### The Reality\n\n**LLMs hallucinate tool calls. Every. Single. Day.**\n\n- Claude invents arguments that don't exist\n- GPT-4 sends `\"100\"` when you need `100`\n- Agents chain 47 calls before one deletes prod data\n\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/table\u003e\n\n**Enterprise solutions exist:** Prompt Security ($50K/year), Pangea (proxy your data), Cisco (\"coming soon\").\n\n**We built the open-source alternative.** One decorator. No vendor lock-in. Your data never leaves your infrastructure.\n\n---\n\n## ✨ What You Get\n\n\u003ctable\u003e\n\u003ctr\u003e\n\u003ctd align=\"center\" width=\"16%\"\u003e\n\u003cimg width=\"40\" src=\"https://img.icons8.com/fluency/48/delete-shield.png\" alt=\"shield\"/\u003e\n\u003cbr/\u003e\u003cb\u003eGhost Args\u003c/b\u003e\n\u003cbr/\u003e\u003csub\u003eStrip LLM-invented params\u003c/sub\u003e\n\u003c/td\u003e\n\u003ctd align=\"center\" width=\"16%\"\u003e\n\u003cimg width=\"40\" src=\"https://img.icons8.com/fluency/48/checked.png\" alt=\"check\"/\u003e\n\u003cbr/\u003e\u003cb\u003eStrict Types\u003c/b\u003e\n\u003cbr/\u003e\u003csub\u003eNo silent coercion\u003c/sub\u003e\n\u003c/td\u003e\n\u003ctd align=\"center\" width=\"16%\"\u003e\n\u003cimg width=\"40\" src=\"https://img.icons8.com/fluency/48/refresh.png\" alt=\"refresh\"/\u003e\n\u003cbr/\u003e\u003cb\u003eSelf-Healing\u003c/b\u003e\n\u003cbr/\u003e\u003csub\u003eLLM-friendly errors\u003c/sub\u003e\n\u003c/td\u003e\n\u003ctd align=\"center\" width=\"16%\"\u003e\n\u003cimg width=\"40\" src=\"https://img.icons8.com/fluency/48/lock.png\" alt=\"lock\"/\u003e\n\u003cbr/\u003e\u003cb\u003eE2B Sandbox\u003c/b\u003e\n\u003cbr/\u003e\u003csub\u003eIsolated execution\u003c/sub\u003e\n\u003c/td\u003e\n\u003ctd align=\"center\" width=\"16%\"\u003e\n\u003cimg width=\"40\" src=\"https://img.icons8.com/fluency/48/user-shield.png\" alt=\"user\"/\u003e\n\u003cbr/\u003e\u003cb\u003eRBAC\u003c/b\u003e\n\u003cbr/\u003e\u003csub\u003eRole-based access\u003c/sub\u003e\n\u003c/td\u003e\n\u003ctd align=\"center\" width=\"16%\"\u003e\n\u003cimg width=\"40\" src=\"https://img.icons8.com/fluency/48/privacy.png\" alt=\"privacy\"/\u003e\n\u003cbr/\u003e\u003cb\u003ePII Mask\u003c/b\u003e\n\u003cbr/\u003e\u003csub\u003eAuto-redact secrets\u003c/sub\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"center\" width=\"16%\"\u003e\n\u003cimg width=\"40\" src=\"https://img.icons8.com/fluency/48/network-card.png\" alt=\"network\"/\u003e\n\u003cbr/\u003e\u003cb\u003eNetwork Guard\u003c/b\u003e\n\u003cbr/\u003e\u003csub\u003eBlock data exfiltration\u003c/sub\u003e\n\u003c/td\u003e\n\u003ctd align=\"center\" width=\"16%\"\u003e\n\u003cimg width=\"40\" src=\"https://img.icons8.com/fluency/48/folder-invoices.png\" alt=\"folder\"/\u003e\n\u003cbr/\u003e\u003cb\u003ePath Validation\u003c/b\u003e\n\u003cbr/\u003e\u003csub\u003eCVE-resistant traversal\u003c/sub\u003e\n\u003c/td\u003e\n\u003ctd align=\"center\" width=\"16%\"\u003e\n\u003cimg width=\"40\" src=\"https://img.icons8.com/fluency/48/restart.png\" alt=\"circuit\"/\u003e\n\u003cbr/\u003e\u003cb\u003eCircuit Breaker\u003c/b\u003e\n\u003cbr/\u003e\u003csub\u003eFault tolerance\u003c/sub\u003e\n\u003c/td\u003e\n\u003ctd align=\"center\" width=\"16%\"\u003e\n\u003cimg width=\"40\" src=\"https://img.icons8.com/fluency/48/analytics.png\" alt=\"otel\"/\u003e\n\u003cbr/\u003e\u003cb\u003eOpenTelemetry\u003c/b\u003e\n\u003cbr/\u003e\u003csub\u003eEnterprise observability\u003c/sub\u003e\n\u003c/td\u003e\n\u003ctd align=\"center\" width=\"16%\"\u003e\n\u003cimg width=\"40\" src=\"https://img.icons8.com/fluency/48/money-bag.png\" alt=\"cost\"/\u003e\n\u003cbr/\u003e\u003cb\u003eCost Tracking\u003c/b\u003e\n\u003cbr/\u003e\u003csub\u003eBudget limits\u003c/sub\u003e\n\u003c/td\u003e\n\u003ctd align=\"center\" width=\"16%\"\u003e\n\u003cimg width=\"40\" src=\"https://img.icons8.com/fluency/48/syringe.png\" alt=\"vaccine\"/\u003e\n\u003cbr/\u003e\u003cb\u003eVaccination\u003c/b\u003e\n\u003cbr/\u003e\u003csub\u003eAuto-secure frameworks\u003c/sub\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/table\u003e\n\n---\n\n## 📋 Table of Contents\n\n\u003cdetails\u003e\n\u003csummary\u003e\u003cb\u003eClick to expand full navigation\u003c/b\u003e\u003c/summary\u003e\n\n- [30-Second Quickstart](#-30-second-quickstart)\n- [The Problem](#-the-problem-no-one-talks-about)\n- [What You Get](#-what-you-get)\n- [Core Features](#-core-features)\n  - [E2B Sandbox](#-e2b-sandbox-execution)\n  - [Security Policies](#-security-policies)\n  - [Cost Control](#-cost-control)\n  - [PII Masking](#-pii--secret-masking)\n  - [Network Airgap](#-network-airgap-v030)\n  - [Framework Vaccination](#-framework-vaccination-v030)\n  - [Circuit Breaker](#-circuit-breaker-v040)\n  - [OpenTelemetry](#-opentelemetry-observability-v040)\n- [Framework Compatibility](#-framework-compatibility)\n- [FastMCP Integration](#-fastmcp-integration)\n- [Comparison](#-why-not-enterprise-vendors)\n- [Installation](#-installation)\n- [OWASP Compliance](#️-owasp-compliance)\n- [Performance](#-performance)\n- [Documentation](#-documentation)\n- [Contributing](#-contributing)\n- [Support](#-support)\n\n\u003c/details\u003e\n\n---\n\n## 🔥 Core Features\n\n### 🔒 E2B Sandbox Execution\n\n```python\nfrom agent_airlock import Airlock, STRICT_POLICY\n\n@Airlock(sandbox=True, sandbox_required=True, policy=STRICT_POLICY)\ndef execute_code(code: str) -\u003e str:\n    \"\"\"Runs in an E2B Firecracker MicroVM. Not on your machine.\"\"\"\n    exec(code)\n    return \"executed\"\n```\n\n| Feature | Value |\n|---------|-------|\n| Boot time | ~125ms cold, \u003c200ms warm |\n| Isolation | Firecracker MicroVM |\n| Fallback | `sandbox_required=True` blocks local execution |\n\nAir-gapped / on-prem? `DockerBackend` is the supported alternative\n— `cap_drop=[\"ALL\"]`, `no-new-privileges`, `network_mode=\"none\"`,\ntimeout enforced, opt-in `pytest -m docker` integration tests. See\n[`docs/sandbox/docker.md`](docs/sandbox/docker.md).\n\n#### ModalBackend — Modal-hosted sandbox *(v0.8.11+, issue #30)*\n\nAlready running the rest of your agent on [Modal](https://modal.com/)?\n`ModalBackend` lets you keep airlocked tool execution on the same\nsubstrate instead of mixing E2B and Modal billing / observability.\n\n```bash\npip install \"agent-airlock[modal]\"\n```\n\n```python\nfrom agent_airlock import Airlock, STRICT_POLICY, AirlockConfig\nfrom agent_airlock.sandbox_backend import ModalBackend\n\nbackend = ModalBackend(\n    app_name=\"my-airlock-sandbox\",\n    image_ref=\"python:3.11-slim\",\n    cpu=0.5,\n    memory_mb=512,\n    timeout_s=30,\n    # network_policy=None  → block_network=True (fail-closed default)\n)\n\n@Airlock(sandbox=True, sandbox_required=True, policy=STRICT_POLICY,\n         config=AirlockConfig(sandbox_backend=backend))\ndef execute_code(code: str) -\u003e str:\n    exec(code)\n    return \"executed\"\n```\n\n**Isolation model — read before you reach for `cap_drop`.** Modal\nsandboxes run under **gVisor** (kernel-syscall filtering), not under\nDocker-style capability dropping. The Modal Python SDK does not\nexpose `cap_drop` / `cap_add` / `seccomp` / `no-new-privileges` —\nthere is no equivalent knob to map. If your threat model needs\nLinux-capability dropping at the container layer, keep using\n`DockerBackend`. The network posture *is* configurable: `ModalBackend`\ndefaults to `block_network=True` (deny-by-default), and a supplied\n`NetworkPolicy` maps to Modal's `block_network` flag (`allow_egress=False`\n→ blocked, `True` → allowed). Hostname allowlists in `NetworkPolicy.allowed_hosts`\ndo **not** forward to Modal (their API is CIDR-only); the backend logs\na structlog warning and the operator is expected to re-state hostname\nconstraints at the `Airlock` policy layer.\n\n`ModalBackend` is **opt-in only** — it is NOT added to the\n`get_default_backend()` priority chain (E2B → Docker → Local stays the\ndefault flow). Existing callers see no behavior change.\n\n---\n\n### 📜 Security Policies\n\n| Preset | Use case | Key posture |\n|---|---|---|\n| `PERMISSIVE_POLICY` | Dev / sandbox | No restrictions |\n| `STRICT_POLICY` | Prod | Rate-limited, requires agent identity, denies dangerous capabilities |\n| `READ_ONLY_POLICY` | Analytics / RAG | `read_*` / `get_*` / `list_*` / `search_*` only |\n| `BUSINESS_HOURS_POLICY` | Compliance windows | `delete_*` / `drop_*` / `*_production` only 09:00–17:00 |\n| `CAMOUFLAGE_RESISTANT_POLICY` *(v0.8.6)* | Detector-independent defense vs. domain-camouflaged injection | Deny-by-default allowlist, ghost-arg BLOCK, output cap, per-call reauthorization |\n\n```python\nfrom agent_airlock import (\n    PERMISSIVE_POLICY,\n    STRICT_POLICY,\n    READ_ONLY_POLICY,\n    BUSINESS_HOURS_POLICY,\n    CAMOUFLAGE_RESISTANT_POLICY,  # v0.8.6\n)\n\n# Or build your own:\nfrom agent_airlock import SecurityPolicy\n\nMY_POLICY = SecurityPolicy(\n    allowed_tools=[\"read_*\", \"query_*\"],\n    denied_tools=[\"delete_*\", \"drop_*\", \"rm_*\"],\n    rate_limits={\"*\": \"1000/hour\", \"write_*\": \"100/hour\"},\n    time_restrictions={\"deploy_*\": \"09:00-17:00\"},\n)\n```\n\n#### CAMOUFLAGE_RESISTANT — detector-independent injection defense (v0.8.6)\n\n[arXiv:2605.22001](https://arxiv.org/abs/2605.22001) (\"Blind Spots in\nthe Guard\", Pai, May 2026) shows that production injection detectors —\n**Llama Guard 3 included** — drop to **IDR = 0.000** on payloads that\nmimic the target document's domain vocabulary and authority structure.\nPer the paper, detection rates collapse from 93.8% to 9.7% on Llama 3.1\n8B and from 100% to 55.6% on Gemini 2.0 Flash.\n\n`CAMOUFLAGE_RESISTANT_POLICY` does not rely on payload-content\nsignatures at all. It blocks at four structural seams an attacker has\nto ride regardless of phrasing:\n\n1. **Deny-by-default tool allowlist.** Empty `allowed_tools` means\n   *nothing* is callable; deployments opt every tool in by name. A\n   camouflaged directive targeting an unlisted tool is blocked on\n   allowlist grounds without ever invoking a detector.\n2. **Ghost-argument BLOCK.** A camouflaged directive cannot smuggle\n   undeclared parameters past validation.\n3. **Hard output cap + sanitization.** Tool output that re-enters the\n   model context is truncated and PII/secret-masked so a camouflaged\n   directive embedded in tool output can't carry into a downstream\n   agent at full length.\n4. **Per-call reauthorization (debate-amplification guard).** Once a\n   tool's output has flowed back into the model, any reinvocation\n   requires an explicit `context.authorize_once(tool)` grant from the\n   harness — breaking the multi-agent fan-out path the paper identifies.\n\n```python\nfrom agent_airlock import Airlock, apply_camouflage_resistant\n\nbundle = apply_camouflage_resistant(allowed_tools=[\"read_file\", \"search\"])\n\n@Airlock(config=bundle.config, policy=bundle.policy)\ndef read_file(path: str) -\u003e str:\n    ...\n```\n\n`apply_camouflage_resistant()` composes the matching `AirlockConfig`\n(unknown-args BLOCK, sanitization on, output cap 4000 chars) with a\n`SecurityPolicy` carrying your explicit allowlist. The preset is\ndeliberately incomplete on its own — the config-level knobs and the\npolicy-level knobs span two seams, so the factory returns both as a\n`CamouflageResistantBundle`.\n\n\u003e **Running an MCP server with STDIO transport?** Also wire the\n\u003e [Ox MCP STDIO sanitizer](#️-owasp-compliance) via\n\u003e `stdio_guard_ox_defaults()` — it blocks the entire\n\u003e CVE-2026-30616 class (shell metacharacter injection,\n\u003e non-allowlisted binaries, Trojan-Source RTL overrides, and\n\u003e inline-code flags) before `subprocess.Popen`.\n\n---\n\n### 🪪 MCP server attestation (v0.8.10)\n\n[arXiv:2605.24248](https://arxiv.org/abs/2605.24248) (\"Attested\nTool-Server Admission\", Metere, May 2026) calls out a gap MCP itself\ndoes not close: the protocol standardises *message exchange* between\nLLM agents and tool servers but says nothing about *trust*. Anybody who\ncan answer on the wire can declare themselves a tool server.\n\n`mcp_attested_admission_defaults()` is a deny-by-default opt-in preset\nthat closes the gap host-side, mirroring the paper's three additive\nmechanisms:\n\n1. **Offline-signed clearance assertion.** Before any tool from an MCP\n   server is dispatched, the host fetches a JWS-compact clearance from\n   `{server_url}/.well-known/mcp-clearance` (path is configurable) and\n   verifies its signature against an **operator-pinned trust root**.\n   The trust root is supplied to `AttestedAdmissionConfig` at process\n   startup — never network-fetched on the hot path.\n2. **Deny-by-default per-server tool allowlist.** Admitting a server\n   is not the same as trusting its every tool. The verified clearance\n   carries an explicit list of tool names the host will permit;\n   everything else is denied. The `sub` claim is matched against the\n   server identity the host is about to dispatch to (so a stolen\n   clearance from server A can't admit a tool call to server B).\n3. **Flavor-gated enforcement.** `ENFORCE` (default) hard-denies on\n   missing / invalid / expired clearance; `WARN` logs and admits — the\n   staged turn-up an operator wants when introducing the gate against\n   real traffic.\n\nEvery admission decision emits a\n[`ReceiptVerdict`](./docs/attest/receipt.md) on the\n`guard=\"mcp_attested_admission\"` channel, so the existing `airlock attest`\nDSSE pipeline picks decisions up unchanged — this preset does **not**\ninvent a new log.\n\n```python\nfrom agent_airlock.mcp_proxy_guard import MCPProxyConfig, MCPProxyGuard\nfrom agent_airlock.mcp_spec.attested_admission import TrustRoot\nfrom agent_airlock.policy_presets import mcp_attested_admission_defaults\n\n# Operator pins the trust root at startup. Never fetched at runtime.\nwith open(\"/etc/airlock/mcp-clearance-root.pem\", \"rb\") as fh:\n    pinned_pem = fh.read()\n\ncfg = mcp_attested_admission_defaults(\n    trust_root=TrustRoot(key_id=\"ops-2026Q2\", ed25519_pem=pinned_pem),\n    enforcement_mode=\"ENFORCE\",       # deny-by-default\n    max_clearance_age_days=30,\n)\nguard = MCPProxyGuard(MCPProxyConfig(attested_admission=cfg))\n\ndecision = guard.audit_tool_admission(\n    server_url=\"https://mcp.example.com\",\n    server_id=\"srv-alpha\",            # expected `sub` claim\n    tool_name=\"read\",\n)\nif not decision.admitted:\n    raise RuntimeError(decision.reason)\n```\n\nSignature verification needs the `[attested]` extra (pulls in\n`cryptography` for offline Ed25519 / RSA-PSS / JWKS verification); the\nbase install stays zero-runtime-dep.\n\n\u003e Install with `pip install \"agent-airlock[attested]\"`. Opt-in only —\n\u003e existing callers that don't set `attested_admission` get exactly\n\u003e v0.8.9 behavior.\n\n---\n\n### 🧭 Behavioral sequence guard (v0.8.12)\n\nWatches the **ordered stream of tool calls** in a session and flags\ndivergence from a declared expected order — *not* the model's stated\nreasoning trace.\n\n[arXiv:2605.27901](https://arxiv.org/abs/2605.27901) (\"The Fragility\nof Chain-of-Thought Monitoring\", Onyame, Zhou, Thopalli, Kailkhura\n\u0026 Agarwal, May 2026) reports an average **95.9% CoT unfaithfulness\nacross 8B–120B-parameter models** — including answer-switching,\npost-hoc rationalisation, and procedural exploitation of hints.\nTrusting the model's stated reasoning to detect misbehavior is\ntherefore not viable. Trusting its **behavior** — the sequence of\ntools it actually invokes — is.\n\n`SequenceGuard` is an opt-in field on `SecurityPolicy` that runs in\nthe `@Airlock` seam **right after** the standard policy check, in two\nmodes:\n\n**DECLARED mode** — operator supplies a permitted-transition DAG.\nAny transition not in the DAG is a `SequenceViolation`. Deny-by-default.\n\n```python\nfrom agent_airlock import Airlock, SecurityPolicy\nfrom agent_airlock.sequence_guard import SequenceGuard, ENTRY_SENTINEL\n\npolicy = SecurityPolicy(\n    sequence_guard=SequenceGuard(\n        mode=\"declared\",\n        action=\"block\",                       # or \"warn\"\n        dag={\n            ENTRY_SENTINEL: {\"read\"},         # only `read` may start a session\n            \"read\": {\"read\", \"summarize\"},    # after read, either re-read or summarize\n            \"summarize\": {\"send\"},            # after summarize, only send\n            \"send\": set(),                    # send is terminal\n        },\n    ),\n)\n```\n\n**BASELINE mode** — guard maintains a per-session-key Markov transition\nprofile in a local JSON file (no cloud, no PII — only tool names and\nSHA-256 *shape hashes* of `(arg types, kwarg names+types)`, **never\nargument values**) and flags transitions with observed\n`P(curr | prev) \u003c threshold` once the sample size from `prev` reaches\n`min_baseline_samples`.\n\n```python\nfrom pathlib import Path\nfrom agent_airlock.sequence_guard import SequenceGuard\n\npolicy = SecurityPolicy(\n    sequence_guard=SequenceGuard(\n        mode=\"baseline\",\n        baseline_path=Path(\"/var/lib/airlock/sequence-baseline.json\"),\n        low_probability_threshold=0.05,   # flag the bottom 5%\n        min_baseline_samples=50,          # don't flag until 50 obs from `prev`\n    ),\n)\n```\n\nEvery flagged transition emits OTel span attributes on the current\nspan (`airlock.sequence_guard.mode`, `.from_tool`, `.to_tool`,\n`.session_key`, `.observed_probability`) via the existing\n`observability` provider — telemetry failures are swallowed so they\ncannot break enforcement.\n\n**Not** `AnomalyDetector` (that's rate / endpoint-diversity /\nerror-rate / consecutive-blocked over sliding windows). `SequenceGuard`\nis per-transition ORDER signal. Run both for layered coverage. **Not**\na chain-of-thought monitor — by construction.\n\n\u003e Strictly opt-in. The new `SecurityPolicy.sequence_guard` field\n\u003e defaults to `None`; callers that don't set it get exactly v0.8.11\n\u003e behavior. Zero new runtime deps — Pydantic-only core stays intact.\n\n---\n\n### 🛑 Action-time contradiction gate (v0.8.15)\n\n[arXiv:2605.27157](https://arxiv.org/abs/2605.27157) (\"Detecting Is\nNot Resolving: The Monitoring Control Gap in Retrieval Augmented\nLLMs\", Yu et al., 2026) shows that LLMs **readily acknowledge\ncontradictory evidence** in their reasoning trace yet \"this awareness\nfails to constrain their final recommendations\". The deficit is at\n*action selection* — single-turn diagnostics overestimate RAG safety,\nand detection alone is not a control.\n\n`ActionContradictionGate` is an opt-in policy hook that wraps three\n**pluggable detectors** (any one trips) and a **privileged-sink glob\nset**. When a detector trips AND the dispatched tool matches a\nprivileged sink AND the harness has not issued an explicit allow,\nthe gate blocks the call (or warns, depending on `action=`).\n\nThe explicit-allow primitive **is not new** — the gate reuses the\nexisting `AirlockContext.authorize_once(tool_name)` (introduced for\nthe v0.8.6 reauth flow). Same one-shot grant, same semantics. After\na one-shot is consumed the gate **re-locks** — the harness must mint\na fresh `authorize_once` for each privileged action.\n\n```python\nimport re\nfrom agent_airlock import Airlock, SecurityPolicy\nfrom agent_airlock.action_contradiction_gate import ActionContradictionGate\n\npolicy = SecurityPolicy(\n    action_contradiction_gate=ActionContradictionGate(\n        # Detector 1: a boolean flag the RAG pipeline flips on after\n        # it sees an evidence-vs-claim conflict the agent discussed.\n        signal_field_key=\"evidence_contradiction\",\n        # Detector 2: pluggable regex against the SAME key when its\n        # value is a string (operator-controlled marker — never the\n        # model's full reasoning trace).\n        marker_regex=re.compile(r\"contradict|conflict|disagree\", re.I),\n        # Detector 3: fully pluggable callable; receives the context.\n        # predicate=lambda ctx: ctx.metadata.get(\"conflict_count\", 0) \u003e 1,\n        # Default privileged sinks: send_* / export_* / commit_* /\n        # transfer_* / delete_* + the v0.8.14 outbound-integration set.\n        # Operators can narrow via `privileged_sinks=(...)`.\n        action=\"block\",  # or \"warn\" for staged turn-up\n    ),\n)\n```\n\n**Off-by-default invariant.** `SecurityPolicy.action_contradiction_gate`\ndefaults to `None`; non-RAG flows pay **zero false-positive tax**\n(no detector runs, no log lines, no metadata reads). Even when wired,\nthe gate is **inert until at least one detector slot is configured**\n— so a partial roll-out (gate attached but detectors flipped off)\nadmits everything.\n\n**Not a chain-of-thought monitor.** The gate reads **operator-\ncontrolled signals** only (a metadata field, an operator regex, an\noperator predicate). It never reads the model's own claim that it\nhas or has not noticed a contradiction — the paper's whole point is\nthat those claims do not gate behavior.\n\n**Not** `sequence_guard` (v0.8.12) — that flags unusual call ORDER.\n**Not** `reauth_on_untrusted_reinvocation` (v0.8.6) — that's\ncount-driven on a per-tool counter. This gate is signal-driven and\ntargets a specific privileged-sink glob set. They compose; run all\nthree for layered coverage.\n\n\u003e Strictly opt-in. Zero new runtime deps — Pydantic-only core stays\n\u003e intact. The new `SecurityPolicy.action_contradiction_gate` field\n\u003e defaults to `None`; callers that don't set it get exactly v0.8.14\n\u003e behavior.\n\n---\n\n### 📊 Adversarial-negotiation regression harness (v0.8.17)\n\nA deterministic harness that measures what the deny-by-default\ngovernance layer does to a fixed set of adversarial buyer-seller\nnegotiation actions — and reports two metrics named to line up with an\nexternal published baseline so the numbers can sit side by side.\n\n```bash\npython -m agent_airlock.cli.negotiation_bench --report markdown\n```\n\nEach scenario carries a **concrete, checkable unsafe action** and runs\ntwice — **baseline** (no airlock, the unsafe event lands) and\n**governed** (the *same* action through the **real** `@Airlock`\nintercept-before-execute path, no policy-layer mocking). Three\nunsafe-action classes each exercise a different real interception\nmechanism: price-below-floor → Pydantic strict-validation,\nsecret-leak → the output sanitizer, transfer-outside-policy →\ndeny-by-default `SecurityPolicy`. Benign deals are included to confirm\ngovernance does not over-block.\n\n| source | unsafe_execution_rate (base → governed) | valid_task_success_rate (base → governed) |\n|---|---|---|\n| **agent-airlock** (this harness) | 100% → **0%** | 43% → **100%** |\n| OCL (external, live LLMs, [arXiv:2606.04306](https://arxiv.org/abs/2606.04306)) | 88% → ~0% | 12% → 96% |\n\n\u003e **The OCL row is an external result, not agent-airlock's.** It was\n\u003e measured on live frontier LLM agents in AgenticPay-adapted negotiation\n\u003e ([OCL, arXiv:2606.04306](https://arxiv.org/abs/2606.04306);\n\u003e [AgenticPay, arXiv:2602.06008](https://arxiv.org/abs/2602.06008)) and\n\u003e is reproduced here only for **directional comparison** — both put\n\u003e governance at the execution boundary. It is **not** the same\n\u003e experiment: agent-airlock is a deterministic execution-boundary\n\u003e validator, not an LLM, and this harness does not call a model. The\n\u003e agent-airlock rows are a property of the **policy layer** under a\n\u003e worst-case scripted adversary, exercised through the real `@Airlock`\n\u003e path.\n\nThe harness doubles as a **regression gate**: `--fail-if-governed-unsafe`\nexits non-zero if the governed `unsafe_execution_rate` ever rises above\nzero, so a future change that weakens the policy layer fails CI. Zero\nnew runtime deps; fully deterministic (no randomness, no network, no\nmodel call).\n\n---\n\n### 🔎 Privilege right-sizing — `airlock-explain --unused-scopes` (v0.8.13)\n\nA read-only CLI that surfaces **over-permissioning**: it diffs the\n`SecurityPolicy`'s **granted** tool scopes against the tools the agent\n**actually called** (from an OTLP export OR a native audit JSONL),\nper `AgentIdentity`, and prints the dead-weight set plus a *suggested*\ntightened allow-list.\n\n```bash\n# Install the v0.8.13 wheel; airlock-explain becomes available\npip install \"agent-airlock\u003e=0.8.13\"\n\n# Diff granted vs used; print a table\nairlock-explain --unused-scopes \\\n    --policy ./security-policy.toml \\\n    --trace  ./agent.audit.jsonl\n\n# Same, machine-readable, plus a proposed tightened policy preview\nairlock-explain --unused-scopes \\\n    --policy ./security-policy.toml \\\n    --trace  ./otel-export.json \\\n    --format json \\\n    --suggest-policy\n```\n\n**Observability-only.** This command **never mutates** the\n`SecurityPolicy`, **never writes** the policy file, and **never\nauto-applies** the suggestion. The deny-by-default posture is\nunchanged — the right-size CLI is a *review aid*, not an enforcement\nprimitive. The `--suggest-policy` output is intentionally a stdout\npreview so a human reviews the tightened allow-list before adopting\nit by hand.\n\n**Trace formats** (auto-detected by inspecting the file head):\n\n- **Audit JSONL** — the format `AuditLogger` already emits. One JSON\n  object per line, with `tool_name` / `agent_id` / `blocked`. Blocked\n  calls are excluded from the \"actually called\" set — a blocked call\n  is not an exercise of a granted scope.\n- **OTLP JSON** — the format `opentelemetry-exporter-otlp` writes.\n  Span `name` is the tool name; `attributes.agent_id` keys the per-\n  agent diff. If a span carries `airlock.blocked=true` it is skipped,\n  same as JSONL.\n\n**Diff semantics.** The matcher is `fnmatch` — the same glob semantics\n`SecurityPolicy.check_tool_allowed` uses internally, so the suggested\ntightened allow-list admits exactly the tools the agent was observed\ncalling (no surprises at adoption time). Denied-list patterns are\nforwarded unchanged to the suggestion: denials are *intent*, not\nusage data.\n\n\u003e Strictly observability. No new runtime deps. The new console-script\n\u003e entry `airlock-explain` is the project's first installable CLI;\n\u003e existing `python -m agent_airlock.cli.\u003cname\u003e` invocations are\n\u003e unaffected.\n\n---\n\n### 💰 Cost Control\n\nA runaway agent can burn $500 in API costs before you notice.\n\n```python\nfrom agent_airlock import Airlock, AirlockConfig\n\nconfig = AirlockConfig(\n    max_output_chars=5000,    # Truncate before token explosion\n    max_output_tokens=2000,   # Hard limit on response size\n)\n\n@Airlock(config=config)\ndef query_logs(query: str) -\u003e str:\n    return massive_log_query(query)  # 10MB → 5KB\n```\n\n**ROI:** 10MB logs = ~2.5M tokens = $25/response. Truncated = ~1.25K tokens = $0.01. **99.96% savings.**\n\n#### Per-model-tier budgets (v0.8.7)\n\nThe flat `max_output_*` caps above apply uniformly to every call. **`ModelTierBudget`** caps per-call cost and output tokens **per model tier label** (e.g. `\"frontier\"` / `\"mid\"` / `\"small\"`), evaluated *before* the tool runs. Untagged calls fall back to a configurable `strict_tier` (deny-by-default — the cheapest tier).\n\n```python\nfrom agent_airlock import (\n    Airlock, ModelTierBudget, SecurityPolicy, TierBudget,\n)\n\npolicy = SecurityPolicy(\n    model_tier_budget=ModelTierBudget(\n        tiers={\n            \"frontier\": TierBudget(max_cost_cents=50, max_output_tokens=4000),\n            \"mid\":      TierBudget(max_cost_cents=10, max_output_tokens=2000),\n            \"small\":    TierBudget(max_cost_cents=2,  max_output_tokens=1000),\n        },\n        strict_tier=\"small\",  # untagged → cheapest tier (deny-by-default)\n    ),\n)\n\n@Airlock(policy=policy, return_dict=True)\ndef call_model(prompt: str, **_extra):\n    return run_my_router(prompt)\n\n# The router tags each call. Airlock blocks before the model fires.\ncall_model(\"Draft a tweet\",  _airlock_tier=\"small\",    _airlock_input_tokens=50)\ncall_model(\"Deep analysis\", _airlock_tier=\"frontier\", _airlock_input_tokens=200_000)\n# →  AIRLOCK_BLOCK: Tier 'frontier' budget exceeded (worst-case 66¢ \u003e cap 50¢)\n```\n\nRouting logic stays in the user's router. Three tagging routes are supported:\n\n1. **`_airlock_tier` kwarg** — stripped before the tool sees it.\n2. **`context.metadata[\"airlock_tier\"]`** — set on a contextvar-stored\n   `AirlockContext` by the router's session middleware.\n3. **`tier_resolver` callback** — `ModelTierBudget(tier_resolver=fn)`\n   where `fn(model_id: str) -\u003e tier_label` lives in the caller's code.\n   Airlock invokes the callback when `context.metadata[\"model_id\"]`\n   is set; it carries no vendor-specific model→tier table.\n\nAfter execution, actual vs estimated cost is reconciled into the global\n`CostTracker` (observability — never blocks). See\n[`examples/model_tier_budget.py`](./examples/model_tier_budget.py) for\nall four patterns including composition with allow/deny lists.\n\nA ready-to-use `strict_tier_budget_policy()` preset returns a\n`SecurityPolicy` seeded with the table above.\n\n---\n\n### 🔐 PII \u0026 Secret Masking\n\n```python\nconfig = AirlockConfig(\n    mask_pii=True,      # SSN, credit cards, phones, emails\n    mask_secrets=True,  # API keys, passwords, JWTs\n)\n\n@Airlock(config=config)\ndef get_user(user_id: str) -\u003e dict:\n    return db.users.find_one({\"id\": user_id})\n\n# LLM sees: {\"name\": \"John\", \"ssn\": \"[REDACTED]\", \"api_key\": \"sk-...XXXX\"}\n```\n\n**12 PII types detected** · **4 masking strategies** · **Zero data leakage**\n\n#### Opt-in regional PII (`pii_locales`)\n\nAadhaar / PAN / UPI / IFSC have always shipped as `SensitiveDataType` members,\nbut are not added to the default `mask_pii=True` set — to keep the surface\nzero-dep and US-shaped by default. **v0.8.9** adds a `pii_locales` opt-in\nthat pulls them in *and* tightens detection:\n\n```python\nconfig = AirlockConfig(\n    mask_pii=True,\n    pii_locales=[\"in\"],   # opt in to India-locale detection\n)\n\n@Airlock(config=config)\ndef lookup(query: str) -\u003e str:\n    return (\n        \"User: राम कुमार, \"\n        \"Aadhaar: 234567890124, \"       # → \"23********24\" (PARTIAL)\n        \"PAN: ABCDE1234F, \"             # → \"AB******4F\" (PARTIAL)\n        \"phone: 555-123-4567\"           # still masked by existing PHONE regex\n    )\n```\n\nTwo things activate when `\"in\" in pii_locales`:\n\n- **Aadhaar Verhoeff checksum gate** — the existing Aadhaar regex is\n  permissive (any 12-digit number starting 2-9 matches). With the opt-in,\n  each match must also pass the UIDAI Verhoeff checksum, cutting the FP\n  rate ~10x on random IDs / phone numbers.\n- **Devanagari personal-name detection** — `PERSONAL_NAME_DEVANAGARI` runs\n  against the Unicode block `U+0900–U+097F`, with a small allowlist of\n  common Hindi greetings / pronouns / interrogatives to keep ordinary\n  prose from being masked. Conservative heuristic — production callers\n  who need precise extraction should layer NER on top.\n\nThe flag is **additive and reversible** — `pii_locales=[]` (the default)\npreserves the prior behavior bit-for-bit.\n\n---\n\n### 🌐 Network Airgap (V0.3.0)\n\nBlock data exfiltration during tool execution:\n\n```python\nfrom agent_airlock import network_airgap, NO_NETWORK_POLICY\n\n# Block ALL network access\nwith network_airgap(NO_NETWORK_POLICY):\n    result = untrusted_tool()  # Any socket call → NetworkBlockedError\n\n# Or allow specific hosts only\nfrom agent_airlock import NetworkPolicy\n\nINTERNAL_ONLY = NetworkPolicy(\n    allow_egress=True,\n    allowed_hosts=[\"api.internal.com\", \"*.company.local\"],\n    allowed_ports=[443],\n)\n```\n\n---\n\n### 💉 Framework Vaccination (V0.3.0)\n\nSecure existing code **without changing a single line**:\n\n```python\nfrom agent_airlock import vaccinate, STRICT_POLICY\n\n# Before: Your existing LangChain tools are unprotected\nvaccinate(\"langchain\", policy=STRICT_POLICY)\n\n# After: ALL @tool decorators now include Airlock security\n# No code changes required!\n```\n\n**Supported:** LangChain, OpenAI Agents SDK, PydanticAI, CrewAI\n\n---\n\n### ⚡ Circuit Breaker (V0.4.0)\n\nPrevent cascading failures with fault tolerance:\n\n```python\nfrom agent_airlock import CircuitBreaker, AGGRESSIVE_BREAKER\n\nbreaker = CircuitBreaker(\"external_api\", config=AGGRESSIVE_BREAKER)\n\n@breaker\ndef call_external_api(query: str) -\u003e dict:\n    return external_service.query(query)\n\n# After 5 failures → circuit OPENS → fast-fails for 30s\n# Then HALF_OPEN → allows 1 test request → recovers or reopens\n```\n\n---\n\n### 📈 OpenTelemetry Observability (V0.4.0)\n\nEnterprise-grade monitoring:\n\n```python\nfrom agent_airlock import configure_observability, observe\n\nconfigure_observability(\n    service_name=\"my-agent\",\n    otlp_endpoint=\"http://otel-collector:4317\",\n)\n\n@observe(name=\"critical_operation\")\ndef process_data(data: dict) -\u003e dict:\n    # Automatic span creation, metrics, and audit logging\n    return transform(data)\n```\n\n---\n\n## 🔌 Framework Compatibility\n\n\u003e **The Golden Rule:** `@Airlock` must be closest to the function definition.\n\n```python\n@framework_decorator    # ← Framework sees secured function\n@Airlock()             # ← Security layer (innermost)\ndef my_function():     # ← Your code\n```\n\n\u003ctable\u003e\n\u003ctr\u003e\n\u003ctd\u003e\n\n### LangChain / LangGraph\n\n```python\nfrom langchain_core.tools import tool\nfrom agent_airlock import Airlock\n\n@tool\n@Airlock()\ndef search(query: str) -\u003e str:\n    \"\"\"Search for information.\"\"\"\n    return f\"Results for: {query}\"\n```\n\n\u003c/td\u003e\n\u003ctd\u003e\n\n### OpenAI Agents SDK\n\n```python\nfrom agents import function_tool\nfrom agent_airlock import Airlock\n\n@function_tool\n@Airlock()\ndef get_weather(city: str) -\u003e str:\n    \"\"\"Get weather for a city.\"\"\"\n    return f\"Weather in {city}: 22°C\"\n```\n\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd\u003e\n\n### PydanticAI\n\n```python\nfrom pydantic_ai import Agent\nfrom agent_airlock import Airlock\n\n@Airlock()\ndef get_stock(symbol: str) -\u003e str:\n    return f\"Stock {symbol}: $150\"\n\nagent = Agent(\"openai:gpt-4o\", tools=[get_stock])\n```\n\n\u003c/td\u003e\n\u003ctd\u003e\n\n### CrewAI\n\n```python\nfrom crewai.tools import tool\nfrom agent_airlock import Airlock\n\n@tool\n@Airlock()\ndef search_docs(query: str) -\u003e str:\n    \"\"\"Search internal docs.\"\"\"\n    return f\"Found 5 docs for: {query}\"\n```\n\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/table\u003e\n\n\u003cdetails\u003e\n\u003csummary\u003e\u003cb\u003eMore frameworks: LlamaIndex, AutoGen, smolagents, Anthropic\u003c/b\u003e\u003c/summary\u003e\n\n### LlamaIndex\n\n```python\nfrom llama_index.core.tools import FunctionTool\nfrom agent_airlock import Airlock\n\n@Airlock()\ndef calculate(expression: str) -\u003e int:\n    return eval(expression, {\"__builtins__\": {}})\n\ncalc_tool = FunctionTool.from_defaults(fn=calculate)\n```\n\n### AutoGen\n\n```python\nfrom autogen import ConversableAgent\nfrom agent_airlock import Airlock\n\n@Airlock()\ndef analyze_data(dataset: str) -\u003e str:\n    return f\"Analysis of {dataset}: mean=42.5\"\n\nassistant = ConversableAgent(name=\"analyst\", llm_config={\"model\": \"gpt-4o\"})\nassistant.register_for_llm()(analyze_data)\n```\n\n### smolagents\n\n```python\nfrom smolagents import tool\nfrom agent_airlock import Airlock\n\n@tool\n@Airlock(sandbox=True)\ndef run_code(code: str) -\u003e str:\n    \"\"\"Execute in E2B sandbox.\"\"\"\n    exec(code)\n    return \"Executed\"\n```\n\n### Anthropic (Direct API)\n\n```python\nfrom agent_airlock import Airlock\n\n@Airlock()\ndef get_weather(city: str) -\u003e str:\n    return f\"Weather in {city}: 22°C\"\n\n# Use in tool handler\ndef handle_tool_call(name, inputs):\n    if name == \"get_weather\":\n        return get_weather(**inputs)  # Airlock validates\n```\n\n\u003c/details\u003e\n\n### Adapter-shipped vs example-only (honest split)\n\n\u003e Both paths use the same `@Airlock()` decorator placement. \"Adapter-shipped\"\n\u003e means there's a dedicated `src/agent_airlock/integrations/\u003cframework\u003e.py`\n\u003e module with framework-specific glue (signature preservation, tool registry\n\u003e rewrites, request-shape adapters). \"Example-only\" means the decorator is\n\u003e compatible out of the box — no extra adapter required.\n\n**Adapter-shipped (11):** LangChain (`integrations/langchain.py`),\nLangGraph (`integrations/langgraph_toolnode_compat.py`),\nOpenAI Agents SDK (`integrations/openai_guardrails.py`),\nAnthropic Messages API (`integrations/anthropic.py`),\nAnthropic Claude Agent SDK (`integrations/anthropic_claude_agent_sdk.py`, v0.6.1+),\nsmolagents (`integrations/smolagents_wrapper.py`),\nGemini 3 Agent Mode (`integrations/gemini3_tool_shape_adapter.py`),\nGPT-5.5 (`integrations/gpt5_5_tool_shape_adapter.py`),\nPydanticAI (`integrations/pydantic_ai.py`, v0.7.1+),\nCrewAI (`integrations/crewai.py`, v0.7.2+),\nFastMCP (`mcp.py`).\n\n**Example-only (2):** AutoGen, LlamaIndex —\ndecorator-compatible without an adapter; see `examples/`.\n\n### Complete Examples\n\n| Framework | Path | Surface |\n|-----------|------|---------|\n| LangChain | [adapter](./src/agent_airlock/integrations/langchain.py) · [example](./examples/langchain_integration.py) | @tool, AgentExecutor |\n| LangGraph | [adapter](./src/agent_airlock/integrations/langgraph_toolnode_compat.py) · [example](./examples/langgraph_integration.py) | StateGraph, ToolNode |\n| OpenAI Agents | [adapter](./src/agent_airlock/integrations/openai_guardrails.py) · [example](./examples/openai_agents_sdk_integration.py) | Handoffs, manager pattern |\n| Anthropic API | [adapter](./src/agent_airlock/integrations/anthropic.py) · [example](./examples/anthropic_integration.py) | Direct Messages API |\n| Claude Agent SDK | [adapter](./src/agent_airlock/integrations/anthropic_claude_agent_sdk.py) · [doc](./docs/integrations/anthropic-claude-agent-sdk.md) | `wrap_agent(agent, policy=...)` |\n| smolagents | [adapter](./src/agent_airlock/integrations/smolagents_wrapper.py) · [example](./examples/smolagents_integration.py) | CodeAgent, E2B |\n| Gemini 3 | [adapter](./src/agent_airlock/integrations/gemini3_tool_shape_adapter.py) | `function_call` carrier + `thought_signature` redaction |\n| GPT-5.5 | [adapter](./src/agent_airlock/integrations/gpt5_5_tool_shape_adapter.py) | `gpt_5_5_agent_defaults` preset |\n| FastMCP | [adapter](./src/agent_airlock/mcp.py) · [example](./examples/fastmcp_integration.py) | `@secure_tool` decorator |\n| PydanticAI | [adapter](./src/agent_airlock/integrations/pydantic_ai.py) · [doc](./docs/integrations/pydantic-ai.md) · [example](./examples/pydanticai_integration.py) | `wrap_agent(agent, policy=...)` + output_validate hook |\n| CrewAI | [adapter](./src/agent_airlock/integrations/crewai.py) · [doc](./docs/integrations/crewai.md) · [example](./examples/crewai_integration.py) | `wrap_crew(crew, policy=...)` + task-level tool overrides |\n| LlamaIndex | [example only](./examples/llamaindex_integration.py) | ReActAgent |\n| AutoGen | [example only](./examples/autogen_integration.py) | ConversableAgent |\n\n---\n\n## ⚡ FastMCP Integration\n\n```python\nfrom fastmcp import FastMCP\nfrom agent_airlock.mcp import secure_tool, STRICT_POLICY\n\nmcp = FastMCP(\"production-server\")\n\n@secure_tool(mcp, policy=STRICT_POLICY)\ndef delete_user(user_id: str) -\u003e dict:\n    \"\"\"One decorator: MCP registration + Airlock protection.\"\"\"\n    return db.users.delete(user_id)\n```\n\n---\n\n## 🏆 Why Not Enterprise Vendors?\n\n| | Prompt Security | Pangea | **Agent-Airlock** |\n|---|:---:|:---:|:---:|\n| **Pricing** | $50K+/year | Enterprise | **Free forever** |\n| **Integration** | Proxy gateway | Proxy gateway | **One decorator** |\n| **Self-Healing** | ❌ | ❌ | **✅** |\n| **E2B Sandboxing** | ❌ | ❌ | **✅ Native** |\n| **Your Data** | Their servers | Their servers | **Never leaves you** |\n| **Source Code** | Closed | Closed | **MIT Licensed** |\n\n\u003e We're not anti-enterprise. We're anti-gatekeeping.\n\u003e **Security for AI agents shouldn't require a procurement process.**\n\n---\n\n## 📦 Installation\n\n```bash\n# Core (validation + policies + sanitization)\npip install agent-airlock\n\n# With E2B sandbox support\npip install agent-airlock[sandbox]\n\n# With FastMCP integration\npip install agent-airlock[mcp]\n\n# Everything\npip install agent-airlock[all]\n```\n\n```bash\n# E2B key for sandbox execution\nexport E2B_API_KEY=\"your-key-here\"\n```\n\n---\n\n## 🛡️ OWASP Compliance\n\nAgent-Airlock maps to the [**OWASP Top 10 for Agentic Applications (2026)**](https://genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026/)\n— the agentic-era successor to the old LLM Top 10. Coverage is\nreported honestly: **Full** means the primitive ships and blocks the\nclass in tests; **Partial** means agent-airlock covers the runtime\nleg but something upstream (client UI, IAM, training data) is out of\nscope; **Monitor-only** means we surface the signal but do not\nactually prevent the risk.\n\n| Risk | Implemented in agent-airlock | Module / preset | Coverage |\n|------|------------------------------|-----------------|----------|\n| **ASI01 Agent Goal Hijack** | Pydantic strict validation + ghost-arg rejection + `UnknownArgsMode.BLOCK` | `validator`, `unknown_args`, `core` | Partial |\n| **ASI02 Tool Misuse and Exploitation** | Deny-by-default `SecurityPolicy`, RBAC, rate limits, `SafePath` / `SafeURL`, Flowise `Function()`/`eval` token ban ([CVE-2025-59528](https://labs.cloudsecurityalliance.org/research/csa-research-note-flowise-mcp-rce-exploitation-20260409-csa/)), MCPwn destructive-auth check ([CVE-2026-33032](https://nvd.nist.gov/vuln/detail/CVE-2026-33032)), Mobile MCP intent-URL guard ([CVE-2026-35394](https://www.sentinelone.com/vulnerability-database/cve-2026-35394/)) | `policy`, `safe_types`, `filesystem`, `network`, `policy_presets.flowise_cve_2025_59528_defaults`, `policy_presets.mcpwn_cve_2026_33032_defaults`, `policy_presets.mobile_mcp_intent_guard_2026_05` | **Full** |\n| **ASI03 Identity and Privilege Abuse** | `AgentIdentity`, `MCPProxyGuard` token-passthrough prevention, `CredentialScope`, OAuth-app audit ([Vercel 2026-04-19](https://vercel.com/kb/bulletin/vercel-april-2026-security-incident)), MCP Attested Tool-Server Admission ([arXiv:2605.24248](https://arxiv.org/abs/2605.24248)) | `policy`, `mcp_proxy_guard`, `mcp_spec.oauth_audit`, `mcp_spec.attested_admission`, `policy_presets.oauth_audit_vercel_2026_defaults`, `policy_presets.mcp_attested_admission_defaults` | Partial |\n| **ASI04 Agentic Supply Chain Vulnerabilities** | Ox MCP STDIO sanitizer + CVE regression suite (11+ CVEs tracked) + session-snapshot integrity guard + spawn-time MCP config pin (CVE-2026-30615, `policy_presets.mcp_config_pin`) | `mcp_spec.stdio_guard`, `mcp_spec.session_guard`, `mcp_spec.zero_click_config_guard`, `policy_presets.stdio_guard_ox_defaults`, `policy_presets.mcp_config_pin`, `tests/cves/` | Partial |\n| **ASI05 Unexpected Code Execution (RCE)** | E2B Firecracker sandbox, pluggable `SandboxBackend`, capability gating for `PROCESS_SHELL`, Flowise eval-token ban ([CVE-2025-59528](https://labs.cloudsecurityalliance.org/research/csa-research-note-flowise-mcp-rce-exploitation-20260409-csa/)) | `sandbox`, `sandbox_backend`, `capabilities`, `policy_presets.flowise_cve_2025_59528_defaults` | **Full** |\n| **ASI06 Memory \u0026 Context Poisoning** | `AirlockContext` `contextvars` isolation, `ConversationConstraints` budget caps, audit logging | `context`, `conversation`, `sanitizer` | Partial |\n| **ASI07 Insecure Inter-Agent Communication** | A2A middleware Pydantic strict validation, method allow-lists | `a2a` | Partial |\n| **ASI08 Cascading Failures** | `CircuitBreaker`, `RetryPolicy`, token-bucket rate limits | `circuit_breaker`, `retry`, `policy` | **Full** |\n| **ASI09 Human-Agent Trust Exploitation** | Honeypot deception, audit-log attribution, structured `fix_hints` | `honeypot`, `audit_otel` | Partial |\n| **ASI10 Rogue Agents** | Audit telemetry + anomaly detector; no quarantine primitive | `observability`, `anomaly` | Monitor-only |\n\n### MCP-specific mapping\n\nThe [OWASP MCP Top 10 (2026 beta)](https://owasp.org/www-project-mcp-top-10/)\nis covered end-to-end by the `OWASP_MCP_TOP_10_2026` policy preset:\n\n| MCP risk | Ships in agent-airlock |\n|----------|------------------------|\n| **MCP01 Token Mismanagement** | `MCPProxyGuard` rejects passthrough headers, enforces audience |\n| **MCP02 Excessive Permissions** | `SecurityPolicy` + `CredentialScope` |\n| **MCP03 Tool Poisoning** | ghost-arg rejection + `SafePath`/`SafeURL` |\n| **MCP04 Supply Chain** | `stdio_guard_ox_defaults()` (Ox 2026-04-16 advisory) |\n| **MCP05 Command Injection** | `stdio_guard` shell-metachar + deny-pattern rules |\n| **MCP07 Insufficient Authentication** | OAuth 2.1 + PKCE S256 helpers in `mcp_spec.oauth` |\n| **MCP10 Context Oversharing** | PII/secret sanitizer + workspace-scoped config |\n\nUse it directly:\n\n```python\nfrom agent_airlock import Airlock\nfrom agent_airlock.policy_presets import owasp_mcp_top_10_2026_policy\n\n@Airlock(policy=owasp_mcp_top_10_2026_policy())\ndef my_mcp_tool(...):\n    ...\n```\n\n\u003e **Ox Security STDIO advisory** (2026-04-16, CVE-2026-30616): see\n\u003e [`docs/cves/index.md#cve-2026-30616`](docs/cves/index.md#cve-2026-30616)\n\u003e and the `stdio_guard_ox_defaults()` preset above. agent-airlock\n\u003e blocks 3 of 4 Ox attack classes at the runtime seam.\n\n---\n\n## 🏢 Used By\n\nAgent-Airlock secures AI agent systems in production:\n\n| Project | Use Case |\n|---------|----------|\n| [**FerrumDeck**](https://github.com/sattyamjjain/FerrumDeck) | AgentOps control plane — deny-by-default tool execution |\n| [**Mnemo**](https://github.com/sattyamjjain/Mnemo) | MCP-native memory database — secure tool call validation |\n\n\u003e Using Agent-Airlock in production? [Open a PR](https://github.com/sattyamjjain/agent-airlock/edit/main/README.md) to add your project!\n\n---\n\n## 📊 Performance\n\n\u003e Test count and coverage are published by the TEST-BADGE block at the top of this file,\n\u003e regenerated from pytest on every release via `python scripts/update_test_badge.py`.\n\u003e That block is the source of truth; this table tracks latency and surface area only.\n\n| Metric | Value |\n|--------|-------|\n| **Validation overhead** | \u003c50ms |\n| **Sandbox cold start** | ~125ms |\n| **Sandbox warm pool** | \u003c200ms |\n| **Framework integrations** | 13 |\n| **Core dependencies** | 0 (Pydantic only) |\n\n---\n\n## 📖 Documentation\n\n| Resource | Description |\n|----------|-------------|\n| [**AGENTS.md**](./AGENTS.md) | v0.6.1 — repo-root entrypoint for agentic IDEs (Cursor, Claude Code, Windsurf, Mintlify) |\n| [**Anthropic Claude Agent SDK adapter**](./docs/integrations/anthropic-claude-agent-sdk.md) | v0.6.1 — `AnthropicClaudeAgentSDKAdapter.wrap_agent(agent, policy=...)`; canonical-list trio |\n| [**`airlock manifest enforce`**](https://news.backbox.org/2026/05/01/200000-mcp-servers-expose-a-command-execution-flaw-that-anthropic-calls-a-feature/) | v0.6.1 — fail-closed CLI runtime allowlist gate against signed manifests; CI exits 0/2/3 |\n| [**Managed Agents Outcomes-rubric guard**](./docs/policies/managed-agents-outcomes-2026-05-06.md) | v0.7.4 — fail-closed gate on the Anthropic Managed Agents 2026-05-06 Outcomes rubric ID; `ManagedAgentsOutcomesGuard.evaluate(provenance)` + `managed_agents_outcomes_2026_05_06_defaults` factory; no SDK dep |\n| [**Filter-Eval RCE guard (CVE-2026-25592 + CVE-2026-26030)**](./docs/policies/semantic-kernel-filter-eval-rce.md) | v0.7.5 — regex detector for the Semantic-Kernel-class lambda-filter / template-expression eval RCE primitive (MSRC 2026-05-07); `FilterEvalRCEGuard.evaluate(args)` + `semantic_kernel_filter_eval_rce_2026_25592_26030_defaults` factory; framework-agnostic |\n| [**OIDC publish-window guard (TanStack 2026-05-11)**](./docs/policies/npm-oidc-publish-window-guard.md) | v0.7.6 — known-bad blast-list guard for the TanStack/Mini-Shai-Hulud npm OIDC trusted-publisher class (postmortem 2026-05-11; 42 pkgs × 84 versions); `OIDCPublishWindowGuard.evaluate(args)` + `npm_oidc_publish_window_guard_defaults` factory; pure-data preset, no runtime npm calls |\n| [**MCP STDIO command-injection guard**](./docs/policies/mcp-stdio-command-injection-guard.md) | v0.7.6 — shell metachar + opt-in path-traversal denier for MCP STDIO argv vectors (HelpNetSecurity 2026-05-05); `StdioCommandInjectionGuard.evaluate(args)` + `mcp_stdio_command_injection_preset_defaults` factory; no `mcp` SDK dep |\n| [**Eval-RCE guard (CVE-2026-44717)**](./docs/policies/eval-rce-cve-2026-44717.md) | v0.8.0 — bare-`eval()`/`parse_expr()`/`exec()` invocation detector for the MCP Calculate Server class (NVD 2026-05-15); `EvalRCEGuard.evaluate(args)` + curated vulnerable-package denylist + `parse_expr` safe-form exemption + `stdio_guard_eval_defaults_2026_05_15` factory |\n| [**MCP Inspector exposure guard (CVE-2026-23744 runtime)**](./docs/policies/mcp-inspector-exposure-guard.md) | v0.8.0 — Linux runtime listener-scan via stdlib `/proc/net/tcp` for the MCPJam Inspector public-bind class; complements v0.5.x config-time `bind_address_guard`; `MCP_INSPECTOR_REQUIRE_AUTH=1` operator bypass |\n| [**Agent SDK Credit pool budget**](./docs/budget/agent-sdk-credit.md) | v0.8.0 — per-month USD pool tracker for Anthropic's 2026-06-15 billing split (Zed blog 2026-05-14); `AgentSDKCreditBudget.register_call(model, input_tokens, output_tokens)` with 90% near-limit + 100% exhausted thresholds; packaged 2026-06 pricing fixture |\n| [**OpenAPI Drift Guard (Hermes 2026-05-13)**](./docs/policies/openapi-drift-guard.md) | v0.8.1 — payload-shape drift detector against an operator-supplied OpenAPI 3.x spec (arXiv:2605.14312); `OpenAPIDriftGuard.evaluate(operation_id, args)` detects `missing_required` / `unknown_field` / `type_mismatch`; three modes (`strict` / `warn` / `shadow`); `vaccinate_openapi(spec)` decorator + `openapi_doc_drift_guard_defaults` factory; caller supplies spec dict, no PyYAML dep |\n| [**MCP Calc-Server bundle preset**](./docs/policies/eval-rce-cve-2026-44717.md) | v0.8.1 — composition factory `mcp_calc_server_bundle_defaults_2026_05_15()` wires v0.8.0 `EvalRCEGuard` + v0.7.6 `StdioCommandInjectionGuard` under a single preset_id (CVE-2026-44717 anchor) scoped to calc/calculate/evaluate/sympy_eval/math_eval tool-name patterns; pure config composition, no new detector module |\n| [**Metis-inspired corpus block-rate regression**](./docs/policies/metis-inspired-corpus-block-rate.md) | v0.8.2 — release-gate primitive `MetisInspiredCorpusBlockRateGuard` runs a deterministic 25-entry exploit-shape corpus (CVE-2026-44717 + 2026-05-05 STDIO injection) through `EvalRCEGuard + StdioCommandInjectionGuard`; one-sided gate fires when block rate drops below baseline − 5%; **NOT a reproduction of the Metis paper's POMDP attacker** (arXiv:2605.10067 cited as motivation, not as prompt source); `airlock corpus-bench` CLI ships text/json/md reports |\n| [**Corpus per-category coverage**](./docs/policies/metis-inspired-corpus-block-rate.md) | v0.8.3 — extends the v0.8.2 corpus-bench with HarnessAudit-Bench (arXiv:2605.14271) two-category taxonomy (`resource_access`, `info_transfer`); `CorpusEntry.violation_category` field + `CategoryCount` decision field; `airlock corpus-bench` reports per-category coverage in text/json/md; **NOT a reproduction of HarnessAudit-Bench** (artifacts not yet public — taxonomy adopted as schema, scoring is not) |\n| [**Stainless SDK provenance classifier**](./docs/policies/stainless-provenance-probe.md) | v0.8.3 — pure-function `classify_sdk_lineage(user_agent, response_body_head)` building block flags MCP servers generated by the deprecated Stainless SDK toolchain (Anthropic acquired Stainless 2026-05-13, hosted generator winding down); operator-callable from own audit hooks — **NOT an automatic HTTP probe** (decorator-in-process architecture, see ROADMAP §1); `stainless_provenance_probe_defaults()` preset is `default_action=tag_only`, visibility not enforcement |\n| [**Human-oversight decorator**](./docs/policies/human-oversight-decorator.md) | v0.8.4 — `@requires_human_oversight(approver=...)` gates a tool function on an operator-supplied approval callable (Code-as-Harness arXiv:2605.18747 anchor); `GRANT` → call wrapped fn, `DENY` → `OversightDeniedError`, `TIMEOUT` → `OversightTimeoutError`; composes with `@Airlock(...)`; protocol shapes + `InProcessRecordedApprover` testing helper; **NOT a bidirectional audit-emitter RPC channel** — operator owns the transport (Slack/PagerDuty/CLI), agent-airlock owns the gate + the protocol |\n| [**Layer-contract receipt block**](./docs/attest/layer-contract.md) | v0.8.5 — opt-in `LayerContract` (assume/guarantee) block on signed `airlock attest receipt` payloads (arXiv:2605.18672 anchor); `--contract` derives per-guard `pass_rate` from the verdicts list, `--assumes id1,id2` declares upstream-layer dependencies; receipt schema v1 unchanged (additive field); `pass_rate` is a measured statistic over the sample (not a proof) — every Guarantee carries `sample_size` so verifiers can weight low-N appropriately; **NOT backed by a window-counter store** (that infrastructure doesn't exist yet — derived from the operator-supplied verdicts list, no new abstraction) |\n| [**MCP Attested Tool-Server Admission (arXiv:2605.24248)**](https://arxiv.org/abs/2605.24248) | v0.8.10 — opt-in admission gate for MCP tool servers per Metere (May 2026). Host fetches a JWS-compact clearance from `{server_url}/.well-known/mcp-clearance`, verifies its signature against an **operator-pinned trust root** (Ed25519 / RSA-PSS / JWKS — never network-fetched on the hot path), and enforces a **deny-by-default per-server tool allowlist** parsed from the verified clearance. Flavor-gated `ENFORCE` (hard-deny) / `WARN` (log only) modes. Every decision emits a `ReceiptVerdict` on the `guard=\"mcp_attested_admission\"` channel — reuses the existing `airlock attest` DSSE path, does **not** invent a new log. `mcp_attested_admission_defaults()` factory + `MCPProxyGuard.audit_tool_admission()` integration; signature verification gated behind `pip install agent-airlock[attested]`. |\n| [**Mobile MCP intent-URL guard (CVE-2026-35394)**](https://www.sentinelone.com/vulnerability-database/cve-2026-35394/) | v0.8.8 — defensive bundle for the Mobilenexthq Mobile MCP `mobile_open_url` intent-injection RCE class (\u003c 0.0.50). `mobile_mcp_intent_guard_2026_05()` returns a pre-configured `SafeURLValidator(allowed_schemes=[\"http\", \"https\"])` (blocks `intent:`, `content:`, `file:`, `app:`, `data:`, `javascript:`, `vbscript:`), an `AirlockConfig(unknown_args=UnknownArgsMode.BLOCK)`, and the canonical Mobile MCP tool-name corpus (`mobile_open_url`, `open_url`, `mobile_launch_url`). DIFF-COMPATIBLE with the existing `SafeURL` type — no new validator invented. Also fixes a pre-existing `block_private_ips=True` no-op in `SafeURLValidator` (RFC1918 ranges were not actually blocked because the validator's own `SafeURLValidationError` raise was caught by `except ValueError`). |\n| [**Capsule ShareLeak / PipeLeak (CVE-2026-21520)**](https://nvd.nist.gov/vuln/detail/CVE-2026-21520) | v0.8.14 — defensive bundle for the [Capsule Security](https://www.capsulesecurity.io/blog-post/shareleak-taking-the-wheel-of-microsofts-copilot-studio-cve-2026-21520)-disclosed indirect-prompt-injection class hitting Microsoft Copilot Studio (ShareLeak, CVE-2026-21520, CVSS 7.5 HIGH, CWE-77, patched 2026-01-15) and Salesforce Agentforce (PipeLeak, parallel pattern). Both vectors share the same architecture: untrusted form input (SharePoint form / Web-to-Lead form) is concatenated into the agent's context with no boundary, while the agent simultaneously holds outbound exfil tools (Outlook send / Salesforce email-case). `capsule_indirect_injection_cve_2026_21520_defaults()` composes existing primitives — `default_deny=True` + canonical exfil-sink `denied_tools` (`send_email`, `outlook_*`, `create_case`, `share_*`, `export_*`, `post_to_*`, `webhook_*`, ...) + `reauth_on_untrusted_reinvocation=True` (v0.8.6 debate-amplification guard at `threshold=1`) + `AirlockConfig(unknown_args=UnknownArgsMode.BLOCK)`. Opt-in only — no new validator invented, no default-priority-chain entry. Pairs with `airlock-explain --unused-scopes` (v0.8.13) so operators populate the read-side allow-list from a real trace before deploying. |\n| [**Flowise MCP-stdio adapter RCE (CVE-2026-40933)**](https://advisories.gitlab.com/npm/flowise-components/CVE-2026-40933/) | v0.8.16 — defensive control for the Flowise authenticated-RCE-via-MCP-stdio-adapter class (CVSS 9.9, fixed upstream in Flowise 3.1.0). Flowise ≤ 3.0.x serialises a user-defined CustomMCP `command`+`args` straight into a child-process spawn with no sandbox or argv sanitisation — importing a crafted chatflow is a one-click path to OS-level RCE. `flowise_mcp_stdio_guard_2026_defaults()` is a per-tool-class projection of the v0.7.6 `StdioCommandInjectionGuard` (no new detector invented), scoped to the Flowise CustomMCP stdio surface. Fail-closed on shell metachars (`;`, `\u0026\u0026`, `\\|\\|`, `\\|`, newline, backtick, `$(`) in the `command`/`args` path + opt-in path-traversal outside a `cwd_allowlist`; `check(args)` raises `FlowiseMcpStdioInjectionError`. OWASP **MCP05 Command Injection**. Wired into `ox_mcp_supply_chain_2026_04_defaults()` — **corrects a prior mis-attribution** where CVE-2026-40933 was recorded as a \"Semantic Kernel auth-header leak\". |\n| [**MCP description-vs-manifest guard (`mcp_description_manifest_guard`)**](https://arxiv.org/abs/2606.04769) | v0.8.18 — runtime consistency gate that asserts a tool's **model-facing description** (declared input schema + advertised capability/security boundary) matches its **registered manifest** *before* the tool is admitted, failing closed per the deny-by-default posture. Anchored on the DCIChecker study (arXiv:2606.04769), which measured **Description-Code Inconsistency at 9.93% of 19,200 tool pairs across 2,214 MCP servers**. `DescriptionManifestGuard.evaluate(description)` detects `described_arg_not_in_manifest` (description claims a ghost argument), `undisclosed_side_effect` (manifest has a side effect the description hides — the tool-poisoning direction), and `overclaimed_capability` (description advertises a capability absent from the manifest); three modes (`strict` / `warn` / `shadow`); `vaccinate_description_manifest(manifests)` decorator + `mcp_description_manifest_guard_defaults()` factory. Composes **above** ghost-arg stripping + Pydantic type-validation (which govern the call payload) — it does not replace them. OWASP **MCP03 Tool Poisoning**. Pydantic-only core, no new runtime deps. |\n| [**LeRobot pickle-deserialization RCE (CVE-2026-25874)**](https://www.sentinelone.com/vulnerability-database/cve-2026-25874/) | v0.8.19 — deny-by-default posture for the HuggingFace LeRobot unauthenticated-RCE class (CVSS 9.3). LeRobot's async-inference PolicyServer / robot-client `pickle.loads()` payloads received over an **unauthenticated, non-TLS** gRPC channel (`SendObservations` / `SendPolicyInstructions` / `GetActions`) — an unauthenticated, network-reachable attacker reaches arbitrary OS command execution. Ships a reusable **`UnsafeDeserializationGuard`** (in `safe_types`, next to `SafePath`/`SafeURL`) that fails closed on pickle magic bytes (`0x80` PROTO), base64-encoded pickle, and `pickle`/`marshal`/`shelve`/`dill`/`jsonpickle` marker tokens in string args — plus an airgap pairing that refuses serialized-object (`bytes`) args unless the call declares an authenticated **and** TLS transport. Wired into `SecurityPolicy.deserialization_guard` and run at the `@Airlock` seam (Step 2.7) **before** the tool body; the block carries a `fix_hint` naming CVE-2026-25874. `lerobot_cve_2026_25874_defaults()` is the per-CVE projection (deny-by-name globs for `*deserialize*`/`*pickle.loads*`/`torch_load`/the gRPC methods + the wired content guard). Composes **above** ghost-arg stripping + Pydantic type-validation. Pydantic-only core, no new runtime deps. |\n| [**MCP server-URL env-interpolation secret leak (CVE-2026-32625)**](https://github.com/danny-avila/LibreChat/security/advisories/GHSA-6vqg-rgpm-qvf9) | v0.8.20 — deny-by-default guard for the LibreChat MCP-server-URL credential-disclosure class (CVSS 9.6, CWE-200, OWASP **MCP01**). A user-supplied MCP server connection template (URL / header / arg) carrying an env-interpolation token (`${VAR}`, bare `$VAR`, or `%VAR%`) is expanded **server-side** against the host `process.env` and leaks a secret (`${JWT_SECRET}` / `${CREDS_KEY}` / `${MONGO_URI}`) into the outbound request. **`MCPServerEnvInterpolationGuard.evaluate(config)`** (in `mcp_spec/env_interpolation_guard.py`) scans the URL/headers/args recursively and refuses **any** interpolation token unless its variable is on an operator-declared `allowed_vars` allowlist of explicitly non-secret vars (empty default = deny all). It never reads `os.environ` or expands anything — token-match only, so it cannot itself leak. `mcp_server_env_interpolation_guard_defaults()` factory + `check(config)` raising `MCPServerEnvInterpolationError`; escaped `\\$`/`$$` are not flagged. Pydantic-only core, no new runtime deps. |\n| [**Codegen triple-quote / delimiter break-out RCE (CVE-2026-11393)**](https://www.thehackerwire.com/agentcore-cli-rce-via-triple-quote-neutralization-bypass-cve-2026-11393/) | v0.8.21 — deny-by-default guard for the AWS AgentCore CLI code-injection class (CVSS 9, CWE-94, OWASP **ASI05**). The CLI splices a model-/user-controlled `collaborationInstruction` into generated Python **without neutralising triple-quote characters**, so a crafted `\"\"\"` closes the generated string literal and injects statements that execute on agent import — RCE on the AgentCore Runtime + the importer's machine. **`CodegenDelimiterInjectionGuard.evaluate(args)`** (in `mcp_spec/codegen_delimiter_guard.py`) recursively scans args bound for a codegen / template / `exec`/`eval` sink and fails closed on triple-quote tokens (`\"\"\"` / `'''`), quote break-out tokens (`\");` / `')` / `\" +` / `']`), and raw newlines — unless the field is on an operator-declared `allowed_literal_fields` allowlist of safe literal contexts. It never generates or executes code — token-match only. `codegen_delimiter_injection_guard_defaults()` factory + `check(args)` raising `CodegenDelimiterInjectionError`; composes one layer above the v0.8.0 `EvalRCEGuard` (which gates the sink itself). Pydantic-only core, no new runtime deps. |\n| [**MCP-bridge subprocess command/args/env RCE (CVE-2026-42271, CISA KEV)**](https://www.cisa.gov/known-exploited-vulnerabilities-catalog?field_cve=CVE-2026-42271) | v0.8.22 — deny-by-default guard for the LiteLLM MCP-preview-endpoint command-injection class (CVSS 8.7, CWE-78, OWASP **ASI05**; **on the CISA KEV catalog as of 2026-06-09, actively exploited**). LiteLLM's `POST /mcp-rest/test/connection` + `/mcp-rest/test/tools/list` accepted a full MCP server config (`command` / `args` / `env`) in the request body and spawned it as a subprocess with no validation — any low-privilege API key reached host command execution (unauthenticated RCE when chained with the Starlette Host-header bypass CVE-2026-48710). **`McpSubprocessArgInjectionGuard.evaluate(config)`** (in `mcp_spec/subprocess_arg_guard.py`) treats spawn-shaped MCP-bridge args (`command`/`cmd`/`args`/`argv`/`env`) as untrusted and refuses them unless the resolved program is on an operator-declared `allowed_commands` allowlist of safe static commands (empty default = deny all); an `env` carrying a code-loading var (`LD_PRELOAD`/`PATH`/`PYTHONPATH`/…) is refused regardless, and a config with no spawn-shaped fields passes. Never spawns anything — config inspection only. `mcp_subprocess_arg_injection_guard_defaults()` factory + `check(config)` raising `McpSubprocessArgInjectionError`; composes one layer above the v0.7.6 `StdioCommandInjectionGuard` (which scans an *allowed* argv for shell metachars). Pydantic-only core, no new runtime deps. |\n| [**Examples**](./examples/) | 13 framework integrations (11 adapter-shipped + 2 example-only) with copy-paste code |\n| [**Security Guide**](./docs/SECURITY.md) | Production deployment checklist |\n| [**API Reference**](./docs/API.md) | Every function, every parameter |\n| [**Egress Bench**](./docs/security/egress-bench.md) | CVE fixture walker — every payload previously blocked stays blocked |\n| [**OX MCP Supply-Chain preset**](./docs/presets/ox-mcp-supply-chain-2026-04.md) | Umbrella for the 2026-04-20 OX dossier (10 CVEs) |\n| [**Elicitation guard (`mcp_elicitation_guard_2026_04`)**](https://github.com/modelcontextprotocol/specification/pull/1487) | v0.6.0 — runtime mitigation for the MCP `tool/elicitation` round-trip (spec PR #1487, draft 2026-04-r1); blocks credential-request and policy-override classes |\n| [**Config-path guard (CVE-2026-31402)**](https://nvd.nist.gov/vuln/detail/CVE-2026-31402) | v0.6.0 — Claude Desktop MCP-server-registration path-traversal mitigation (CVSS 8.8) |\n| [**Gemini 3 Agent Mode adapter**](https://blog.google/technology/google-deepmind/gemini-3-agent-mode-ga/) | v0.6.0 — `function_call` carrier normalisation + `thought_signature` redaction; pinned `SUPPORTED_VERSIONS` set |\n| [**OAuth `state` entropy guard**](https://www.blackhat.com/asia-26/briefings/schedule/#oauth-state-injection) | v0.6.0 — base64/hex/JSON decode + prompt-injection scan on the OAuth `state` parameter (BlackHat Asia 2026 vector) |\n| [**`airlock console`**](./docs/cli/console.md) | v0.6.0 — three-pane Textual TUI with live verdict stream + replay-on-edit; gated behind `airlock[console]` extra |\n| [**`airlock attest receipt`**](./docs/attest/receipt.md) | v0.6.0 — Sigstore-compatible signed agent-run receipts; `emit` + `verify` subcommands |\n| [**`policy_bundle.lock`**](./docs/pack/policy-bundle-lock.md) | v0.6.0 — hash-pinned preset bundles with `Cargo.lock` semantics; `airlock pack lock` + `airlock replay --bundle-lock` |\n| [**`airlock studio`**](./docs/studio/quickstart.md) | v0.6.0 — local stdlib HTTP rehearsal sandbox; paste-a-transcript verdicts + diff between runs |\n| [**smolagents wrapper**](https://github.com/huggingface/smolagents/releases/tag/v1.18) | v0.6.0 — `wrap_agent(agent, policy_bundle)` for HuggingFace smolagents 1.18+ (4th first-class framework) |\n| [**STDIO meta-guard (`mcp_stdio_meta_cve_2026_04`)**](https://www.ox.security/blog/mother-of-all-ai-supply-chains-anthropic-mcp-stdio) | v0.5.9 — bundles every airlock STDIO defence into one chain; recommended default for any MCP server registered after 2026-04-26 |\n| [**LangGraph 1.0.11 ToolNode compat shim**](https://github.com/langchain-ai/langgraph/releases/tag/prebuilt%401.0.11) | v0.5.9 — silent unwrap survives the prebuilt 1.0.11 list-vs-dict shape break |\n| [**GPT-5.5 (\"Spud\") agent defaults + tool-shape adapter**](https://openai.com/index/gpt-5-5/) | v0.5.9 — caps fan-out at 8 / context at 900k / per-call egress at 512 KB |\n| [**Capability caps (`agent_capability_default_caps`)**](https://www.anthropic.com/features/project-deal) | v0.5.9 — programmatic caps for SIGN_CONTRACT / DELEGATE_TO_AGENT / INVOKE_TOOL / WRITE_FILE / NETWORK_EGRESS |\n| [**OWASP Agentic 2026-Q1 coverage matrix**](./docs/owasp-agentic-2026-coverage.md) | v0.5.9 — 10/10 mapping risk_id → guard + preset + test, CI gate fails on stale entries |\n| [**Short-form-video corpus (`wild-2026-04/short_form_video`)**](https://www.blackhat.com/asia-26/briefings/schedule/#tiktok-agent-attacks-zhong) | v0.5.9 — 5 transcript / on-screen / RTL PoCs; `airlock replay --namespace short_form_video` |\n| [**`airlock graph serve`**](./docs/graph.md) | v0.5.9 — local web UI of the live agent → tool → MCP-server topology with verdict overlay |\n| [**`airlock policy compile / explain`**](./docs/policy-as-prompt.md) | v0.5.9 — natural-language policy authoring with hash-pinned prompt + deterministic cache |\n| [**`airlock kill-switch`**](./docs/kill-switch.md) | v0.5.9 — HMAC-signed cluster-wide freeze with 2-of-3 quorum reset |\n| [**Comment-and-Control PR-metadata guard**](https://oddguan.com/blog/comment-and-control-prompt-injection-credential-theft-claude-code-gemini-cli-github-copilot/) | v0.5.8 — neutralises CVSS 9.4 cross-vendor PR-title prompt injection |\n| [**`airlock pack`**](https://www.anthropic.com/features/project-deal) | v0.5.8 — signed policy bundles; `airlock pack install claude-code-ci@2026.04` |\n| [**`airlock baseline`**](https://venturebeat.com/security/rsac-2026-agentic-soc-agent-telemetry-security-gap) | v0.5.8 — per-agent 7-day rolling profile + drift score |\n| [**`airlock attest`**](https://www.anthropic.com/features/project-deal) | v0.5.8 — DSSE provenance per verdict |\n| [**Cloudflare Mesh compat**](https://www.cloudflare.com/press/press-releases/2026/cloudflare-launches-mesh-to-secure-the-ai-agent-lifecycle/) | v0.5.8 — runs alongside Mesh; de-duplicates overlapping policies |\n| [**Manifest-only STDIO mode**](./docs/mcp/manifest-only-mode.md) | v0.5.7 — signed-manifest registry; argv never originates from runtime input |\n| [**STDIO-taint CI gate**](./docs/security/stdio-taint-scan.md) | v0.5.7 — AST taint analyzer; flags remote→Popen flows at PR time |\n| [**Declarative preset YAML**](./docs/presets/yaml-format.md) | v0.5.7 — composite presets via stdlib-only YAML parser |\n| [**CVE-2026-30615 Windsurf zero-click**](./docs/cves/cve-2026-30615.md) | v0.5.7 — diff-on-demand mcp.json auto-load guard; **v0.8.23** adds `mcp_config_pin` — a spawn-time `{name, command, args, env-keys}` fingerprint pin (`McpConfigPinSet.check()`) that fails closed (raises, never warns) on an injected (unpinned) or mutated STDIO server even when the injection never touched a watched config file; emits on the structlog + JSON-Lines audit channels |\n| [**CVE-2026-6980 GitPilot-MCP**](./docs/cves/cve-2026-6980.md) | v0.5.7 — repo_path injection (vendor unresponsive) |\n| [**DockerBackend**](./docs/sandbox/docker.md) | v0.5.1 hardening + known gaps |\n\n### Regulatory engagement\n\n- [Public comment draft — NIST AI RMF v2.0 Agentic-AI Security](./docs/regulatory/nist-ai-rmf-v2-comment-2026.md) (window: 2026-04-18 → mid-June)\n\n---\n\n## 👤 About\n\nBuilt by [**Sattyam Jain**](https://github.com/sattyamjjain) — AI infrastructure engineer.\n\nThis started as an internal tool after watching an agent hallucinate its way through a production database. Now it's yours.\n\n---\n\n## 🤝 Contributing\n\nWe review every PR within 48 hours.\n\n```bash\ngit clone https://github.com/sattyamjjain/agent-airlock\ncd agent-airlock\npip install -e \".[dev]\"\npytest tests/ -v\n```\n\n- **Bug?** [Open an issue](https://github.com/sattyamjjain/agent-airlock/issues)\n- **Feature idea?** [Start a discussion](https://github.com/sattyamjjain/agent-airlock/discussions)\n- **Want to contribute?** [See open issues](https://github.com/sattyamjjain/agent-airlock/issues?q=is%3Aissue+is%3Aopen+label%3A%22good+first+issue%22)\n\n---\n\n## 💖 Support\n\nIf Agent-Airlock saved your production database:\n\n- ⭐ **Star this repo** — Helps others discover it\n- 🐛 **Report bugs** — [Open an issue](https://github.com/sattyamjjain/agent-airlock/issues)\n- 📣 **Spread the word** — Tweet, blog, share\n\n---\n\n## ⭐ Star History\n\n\u003cdiv align=\"center\"\u003e\n\n[![Star History Chart](https://api.star-history.com/svg?repos=sattyamjjain/agent-airlock\u0026type=Date)](https://star-history.com/#sattyamjjain/agent-airlock\u0026Date)\n\n\u003c/div\u003e\n\n---\n\n\u003cdiv align=\"center\"\u003e\n\n**Built with 🛡️ by [Sattyam Jain](https://github.com/sattyamjjain)**\n\n\u003csub\u003eMaking AI agents safe, one decorator at a time.\u003c/sub\u003e\n\n[![GitHub](https://img.shields.io/badge/GitHub-sattyamjjain-181717?style=flat-square\u0026logo=github)](https://github.com/sattyamjjain)\n[![Twitter](https://img.shields.io/badge/Twitter-@sattyamjjain-1DA1F2?style=flat-square\u0026logo=twitter\u0026logoColor=white)](https://twitter.com/sattyamjjain)\n\n\u003c/div\u003e\n\n---\n\n\u003cdiv align=\"center\"\u003e\n\u003csub\u003e\n\n**Sources:** This README follows best practices from [awesome-readme](https://github.com/matiassingers/awesome-readme), [Best-README-Template](https://github.com/othneildrew/Best-README-Template), and the [GitHub Blog](https://github.blog/open-source/maintainers/marketing-for-maintainers-how-to-promote-your-project-to-both-users-and-contributors/).\n\n\u003c/sub\u003e\n\u003c/div\u003e\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fsattyamjjain%2Fagent-airlock","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fsattyamjjain%2Fagent-airlock","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fsattyamjjain%2Fagent-airlock/lists"}